Where It Does the Work
Solving a series reasoning item runs several operations in parallel. You break a figure into its attributes and find which of them is changing, form a hypothesis about how it changes, and test that hypothesis against the next element. If the hypothesis fails, attention shifts to a different attribute. Throughout this the hypotheses you have formed must be held in mind, which is why working memory capacity bears heavily on performance. Items in which several rules operate at once are hard because the number of hypotheses to hold grows and presses against that capacity. The correlation between performance on visual reasoning and on memory tasks is thought to arise from this shared foundation.
Measurement and the Influence of Culture
Tasks intended to measure fluid intelligence are expected to depend as little as possible on language and culture. Problems built out of words turn differences in native language and schooling directly into differences in score, so visual tasks dealing with changes in attributes such as form, count, orientation, and fill have been used instead. No task is entirely free of culture, but the influence of background is smaller than it is for a vocabulary item. This property is why Raven's Progressive Matrices came to be so widely used as an index of non-verbal reasoning.
Reading Changes That Come from Training
Repeat a task in the same format and the score will rise. The procedure for breaking a figure into attributes becomes second nature and frequently recurring rules get committed to memory. Reading this as an improvement in fluid intelligence itself calls for caution. Whether the benefits of cognitive training spread beyond the trained task has been debated for a long time, and many reports find such transfer to be limited. The broad view at present is that people get better at the task they trained on while transfer to different kinds of problem is small. Treating a rising score as evidence of fluency with the task is the reasonable reading.