Rayleigh inference test generative solution method and system of multi-stage deep learning model
Through the generative solution method of Raven inference test of multi-stage deep learning model, the problem of relying on candidate answer sets and prior knowledge in the existing technology is solved, and the answer images in the Raven matrix are learned and generated without auxiliary supervision information, which improves the generalization ability and robustness of the model.
Patent Information
- Application Number
- CN202510235468.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
The generative solvers of existing Riven inference tests are insufficient in terms of interpretable properties and their variation rules in the analytical matrix, and usually rely on the candidate answer set and manually set prior knowledge, affecting the generalization ability and robustness of the model.
The generative solution method of Raven inference test using a multi-stage deep learning model is used to predict the potential concept of the target position image through the interaction between the generation process and the inference process, the implicit rules of the inference matrix, and finally generate the answer image. The model is able to learn target images at any location without auxiliary supervision information.
It realizes learning and generating answer images in the Raven matrix without auxiliary supervision information, showing excellent generative abstract reasoning ability, and improving the generalization and robustness of the model.
Smart Images

Figure CN120163247A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer abstract visual reasoning, and particularly to a method and system for generatively solving Raven's Progressive Matrices of a multi-stage deep learning model. Background Art
[0002] Abstract reasoning ability is the key to extracting potential rules from observations and quickly adapting to new situations, and is crucial for cognitive processes such as number sense, spatial reasoning, and physical reasoning. When an intelligent system processes unknown tasks, if it can possess abstract reasoning ability similar to that of humans, its adaptability and generalization ability will be significantly improved. Therefore, endowing intelligent systems with abstract reasoning ability is not only the cornerstone of building higher intelligent systems, but also an important research direction in the field of artificial intelligence for a long time.
[0003] Raven's Progressive Matrix (RPM) is a classic test for evaluating the abstract reasoning ability of humans and intelligent systems. The test requires the subject to select one from eight candidate images to fill the blank position in the lower right corner of a 3×3 image matrix. Existing research has shown that human subjects demonstrate strong reasoning ability in this test and can imagine the missing image in the matrix by inferring and combining attribute change rules, and this ability is an important part of human abstract reasoning.
[0004] To solve the answer selection problem of Raven's reasoning, many solvers fill each candidate answer into the matrix for score estimation, but these models are difficult to imagine the answer from the given context. In contrast, the answer generation task can more accurately reflect the model's understanding of potential rules. In recent years, some studies have proposed generative solvers based on neural networks, which can generate the image missing in the lower right corner of the Raven's matrix and select the final answer by comparing the generated image with the candidate answers. However, such generative solvers perform poorly in parsing the interpretable attributes and their change rules in the Raven's problem matrix, and usually introduce artificially set prior knowledge in the representation learning or abstract reasoning process. In addition, most existing generative solvers rely on the candidate answer set and the rule annotations provided in the training set for auxiliary training, and this method may lead to potential shortcut learning risks, thus affecting the generalization ability and robustness of the model.
[0005] Deep latent variable models (DLVMs) can capture the underlying structure of noisy observed data through an interpretable latent space. Previous studies have addressed the generative matrix reasoning problem by treating attributes and their change rules as latent concepts, which can generate answers by performing a prediction process for specific attributes. By considering the conditional answer generation process for the underlying structure of matrix reasoning images, deep latent variable models can be trained without relying on distractors. Although previous studies have achieved answer generation in matrix reasoning problems with continuous attributes, it remains challenging for deep latent variable models to understand complex discrete rules and abstract attribute change rules in real-world datasets. Summary of the Invention
[0006] Based on the technical problems existing in the background art, the present invention proposes a method and system for generatively solving Raven's Progressive Matrices (RPM) using a multi-stage deep learning model to predict the latent concepts of the target location image.
[0007] The method for generatively solving Raven's Progressive Matrices using a multi-stage deep learning model proposed by the present invention inputs the RPM problem matrix into a trained generative solution model and outputs a predicted answer image for the RPM problem matrix.
[0008] The training process of the generative solution model is as follows:
[0009] Obtain the RPM problem matrix and candidate answers and perform feature extraction on both to obtain the context image and the latent concepts of the candidate answers.
[0010] Parse the latent concepts of the context image into the rules of the target concept through a generative process rule parser, and parse the latent concepts of the context image and the target answer into the rules of the target concept through an inference process rule parser, and the two rules tend to be consistent.
[0011] Input the predicted rules output by the generative process rule parser and the latent concepts of the context image into a generative process target predictor to obtain a prediction of the latent concepts of the target answer.
[0012] Input the latent concepts of the target answer predicted by the generative process into a decoder to generate a predicted target image with the same size as the original input image.
[0013] Compare the latent concepts of the target answer predicted by the generative process with the latent concepts of the candidate answers, and select the candidate answer that is closest to the latent concepts of the target answer as the predicted answer for the RPM problem matrix.
[0014] Set an optimizer to iteratively train the learnable parameters in the generative solution model.
[0015] Further, the RPM problem matrix includes N images, where N - 1 are context images and 1 is an unknown target image;
[0016] The candidate answers include a target image and multiple interference images, and the target image is the answer image of the unknown target image in the RPM problem matrix.
[0017] Further, among the rules for parsing potential concepts into target concepts by the rule parser, it specifically includes:
[0018] The rule parser constructs a 3×3 feature matrix for the obtained potential concepts, and through two convolutional layers and respectively solve the representations in the row direction and column direction of the feature matrix;
[0019] After connecting the representations in the row direction and column direction and passing through a convolutional layer obtain the rule representation of the overall matrix as the prediction of the implicit rule information;
[0020] Process the unknown concepts of the unknown target image in the generated RPM problem matrix by means of zero padding.
[0021] Further, when inputting the potential concepts of the context images and the prediction output by the generative process rule parser into the generative process target predictor to obtain the prediction of the potential concepts of the target answer, it specifically includes:
[0022] Error calculation: Connect the potential concepts of the context images and the implicit rule information and obtain the prediction error through convolutional operation;
[0023] Convolutional residual calculation: Connect the potential concepts of the context images and the prediction error and input them into two convolutional layers, and add the potential concepts of the context images and the result after convolution through residual connection to obtain the first prediction result;
[0024] Repeat the above error calculation and convolutional residual calculation, and use the obtained second prediction result as the prediction of the potential concepts of the target answer.
[0025] Further, when comparing the potential concepts of the target answer with the potential concepts of the candidate answers and selecting the candidate answer closest to the potential concepts of the target answer as the predicted answer of the RPM problem matrix, specifically:
[0026] Calculate the mean square error between the potential concepts of the target answer and the potential concepts of the candidate answers, and use the candidate answer with the minimum mean square error as the predicted answer of the unknown target image in the RPM problem matrix.
[0027] Further, set a total loss function to train the generative solution model, and the construction process of the total loss function is as follows:
[0028]
[0029] Among them, is the total loss function, is the loss of the reconstructed quality of the target image, is the loss of measuring the reconstructed quality of the context image, is the regularization term of the rule, is the target regularization term, β rc , β r , β t , β cls are hyperparameters, is the regularization term of rule parsing.
[0030] Furthermore, the generation process of the generative solution model interacts with the inference process, infers the implicit rules of the RPM problem matrix, predicts the potential concepts of the target image, and finally generates the target answer image.
[0031] The generative solution system of the Raven's Progressive Matrices (RPM) problem in the multi-stage deep learning model inputs the RPM problem matrix into the trained generative solution model and outputs the predicted answer image of the RPM problem matrix;
[0032] The training process of the generative solution model includes an encoder, a rule parser in the generation process, a rule parser in the inference process, a target predictor in the generation process, a decoder, an answer selector, and an optimizer;
[0033] The encoder is used to extract features from the RPM problem matrix and candidate answers to obtain the potential concepts of the context image and candidate answers;
[0034] The rule parser in the generation process is used to parse the potential concepts of the context image into the rules of the target concept;
[0035] The rule parser in the inference process is used to parse the potential concepts of the context image and the target answer into the rules of the target concept, and the predicted rules output by the rule parser in the generation process tend to be consistent with the predicted rules output by the rule parser in the inference process;
[0036] The target predictor takes the predicted rules output by the rule parser in the generation process and the potential concepts of the context image as inputs to obtain the prediction of the potential concepts of the target answer;
[0037] The decoder takes the potential concepts of the target answer predicted in the generation process as inputs and generates a predicted target image with the same size as the original input image;
[0038] The answer selector is used to compare the potential concepts of the target answer with the potential concepts of the candidate answers, and select the candidate answer that is closest to the potential concept of the target answer as the predicted answer for the RPM question matrix;
[0039] The optimizer is used to iteratively train the learnable parameters in the generative solution model.
[0040] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the Ravens Progressive Matrices generative solution method of the multi-stage deep learning model as described above.
[0041] A computer-readable storage medium stores a number of classification programs, and the number of classification programs are used to be called by a processor and execute the Ravens Progressive Matrices generative solution method of the multi-stage deep learning model as described above.
[0042] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0043] Existing RPM generative solution models generally have several limitations: their training processes often rely on auxiliary supervision information, such as rule annotations or candidate answer sets; they are mainly applicable to processing simple denoising RPM problems; they cannot automatically learn interpretable concepts and global rules, etc. The advantages of the Ravens Progressive Matrices generative solution method and system of the multi-stage deep learning model provided by the present invention are as follows: This embodiment aims to solve the deficiencies of the existing RPM generative solver model technology. Through the interaction between the generation process and the inference process, infer the implicit rules of the inference matrix, predict the potential concepts of the target position image, and finally generate the answer image. The generative solution model can learn the target image at any position without auxiliary supervision information, demonstrating excellent generative abstract reasoning ability. Description of the Drawings
[0044] Figure 1 It is a schematic structural diagram of an RPM question matrix;
[0045] Figure 2 It is a schematic framework diagram of a generative solution model;
[0046] Figure 3 It is a schematic structural diagram of a rule parser;
[0047] Figure 4It is a structural diagram of a target predictor;
[0048] Figure 5 It is a schematic diagram of the generation result of the target image. Specific implementation manner
[0049] Next, the technical solution of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0050] As Figures 1 to 5 shown, in the Raven's Progressive Matrices (RPM) generative solution method of the multi-stage deep learning model proposed by the present invention, the RPM problem matrix is input into the trained generative solution model to output the predicted answer of the RPM problem matrix.
[0051] Existing RPM generative solution models generally have several limitations: their training processes often rely on auxiliary supervision information, such as regular annotations or candidate answer sets; they are mainly applicable to processing simple denoising RPM problems; they cannot automatically learn interpretable concepts and global rules, etc. This embodiment aims to solve the deficiencies of the existing RPM generative solver model technology. By interacting the generation process and the inference process, the implicit rules of the inference matrix are inferred, the potential concepts of the target position image are predicted, and finally the answer image is generated. The generative solution model can learn the target image at any position without auxiliary supervision information, demonstrating excellent generative abstract reasoning ability.
[0052] The generation and inference processes of the generative solution model are as follows:
[0053] (1) Initial feature encoding;
[0054] The RPM problem matrix and candidate answers are obtained and feature extraction is performed on both to obtain the context image and the potential concepts of the candidate answers.
[0055] The RPM problem matrix includes N images, where N - M are context images and 1 is an unknown target image; the candidate answers include one target image and multiple interference images, and this target image is the answer image of the unknown target image in the RPM problem matrix.
[0056] Taking the RPM problem matrix with 9 images as an example for illustration, in the feature encoding process, the encoder uses a total of 9 images, namely 8 context images and 1 unknown target image in the input RPM problem matrix, to learn the potential concepts of the matrix and the attribute features possessed by the concepts corresponding to each image.
[0057] The encoder is mainly composed of multiple convolutional layers combined. Each convolutional layer contains a convolutional kernel, a batch normalization layer, and a ReLU activation function. The convolutional operation plays a key role in this process. It gradually extracts the features of the image and aggregates semantic information to achieve a high-level abstraction of the input data. The batch normalization operation standardizes the data, restricting it within a specific range, aiming to eliminate the influence of the dimension and unit of measurement of the data on the model performance, thereby improving the generalization ability of the model. The introduction of the ReLU activation function provides the necessary non-linear characteristics for the neural network, enabling the network to capture and learn more complex data patterns and function mappings.
[0058] (2) The inference process learns the implicit rule information of each concept;
[0059] The inference process uses the information of the entire RPM problem matrix extracted in step (1) to parse the rules of each concept. The variational distribution q(h|x) derives the latent variables from the context image and the target image. The latent variables contain the latent concepts extracted by the encoder and the latent rule information parsed by the rule parser:
[0060]
[0061] where x represents the input image, h is the latent variable, x n is the nth image, represents the mth latent concept of the nth image in the RPM problem matrix, z m is the mth latent concept in the RPM problem matrix, r m represents the rule information corresponding to the mth latent concept, and M is the number of all latent concepts. represents the process of converting the image into latent concepts by the encoder in step (1). Connect the concepts of different images, and q(r m |z m ) parses the concept-specific rules through the rule parser.
[0062] The structure of the rule parser is as Figure 3 shown. First, the rule parser constructs a 3×3 feature matrix Z m for each latent concept. Through two convolutional layers and respectively solve the row-wise and column-wise representations:
[0063]
[0064] where, through represents the mth latent concept of the image in the ith row and jth column, and its physical meaning is basically the same as that of , both representing the mth latent feature of a certain image in the RPM problem matrix. is the representation of the i-th row, is the representation of the j-th column.
[0065] Take the mean of the row representations and the mean of the column representations to obtain an aggregated representation of the row-column rule. Then concatenate and through a convolutional layer to obtain the rule representation of the overall matrix.
[0066]
[0067] Among them, is the mean value of the rule representation of the overall matrix, is the standard deviation of the rule representation of the overall matrix.
[0068] Since the inference process directly obtains the concept of the target image from the target image, the inference process does not involve the part of target prediction.
[0069] (3) Generate the implicit rule information of the process prediction matrix
[0070] The process of generating the implicit rule information of the process prediction matrix is similar to step (2). The variational distribution p(h|x) derives the latent variable only from the context image of the RPM problem matrix:
[0071]
[0072] Among them, x represents the input context image, h is the latent variable, and x c is the c-th image in the context, represents the m-th latent concept among all the context images of the RPM problem matrix, C is the number of context images of the RPM problem matrix, and r m represents the rule information corresponding to the m-th latent concept, and M is the number of all latent concepts. represents the process of converting the image into a latent concept by the encoder in step (1). Different from the inference process, the generation process only encodes the latent concepts of the context. Since the inference process and the generation process use the same encoder, p(z|x) and q(z|x) are the same distribution. Concatenate the concepts of different images, through Figure 3 the rule parser shown to parse the concept-specific rules, and use zero-padding to handle the unknown concepts of the unknown target images.
[0073] (4) Generate the latent concepts of the target image in the generation process;
[0074] Take the latent concepts of the context images extracted in step (1) and the matrix implicit rule information predicted in step (3) as the input of the target predictor network structure, and output the prediction of the latent concepts of the target answer. The target predictor designed in this embodiment draws on the concept of prediction error in neuroscience. In the RPM task, the subject should first examine all 8 context images to learn the implicit rules, then predict what attributes the unknown target image should have based on the learned rules, and then compare the actual target latent concepts with them to calculate the prediction error. Based on this, the predicted latent concepts of the target answer are continuously updated. The prediction error can further refine the learned rules, and the learned rules and predicted concepts will be continuously updated in this iteration. In this case, the prediction error is the key clue to correct reasoning.
[0075] The structure of the target predictor is as Figure 4 shown. First, convert the input concept tensor of size [9, M×D] into a tensor of size [M×D, 3, 3], where M is the number of latent concepts and D is the dimension of each latent concept vector. The unknown concepts at the target positions are processed by zero-padding, which allows the use of 2D convolution operations to predict the concepts of the target according to the rules of rows and columns. Specifically, first, the latent concepts of the input context images and the matrix implicit rule information r m are concatenated to obtain a new tensor The prediction error is calculated by the following formula:
[0076]
[0077] where Conv represents the convolution operation, BN is the batch normalization operation, ReLU is the activation function, and PE m represents the prediction error between the latent concepts of the target answer predicted from the context and the initially predicted target latent concepts . Then, connect the prediction error PE m with the latent concepts to obtain a new tensor and input the combined new tensor into another two convolutional layers and add the original concept Z m to the result of the convolution with a residual connection. The specific process is as follows:
[0078]
[0079] However, the process of single prediction and matching cannot accurately identify the latent concepts of the target answer. Therefore, the target predictor repeats this process twice and takes the concept vector output in the second time as the latent concepts of the target answer. Such a process allows the target predictor to gradually test and refine the predicted target concepts.
[0080] (5) Generation process: Generate the target image;
[0081] Use the target answer latent concepts predicted in step (4) as the input of the decoder. Through a series of upsampling and feature fusion operations, an image with the same size as the original input image is reconstructed. The decoder structure is mainly composed of a series of transposed convolutional layers cascaded. Each transposed convolutional layer consists of a transposed convolutional kernel, a batch normalization layer, and a LeakyReLU activation function. The transposed convolutional operation, also known as deconvolution, its core role is to upsample the feature map output by the encoder to restore the spatial dimension of the image. This process not only helps to reconstruct the detailed information of the image but also can retain the deep features extracted in the encoding stage to a certain extent. The LeakyReLU activation function is an improvement of ReLU. It allows negative input values to have non-zero gradients, thus alleviating the problem of "dead neurons" in ReLU, that is, when the input is negative, ReLU outputs zero, resulting in some neurons being unable to update their weights. LeakyReLU provides a non-zero slope for negative input values, enabling the network to still learn when facing negative inputs, thereby enhancing the flexibility and robustness of the model. In this way, the decoder can more effectively map the compressed features of the encoder back to the high-resolution output space to achieve accurate image reconstruction or generation tasks.
[0082] (6) Answer selection;
[0083] Use the encoder in step (1) to extract the latent concepts of all candidate answers. Compare the latent concepts of the target image predicted in the generation process with the latent concepts of the candidate answers, and select the candidate answer with the smallest mean squared error (MSE) between the latent concepts of the predicted target image as the predicted answer of the question matrix:
[0084]
[0085] where M represents the number of latent concept vectors, m is the index of the latent concept vector, taking values between 1 and M, represents the predicted target answer latent concept, z d represents the latent concept of the d-th candidate answer.
[0086] (7) Training of the generative solution model;
[0087] This model is based on the VAE (Variational Autoencoder) theory. The optimization goal is to maximize the evidence lower bound (ELBO), which can be expressed as:
[0088]
[0089] where:
[0090]
[0091] Among them, is the total loss function, is the loss of the target image reconstruction quality, is the loss for measuring the context image reconstruction quality, is the regularization term of the rule, is the target regularization term, β rc , β r , β t , β cls are hyperparameters, is the regularization term of rule parsing.
[0092] Measures the quality of the target image reconstruction, As a measure of the quality of the context image reconstruction, it enables the VAE to consider more samples and assist in the training of the generative solution model to improve its feature extraction ability. The regularization term of the rule Indicates the conceptual consistency of the parsing rule through the Kullback-Leibler (KL) divergence between the distribution q(r m |z m ) and . Minimizing encourages the generative solution model to infer the same rule on different context images of the RPM. The target regularization term Measures the distance between the latent concept encoded from the target image (this is the target image latent concept obtained through step (2)) and the target answer latent concept predicted from the context (this is the target latent concept obtained through step (4)). The regularization term of rule parsing will guide the distribution q(r m ) to approach the prior p(r m ). Hyperparameters β rc , β r , β t , β cls are introduced into the loss function to control the importance of each regularization term.
[0093] (8) Model estimation;
[0094] In this embodiment, the RAVEN dataset (a relational and analogical visual reasoning dataset) is used to evaluate the generative solving model. RAVEN is a structured RPM problem dataset that contains 70,000 problem sets. Among them, the training set contains 42,000 problem sets, and the validation set and the test set each contain 14,000 problem sets. Each problem set includes a problem image of a 3×3 matrix, where the last image is missing, and eight candidate answer images. The dataset is evenly distributed in 7 configurations: Center (C), 2*2Grid, 3*3Grid, Left-Right (L-R), Up-Down (U-D), Out-InCenter (O-IC) and Out-InGrid (O-IG). Each problem contains six visual attributes (angle, quantity, position, type, size, and color) and four potential rules (constant, progressive, arithmetic, and distribution three). In addition, to increase the challenge of the problems, the dataset introduces additional noise in the attributes.
[0095] Table 1 below shows the comparison results of the accuracies of different models on the RAVEN dataset. The generative solving model of this embodiment has the highest average accuracy among all models trained without auxiliary information, and is only slightly lower than the current optimal results in the 2×2Grid and 3×3Grid graphic configurations, where Grid is the grid.
[0096] Table 1
[0097]
[0098] Figure 5 Shows the answer generation results at any position on the RAVEN dataset in this embodiment. It can be clearly seen from the figure that this embodiment demonstrates strong answer generation capabilities in most graphic configurations. It should be noted that for the 2×2Grid, 3×3Grid, and O-IG graphic configurations, although there are differences between the generated images and the actual correct answers, it can be clearly seen that the present invention has captured the logical rules implicit in the matrix. For some noise attributes, the present invention cannot guarantee that the generated results are consistent with the actual answers, such as the shape attribute in the 2×2Grid and 3×3Grid example diagrams and the graphic position attribute in the O-IG example diagram.
[0099] Therefore, the RPM generative solving model based on the variational autoencoder proposed in this embodiment aims to solve RPM problems containing noise interference, can generate images with missing positions at any location without auxiliary supervision, and achieves an accuracy comparable to that of models trained under auxiliary supervision. In addition, the model can abstract global rules from the encoded latent concepts, thereby achieving higher accuracy and robustness in the RPM task.
[0100] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A generative solution method for the Raven's Reasoning Test based on a multi-stage deep learning model, characterized by: Input the RPM question matrix into the trained generative solving model and output the predicted answer image of the RPM question matrix; The training process of the generative solution model is as follows: Obtain the RPM question matrix and candidate answers and perform feature extraction on them to obtain the context image and potential concepts of the candidate answers; The rule parser of the generation process parses the latent concepts of the context image into the rules of the target concept, and the rule parser of the inference process parses the latent concepts of the context image and the target answer into the rules of the target concept, and the rules of the two tend to be consistent; The prediction rules output by the generation process rule parser and the latent concepts of the context image are input into the generation process target predictor to obtain the prediction of the latent concept of the target answer; The target answer latent concept predicted by the generation process is input into the decoder to generate a predicted target image with the same size as the original input image; Compare the target answer potential concept predicted by the generation process with the potential concept of the candidate answer, and select the candidate answer that is closest to the target answer potential concept as the predicted answer of the RPM question matrix; Set up an optimizer to iteratively train the learnable parameters in the generative solver model.
2. The method for solving the Raven's Reasoning Test using a multi-stage deep learning model according to claim 1, characterized in that: The RPM problem matrix includes N images, including N-1 context images and 1 unknown target image; The candidate answers include a target image and multiple interference images, and the target image is the answer image of the unknown target image in the RPM problem matrix.
3. The method for solving the Raven's Reasoning Test using a multi-stage deep learning model according to claim 1, characterized in that: In the rules for parsing potential concepts into target concepts through the rule parser, specifically including: The rule parser constructs a 3×3 feature matrix for the latent concept obtained, passing through two convolutional layers and Solve the row-wise and column-wise representations of the characteristic matrix respectively; The row and column representations are connected and passed through the convolution layer Obtain the rule representation of the overall matrix as a prediction of implicit rule information; Zero filling is used to handle the unknown concept of unknown target images in the RPM problem matrix of the generation process.
4. The method for solving the Raven's Reasoning Test generative formula using a multi-stage deep learning model according to claim 1, characterized in that: The latent concepts of the context image and the predictions output by the generation process rule parser are input into the generation process target predictor to obtain the predictions of the latent concepts of the target answer, specifically including: Error calculation: The potential concepts and implicit rule information of the context image are connected and then the prediction error is obtained through convolution operation; Convolution residual calculation: The potential concept of the context image is connected to the prediction error and then input into two convolution layers. The potential concept of the context image is added to the convolution result through the residual connection to obtain the first prediction result. Repeat the above error calculation and convolution residual calculation, and use the second prediction result as the prediction of the potential concept of the target answer.
5. The method for solving the Raven's Reasoning Test generative formula using a multi-stage deep learning model according to claim 1, characterized in that: Compare the potential concept of the target answer with the potential concept of the candidate answer, and select the candidate answer that is closest to the potential concept of the target answer as the predicted answer in the RPM question matrix, specifically: The mean square error between the target answer potential concept and the candidate answer potential concept is calculated, and the candidate answer with the minimum mean square error is used as the predicted answer for the unknown target image in the RPM problem matrix.
6. The method for solving the Raven's Reasoning Test generative formula using a multi-stage deep learning model according to claim 1, characterized in that: Set the total loss function to train the generative solver model. The total loss function construction process is as follows: in, is the total loss function, Reconstruct quality loss for the target image, To measure the quality loss of context image reconstruction, is the rule regularization term, is the target regularization term, β rc ,β r ,β t ,β cls is a hyperparameter, Regularization term for rule parsing.
7. The method for solving the Raven's Reasoning Test generative formula using a multi-stage deep learning model according to claim 1, characterized in that: The generation process of the generative solving model interacts with the inference process to infer the implicit rules of the RPM problem matrix, predict the latent concepts of the target image, and finally generate the target answer image.
8. A generative solution system for the Raven's Reasoning Test based on a multi-stage deep learning model, characterized by: Input the RPM question matrix into the trained generative solving model and output the predicted answer image of the RPM question matrix; The training process of the generative problem-solving model includes an encoder, a generative process rule parser, an inference process rule parser, a generative process target predictor, a decoder, an answer selector, and an optimizer; The encoder is used to extract features from the RPM question matrix and candidate answers to obtain the potential concepts of the context image and candidate answers; The generation process rule parser is used to parse the latent concepts of the context image into the rules of the target concept; The inference process rule parser is used to parse the potential concepts of the context image and the target answer into the rules of the target concept, and the prediction rules output by the generation process rule parser tend to be consistent with the prediction rules output by the inference process rule parser; The target predictor takes the prediction rules output by the generative process rule parser and the matrix concept of the context image as input to obtain the prediction of the potential concept of the target answer; The decoder takes as input the target answer latent concept predicted by the generation process and generates a predicted target image of the same size as the original input image; The answer selector is used to compare the potential concept of the target answer with the potential concept of the candidate answer, and select the candidate answer that is closest to the potential concept of the target answer as the predicted answer of the RPM question matrix; The optimizer is used to iteratively train the learnable parameters in the generative solver model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the generative solution method of the Raven's Reasoning Test based on the multi-stage deep learning model as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of classification programs, which are used to be called by a processor and execute the Raven's Reasoning Test generative solution method of the multi-stage deep learning model as described in any one of claims 1-7.