Image recognition method, apparatus and device, storage medium and product
Through the multimodal alignment network combined with visual and language processing models, the problem of insufficient universality of existing image recognition models in new category recognition is solved, and efficient image recognition universality and adaptability are achieved.
Patent Information
- Application Number
- PCT/CN2024/141101
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-17
AI Technical Summary
Existing image recognition models are less versatile in category recognition that are not clearly defined in training data, and require re-collecting data and annotation, resulting in high time and labor costs and inability to adapt to business rules changes.
By inputting the image to be identified, problem text and description rule text into the multimodal alignment network completed by training, the image feature vector is obtained using the visual processing model, and the text feature vector is obtained using the language processing model, and the target label is determined based on the setting query category, so as to avoid re-collecting sample data to train a new recognition model.
It improves the versatility of image recognition, reduces the time and labor cost for new categories, and enhances the adaptability of the model.
Smart Images

Figure CN2024141101_17072025_PF_FP_ABST
Abstract
Description
Image recognition method, device, equipment, storage medium and product
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 11, 2024, with application number 202410048426.8, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of image processing technology, and in particular to an image recognition method, apparatus, device, storage medium, and product. Background Art
[0003] At present, the review of images is generally carried out through the recognition results of image recognition models, that is, image recognition is performed by image recognition models trained through supervised learning based on a large number of samples and annotated labels.
[0004] Image recognition models require a large amount of training data to ensure good generalization and practical application value. They must obtain the desired positive examples from a large amount of data. When applied to large-scale audit data, these positive examples often require hundreds of thousands or even tens of thousands of cumulative examples. However, for identifying categories that are not clearly defined in the training data, the current common solution is to manually collect more data on the desired categories to alleviate the data shortage, while further training a separate image recognition model to effectively identify the new categories. However, retraining an image recognition model is extremely time-consuming, and because audit rules change with business and scale, data recollection and annotation are often required, resulting in low versatility in image recognition. Summary of the Invention
[0005] The embodiments of the present application provide an image recognition method, apparatus, device, storage medium and product to solve the technical problem of low versatility of existing image recognition solutions and effectively improve the versatility of image recognition.
[0006] In a first aspect, an embodiment of the present application provides an image recognition method, comprising:
[0007] Obtaining an image to be recognized, a question text, and a description rule text, wherein the question text and the description rule text are configured based on a set query category;
[0008] The image to be identified, the question text and the description rule text are input into a trained multimodal alignment network, and the image feature vector of the image to be identified is obtained by using a visual processing model through the multimodal alignment network, and the text feature vectors of the question text and the description rule text are obtained by using a language processing model, and the target label corresponding to the image to be identified in the set query category is determined by using the language processing model based on the image feature vector and the text feature vector.
[0009] In a second aspect, an embodiment of the present application provides an image recognition device, including an information acquisition module and an image recognition module, wherein:
[0010] The information acquisition module is configured to acquire an image to be recognized, a question text, and a description rule text, wherein the question text and the description rule text are configured based on a set query category;
[0011] The image recognition module is configured to input the image to be recognized, the question text and the description rule text into a trained multimodal alignment network, obtain the image feature vector of the image to be recognized by using a visual processing model through the multimodal alignment network, and obtain the text feature vectors of the question text and the description rule text by using a language processing model, and use the language processing model to determine the target label corresponding to the image to be recognized in the set query category based on the image feature vector and the text feature vector.
[0012] In a third aspect, an embodiment of the present application provides an image recognition device, including: a memory and one or more processors;
[0013] The memory is used to store one or more programs;
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the image recognition method as described in the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a non-volatile storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the image recognition method as described in the first aspect.
[0016] In the fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads and executes the computer program from the computer-readable storage medium, so that the device performs the image recognition method described in the first aspect.
[0017] The embodiment of the present application inputs the image to be identified, the question text and the description rule text into a trained multimodal alignment network, obtains the image feature vector of the image to be identified by using a visual processing model through the multimodal alignment network, and obtains the text feature vectors of the question text and the description rule text by using a language processing model, and determines the target label corresponding to the image to be identified in the set query category based on the image feature vector and the text feature vector using the language processing model. There is no need to re-collect sample data to train a new recognition model. The target label corresponding to the image to be identified in the set query category can be obtained by configuring the question text and the description rule text based on the set query category, thereby effectively improving the versatility of image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] FIG1 is a flow chart of an image recognition method provided by an embodiment of the present application;
[0019] FIG2 is a flow chart of another image recognition method provided by an embodiment of the present application;
[0020] FIG3 is a schematic diagram of a neural network structure based on a self-attention mechanism provided in an embodiment of the present application;
[0021] FIG4 is a schematic diagram of data flow during the training process of a multimodal alignment network provided in an embodiment of the present application;
[0022] FIG5 is a schematic diagram of data flow during the training process of a multimodal alignment network provided in an embodiment of the present application;
[0023] FIG6 is a schematic structural diagram of an image recognition device provided in an embodiment of the present application;
[0024] FIG7 is a schematic structural diagram of an image recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only some, but not all, of the contents related to the present application are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The above process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The above process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0026] The image recognition method provided in this application can be applied to image review and labeling scenarios. It aims to configure question text and description rule text based on the set query category, analyze and process the image to be recognized, question text and description rule text through the visual processing model and language processing model in the multimodal alignment network, and obtain the target label corresponding to the set query category of the image to be recognized, thereby effectively improving the versatility of image recognition.
[0027] In existing image recognition solutions, when it is necessary to identify categories that have not been clearly defined in the training data, it is generally necessary to re-collect positive sample data of the category and re-train a new recognition model, that is, to re-collect data and label it. The process of collecting and labeling data often consumes a huge amount of time and manpower. Collecting target data from raw data requires a very large manpower cost. At the same time, since there is no dedicated algorithm model to identify the target category, it is necessary to further develop a new visual recognition depth model, and the image recognition versatility is poor. Based on this, an image recognition method of an embodiment of the present application is provided to solve the technical problem of poor versatility of image recognition in existing image recognition solutions.
[0028] Figure 1 shows a flowchart of an image recognition method provided in an embodiment of the present application. The image recognition method provided in an embodiment of the present application can be executed by an image recognition device, which can be implemented by hardware and / or software and integrated into an image recognition device.
[0029] The following description is based on an example of an image recognition method performed by an image recognition device. Referring to FIG1 , the image recognition method includes:
[0030] S110: Obtain an image to be recognized, a question text, and a description rule text, wherein the question text and the description rule text are configured based on a set query category.
[0031] For example, an image to be identified that needs to be reviewed or labeled, as well as a question text and a description rule text corresponding to the image to be identified, are obtained. The solution may provide one or more images to be identified, and the question text and description rule text corresponding to each image to be identified may be the same or different.
[0032] In one embodiment, the question text and description rule text provided by this solution are configured based on a set query category. The question text can be used to describe the type of answer that the multimodal alignment network needs to output. For example, the question text can be used to inquire about the label corresponding to the to-be-identified object, or the label corresponding to a specific person or object in the image to be identified. The description rule text can be used to describe the form of the answer output by the multimodal alignment network.
[0033] S120: Input the image to be identified, the question text, and the description rule text into the trained multimodal alignment network, obtain the image feature vector of the image to be identified by using the visual processing model through the multimodal alignment network, and obtain the text feature vectors of the question text and the description rule text by using the language processing model, and use the language processing model to determine the target label corresponding to the image to be identified in the set query category based on the image feature vector and the text feature vector.
[0034] Exemplarily, the image to be identified, the question text, and the description rule text are input into the trained multimodal alignment network, and the multimodal alignment network analyzes and processes the image to be identified, the question text, and the description rule text and outputs the target label corresponding to the set query category of the image to be identified.
[0035] The multimodal alignment network provided by this solution is configured with a trained visual processing model and a language processing model. After receiving the image to be identified, the question text, and the description rule text, the multimodal alignment network uses the visual processing model to obtain the image feature vector of the image to be identified, and uses the language processing model to obtain the text feature vectors of the question text and the description rule text. The language processing model is then used to determine the target label corresponding to the image to be identified in the set query category based on the image feature vector and the text feature vector.
[0036] Optionally, the visual processing model provided by this solution can be a large vision model (LVM), and the language processing model provided by this solution can be a large language model (LLM). The visual processing model and language processing model provided by this solution can be built based on a Transformer network (a neural network based on a self-attention mechanism), an RNN network (recurrent neural network), or a CNN network (convolutional neural network).
[0037] In the above, by inputting the image to be identified, the question text and the description rule text into the trained multimodal alignment network, the multimodal alignment network uses the visual processing model to obtain the image feature vector of the image to be identified, and uses the language processing model to obtain the text feature vectors of the question text and the description rule text, and uses the language processing model to determine the target label corresponding to the image to be identified in the set query category based on the image feature vector and the text feature vector. There is no need to re-collect sample data to train a new recognition model. The target label corresponding to the image to be identified in the set query category can be obtained by configuring the question text and the description rule text based on the set query category, thereby effectively improving the versatility of image recognition.
[0038] Based on the above embodiment, FIG2 shows a flowchart of another image recognition method provided by an embodiment of the present application. The image recognition method is a specific embodiment of the above image recognition method. Referring to FIG2, the image recognition method includes:
[0039] S210: Acquire an image-text data pair, where the image-text data pair includes a sample image, a sample description text describing the sample image, and sample object coordinate information.
[0040] In one embodiment, before using a multimodal alignment network to determine the target label corresponding to the set query category of the image to be identified, image-text data pairs can be collected first, and the multimodal alignment network can be optimized and trained using the image-text data pairs. The optimized multimodal alignment network can then be configured in the image recognition device so that the image recognition device can determine the target label corresponding to the set query category of the image to be identified based on the trained multimodal alignment network.
[0041] For example, the data sources for the audit business can be organized, and the sample images corresponding to the data sources can be annotated, such as the sample description and object coordinates in the sample images. The corresponding sample images and the corresponding annotated data such as the sample description and object coordinates can be used as image-text data pairs. An image-text data pair includes a sample image, a sample description text describing the sample image, and sample object coordinate information (i.e., the annotated sample description and object coordinates).
[0042] Optionally, corresponding sample question-and-answer data can be configured for each image-text pair as a supervisory signal during the multimodal alignment network training process. The sample question-and-answer data includes sample questions about the content in the sample image and sample answers to the sample questions based on the sample image. Optionally, the sample answers in the sample question-and-answer data can serve as the true labels corresponding to the sample images.
[0043] S220: Using image-text data to train a multimodal alignment network, wherein the multimodal alignment network uses a visual processing model to obtain a sample image vector of a sample image, and uses a language processing model to obtain a sample text vector of sample description text and sample object coordinate information, and uses the language processing model to determine a predicted label corresponding to the sample image based on the sample image vector and the sample text vector, and optimizes network parameters of the visual processing model and the language processing model based on the predicted label and the true label corresponding to the sample image vector.
[0044] Exemplarily, the collected image-text data pairs are used to train the multimodal alignment network, wherein during the training process, the multimodal alignment network provided by the present solution can process the image-text data pairs through the visual processing model and the language processing model configured by the multimodal alignment network. For example, the multimodal alignment network analyzes and processes the sample images in the image-text data pairs through the visual processing model to determine the sample image vector corresponding to the sample image, and analyzes and processes the sample description text and sample object coordinate information in the image-text data pairs through the Q-Benz model to determine the sample text vector of the sample description text and the sample object coordinate information.
[0045] In one embodiment, after obtaining a sample image vector and a sample text vector, the sample image vector and the sample text vector are analyzed and processed using a language processing model to obtain a predicted label corresponding to the sample image, and network parameters of the visual processing model and the language processing model are optimized based on the predicted label and the true label corresponding to the sample image vector (for example, the sample answer data in the sample question and answer data, and the true label that is separately annotated). This solution uses image-text data to train a multimodal alignment network, and uses a large amount of image-text data in business scenarios to obtain a multimodal alignment network that semantically aligns text and image information. This enables the multimodal alignment network to have the ability to recognize images under specified questions and description rules, without the need to collect sample data and train new recognition models separately for different recognition types, effectively improving the versatility of image recognition.
[0046] In one possible embodiment, the multimodal alignment network provided by this solution, when using a visual processing model to obtain a sample image vector of a sample image, and using a language processing model to obtain a sample text vector of sample description text and sample object coordinate information, and using the language processing model to determine a predicted label corresponding to the sample image based on the sample image vector and the sample text vector, includes:
[0047] S221: Obtain a sample image vector of the sample image using the image encoding module in the visual processing model.
[0048] S222: Utilize the text decoding module in the language processing model to obtain a sample text vector of the sample description text and the sample object coordinate information.
[0049] S223: Analyze and process the sample image vector and the sample text vector using the prediction output layer in the language processing model to obtain a prediction label corresponding to the sample image.
[0050] Exemplarily, when the multimodal alignment network uses a visual processing model to obtain a sample image vector of a sample image, it may use an image encoding module in the visual processing model to analyze and process the sample image and obtain a sample image vector of the sample image.
[0051] When the multimodal alignment network uses a language processing model to obtain sample text vectors of sample description text and sample object coordinate information, it can use the text decoding module in the language processing model to analyze and process the sample description text and sample object coordinate information, and obtain sample text vectors of the sample description text and sample object coordinate information.
[0052] After determining the sample image vector and sample text vector, the sample image vector and sample text vector are input into the prediction output layer of the language processing model. The prediction output layer of the language processing model is used to analyze and process the sample image vector and sample text vector to obtain the predicted label corresponding to the sample image. This solution uses the image encoding module in the visual processing model to obtain the sample image vector of the sample image and the text decoding module in the language processing model to obtain the sample text vector of the sample description text and sample object coordinate information. The prediction output layer of the language processing model is used to analyze and process the sample image vector and sample text vector to obtain the predicted label. This allows the multimodal alignment network to accurately learn the ability to recognize images under specified problems and description rules, effectively improving the versatility of image recognition.
[0053] Figure 3 shows a schematic diagram of a neural network structure based on a self-attention mechanism. Taking the self-attention neural network as an example, the self-attention neural network is a 12-layer self-attention and residual neural network, consisting of 16 layers in five stages. ① is the input layer for receiving images or text, and ② is the normalization layer for normalizing the images or text. ③ in Figure 3 shows the core structure of the self-attention neural network, which primarily extracts information through multi-head self-attention and short-cut residual links. It also includes a learnable feedforward network. ③ in Figure 3 forms the encoder block of the self-attention neural network. The input and output of each encoder block can be fixed matrices with 768 dimensions, maintaining the same dimensionality. A self-attention neural network is configured with 12 cascaded encoder blocks, effectively extracting high-level semantic information from images and text. The fully connected layer network shown in ④ in Figure 3 encodes text or visual information. ①-④ in Figure 3 form the encoding module of the self-attention neural network.
[0054] In one possible embodiment, the multimodal alignment network provided by this solution, when using the text decoding module in the language processing model to obtain the sample text vector of the sample description text and the sample object coordinate information, includes:
[0055] S2221: Integrate the sample description text and the sample object coordinate information text to obtain a sample integration matrix.
[0056] S2222: Utilize the text decoding module in the language processing model to obtain the text feature vector of the sample integration matrix.
[0057] Exemplarily, the corresponding text in each image-text data pair is integrated, for example, the sample description text and the sample object coordinate information text in the image-text data pair are integrated to obtain a sample integration matrix. This sample integration matrix is then sent to a text decoding module in the language processing model, which analyzes and processes the sample integration matrix to obtain a text feature vector for the sample integration matrix. This solution improves the efficiency of extracting sample text vectors and the training and data processing efficiency of the multimodal alignment network by integrating the sample description text and the sample object coordinate information text into a sample integration matrix and using the text decoding module in the language processing model to obtain the text feature vector for the sample integration matrix.
[0058] In one possible embodiment, the image recognition method provided by this solution can optimize the network parameters of the visual processing model and the language processing model based on the predicted labels and the true labels corresponding to the sample image vectors. A loss function can be determined based on the predicted labels and the true labels corresponding to the sample image vectors, and the network parameters of the visual processing model and the language processing model can be optimized based on the loss function. Optionally, the network parameters can be optimized using stochastic gradient descent or a non-heuristic optimization algorithm to improve the network convergence speed.
[0059] Exemplarily, during the training of the multimodal alignment network, the loss function is determined based on the predicted label and the true label corresponding to the sample image vector, and the network parameters of the visual processing model and the language processing model are optimized based on the loss function, and the network parameters of the visual processing model and the language processing model are continuously updated (the network parameters corresponding to the image encoding module in the visual processing model, the text decoding module in the language processing model, and the prediction output layer are continuously updated) so that the loss function of the converged multimodal alignment network reaches the minimum or is less than the set loss threshold.
[0060] In one embodiment, the loss function corresponding to the predicted label and the true label can be determined based on the predicted probability corresponding to the predicted label and the true probability corresponding to the true label. The true probability can be a numerical representation of the true answer, such as a one-hot vector. Optionally, the loss function corresponding to the predicted label and the true label can be expressed as: L(P(A), y) = -ylog(P(A))
[0061] Where P(A) is the predicted probability corresponding to the predicted label, and y is the true probability corresponding to the true label. This solution optimizes the network parameters of the visual processing model and language processing model by determining the loss function based on the predicted and true labels, effectively improving the image recognition accuracy of the multimodal alignment network.
[0062] As shown in the data flow diagram of a multimodal alignment network during the training process provided in Figure 4, in Figure 4, ① is a sample image corresponding to the image-text data pair, ② and ③ are respectively the sample description text and sample object coordinate information corresponding to the image-text data pair. For example, the content corresponding to the sample description text can be "four dogs set on a stone in a wild. from left to right, a white dog, a black dog. a head black-striped white dog and the right is a brown-color dog", which describes the four dogs in the sample image and the corresponding colors. The content corresponding to the object coordinate information can be "tag:dog bbox:(194, 43, 624, 1055); tag:dog bbox:(1206, 176, 1705, 1033); tag:dog bbox:(776, 399, 1245, 1025); tag:dog bbox:(560, 479, 848, 1014); tag:grass bbox:(728, 997, 1022, 1078); tag: tree bbox:(1739, 0, 1848, 718)", which respectively record the coordinates of the positioning boxes corresponding to the four dogs, grass, and trees in the sample image.
[0063] Optionally, the sample image can be normalized into a tensor matrix of set size and dimension, and then the tensor matrix obtained after normalization is sent to the visual processing model to extract the sample image vector. For example, the sample image is normalized into a tensor matrix of 224*224*3d, and then processed by the image encoding module (Image-Encoder) of the visual processing model shown in Figure 4 to obtain a sample image vector x of 192*512 dimension. I =P(x;θ vision )=E I (I)∈R 192×512 , which is the sample image vector (Image Embedding) shown in ⑥ in Figure 4.
[0064] Optionally, the sample description text and sample object coordinate information can be standardized and normalized into a matrix of set dimensions, and then the matrix obtained after normalization is sent to the language processing model to extract the sample text vector. For example, the sample description text and sample object coordinate information are aggregated and normalized into a matrix of 1024*512 dimensions, and processed by the text decoding module (Language-Decoder) of the language processing model shown in Figure ⑤ to obtain a 1024*512-dimensional sample text vector E T (T)∈R 1024×512 , which is the sample text vector (Text Embedding) shown in Figure 3.
[0065] After obtaining the sample image vector and the sample text vector, the sample image vector and the sample text vector are input into the prediction output layer (Generation Layer shown in ⑧ in Figure 4) in the language processing model for analysis and processing, and the vocabulary probability distribution (the predicted label corresponding to the sample image and the corresponding predicted probability) is output. The network parameters of the visual processing model and the language processing model can be optimized based on the loss function corresponding to the predicted label and the true label corresponding to the sample image vector. Optionally, the vocabulary probability distribution can be expressed as: P(A)=P(x answer |xdescribtion,x objects ,x Question ,x I θ LLM ,θ vision )
[0066] Where P(A) is the probability of the answer output by the language processing model (predicted probability), which is calculated as a probability distribution over the speech processing model dictionary (e.g., the language model dictionary), that is, a non-negative vector whose sum is 1 and whose length is the length of the dictionary. LLM and θ vision This section respectively represents language processing model and visual processing model.
[0067] [xdescribtion,x objects ,x Question ,x image ]=concat([E T (Descrbtion),E T (Objects),E T (Question)],E I(Image) represents the concatenation of the text vectors and image vectors corresponding to the sample description text, sample object coordinate information, the set question text, and the sample image. After training the multimodal alignment network using image-text data and continuously updating the network parameters of the visual processing model and language processing model, the network weights are no longer updated after convergence to the optimal network. Specifically, the network parameters corresponding to ④, ⑤, and ⑧ in Figure 4, for the image encoding module, text decoding module, and prediction output layer, are fixed and used as the feature extractors for images and text, respectively, and the parameters for generating predicted text.
[0068] S230: Acquire an image to be recognized, a question text, and a description rule text, wherein the question text and the description rule text are configured based on a set query category.
[0069] S240: Input the image to be identified, the question text, and the description rule text into the trained multimodal alignment network, obtain the image feature vector of the image to be identified by using the visual processing model through the multimodal alignment network, and obtain the text feature vectors of the question text and the description rule text by using the language processing model, and determine the target label corresponding to the image to be identified in the set query category based on the image feature vector and the text feature vector using the language processing model.
[0070] In one possible embodiment, the multimodal alignment network provided by this solution uses a visual processing model to obtain an image feature vector of an image to be identified, uses a language processing model to obtain text feature vectors of question text and rule description text, and uses the language processing model to determine a target label corresponding to a query category for the image to be identified based on the image feature vector and the text feature vector, including:
[0071] S241: Utilize the image encoding module in the visual processing model to obtain the image feature vector of the image to be identified.
[0072] S242: Utilize the text decoding module in the language processing model to obtain the text feature vectors of the question text and the text describing the rule.
[0073] S243: Analyze and process the image feature vector and the text feature vector using the prediction output layer in the language processing model to determine the target label corresponding to the set query category for the image to be identified.
[0074] Exemplarily, when the multimodal alignment network uses a visual processing model to obtain the image feature vector of the image to be identified, it can use the image encoding module in the visual processing model to analyze and process the image to be identified and obtain the image feature vector of the image to be identified.
[0075] When the multimodal alignment network uses the language processing model to obtain the text feature vectors of the question text and the description rule text, it can use the text decoding module in the language processing model to analyze and process the question text and the description rule text, and obtain the text feature vectors of the question text and the description rule text.
[0076] After determining the image feature vector and text feature vector, the image feature vector and text feature vector are input into the prediction output layer of the language processing model. The prediction output layer of the language processing model is used to analyze and process the image and text feature vectors to be identified, and obtain the target label corresponding to the set query category of the image to be identified. This solution uses the image encoding module in the visual processing model to obtain the image feature vector of the image to be identified and the text decoding module in the language processing model to obtain the text feature vectors of the question text and the text describing the rule. The prediction output layer of the language processing model is used to analyze and process the image and text feature vectors to be identified, and obtain the target label corresponding to the set query category. This allows the multimodal alignment network to accurately output the target label corresponding to the set query category, thereby improving image recognition accuracy.
[0077] In one possible embodiment, the multimodal alignment network provided by this solution, when using the text decoding module in the language processing model to obtain the text feature vectors of the question text and the description of the rule text, includes:
[0078] S2421: Integrate the question text and the rule description text to obtain a text integration matrix.
[0079] S2422: Utilize the text decoding module in the language processing model to obtain the text feature vector of the text integration matrix.
[0080] For example, the question text and the rule description text are integrated to obtain a text integration matrix, which is then sent to the text decoding module in the language processing model. The text decoding module analyzes and processes the text integration matrix to obtain the text feature vector of the text integration matrix. This solution improves the efficiency of extracting sample text vectors and image recognition by integrating the question text and the rule description text into a text integration matrix and using the text decoding module in the language processing model to obtain the text feature vector of the text integration matrix.
[0081] Figure 5 shows a diagram of the data flow during the training process of a multimodal alignment network. In Figure 5, ① is the descriptive rule text, ② is the image to be identified, and ③ is the question text. For example, the descriptive rule text might be "Tongue Out: by judging the person's tongue, if the tongue is out of the mouth, then the image can be labeled as tongue out." This describes the descriptive rule for labeling an image as tongue out. The image corresponding to the descriptive rule text in the figure is an example of "tongue out" and is not input as the descriptive rule text. The question text might be "What is the man's label?" This asks the question about the label of the man in the image to be identified.
[0082] The image to be identified is input into the image encoding module (Image-Encoder) of the visual processing model shown in Figure 5 (4). The image encoding module analyzes and processes the image feature vector corresponding to the image to be identified (as shown in Figure 5 (6)). The question text and the description rule text are integrated and sent to the text decoding module (Language-Decoder) of the language processing model shown in Figure 5 (5). The text decoding module analyzes and processes the text feature vectors corresponding to the question text and the description rule text (as shown in Figure 5 (7)). The image feature vector and the text feature vector are input into the prediction output layer (Generation Layer) of the language processing model shown in Figure 5 (8). The prediction output layer analyzes and processes the image feature vector and the text feature vector to output the corresponding target label. As shown in Figure 5 (9), the target label is expressed as "by judging the person's tongue, the man tongue out, so this picture's label is "tongue out"." This means that after analyzing the tongue of the person in the image to be identified, the man in the image to be identified sticks out his tongue, so the target label for the image to be identified can be "tongue out."
[0083] In the above, by inputting the image to be recognized, the question text, and the description rule text into the trained multimodal alignment network, the multimodal alignment network uses a visual processing model to obtain the image feature vector of the image to be recognized, and uses a language processing model to obtain the text feature vectors of the question text and the description rule text, and uses the language processing model to determine the target label corresponding to the image to be recognized in the set query category based on the image feature vector and the text feature vector. There is no need to collect sample data again to train a new recognition model. The target label corresponding to the image to be recognized in the set query category can be obtained by configuring the question text and the description rule text based on the set query category, effectively improving the versatility of image recognition. By using image-text data to train the multimodal alignment network, and using the massive amount of image-text pair data in business scenarios to obtain a multimodal alignment network that semantically aligns text and image information, the multimodal alignment network has the ability to recognize images under specified questions and description rules, without the need to collect sample data and train new recognition models separately for different recognition types, effectively improving the versatility of image recognition.
[0084] FIG6 is a schematic diagram of the structure of an image recognition device provided in an embodiment of the present application. Referring to FIG6 , the image recognition device includes an information acquisition module 61 and an image recognition module 62 .
[0085] Among them, the information acquisition module 61 is configured to obtain the image to be identified, the question text and the description rule text, and the question text and the description rule text are configured based on the set query category; the image recognition module 62 is configured to input the image to be identified, the question text and the description rule text into the trained multimodal alignment network, and obtain the image feature vector of the image to be identified by using the visual processing model through the multimodal alignment network, and obtain the text feature vector of the question text and the description rule text by using the language processing model, and determine the target label corresponding to the image to be identified in the set query category based on the image feature vector and the text feature vector using the language processing model.
[0086] In one possible embodiment, the multimodal alignment network uses a visual processing model to obtain an image feature vector of an image to be identified, uses a language processing model to obtain text feature vectors of question text and rule description text, and uses the language processing model to determine the target label corresponding to the query category for the image to be identified based on the image feature vector and the text feature vector. The configuration is as follows:
[0087] Using the image encoding module in the visual processing model to obtain the image feature vector of the image to be identified;
[0088] Use the text decoding module in the language processing model to obtain the text feature vectors of the question text and the rule text;
[0089] The prediction output layer in the language processing model is used to analyze and process the image feature vector and the text feature vector to determine the target label corresponding to the set query category of the image to be identified.
[0090] In one possible embodiment, when the multimodal alignment network uses the text decoding module in the language processing model to obtain the question text and the text feature vector describing the rule text, the configuration is as follows:
[0091] Integrate the question text and the rule description text to obtain a text integration matrix;
[0092] The text decoding module in the language processing model is used to obtain the text feature vector of the text integration matrix.
[0093] In a possible embodiment, the image recognition device further includes a sample acquisition module and a model training module.
[0094] The sample acquisition module is configured to acquire an image-text data pair, the image-text data pair including a sample image, a sample description text describing the sample image, and sample object coordinate information;
[0095] The model training module is configured to train a multimodal alignment network using image-text data, wherein the multimodal alignment network uses a visual processing model to obtain sample image vectors of sample images, and uses a language processing model to obtain sample text vectors of sample description text and sample object coordinate information, and uses the language processing model to determine the predicted label corresponding to the sample image based on the sample image vector and the sample text vector, and optimizes the network parameters of the visual processing model and the language processing model based on the predicted label and the true label corresponding to the sample image vector.
[0096] In one possible embodiment, when the multimodal alignment network uses a visual processing model to obtain a sample image vector of a sample image, and uses a language processing model to obtain a sample text vector of sample description text and sample object coordinate information, and uses the language processing model to determine a predicted label corresponding to the sample image based on the sample image vector and the sample text vector, the multimodal alignment network is configured as follows:
[0097] Obtaining a sample image vector of the sample image using an image encoding module in a visual processing model;
[0098] Utilize the text decoding module in the language processing model to obtain the sample text vector of the sample description text and the sample object coordinate information;
[0099] The prediction output layer in the language processing model is used to analyze and process the sample image vector and the sample text vector to obtain the predicted label corresponding to the sample image.
[0100] In one possible embodiment, when the multimodal alignment network uses the text decoding module in the language processing model to obtain the sample text vector of the sample description text and the sample object coordinate information, the configuration is as follows:
[0101] Integrate the sample description text and the sample object coordinate information text to obtain a sample integration matrix;
[0102] The text decoding module in the language processing model is used to obtain the text feature vector of the sample integration matrix.
[0103] In one possible embodiment, when the model training module optimizes the network parameters of the visual processing model and the language processing model based on the predicted labels and the true labels corresponding to the sample image vectors, the model training module is configured as follows:
[0104] The loss function is determined according to the predicted label and the true label corresponding to the sample image vector, and the network parameters of the visual processing model and the language processing model are optimized based on the loss function.
[0105] It is worth noting that in the embodiment of the above-mentioned image recognition device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the embodiments of this application.
[0106] The embodiment of the present application also provides an image recognition device, which can integrate the image recognition device provided by the embodiment of the present application. Figure 7 is a structural diagram of an image recognition device provided by the embodiment of the present application. Referring to Figure 7, the image recognition device includes: an input device 73, an output device 74, a memory 72 and one or more processors 71; the memory 72 is used to store one or more programs; when the one or more programs are executed by one or more processors 71, the one or more processors 71 implement the image recognition method provided by the above embodiment. The above-mentioned image recognition device, equipment and computer can be used to execute the image recognition method provided by any of the above embodiments, and have corresponding functions and beneficial effects.
[0107] The embodiments of the present application also provide a non-volatile storage medium that stores computer-executable instructions, which are used to execute the image recognition method provided in the above embodiments when executed by a computer processor. Of course, the non-volatile storage medium that stores computer-executable instructions provided in the embodiments of the present application, whose computer-executable instructions are not limited to the image recognition method provided above, can also execute the related operations in the image recognition method provided in any embodiment of the present application. The image recognition apparatus, device and storage medium provided in the above embodiments can execute the image recognition method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the image recognition method provided in any embodiment of the present application.
[0108] Based on the above embodiments, the embodiments of the present application also provide a computer program product. The technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes a number of instructions for enabling a computer device, a mobile terminal or the processor therein to execute all or part of the steps of the image recognition method provided in each embodiment of the present application.
Claims
1. An image recognition method, wherein, Including: Obtain the image to be recognized, the question text, and the description rule text, where the question text and the description rule text are configured based on a set query category; Input the image to be recognized, the question text, and the description rule text into the trained multi-modal alignment network. Through the multi-modal alignment network, use the visual processing model to obtain the image feature vector of the image to be recognized, and use the language processing model to obtain the text feature vectors of the question text and the description rule text, and use the language processing model to determine the target label of the image to be recognized corresponding to the set query category according to the image feature vector and the text feature vector.
2. The image recognition method according to claim 1, wherein When the multi-modal alignment network uses the visual processing model to obtain the image feature vector of the image to be recognized, and uses the language processing model to obtain the text feature vectors of the question text and the description rule text, and uses the language processing model to determine the target label of the image to be recognized corresponding to the set query category, it includes: Use the image encoding module in the visual processing model to obtain the image feature vector of the image to be recognized; Use the text decoding module in the language processing model to obtain the text feature vectors of the question text and the description rule text; Use the prediction output layer in the language processing model to analyze and process the image feature vector and the text feature vector, and determine the target label of the image to be recognized corresponding to the set query category.
3. The image recognition method according to claim 2, wherein, When the multi-modal alignment network uses the text decoding module in the language processing model to obtain the text feature vectors of the question text and the description rule text, it includes: Integrate the question text and the description rule text to obtain a text integration matrix; Use the text decoding module in the language processing model to obtain the text feature vector of the text integration matrix.
4. The image recognition method according to claim 1, wherein, The training process of the multi-modal alignment network includes: Obtain an image-text data pair, where the image-text data pair includes a sample image, and a sample description text and sample object coordinate information for describing the sample image; Use the image-text data pair to train the multi-modal alignment network. Among them, the multi-modal alignment network uses the visual processing model to obtain the sample image vector of the sample image, and uses the language processing model to obtain the sample text vector of the sample description text and the sample object coordinate information, and uses the language processing model to determine the prediction label corresponding to the sample image according to the sample image vector and the sample text vector, and optimize the network parameters of the visual processing model and the language processing model based on the prediction label and the true label corresponding to the sample image vector.
5. The image recognition method according to claim 4, wherein, When the multi-modal alignment network obtains the sample image vector of the sample image by using the visual processing model, and obtains the sample text vector of the sample description text and the sample object coordinate information by using the language processing model, and uses the language processing model to determine the prediction label corresponding to the sample image according to the sample image vector and the sample text vector, it includes: Obtaining the sample image vector of the sample image by using the image encoding module in the visual processing model; Obtaining the sample text vector of the sample description text and the sample object coordinate information by using the text decoding module in the language processing model; Analyzing and processing the sample image vector and the sample text vector by using the prediction output layer in the language processing model to obtain the prediction label corresponding to the sample image.
6. The image recognition method according to claim 5, wherein, When the multi-modal alignment network obtains the sample text vector of the sample description text and the sample object coordinate information by using the text decoding module in the language processing model, it includes: Integrating and processing the sample description text and the sample object coordinate information text to obtain a sample integration matrix; Obtaining the text feature vector of the sample integration matrix by using the text decoding module in the language processing model.
7. The image recognition method according to claim 4, wherein, Optimizing the network parameters of the visual processing model and the language processing model based on the prediction label and the true label corresponding to the sample image vector includes: Determining a loss function according to the prediction label and the true label corresponding to the sample image vector, and optimizing the network parameters of the visual processing model and the language processing model based on the loss function.
8. An image recognition device, wherein, Including an information acquisition module and an image recognition module, where: The information acquisition module is configured to acquire an image to be recognized, a question text, and a description rule text, where the question text and the description rule text are configured based on a set query category; The image recognition module is configured to input the image to be recognized, the question text, and the description rule text into the trained multi-modal alignment network, and use the visual processing model in the multi-modal alignment network to obtain the image feature vector of the image to be recognized, and use the language processing model to obtain the text feature vectors of the question text and the description rule text, and use the language processing model to determine the target label corresponding to the image to be recognized in the set query category according to the image feature vector and the text feature vector.
9. An image recognition device, wherein, Including: A memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image recognition method according to any one of claims 1-7.
10. A non-volatile storage medium storing computer-executable instructions, wherein, The computer-executable instructions are used to execute the image recognition method according to any one of claims 1-7 when executed by a computer processor.
11. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the image recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Visual common sense reasoning method and system based on graph attention network
CN115759261A
Text information extraction method and device, storage medium and computer equipment
CN116503877A
Loss assessment content generation method, multi-modal model, device, equipment and storage medium
CN117196858A
Image recognition method and device, equipment, storage medium and product
CN118038201A
Cited By
Regional safety detection method, access control equipment and storage medium
CN121392755A
Aerospace material crack image detection method based on CA-MLL two-stage reasoning
CN121437502A
Image processing method and device, computer equipment and storage medium
CN121527596A
Fashion preference prediction method and device
CN121640486A