Image recognition method and device, electronic equipment and readable storage medium
Patent Information
- Application Number
- CN202311237417.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-09-25
AI Technical Summary
然而,现有的图像识别算法具有误报、判断能力弱、应用局限大等缺陷,生成的图像识别结果不够准确
[0010]本申请实施例提供的图像识别方法,获取第一图像;通过训练后的图像识别模型处理第一图像,得到第一图像的图像识别信息和图像识别信息的置信度;在置信度大于第一阈值的情况下,根据图像识别信息确定图像识别结果,在置信度小于或等于第一阈值的情况下,对图像识别信息进行分类处理,并根据分类处理结果确定图像识别结果。通过上述图像识别方法,基于训练后的图像识别模型得到第一图像的图像识别信息和图像识别信息的置信度,在置信度大于第一阈值的情况下,直接根据图像识别信息确定图像识别结果,而在置信度小于或等于第一阈值的情况下,则结合分类算法处理图像识别信息以得到准确的图像识别结果。这样,在图像识别信息的置信度较大时,直接基于图像识别信息确定图像识别结果,在图像识别信息的置信度较小时,结合分类算法进行图像识别得到图像识别结果,在各种情况下均能够得到准确的图像识别结果,提升了图像识别方法的判断能力,降低了图像识别方法的误报概率,提升了图像识别结果的准确性,并且,基于图像识别模型进行图像识别,降低了图像识别方法的应用局限性。
Smart Images

Figure CN117315345B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an image recognition method and apparatus, electronic device and readable storage medium. Background Technology
[0002] Currently, image recognition algorithms are mainly divided into two types: rule-based algorithms and machine learning-based algorithms. Machine learning-based image recognition algorithms are further divided into two categories: feature-based algorithms and deep learning-based algorithms. However, existing image recognition algorithms suffer from drawbacks such as false alarms, weak judgment capabilities, and limited applications, resulting in inaccurate image recognition results. Summary of the Invention
[0003] The purpose of this application is to provide an image recognition method, apparatus, electronic device, and readable storage medium that can improve the accuracy of image recognition results.
[0004] In a first aspect, embodiments of this application provide an image recognition method, the method comprising: acquiring a first image; processing the first image through a trained image recognition model to obtain image recognition information and confidence level of the image recognition information; determining an image recognition result based on the image recognition information when the confidence level is greater than a first threshold, and classifying the image recognition information when the confidence level is less than or equal to the first threshold, and determining the image recognition result based on the classification result.
[0005] Secondly, embodiments of this application provide an image recognition device, which includes: an acquisition unit for acquiring a first image; a processing unit for processing the first image through a trained image recognition model to obtain image recognition information and confidence level of the image recognition information; the processing unit is further configured to determine an image recognition result based on the image recognition information when the confidence level is greater than a first threshold, and to classify the image recognition information when the confidence level is less than or equal to the first threshold, and to determine the image recognition result based on the classification result.
[0006] Thirdly, embodiments of this application provide an electronic device including a processor and a memory. The memory stores programs or instructions that can run on the processor, and when the programs or instructions are executed by the processor, they implement the steps of the image recognition method as described in the first aspect.
[0007] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the image recognition method as described in the first aspect.
[0008] Fifthly, embodiments of this application provide a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the image recognition method as described in the first aspect.
[0009] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the image recognition method as described in the first aspect.
[0010] The image recognition method provided in this application involves acquiring a first image; processing the first image using a trained image recognition model to obtain image recognition information and a confidence level of the image recognition information; determining an image recognition result based on the image recognition information when the confidence level is greater than a first threshold; and classifying the image recognition information based on the classification result when the confidence level is less than or equal to the first threshold. Through this image recognition method, image recognition information and a confidence level of the first image are obtained based on a trained image recognition model. When the confidence level is greater than the first threshold, the image recognition result is directly determined based on the image recognition information; while when the confidence level is less than or equal to the first threshold, a classification algorithm is used to process the image recognition information to obtain an accurate image recognition result. In this way, when the confidence level of the image recognition information is high, the image recognition result can be determined directly based on the image recognition information. When the confidence level of the image recognition information is low, the image recognition result is obtained by combining the classification algorithm. Accurate image recognition results can be obtained in all situations, which improves the judgment ability of the image recognition method, reduces the false alarm probability of the image recognition method, and improves the accuracy of the image recognition result. Furthermore, image recognition based on the image recognition model reduces the application limitations of the image recognition method. Attached Figure Description
[0011] Figure 1 A schematic flowchart illustrating the image recognition method provided in this application embodiment;
[0012] Figure 2 One of the schematic diagrams of the image recognition method provided in the embodiments of this application;
[0013] Figure 3 The second schematic diagram of the image recognition method provided in the embodiments of this application;
[0014] Figure 4 The third schematic diagram of the image recognition method provided in the embodiments of this application;
[0015] Figure 5 The fourth schematic diagram of the image recognition method provided in the embodiments of this application;
[0016] Figure 6 The fifth schematic diagram of the image recognition method provided in the embodiments of this application;
[0017] Figure 7 This is a structural block diagram of the image recognition device provided in the embodiments of this application;
[0018] Figure 8 A structural block diagram of the electronic device provided in the embodiments of this application;
[0019] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0021] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0022] The image recognition method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0023] like Figure 1 As shown, this application provides an image recognition method, which may include the following steps S102 to S106:
[0024] S102: Obtain the first image.
[0025] The image recognition method proposed in this application is executed by an electronic device, which may be a smartphone, tablet computer, laptop computer, or smartwatch, etc., and is not specifically limited to these devices.
[0026] The first image mentioned above can be a single image selected by the user.
[0027] Furthermore, the first image mentioned above can also be a frame from a video selected by the user.
[0028] Specifically, the first image mentioned above may be a clear image selected by the electronic device from the video stream data selected by the user based on a fuzzy judgment algorithm.
[0029] Specifically, in the image recognition method provided in this application embodiment, the electronic device performs fuzzy judgment based on a binary classification network. The network structure of the binary classification network is as follows: Figure 2 As shown, the binary classification network is trained using both clear and blurry images, and outputs a binary classification result of 0 or 1. Here, 0 indicates that the image input to the binary classification network is blurry, and 1 indicates that the image input to the binary classification network is clear. During the process of the electronic device selecting a clear image frame from the user-selected video stream data based on the binary classification network, each frame of the video stream data is sequentially input into the binary classification network. If the electronic device determines that the image input to the binary classification network is blurry based on the binary classification result, it will continue to perform blur judgment on the next frame of the video stream data until a clear image is selected. Furthermore, if the electronic device determines that the first image is blurry, it will return "a blurry image" as the image recognition result.
[0030] S104: Process the first image using the trained image recognition model to obtain the image recognition information and the confidence level of the image recognition information.
[0031] The image recognition model is the MIC (Multi-modal Image Caption) model, which can dynamically generate corresponding image recognition information for the input image.
[0032] Furthermore, the image recognition model is a small-volume model.
[0033] Furthermore, the image recognition information mentioned above includes at least one recognition word for the first image, which is used to identify the image content of the first image.
[0034] Furthermore, the confidence level of the aforementioned image recognition information is used to indicate the degree of credibility of the match between the image recognition information and the first image.
[0035] S106: If the confidence level is greater than the first threshold, determine the image recognition result based on the image recognition information; if the confidence level is less than or equal to the first threshold, classify the image recognition information and determine the image recognition result based on the classification result.
[0036] In the process of recognizing the first image using the trained image recognition model, the model also outputs a score for each recognized word in the first image. Based on this, the aforementioned first threshold is dynamically generated based on the score of each recognized word in the first image during the recognition process.
[0037] Furthermore, the first threshold serves as the criterion for determining the reliability of the image recognition information. If the confidence level of the image recognition information is greater than the first threshold, it indicates a high degree of reliability, meaning a high degree of matching between the image recognition information and the first image. Conversely, if the confidence level is less than or equal to the first threshold, it indicates a low degree of reliability, meaning a low degree of matching between the image recognition information and the first image.
[0038] Furthermore, the above image recognition result is a complete sentence information that recognizes the image content of the first image, such as "a man standing on the grass" or "an image containing people, a beach, and the sky" and other image recognition sentences.
[0039] Specifically, in the image recognition method provided in this application embodiment, after performing image recognition on the first image using a trained image recognition model to obtain image recognition information and confidence level of the image recognition information, the confidence level is compared with a dynamically generated first threshold. Further, if the confidence level is greater than the first threshold, the image recognition result is directly determined based on the image recognition information; if the confidence level is less than or equal to the first threshold, the image recognition information is classified using a classification model based on a classification algorithm, and the classification results output by the classification model are combined into a complete image recognition statement, which is then used as the image recognition result for the first image.
[0040] The aforementioned classification model is based on a multi-label classification network, which employs the MobileNet architecture. In practical applications, the multi-label classification network annotates a large number of images using multi-label image annotation and trains the network with these images, thus enabling it to label images. Based on this, the output of the classification model is at least one label for the first image. By integrating this at least one label from the model's output, a complete image recognition statement can be obtained, which serves as the image recognition result for the first image.
[0041] For example, the classification model based on a multi-label classification network outputs labels "person," "beach," and "sky." By integrating the label information output by the classification model, the image recognition result of the first image can be obtained: "an image containing a person, a beach, and a sky." Based on this, the electronic device can display the image recognition result of the first image to the user through text display or voice playback.
[0042] The image recognition method provided in this application involves acquiring a first image; processing the first image using a trained image recognition model to obtain image recognition information and a confidence level of the image recognition information; determining an image recognition result based on the image recognition information when the confidence level is greater than a first threshold; and classifying the image recognition information based on the classification result when the confidence level is less than or equal to the first threshold. In this image recognition method, image recognition information and a confidence level of the first image are obtained based on a trained image recognition model. When the confidence level is greater than the first threshold, the image recognition result is directly determined based on the image recognition information. When the confidence level is less than or equal to the first threshold, a classification algorithm is used to process the image recognition information to obtain an accurate image recognition result. In this way, when the confidence level of the image recognition information is high, the image recognition result can be determined directly based on the image recognition information. When the confidence level of the image recognition information is low, the image recognition result is obtained by combining the classification algorithm. Accurate image recognition results can be obtained in all situations, which improves the judgment ability of the image recognition method, reduces the false alarm probability of the image recognition method, and improves the accuracy of the image recognition result. Furthermore, image recognition based on the image recognition model reduces the application limitations of the image recognition method.
[0043] In this embodiment of the application, S104 may specifically include the following S104a to S104c:
[0044] S104a: Adjust the image size of the first image to obtain the second image.
[0045] The image recognition model mentioned above includes a size adjustment module.
[0046] Specifically, in the image recognition method provided in the embodiments of this application, after inputting the first image into the image recognition model, the image size of the first image is adjusted by the size adjustment module in the image recognition model to obtain the second image.
[0047] For example, the size adjustment module in the image recognition model can reduce the size of the first image input to the model from its original size to a second image with image size parameters of (256, 256). The image size parameters of the second image are variable, and those skilled in the art can set them according to actual conditions; no specific limitations are imposed here.
[0048] S104b: Extract the feature matrix of the second image using the trained image recognition model.
[0049] The image recognition model mentioned above includes a feature extraction module, namely a visual encoder.
[0050] In practical applications, the aforementioned feature extraction module may specifically be a ResNet visual feature extraction module, a MobileVirt feature extraction module, or other modules. Those skilled in the art can select the specific type of the aforementioned feature extraction module according to the actual situation, and no specific restrictions are imposed here.
[0051] Specifically, in the image recognition method provided in this application embodiment, after adjusting the image size of the first image through the size adjustment module in the trained image recognition model to obtain the second image, the second image is input into the feature extraction module in the trained image recognition model, and the feature matrix of the second image is extracted through the feature extraction module.
[0052] The input to the feature extraction module is the second image output by the size adjustment module, and the output of the feature extraction module is the encoded feature matrix of the second image.
[0053] For example, the input to the feature extraction module is a second image of size [batch, 3, 256, 256], where batch represents the number of input images, typically 1. The output of the feature extraction module is a feature matrix of size [batch, Ne = 256, dim = 768], where Ne represents the number of features, typically 256 (other values are also possible), and dim represents the feature dimension, typically 768 (other values are also possible).
[0054] S104c: Based on the trained image recognition model, determine image recognition information and confidence level according to the feature matrix.
[0055] The image recognition model mentioned above includes a recognition generation module.
[0056] Specifically, in the image recognition method provided in this application embodiment, after the feature matrix of the second image is extracted by the feature extraction module, the feature matrix of the second image is used as an input to the recognition generation module. The recognition generation module, based on the trained image recognition model, determines the image recognition information and confidence level according to the feature matrix.
[0057] In practical applications, during the process of determining image recognition information and confidence levels based on the feature matrix in the image recognition model-based recognition generation module, the module also has another input: a fixed-length buffer queue. Based on this, the recognition generation module iteratively generates recognition words for the second image using the feature matrix, and sequentially fills the buffer queue with the feature vectors corresponding to the generated recognition words, until the number of characters in the recognition words output by the module exceeds the set maximum number of characters or the module outputs a stop symbol.
[0058] For example, another input to the recognition generation module is a buffer queue of length [batch, Nd = 25, dim = 768], where Nd represents the maximum number of characters that the recognition generation module can output, typically 25, indicating that the recognition generation module supports a maximum of 25 characters for input and output. Based on this, the recognition generation module iteratively generates recognition words for the second image based on the feature matrix of the second image and the set vocabulary, and obtains the order of the recognition words in the vocabulary, i.e., the character number. Then, Word2Embed is used to convert the character number of the recognition word into the corresponding feature vector, and the feature vector corresponding to the recognition word is filled into the aforementioned buffer queue, until the number of characters in the recognition words output by the recognition generation module exceeds the set maximum number of characters, i.e., 25, or the recognition generation module outputs a stop symbol.
[0059] The embodiments provided in this application, during the process of processing a first image using a trained image recognition model to obtain image recognition information and confidence levels for the first image, adjust the image size of the first image to obtain a second image; extract the feature matrix of the second image using the trained image recognition model; and determine the image recognition information and confidence levels based on the feature matrix using the trained image recognition model. In this way, determining image recognition information and confidence levels based on a small-volume image recognition model and the image's feature matrix ensures the accuracy of the image recognition results. Furthermore, the small-volume image recognition model is easy to implement in mobile devices, reducing the application limitations of image recognition methods.
[0060] In this embodiment of the application, the feature matrix includes N feature vectors, and the N feature vectors correspond to multiple recognition words in the vocabulary. Based on this, the steps of determining image recognition information and confidence level according to the feature matrix may specifically include the following steps S108 to S112:
[0061] S108: Construct a cache queue with a length of the second threshold, and fill the first position of the cache queue with the feature vector corresponding to the first recognized word in the vocabulary.
[0062] The first position is the initial position of the cache queue.
[0063] Furthermore, the length of the buffer queue is a second threshold, which represents the maximum number of characters that the recognition generation module of the image recognition model can output. For example, the length of the buffer queue is [batch, Nd = 25, dim = 768], where Nd is the aforementioned second threshold, indicating that the recognition generation module supports a maximum of 25 characters for input and output.
[0064] Furthermore, the aforementioned feature matrix includes N feature vectors, each corresponding to one of the N recognized words in a predefined vocabulary.
[0065] Specifically, in the image recognition method provided in this application embodiment, in the process of determining image recognition information and confidence based on the feature matrix in the recognition generation module based on the trained image recognition model, a cache queue with a length of a second threshold is constructed, and the feature vector corresponding to the first recognition word in the vocabulary is filled into the initial position of the cache queue, i.e., the first position.
[0066] S110: Based on the cache queue and feature matrix, determine the confidence level and the first ranking order, and fill the feature vector corresponding to the recognition word in the first ranking order in the vocabulary into the second position of the cache queue to update the cache queue.
[0067] Specifically, in the image recognition method provided in this application embodiment, after constructing a cache queue of length two thresholds and filling the initial position of the cache queue with the feature vector corresponding to the first recognition word in the vocabulary, the cache queue and the feature matrix are input into the recognition generation module of the trained image recognition model. Based on the cache queue and the feature matrix, the recognition generation module outputs the text sequence number (i.e., the first arrangement order) of the predicted recognition word in the vocabulary and the confidence score of the predicted recognition word matching the second image. On this basis, Word2Embed is used to convert the text sequence number (i.e., the first arrangement order) of the predicted recognition word in the vocabulary into a corresponding feature vector, and this feature vector is filled into the second position of the cache queue, which is the next position after the first position, so that this feature vector can be used as the module input for the next predicted recognition word.
[0068] S112: Based on the feature matrix and the updated cache queue, update the confidence and determine the second arrangement order until the determined arrangement order meets the first condition or the number of confidence updates reaches the second threshold. Based on the recognition words corresponding to each feature vector in the cache queue after the last update, determine the image recognition information.
[0069] Specifically, in the image recognition method provided in this application embodiment, after filling the feature vector corresponding to the first predicted recognition word into the second position of the cache queue, the feature vector is used as the module input for the next predicted recognition word. The recognition generation module, based on the feature matrix and the updated cache queue, continues to output the text sequence number (second arrangement order) of the predicted next recognition word in the word list, as well as the confidence score of the predicted recognition word matching the second image. The confidence score previously output by the recognition generation module is then updated based on this confidence score. Further, Word2Embed converts the text sequence number (second arrangement order) of the newly predicted recognition word in the word list into a corresponding feature vector, and this feature vector is filled into the third position of the cache queue, which is the next position after the second position, so that the feature vector can be used as the module input for the next predicted recognition word. This process is repeated until the order of the outputs from the recognition generation module meets the first condition, or until the number of confidence updates, i.e., the number of iterations, reaches the second threshold mentioned above. Based on the recognition words corresponding to each feature vector in the cache queue after the last update, the image recognition information is determined, and the confidence level output by the recognition generation module at the end is taken as the confidence level of the image recognition information.
[0070] The first condition mentioned above is that the order of the output of the recognition generation module corresponds to the stop symbol in the vocabulary.
[0071] For example, in the image recognition process, a batch of images is acquired. After preprocessing, the image input I is obtained, with a size of [batch, 3, 256, 256]. The image input I is processed by the ResNet feature extraction module to obtain a feature matrix with a length of [batch, Ne = 256, dim = 768]. This feature matrix is denoted as Iencoder. Further, a buffer queue of length [batch, Nd = 25, dim = 768] is acquired, denoted as Idecoder. The buffer queue is initially filled with 0s, and the feature vector corresponding to Wbos = 0 is filled at the 0th index of the buffer queue, i.e., Idecoder[:, 0,:] = Word2Embed(Wbos). Here, Word2Embed is a matrix of size [vocab_size, dim = 768], where vocab_size is the size of the vocabulary, and dim is the feature dimension of each recognized word. Further, Wbos represents the text index of the starting recognized word in the vocabulary, which is assumed to be 0 here. That is, the feature vector corresponding to the first recognized word in the vocabulary is filled into the initial position of the cache queue. Further, the feature vector corresponding to Wcls=1 is filled into the Nd=25th position of the cache queue, i.e., Idecoder[:,24,:]=Word2Embed(Wcls). Here, Wcls represents the text index of the recognized word in the vocabulary, which is assumed to be 1 here.
[0072] Furthermore, Iencoder and Idecoder are used as inputs to the recognition generation module, yielding two outputs: one is the character sequence number of the predicted recognition word in the vocabulary, and the other is the confidence score of the predicted recognition word matching the second image. Further, Word2Embed converts the output character sequence number into a feature vector, which is then used as input to the next predicted recognition word, filling the next position in Idecoder. Based on this, the recognition generation module is repeatedly called to generate the character sequence number of the next predicted recognition word and the confidence score of the recognition word matching the second image, until the character sequence number output by the recognition generation module corresponds to a stop symbol or the number of iterations exceeds Nd (25).
[0073] For example, for the recognition-generating single module with two inputs, Iencoder and Idecoder, such as Figure 3 As shown in (a), at the initial time step, the beginning and end of the Idecoder are filled with the feature vectors corresponding to the recognized words with text indices Wbos and Wcls, respectively, while the rest are filled with 0. Further, the Iencoder and Idecoder are used as inputs to the recognition generation module to predict the output of the next recognized word, as follows... Figure 3As shown in (b), the first output result is predicted at the position where the Next part is located. On this basis, if the predicted value of the next recognition word output by the recognition generation module is 10, it indicates that the next recognition word corresponds to the recognition word with the character sequence number 10 in the vocabulary. Based on this, by querying the vocabulary, it can be found that the recognition word with the character sequence number 10 in the vocabulary is "I".
[0074] Further, the character sequence number 10 is converted into the corresponding feature vector through Word2Embed, and the feature vector is filled into the next position of Wbos in the cache queue Idecoder, that is, Idecoder[batch, 1, :] = Word2Embed(10). At this time, the recognition generation module is run based on the updated cache queue Idecoder, and the predicted value of the next recognition word output by the recognition generation module is 11. By querying the vocabulary, it can be found that the recognition word with the character sequence number 11 in the vocabulary is "am". Further, the feature vector of the word "am" is continuously filled into the next position of the cache queue Idecoder, that is, Idecoder[batch, 2, :] = Word2Embed(11). At this time, the recognition generation module is run based on the updated cache queue Idecoder, and the predicted value of the next recognition word output by the recognition generation module is 12. By querying the vocabulary, it can be found that the recognition word with the character sequence number 12 in the vocabulary is "human". Further, the feature vector of the word "human" is continuously filled into the next position of the cache queue Idecoder, that is, Idecoder[batch, 3, :] = Word2Embed(12). Further, the recognition generation module continues to run based on the updated cache queue Idecoder. At this time, if the predicted value output by the recognition generation module is 102, and 102 corresponds to a stop symbol, the operation of the recognition generation module is stopped, the image recognition information "I am human" is obtained, and the confidence finally output by the recognition generation module is taken as the confidence of the image recognition information.
[0075] The embodiments provided in this application include a feature matrix comprising N feature vectors, each corresponding to one of the N recognition words in a vocabulary. During the process of determining image recognition information and confidence levels based on the feature matrix, a cache queue of length denoted by a second threshold is constructed. The feature vector corresponding to the first recognition word in the vocabulary is then filled into the first position of the cache queue, where the first position is the initial position of the cache queue. Based on the cache queue and the feature matrix, the confidence level and a first arrangement order are determined. The feature vector corresponding to the recognition word in the first arrangement order is then filled into the second position of the cache queue to update the cache queue, where the second position is the next in order after the first position. Based on the feature matrix and the updated cache queue, the confidence level is updated and a second arrangement order is determined until the determined arrangement order meets the first condition or the number of confidence level updates reaches the second threshold. Finally, image recognition information is determined based on the recognition word corresponding to each feature vector in the last updated cache queue. In this way, recognition words are predicted based on the image's feature vectors, and a fixed-length cache queue is used to output image recognition information, improving the accuracy of the image recognition results.
[0076] In this embodiment of the application, prior to S110, the image recognition method may further include the following S114 to S118:
[0077] S114: Add a mask to the cache queue.
[0078] The aforementioned recognition and generation module is used to fuse the input image code and the generated text code, extract relevant feature information from the image code, and convert it into the corresponding text prediction result. In the operation of the image recognition model, such as... Figure 4 As shown, the recognition generation module is used as a basic processing module in the image recognition model and is stacked together.
[0079] Furthermore, such as Figure 4 As shown, the recognition generation module includes a Mask-M-Attention module, which is a mask-based attention learning module. Based on this, after inputting the cache queue and feature matrix into the trained image recognition model's recognition generation module, the Mask-M-Attention module adds a mask to the cache queue.
[0080] The mask is a lower triangular matrix of all 1s. By multiplying the mask by the corresponding elements of the cache queue, the input part of the recognition and generation module is filtered to prevent the recognition and generation module from seeing the input part after the current text number.
[0081] S116: Perform self-attention learning on the cache queue and feature matrix after adding the mask to obtain the first weight matrix, and perform weighted processing on the feature matrix according to the first weight matrix.
[0082] Among them, such as Figure 4 As shown, the recognition and generation module includes an M-Attention module, which is an attention learning module. Based on this, after masking the cache queue using the Mask-M-Attention module, the output of the Mask-M-Attention module and the aforementioned feature matrix are used as two inputs to the M-Attention module. The M-Attention module performs self-attention learning on the masked cache queue and the feature matrix. Based on the attention learning mechanism, the feature vectors in the feature matrix are fused to obtain a first weight matrix. The feature matrix is then weighted according to this first weight matrix, enabling the trained image recognition model to focus on more useful information, thereby improving the accuracy of image recognition.
[0083] Among them, such as Figure 5 As shown, the M-Attention module includes a Linear layer, a Layernorm & Activation layer, a BMM layer, a Normalize layer, and a Reweight layer. The Linear layer is a fully connected layer that accepts an input tensor and transforms it into an output tensor, with the mathematical formula y = x·A. T +b, where x is the input, and A is the linear weight matrix. T Let be the transpose of A, b be the bias matrix of Linear, and y be the output of Linear.
[0084] Furthermore, the Layernorm & Activation layer, also known as the layer normalization and activation layer, is used to normalize the Nd dimension of the input [batch, Nd, dim]. Specifically, in the Nd dimension, the mean is subtracted from each pixel and then divided by the variance. Then, an activation function such as the ReLU activation function is used to process the data, e.g., ReLU(x) = max(0, x).
[0085] Furthermore, the BMM layer is used to implement BMM operations between matrices, which are three-dimensional matrix multiplication operations. For example, A( b×n1×n2 )×B( b×n2×n3 ) = C( b×n1×n3 This means that multiplying a three-dimensional matrix A by a three-dimensional matrix B results in a new three-dimensional matrix C, where the third dimension of A and the second dimension of B must be the same.
[0086] Furthermore, the Normalize layer is used to normalize a three-dimensional matrix row by row. For example, normalizing the position of the i-th row and j-th column of a two-dimensional matrix A yields:
[0087] Furthermore, the Reweight layer adjusts the weights of the input feature matrix V based on a three-dimensional weight matrix W, which is the output of the Normalize layer. Specifically, the Reweight layer calls the BMM operation to perform matrix multiplication between the weight matrix W and the feature matrix V.
[0088] Based on this, during the operation of the M-Attention module, such as Figure 5 As shown, for inputs Q, K, V, and mask, the output is The entire calculation process of the M-Attention module can specifically include the following formulas (1) to (7):
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096] Where Q, K, and V are all feature vectors. and The feature vector after LayerNorm & Activation processing. K is the feature vector after linear processing. T Let K be the transpose of the eigenvector K. For feature vectors The transpose of W, where W is a three-dimensional matrix. b , i , j This represents the matrix element in the j-th row, i-th column, and b-th layer of the three-dimensional matrix W, where mask is the mask.
[0097] S118: If the dimension of the weighted feature matrix meets the second condition, perform a non-linear mapping on the feature matrix to adjust its dimension.
[0098] Among them, such as Figure 4As shown, the recognition and generation module includes an FNN module, i.e., a nonlinear spatial mapping module. Based on this, and provided the dimension of the weighted feature matrix meets the second condition, a nonlinear mapping is performed on the feature matrix to adjust its dimension. This increases the complexity of the content expressed by the image recognition model, enabling it to learn more complex textual expressions and improving the accuracy of image recognition based on the feature matrix.
[0099] The second condition mentioned above is that the dimension of the feature matrix is [batch, Nd, dim]. That is, the FNN module is used to perform non-linear spatial mapping on the feature matrix with dimension [batch, Nd, dim] output by the M-Attention module in order to adjust the dimension of the feature matrix.
[0100] In practical applications, the FNN module can operate based on the following formula (8):
[0101] FFN(x)=max(0,x×W1+b1)×W2+b2,(8)
[0102] Where x is the feature matrix output by the M-Attention module, and W1, W2, b1, and b2 are the module parameters of the FNN module.
[0103] Based on this, such as Figure 6 As shown, in the image recognition model's operation, the input index value is converted into a feature vector x. Then, x is used as the input to K, Q, and V of the first M-Attention module, resulting in an output vector with the same dimension as V, which is then used as the input to V of the second M-Attention module. Further, the feature vector extracted by the visual feature extractor is used as the input to Q and K of the second M-Attention module, and subsequently, all features located at... Figure 6 In the second row, the Q and K inputs of the M-Attention module are the feature vectors extracted by the visual feature extractor and remain unchanged. Further, the output of the second M-Attention module after FFN transformation is used as the Q, K, and V inputs of the third M-Attention module, and so on. Further, the output of the last FFN module is mapped to the vocabulary to predict the next recognized word x, and the above process is repeated until the image recognition model outputs a stop symbol or the length of the recognized word output by the image recognition model reaches its maximum length.
[0104] The aforementioned vocabulary is a small-scale Chinese vocabulary obtained by training Wordembed on a large-scale image-text corpus. Further, in this embodiment, for... Figure 6The linear module in "nextindex:x" is reduced, so that the output of the final linear layer is reduced from 21128 to 8000, which reduces the inference time of the image recognition model by 60%.
[0105] The embodiments provided in this application, before determining the confidence level and the first ranking order based on the cache queue and the feature matrix, add a mask to the cache queue; perform self-attention learning on the masked cache queue and the feature matrix to obtain a first weight matrix, and then perform weighted processing on the feature matrix according to the first weight matrix; if the dimension of the weighted feature matrix meets the second condition, perform nonlinear mapping on the feature matrix to adjust the dimension of the feature matrix. In this way, on the one hand, by performing a masking operation on the cache queue, the input part of the recognition generation module is filtered to prevent the recognition generation module from seeing the subsequent input parts, thus improving the accuracy of image recognition; on the other hand, by performing weighted processing on the feature matrix, the image recognition model can focus on more useful information, thereby improving the accuracy of image recognition; furthermore, by performing nonlinear mapping on the feature matrix, the complexity of the content expressed by the image recognition model can be increased, enabling the image recognition model to learn more complex textual expressions, thus improving the accuracy of image recognition based on the feature matrix.
[0106] In this embodiment of the application, before determining the image recognition result based on the image recognition information, the above image recognition method may further include the following steps S120 to S126:
[0107] S120: Determine the first vector based on the confidence obtained each time during the process of determining or updating the confidence.
[0108] In the image recognition method provided in this application embodiment, during the process of predicting the recognition word and determining the image recognition information by running the above-mentioned recognition generation module in a loop, for each loop, the recognition generation module will output the confidence level of the predicted recognition word matching the second image.
[0109] Based on this, in the image recognition method provided in the embodiments of this application, the electronic device records the confidence level corresponding to the recognition word output in each loop, and determines the corresponding first vector according to the confidence level corresponding to each recognition word.
[0110] S122: Determine the second vector based on each first vector.
[0111] Specifically, in the image recognition method provided in this application embodiment, after determining a first vector based on the confidence level corresponding to each recognition word, the position index value of the maximum value of each first vector is obtained, and a second vector corresponding to the first vector is determined based on the position index value corresponding to each first vector. The first vector and the second vector have the same magnitude.
[0112] For example, the confidence level is [batch = 1, 1, vocab_size], denoted as the first vector q. Further, the index of the maximum value is calculated for the first vector q, denoted as index. Further, a second vector p is constructed based on the index. In constructing the second vector p, p is a vector of the same size as the first vector q, with a value of 1 at the index position and 0 at all other positions.
[0113] S124: Determine the divergence value between each first vector and its corresponding second vector to obtain multiple divergence values.
[0114] In this case, one first vector corresponds to one second vector.
[0115] Based on this, in the image recognition method provided in the embodiments of this application, after determining the second vector corresponding to each first vector, for the corresponding first vector and second vector, the divergence value between each first vector and the corresponding second vector is calculated, thereby obtaining multiple divergence values.
[0116] In practical applications, the aforementioned divergence value can specifically be the KL divergence value, which is used to evaluate the similarity between the first and second vectors. The KL divergence value can be calculated using the following formula (9):
[0117] D KL (p||q)=-∑ x p(x)logq(x)+∑ x p(x)logp(x), (9)
[0118] Where p and q are two vectors, D KL (p||q) represents the KL divergence between vectors p and q, where p(x) represents the position index of vector p and q(x) represents the position index of vector q.
[0119] S126: Determine the first threshold based on multiple divergence values.
[0120] Specifically, in the image recognition method provided in this application embodiment, after calculating the divergence value between each first vector and the corresponding second vector to obtain multiple divergence values, the variance of the multiple divergence values is calculated, and the variance of the multiple divergence values is determined as the aforementioned first threshold.
[0121] The embodiments provided in this application, before determining the image recognition result based on image recognition information, determine a first vector based on the confidence level obtained each time during the determination or update of confidence; determine a second vector based on each first vector; determine the divergence value between each first vector and the corresponding second vector to obtain multiple divergence values; and determine a first threshold based on the multiple divergence values. In this way, the first threshold is dynamically determined based on the confidence level of the predicted recognition words during the image recognition process, improving the accuracy of the image recognition result when subsequently determining the image recognition result based on the comparison between the first threshold and the confidence level of the image recognition information.
[0122] In this embodiment of the application, after determining the image recognition result, the above image recognition method may further include the following S128:
[0123] S128: If the image recognition result includes a first type of recognition word, replace the first type of recognition word with a second type of recognition word.
[0124] The second type has a wider range of types than the first type.
[0125] Furthermore, the first type of identification words can specifically be identification words containing modifiers, i.e., adjectives, such as man, woman, big tree, red flower, etc. Based on this, the second type of identification words mentioned above can be the identification words of the first type of identification words after removing the modifiers, such as person, tree, flower, etc.
[0126] In practical applications, the specific types of the first type of recognition words and the second type of recognition words, as well as the correspondence between them, can be determined through a preset recognition word replacement template.
[0127] Specifically, in the image recognition method provided in this application embodiment, after obtaining the image recognition result, if the image recognition result includes a first type of recognition word, the first type of recognition word is replaced with a second type of recognition word based on the recognition word replacement template.
[0128] The embodiments provided in this application, after determining the image recognition result, replace the first type of recognition word with a second type of recognition word if the image recognition result includes a first type of recognition word; wherein the range of the second type is greater than the range of the first type. This improves the simplicity of the image recognition result and makes it easier for users to understand.
[0129] In this embodiment of the application, prior to S104, the image recognition method may further include the following S130 to S138:
[0130] S130: Obtain the sample dataset.
[0131] The sample dataset includes multiple sample images and corresponding annotation information for each sample image, with each sample image and annotation information corresponding one-to-one.
[0132] S132: Determine the feature vector of each annotation information based on the order of the character information in the vocabulary.
[0133] Specifically, in the image recognition method provided in this application embodiment, after obtaining the sample dataset, for each annotation information in the sample dataset, each annotation information is split into individual characters, and then the feature vector of each annotation information is determined according to the arrangement order of the character information in the vocabulary.
[0134] For example, given a vocabulary of {0: I, 1: is, 2: one, 3: each, 4: small, 5: child, 6: child}, in the case of the annotation "I am a small child", the order of the characters in the annotation, i.e., the text number, is ground_truth = [0, 1, 2, 3, 4, 5, 6]. Based on this, using the two-dimensional matrix Word2Embed, the row vector of each character is obtained by querying the order of each character in the annotation, i.e., the text number. This row vector is then used as the feature vector of that character, thus obtaining the feature vector of the entire annotation.
[0135] S134: Based on the image recognition model, process the feature vector of each sample image and the corresponding annotation information to obtain the model output information.
[0136] Specifically, in the image recognition method provided in this application embodiment, after determining the feature vector of each annotation information in the sample dataset, the feature vector of each annotation information and the sample image corresponding to each annotation information are input into the image recognition model, so as to process the feature vector of each sample image and the annotation information corresponding to each sample image through the image recognition model and obtain the model output information.
[0137] S136: Determine the loss value based on the order of the characters in the vocabulary of the model output information and each labeled information.
[0138] The aforementioned loss value can be determined using the following formula (10):
[0139] H(x, ground) truth( id ))=-∑ batch , id , k p(k)·logx(batch,0,k),(10)
[0140] Where H(x, ground) truth ( id )) represents the loss value, ground truth ( id ) represents the order of all characters in the annotation information, i.e., the character number. id represents the order of the characters in the annotation information, i.e., the character number. When k = id, p(k) = 1, and in other cases p(k) = 0.
[0141] Furthermore, in practical applications, the labeled information in the aforementioned sample dataset includes positive and negative samples. Positive samples are the original labeled information, while negative samples are obtained by shuffling the order of the characters in the labeled information. For example, swapping the third and first characters in the labeled information yields negative sample 1; swapping the second and first characters yields negative sample 2. In this way, the positive and negative samples of the labeled information in the sample dataset, along with their corresponding ground truth values, can be constructed.
[0142] In this case, when the constructed sample pairs are Q[batch×3, ground_truth], the sample data in the sample pairs are arranged according to [positive sample; negative sample; negative sample]. Based on this, the target value corresponding to the sample pair is [ones(batch, 2); zeros(batch, 2); zeros(batch, 2)], where ones(batch, 2) represents constructing a length of [batch, 2] of all values being 1, and zeros(batch, 2) represents constructing a length of [batch, 2] of all values being 0.
[0143] Based on this, the aforementioned loss value can be specifically determined using the following formula (11):
[0144] H(P, Q) = -∑ i P(i)logQ(i), (11)
[0145] Where H(P, Q) represents the loss value, Q(i) represents the i-th sample pair, and P(i) represents the target value corresponding to sample pair Q(i).
[0146] S138: Iteratively train the image recognition model based on the loss value until the model output information of the image recognition model converges.
[0147] Specifically, in the image recognition method provided in this application embodiment, after processing each group of sample images and annotation information in the sample dataset through the image recognition model and obtaining the corresponding loss value, the image recognition model is iteratively trained according to the loss value, and the next group of sample images and annotation information in the sample dataset is processed through the image recognition model until the model output information of the image recognition model is in a convergent state, and the training of the image recognition model ends.
[0148] The embodiments provided in this application, before processing the first image using the trained image recognition model, acquire a sample dataset, which includes multiple sample images and corresponding annotation information for each sample image; determine the feature vector of each annotation based on the order of character information in the vocabulary; process the feature vectors of each sample image and its corresponding annotation based on the image recognition model to obtain model output information; determine the loss value based on the model output information and the order of character information in the vocabulary; iteratively train the image recognition model based on the loss value until the model output information converges. In this way, before performing image recognition using the trained image recognition model, the ability to judge image-text consistency is trained, ensuring the accuracy of subsequent image recognition using the image recognition model.
[0149] The image recognition method provided in this application can be executed by an image recognition device. This application uses an image recognition device executing the above-described image recognition method as an example to illustrate the image recognition device provided in this application.
[0150] like Figure 7 As shown, this application embodiment provides an image recognition device 700, which may include the acquisition unit 702 and processing unit 704 described below.
[0151] Acquisition unit 702 is used to acquire the first image;
[0152] The processing unit 704 is used to process the first image through the trained image recognition model to obtain the image recognition information of the first image and the confidence level of the image recognition information.
[0153] The processing unit 704 is further configured to determine the image recognition result based on the image recognition information when the confidence level is greater than the first threshold, and to perform classification processing on the image recognition information when the confidence level is less than or equal to the first threshold, and to determine the image recognition result based on the classification processing result.
[0154] The image recognition device 700 provided in this application embodiment acquires a first image; processes the first image using a trained image recognition model to obtain image recognition information and a confidence level of the image recognition information; when the confidence level is greater than a first threshold, it determines an image recognition result based on the image recognition information; when the confidence level is less than or equal to the first threshold, it performs classification processing on the image recognition information and determines the image recognition result based on the classification processing result. Through the above-described image recognition device 700, image recognition information and a confidence level of the first image are obtained based on a trained image recognition model. When the confidence level is greater than the first threshold, the image recognition result is directly determined based on the image recognition information; when the confidence level is less than or equal to the first threshold, a classification algorithm is used to process the image recognition information to obtain an accurate image recognition result. In this way, when the confidence level of the image recognition information is high, the image recognition result can be determined directly based on the image recognition information. When the confidence level of the image recognition information is low, the image recognition result is obtained by combining the classification algorithm. Accurate image recognition results can be obtained in all situations, which improves the judgment ability of the image recognition method, reduces the false alarm probability of the image recognition method, and improves the accuracy of the image recognition result. Furthermore, image recognition based on the image recognition model reduces the application limitations of the image recognition method.
[0155] In this embodiment, the processing unit 704 is specifically used to: adjust the image size of the first image to obtain the second image; extract the feature matrix of the second image through the trained image recognition model; and determine the image recognition information and confidence level based on the trained image recognition model and the feature matrix.
[0156] The embodiments provided in this application, during the process of processing a first image using a trained image recognition model to obtain image recognition information and confidence levels for the first image, adjust the image size of the first image to obtain a second image; extract the feature matrix of the second image using the trained image recognition model; and determine the image recognition information and confidence levels based on the feature matrix using the trained image recognition model. In this way, determining image recognition information and confidence levels based on a small-volume image recognition model and the image's feature matrix ensures the accuracy of the image recognition results. Furthermore, the small-volume image recognition model is easy to implement in mobile devices, reducing the application limitations of image recognition methods.
[0157] In this embodiment, the feature matrix includes N feature vectors, which correspond to N recognition words in the vocabulary. The processing unit 704 is specifically used to: construct a cache queue with a length of a second threshold, and fill the feature vector corresponding to the first recognition word in the vocabulary into the first position of the cache queue, wherein the first position is the initial position of the cache queue; determine the confidence level and a first sorting order according to the cache queue and the feature matrix, and fill the feature vector corresponding to the recognition word in the vocabulary in the first sorting order into the second position of the cache queue to update the cache queue, wherein the second position is the next in order after the first position; update the confidence level and determine the second sorting order according to the feature matrix and the updated cache queue, until the determined sorting order meets the first condition or the number of confidence level updates reaches the second threshold; and determine the image recognition information according to the recognition word corresponding to each feature vector in the cache queue after the last update.
[0158] The embodiments provided in this application include a feature matrix comprising N feature vectors, each corresponding to one of the N recognition words in a vocabulary. During the process of determining image recognition information and confidence levels based on the feature matrix, a cache queue of length denoted by a second threshold is constructed. The feature vector corresponding to the first recognition word in the vocabulary is then filled into the first position of the cache queue, where the first position is the initial position of the cache queue. Based on the cache queue and the feature matrix, the confidence level and a first arrangement order are determined. The feature vector corresponding to the recognition word in the first arrangement order is then filled into the second position of the cache queue to update the cache queue, where the second position is the next in order after the first position. Based on the feature matrix and the updated cache queue, the confidence level is updated and a second arrangement order is determined until the determined arrangement order meets the first condition or the number of confidence level updates reaches the second threshold. Finally, image recognition information is determined based on the recognition word corresponding to each feature vector in the last updated cache queue. In this way, recognition words are predicted based on the image's feature vectors, and a fixed-length cache queue is used to output image recognition information, improving the accuracy of the image recognition results.
[0159] In this embodiment of the application, the processing unit 704 is specifically used to: add a mask to the cache queue; perform self-attention learning on the cache queue and feature matrix after adding the mask to obtain a first weight matrix, and perform weighted processing on the feature matrix according to the first weight matrix; and perform nonlinear mapping on the feature matrix to adjust the dimension of the feature matrix if the dimension of the weighted feature matrix meets the second condition.
[0160] The embodiments provided in this application, before determining the confidence level and the first ranking order based on the cache queue and the feature matrix, add a mask to the cache queue; perform self-attention learning on the masked cache queue and the feature matrix to obtain a first weight matrix, and then perform weighted processing on the feature matrix according to the first weight matrix; if the dimension of the weighted feature matrix meets the second condition, perform nonlinear mapping on the feature matrix to adjust the dimension of the feature matrix. In this way, on the one hand, by performing a masking operation on the cache queue, the input part of the recognition generation module is filtered to prevent the recognition generation module from seeing the subsequent input parts, thus improving the accuracy of image recognition; on the other hand, by performing weighted processing on the feature matrix, the image recognition model can focus on more useful information, thereby improving the accuracy of image recognition; furthermore, by performing nonlinear mapping on the feature matrix, the complexity of the content expressed by the image recognition model can be increased, enabling the image recognition model to learn more complex textual expressions, thus improving the accuracy of image recognition based on the feature matrix.
[0161] In this embodiment of the application, the processing unit 704 is further configured to: determine a first vector based on the confidence obtained each time in the process of determining or updating the confidence; determine a second vector based on each first vector; determine the divergence value between each first vector and the corresponding second vector to obtain multiple divergence values; and determine a first threshold based on the multiple divergence values.
[0162] The embodiments provided in this application, before determining the image recognition result based on image recognition information, determine a first vector based on the confidence level obtained each time during the determination or update of confidence; determine a second vector based on each first vector; determine the divergence value between each first vector and the corresponding second vector to obtain multiple divergence values; and determine a first threshold based on the multiple divergence values. In this way, the first threshold is dynamically determined based on the confidence level of the predicted recognition words during the image recognition process, improving the accuracy of the image recognition result when subsequently determining the image recognition result based on the comparison between the first threshold and the confidence level of the image recognition information.
[0163] In this embodiment of the application, the processing unit 704 is further configured to: replace the first type of recognition word with a second type of recognition word when the image recognition result includes a first type of recognition word; wherein the type range of the second type is greater than the type range of the first type.
[0164] The embodiments provided in this application, after determining the image recognition result, replace the first type of recognition word with a second type of recognition word if the image recognition result includes a first type of recognition word; wherein the range of the second type is greater than the range of the first type. This improves the simplicity of the image recognition result and makes it easier for users to understand.
[0165] In this embodiment, the acquisition unit 702 is further configured to: acquire a sample dataset, the sample dataset including multiple sample images and annotation information corresponding to each sample image; the processing unit 704 is further configured to: determine the feature vector of each annotation information according to the order of the character information in the vocabulary; process the feature vector of each sample image and the annotation information corresponding to each sample image based on the image recognition model to obtain model output information; determine the loss value according to the model output information and the order of the character information in the vocabulary; and iteratively train the image recognition model according to the loss value until the model output information of the image recognition model is in a convergent state.
[0166] The embodiments provided in this application, before processing the first image using the trained image recognition model, acquire a sample dataset, which includes multiple sample images and corresponding annotation information for each sample image; determine the feature vector of each annotation based on the order of character information in the vocabulary; process the feature vectors of each sample image and its corresponding annotation based on the image recognition model to obtain model output information; determine the loss value based on the model output information and the order of character information in the vocabulary; and iteratively train the image recognition model based on the loss value until the model output information converges. In this way, the ability of the image recognition model to judge image-text consistency is trained before image recognition is performed, ensuring the accuracy of subsequent image recognition.
[0167] The image recognition device 700 in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific type of device.
[0168] The image recognition device 700 in this embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this embodiment does not specifically limit its use.
[0169] The image recognition device 700 provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0170] Optionally, such as Figure 8 As shown, this application embodiment also provides an electronic device 800, including a processor 802 and a memory 804. The memory 804 stores a program or instructions that can run on the processor 802. When the program or instructions are executed by the processor 802, they implement the various steps of the above-described image recognition method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0171] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.
[0172] Figure 9 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0173] The electronic device 900 includes, but is not limited to, components such as: radio frequency unit 901, network module 902, audio output unit 903, input unit 904, sensor 905, display unit 906, user input unit 907, interface unit 908, memory 909, and processor 910.
[0174] Those skilled in the art will understand that the electronic device 900 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 910 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 9 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0175] The processor 910 is used to acquire the first image.
[0176] The processor 910 is also used to process the first image through the trained image recognition model to obtain the image recognition information of the first image and the confidence level of the image recognition information.
[0177] The processor 910 is also used to determine the image recognition result based on the image recognition information when the confidence level is greater than the first threshold, and to classify the image recognition information when the confidence level is less than or equal to the first threshold, and to determine the image recognition result based on the classification result.
[0178] In this embodiment, a first image is acquired; the first image is processed by a trained image recognition model to obtain image recognition information and a confidence level of the image recognition information; if the confidence level is greater than a first threshold, an image recognition result is determined based on the image recognition information; if the confidence level is less than or equal to the first threshold, the image recognition information is classified, and the image recognition result is determined based on the classification result. In this embodiment, the image recognition information and the confidence level of the first image are obtained based on a trained image recognition model. If the confidence level is greater than the first threshold, the image recognition result is directly determined based on the image recognition information; if the confidence level is less than or equal to the first threshold, a classification algorithm is used to process the image recognition information to obtain an accurate image recognition result. In this way, when the confidence level of the image recognition information is high, the image recognition result can be determined directly based on the image recognition information. When the confidence level of the image recognition information is low, the image recognition result is obtained by combining the classification algorithm. Accurate image recognition results can be obtained in all situations, which improves the judgment ability of the image recognition method, reduces the false alarm probability of the image recognition method, and improves the accuracy of the image recognition result. Furthermore, image recognition based on the image recognition model reduces the application limitations of the image recognition method.
[0179] Optionally, the processor 910 is specifically used to: adjust the image size of the first image to obtain the second image; extract the feature matrix of the second image through the trained image recognition model; and determine the image recognition information and confidence level based on the feature matrix according to the trained image recognition model.
[0180] The embodiments provided in this application, during the process of processing a first image using a trained image recognition model to obtain image recognition information and confidence levels for the first image, adjust the image size of the first image to obtain a second image; extract the feature matrix of the second image using the trained image recognition model; and determine the image recognition information and confidence levels based on the feature matrix using the trained image recognition model. In this way, determining image recognition information and confidence levels based on a small-volume image recognition model and the image's feature matrix ensures the accuracy of the image recognition results. Furthermore, the small-volume image recognition model is easy to implement in mobile devices, reducing the application limitations of image recognition methods.
[0181] Optionally, the feature matrix includes N feature vectors, each corresponding to one of the N recognition words in the vocabulary. The processor 910 is specifically configured to: construct a cache queue of length equidistant from a second threshold, and fill the first position of the cache queue with the feature vector corresponding to the first recognition word in the vocabulary, where the first position is the initial position of the cache queue; determine the confidence level and a first arrangement order based on the cache queue and the feature matrix, and fill the second position of the cache queue with the feature vector corresponding to the recognition word in the first arrangement order, thereby updating the cache queue, where the second position is the next in order after the first position; update the confidence level and determine a second arrangement order based on the feature matrix and the updated cache queue, until the determined arrangement order meets the first condition or the number of confidence level updates reaches the second threshold; and determine image recognition information based on the recognition word corresponding to each feature vector in the last updated cache queue.
[0182] The embodiments provided in this application include a feature matrix comprising N feature vectors, each corresponding to one of the N recognition words in a vocabulary. During the process of determining image recognition information and confidence levels based on the feature matrix, a cache queue of length denoted by a second threshold is constructed. The feature vector corresponding to the first recognition word in the vocabulary is then filled into the first position of the cache queue, where the first position is the initial position of the cache queue. Based on the cache queue and the feature matrix, the confidence level and a first arrangement order are determined. The feature vector corresponding to the recognition word in the first arrangement order is then filled into the second position of the cache queue to update the cache queue, where the second position is the next in order after the first position. Based on the feature matrix and the updated cache queue, the confidence level is updated and a second arrangement order is determined until the determined arrangement order meets the first condition or the number of confidence level updates reaches the second threshold. Finally, image recognition information is determined based on the recognition word corresponding to each feature vector in the last updated cache queue. In this way, recognition words are predicted based on the image's feature vectors, and a fixed-length cache queue is used to output image recognition information, improving the accuracy of the image recognition results.
[0183] Optionally, the processor 910 is specifically used to: add a mask to the cache queue; perform self-attention learning on the cache queue and feature matrix after adding the mask to obtain a first weight matrix, and perform weighted processing on the feature matrix according to the first weight matrix; and perform non-linear mapping on the feature matrix to adjust the dimension of the feature matrix if the dimension of the weighted feature matrix meets the second condition.
[0184] The embodiments provided in this application, before determining the confidence level and the first ranking order based on the cache queue and the feature matrix, add a mask to the cache queue; perform self-attention learning on the masked cache queue and the feature matrix to obtain a first weight matrix, and then perform weighted processing on the feature matrix according to the first weight matrix; if the dimension of the weighted feature matrix meets the second condition, perform nonlinear mapping on the feature matrix to adjust the dimension of the feature matrix. In this way, on the one hand, by performing a masking operation on the cache queue, the input part of the recognition generation module is filtered to prevent the recognition generation module from seeing the subsequent input parts, thus improving the accuracy of image recognition; on the other hand, by performing weighted processing on the feature matrix, the image recognition model can focus on more useful information, thereby improving the accuracy of image recognition; furthermore, by performing nonlinear mapping on the feature matrix, the complexity of the content expressed by the image recognition model can be increased, enabling the image recognition model to learn more complex textual expressions, thus improving the accuracy of image recognition based on the feature matrix.
[0185] Optionally, the processor 910 is further configured to: determine a first vector based on the confidence obtained each time during the process of determining or updating the confidence; determine a second vector based on each first vector; determine the divergence value between each first vector and the corresponding second vector to obtain multiple divergence values; and determine a first threshold based on the multiple divergence values.
[0186] The embodiments provided in this application, before determining the image recognition result based on image recognition information, determine a first vector based on the confidence level obtained each time during the determination or update of confidence; determine a second vector based on each first vector; determine the divergence value between each first vector and the corresponding second vector to obtain multiple divergence values; and determine a first threshold based on the multiple divergence values. In this way, the first threshold is dynamically determined based on the confidence level of the predicted recognition words during the image recognition process, improving the accuracy of the image recognition result when subsequently determining the image recognition result based on the comparison between the first threshold and the confidence level of the image recognition information.
[0187] Optionally, the processor 910 is further configured to: replace the first type of recognition word with a second type of recognition word when the image recognition result includes a first type of recognition word; wherein the type range of the second type is greater than the type range of the first type.
[0188] The embodiments provided in this application, after determining the image recognition result, replace the first type of recognition word with a second type of recognition word if the image recognition result includes a first type of recognition word; wherein the range of the second type is greater than the range of the first type. This improves the simplicity of the image recognition result and makes it easier for users to understand.
[0189] Optionally, the processor 910 is further configured to: acquire a sample dataset, the sample dataset including multiple sample images and annotation information corresponding to each sample image; the processor 910 is further configured to: determine the feature vector of each annotation information according to the order of the character information in the vocabulary; based on the image recognition model, process the feature vector of each sample image and the annotation information corresponding to each sample image to obtain model output information; determine the loss value according to the model output information and the order of the character information in the vocabulary; and iteratively train the image recognition model according to the loss value until the model output information of the image recognition model converges.
[0190] The embodiments provided in this application, before processing the first image using the trained image recognition model, acquire a sample dataset, which includes multiple sample images and corresponding annotation information for each sample image; determine the feature vector of each annotation based on the order of character information in the vocabulary; process the feature vectors of each sample image and its corresponding annotation based on the image recognition model to obtain model output information; determine the loss value based on the model output information and the order of character information in the vocabulary; and iteratively train the image recognition model based on the loss value until the model output information converges. In this way, the ability of the image recognition model to judge image-text consistency is trained before image recognition is performed, ensuring the accuracy of subsequent image recognition.
[0191] It should be understood that, in this embodiment, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042. The GPU 9041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 906 may include a display panel 9061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include a touch detection device and a touch controller. Other input devices 9072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0192] The memory 909 can be used to store software programs and various data. The memory 909 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 909 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 909 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0193] Processor 910 may include one or more processing units; optionally, processor 910 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 910.
[0194] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image recognition method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0195] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0196] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described image recognition method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0197] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0198] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the image recognition method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0199] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0200] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0201] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image recognition method, characterized in that, The method includes: Get the first image; The first image is processed by the trained image recognition model to obtain the image recognition information of the first image and the confidence level of the image recognition information. If the confidence level is greater than a first threshold, the image recognition result is determined based on the image recognition information; if the confidence level is less than or equal to the first threshold, the image recognition information is classified, and the image recognition result is determined based on the classification result. The step of determining the image recognition result based on the classification processing result includes: The classification results are combined into a complete image recognition statement. The image recognition statement is used as the image recognition result for the first image; The step of processing the first image using a trained image recognition model to obtain image recognition information of the first image and the confidence level of the image recognition information includes: Adjust the image size of the first image to obtain the second image; The feature matrix of the second image is extracted using the trained image recognition model; Based on the trained image recognition model, the image recognition information and the confidence level are determined according to the feature matrix; The feature matrix includes N feature vectors, each corresponding to one of N recognition words in a vocabulary, where N is a positive integer greater than 1. Determining the image recognition information and the confidence level based on the feature matrix includes: Construct a cache queue with a length of a second threshold, and fill the first position of the cache queue with the feature vector corresponding to the first recognized word in the vocabulary, wherein the first position is the initial position of the cache queue; Based on the cache queue and the feature matrix, the confidence level and the first sorting order are determined, and the feature vectors corresponding to the recognition words in the vocabulary that are located in the first sorting order are filled into the second position of the cache queue to update the cache queue, wherein the second position is located in the next order after the first position; Based on the feature matrix and the updated cache queue, the confidence level is updated and a second sorting order is determined until the determined sorting order meets the first condition or the number of times the confidence level is updated reaches the second threshold. Based on the recognition word corresponding to each feature vector in the cache queue after the last update, the image recognition information is determined.
2. The image recognition method according to claim 1, characterized in that, Before determining the confidence level and the first ranking order based on the cache queue and the feature matrix, the image recognition method further includes: Add a mask to the cache queue; Self-attention learning is performed on the cache queue and the feature matrix after adding the mask to obtain a first weight matrix, and the feature matrix is weighted according to the first weight matrix; If the dimension of the weighted feature matrix meets the second condition, a nonlinear mapping is performed on the feature matrix to adjust its dimension.
3. The image recognition method according to claim 1, characterized in that, Before determining the image recognition result based on the image recognition information, the image recognition method further includes: The first vector is determined based on the confidence level obtained each time during the process of determining or updating the confidence level; Determine the second vector based on each of the first vectors; Determine the divergence value between each of the first vectors and the corresponding second vectors to obtain multiple divergence values; The first threshold is determined based on a plurality of said divergence values.
4. The image recognition method according to any one of claims 1 to 3, characterized in that, After determining the image recognition result, the image recognition method further includes: If the image recognition result includes a first type of recognition word, the first type of recognition word is replaced with a second type of recognition word; The type range of the second type is greater than that of the first type.
5. The image recognition method according to any one of claims 1 to 3, characterized in that, Before processing the first image using the trained image recognition model, the image recognition method further includes: Obtain a sample dataset, which includes multiple sample images and annotation information corresponding to each sample image; The feature vector of each annotation is determined based on the order of the characters in the vocabulary. Based on the image recognition model, the feature vectors of each sample image and the corresponding annotation information of each sample image are processed to obtain the model output information; The loss value is determined based on the model output information and the order of the character information in each annotation information in the vocabulary. The image recognition model is trained iteratively based on the loss value until the model output information of the image recognition model converges.
6. An image recognition device, characterized in that, The image recognition device includes: The acquisition unit is used to acquire the first image; The processing unit is configured to process the first image using a trained image recognition model to obtain image recognition information of the first image and the confidence level of the image recognition information. The processing unit is further configured to determine an image recognition result based on the image recognition information when the confidence level is greater than a first threshold, and to perform classification processing on the image recognition information when the confidence level is less than or equal to the first threshold, and to determine the image recognition result based on the classification processing result. The processing unit is specifically used to synthesize the classification processing result into a complete image recognition statement; use the image recognition statement as the image recognition result of the first image; and adjust the image size of the first image to obtain the second image. The feature matrix of the second image is extracted using the trained image recognition model; Based on the trained image recognition model, the image recognition information and the confidence level are determined according to the feature matrix; The feature matrix includes N feature vectors, and the N feature vectors correspond to N recognition words in the vocabulary, where N is a positive integer greater than 1. The processing unit is specifically used to construct a cache queue with a length of a second threshold, and fill the feature vector corresponding to the first recognition word in the vocabulary into the first position of the cache queue, wherein the first position is the initial position of the cache queue. Based on the cache queue and the feature matrix, the confidence level and the first sorting order are determined, and the feature vectors corresponding to the recognition words in the vocabulary that are located in the first sorting order are filled into the second position of the cache queue to update the cache queue, wherein the second position is located in the next order after the first position; Based on the feature matrix and the updated cache queue, the confidence level is updated and a second sorting order is determined until the determined sorting order meets the first condition or the number of times the confidence level is updated reaches the second threshold. Based on the recognition word corresponding to each feature vector in the cache queue after the last update, the image recognition information is determined.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the image recognition method as described in any one of claims 1 to 5.
8. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image recognition method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Image recognition method and device, computer equipment and storage medium
CN111523621A
Tumble behavior detection method and device
CN113642361A
Image recognition method and device, network equipment and computer readable storage medium
CN113780315A
Chest CT image classification model training method, chest CT image classification method, chest CT image classification system and chest CT image classification equipment
CN116630303A