Multi-instruction image recognition method and device, computer equipment and storage medium
Through the pre-trained image processing model combined with voice information, the problem that image recognition in the prior art is difficult to fuse voice information, and more efficient and real-time image area recognition is achieved, which is suitable for small computing devices.
Patent Information
- Application Number
- CN202510379235.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-25
AI Technical Summary
The existing image recognition technology is difficult to effectively integrate voice information, which makes the camera's shooting needs in complex scenarios unable to meet, and the traditional multimodal model has a large amount of computing and excessive memory usage, making it impossible to adapt to small computing devices.
A pre-trained image processing model is used, combined with speech information as a guiding condition, and a mask image is output to determine the desired area by fusing multimodal information of the speech processing layer and the image processing layer.
It realizes more accurate image area recognition, reduces user operation difficulty, improves processing efficiency, meets real-time requirements, and is suitable for small computing devices.
Smart Images

Figure CN120375043A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of neural network models, and particularly to a multi-instruction image recognition method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of computer and Internet technologies, the application of scene recognition has become increasingly widespread. Currently, common control methods include manually controlling the camera direction and voice commands, but there are many problems in actual applications. On the one hand, the diversity of control methods makes it difficult to comprehensively understand the instructions. For example, there may be only manual control without voice commands, or the things in the direction targeted by voice and manual control have no relation at all, which makes it difficult for the camera to clearly identify the image area that the user really expects to pre-edit.
[0003] On the other hand, existing technical means also have limitations. The traditional "salient detection" technology only calculates based on image information and cannot effectively integrate voice information, making it difficult to meet the shooting requirements in complex scenarios. Although there are multi-modal models that can comprehensively process language and images currently, such models have problems of huge computational complexity and excessive memory occupancy, resulting in their inability to adapt to small computing devices such as cameras, severely restricting their applications in camera shooting scenarios. Summary of the Invention
[0004] Based on this, it is necessary to provide a multi-instruction image recognition method, apparatus, computer device, and storage medium for the above technical problems.
[0005] In a first aspect, the present disclosure provides a multi-instruction image recognition method. Applied to a shooting device, the method includes:
[0006] Obtain a panoramic image, and determine a search area in the panoramic image, where the search area includes an expected area;
[0007] In response to detecting voice information associated with the expected area, use the voice information as a guiding condition, and process the search area using a pre-trained image processing model to output a mask image; wherein, the image processing model includes: a voice processing layer and an image processing layer, the voice processing layer is connected to the image processing layer, and is used to recognize the voice information as text information, encode the text information, and then transmit it to the image processing layer; the image processing layer is used to encode the search area, receive the encoded text information, fuse the encoded text information with the encoded search area, and decode the fused result to obtain a mask image;
[0008] Based on the mask image and the search area, determine the expected area.
[0009] In one embodiment, the image processing layer includes: a plurality of sequentially connected encoding layers, a feature fusion layer matching the plurality of encoding layers, and a plurality of sequentially connected decoding layers matching the plurality of encoding layers;
[0010] The encoding layer is used to encode an image, the feature fusion layer matching the plurality of encoding layers is used to receive the text information encoded by the speech processing layer, and fuse the encoded image information and the encoded text information; the decoding layer is used to receive the fused information and output the decoded fused information.
[0011] In one embodiment, the encoding layer is composed of residual blocks, and the residual blocks encode the image using convolutional operations;
[0012] The decoding layer is composed of inverse residual blocks matching the residual blocks, and the inverse residual blocks decode the fused information through transposed convolution and skip connection splicing.
[0013] In one embodiment, the residual blocks in the plurality of encoding layers encode the sequentially received data and then transmit it to the next residual block and the feature fusion layer;
[0014] The feature fusion layer fuses the encoded data and then transmits it to the inverse residual block matching the residual block that performs the encoding;
[0015] The inverse residual blocks in the plurality of decoding layers sequentially decode the received data and then transmit it to the previous inverse residual block, and the inverse residual module of the first decoding layer in the plurality of decoding layers outputs a mask image.
[0016] In one embodiment, the speech processing layer includes a text encoding layer, and the encoding layer and the text encoding layer are CLIP encodings; the image processing model further includes: a class recognition layer connected to the encoding layer, and the class recognition layer is used to recognize the class of the object in the image;
[0017] The method further includes:
[0018] Obtain training parameter data;
[0019] Load the pre-trained parameters of the CLIP encoding into the image processing model to be trained, freeze the weights of the CLIP encoding, and control the CLIP not to participate in the training;
[0020] Use the training parameter data to train the feature fusion layer, decoding layer and class recognition layer in the image processing model to be trained;
[0021] Calculate a loss value based on the result input by the class recognition layer and the text description result in the training parameter data;
[0022] Iteratively adjust the parameters in the feature fusion layer, decoding layer, and class recognition layer of the image processing model to be trained based on the loss value to obtain the image processing model.
[0023] In one embodiment, the CBN (Conditional Batch Normalization) technique is used in the feature fusion layer to link the text information encoded by the speech processing layer to the encoding layer through an MLP (Multilayer Perceptron).
[0024] In one embodiment, the method further includes:
[0025] During the inference process of the image processing model, turn off the class recognition layer in the image processing model.
[0026] In a second aspect, the present disclosure also provides a multi-instruction image recognition device. Applied to a shooting device, the device includes:
[0027] A search area determination module, configured to obtain a panoramic image and determine a search area in the panoramic image, where the search area includes an expected area;
[0028] A model processing module, configured to, in response to detecting speech information associated with the expected area, use the speech information as a guiding condition and process the search area using a pre-trained image processing model to output a mask image; wherein, the image processing model includes: a speech processing layer and an image processing layer, the speech processing layer is connected to the image processing layer, and is configured to recognize the speech information as text information, encode the text information, and transmit it to the image processing layer; the image processing layer is configured to encode the search area, receive the encoded text information, fuse the encoded text information with the encoded search area, and decode the fused result to obtain a mask image;
[0029] An expected area determination module, configured to determine an expected area based on the mask image and the search area.
[0030] In a third aspect, the present disclosure also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in any of the above method embodiments are implemented.
[0031] In a fourth aspect, the present disclosure also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0032] In a fifth aspect, the present disclosure also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the steps in any of the above method embodiments.
[0033] In the above embodiments, by using a pre-trained image processing model and combining voice information as a guiding condition to process the search area. The voice information can provide richer semantic descriptions, enabling the model to more accurately understand the user's intention, thereby more precisely identifying and locating the desired area. Introducing voice information as an interaction method, users can interact with the system through voice commands without the need to manually and tediously set parameters or make selections in the image. This interaction method is more natural and convenient, reducing the operation difficulty for users. The voice processing layer converts the voice information into text information and encodes it, and then the image processing layer fuses and processes it with the image encoding information. The entire process reflects the system's comprehensive understanding and processing ability of two different modal information, namely voice and image. This multi-modal information fusion method enables the system to have a higher level of intelligence, being able to better simulate the human perception and understanding process, and thus more intelligently complete image analysis tasks. By delimiting the search area, it avoids indiscriminate processing of the entire panoramic image, reduces the amount of data processing, and improves the processing efficiency. At the same time, the pre-training of the model enables rapid processing of the input voice and image information and output of results in practical applications, meeting the real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 It is a schematic flowchart of a multi-instruction image recognition method in an embodiment;
[0036] Figure 2 It is a schematic structural diagram of an image processing model in an embodiment;
[0037] Figure 3 It is a schematic structural diagram of an image processing layer in an image processing model in an embodiment;
[0038] Figure 4 It is a schematic structural diagram of an encoding layer, a decoding layer, and a feature fusion layer in an image processing model in an embodiment;
[0039] Figure 5Schematic diagram of the structure of the category recognition layer in the image processing model in one embodiment;
[0040] Figure 6 Schematic flow chart of the training process of the image processing model in one embodiment;
[0041] Figure 7 Block diagram of the structure of the multi-instruction image recognition device in one embodiment;
[0042] Figure 8 Schematic diagram of the internal structure of a computer device in one embodiment. Detailed implementation manners
[0043] In order to make the objectives, technical solutions and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.
[0044] It should be noted that the terms "first", "second", etc. in the specification and claims of this article and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described in this article can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or equipment.
[0045] In this article, the term "and / or" is only a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0046] In one embodiment, as Figure 1 shown, a multi-instruction image recognition method is provided. In this embodiment, it is exemplified that the method is applied to a terminal. The terminal can usually be a shooting device, such as a mobile phone, a panoramic camera, etc. The method includes the following steps:
[0047] S102. Obtain a panoramic image, determine a search area in the panoramic image, and the search area includes a desired area.
[0048] Among them, the search area is usually a specific range delimited in the panoramic image, which is set to search for specific targets, features or information in the panoramic image. The size, shape and position of the search area can be determined according to specific tasks and requirements, and in some embodiments of the present disclosure, there is no limitation on how the search area is determined. The desired area is usually the part of the search area that the user is really interested in, hopes to find or pay attention to. It may be a specific object in the image, a part of the scene or an area with certain specific features.
[0049] S104. In response to detecting voice information associated with the desired area, using the voice information as a guiding condition, and processing the search area by using a pre-trained image processing model to output a mask image; wherein, the image processing model includes: a voice processing layer and an image processing layer, the voice processing layer is connected to the image processing layer, and is configured to recognize the voice information as text information, encode the text information and then transmit it to the image processing layer; the image processing layer is configured to encode the search area, receive the encoded text information, fuse the encoded text information with the encoded search area, and decode the fused result to obtain a mask image.
[0050] Among them, voice information refers to the audio signal collected by a voice input device (such as a microphone), which is processed and then converted into information that can be understood by a computer. In this context, the voice information has a certain association with the desired area, which may be a description or feature explanation of the desired area. The guiding condition in the embodiments of the present disclosure refers to the rule or direction that should be followed when the image processing model processes the search area based on the voice information. The voice information provides clues related to a specific target or task, guiding the model to perform corresponding processing operations. The image processing model is a model constructed based on machine learning or deep learning algorithms. After being trained with a large amount of image data, it can perform various processes on the input image. This model is used to process the search area according to the guidance of the voice information. Mask image: It is a special image, which usually consists of a binary image (only two pixel values of 0 and 1) or a grayscale image, and is used to represent the selection or masking of a specific area in the image. In image processing, the mask image can be used to highlight or hide certain parts of the original image for further analysis or processing. For example Figure 2As shown in the figure, it is a schematic diagram of the structure of an image processing model. The image processing model is a comprehensive model architecture composed of a speech processing layer and an image processing layer, which is used to process the search area in the panoramic image according to the speech information and output a mask image. It integrates the functions of speech processing and image processing to achieve specific image analysis tasks. The speech processing layer is usually part of the image processing model and is responsible for processing the input speech information. Its main function is to convert the speech information into text information and encode the text information so that it can interact and transfer data with the image processing layer in the future. Image processing layer: Another important part of the image processing model, mainly processes the image data of the search area. It includes encoding the image, receiving the encoded text information from the speech processing layer, fusing the two, and finally decoding the fusion result to obtain the mask image. The text information is the text content obtained by the speech processing layer after recognizing the speech information. These text information contain key information such as descriptions and instructions related to the desired area, which are important bases for guiding the image processing layer to perform image analysis. Encoding During the model processing, the text information and image data are converted into a numerical form (such as vectors, tensors, etc.) that the model can understand and process. The purpose of encoding is to convert different types of data (text information corresponding to speech and image data) into a unified format for subsequent processing and fusion operations. Fusion Combines the encoded text information and the encoded search area image data, so that the information of the two is correlated and integrated. Through the fusion operation, the model can use the semantic information in the text information to guide the analysis of the image, improving the accuracy and pertinence of the processing. Decoding Opposite to encoding, it is the process of converting the encoded result obtained after fusion back into the image form (i.e., the mask image). Through the decoding operation, the numerical representation processed by the model is restored to an image with practical significance, thus realizing the processing of the search area and the extraction of the target.
[0051] Specifically, the microphone in the shooting device can be used to collect the speech signal in the environment in real time and convert the collected speech signal into text information. When the information described in the text information is related to the current scene of the panoramic image (for example, if the text information describes a certain scenery or an object in the scenery, and the shooting scene of the current panoramic image is also a scenery scene, it can be considered related). Input the speech information and the search area into the pre-trained image processing model, and use the image processing model to output the mask image.
[0052] S106, determine the desired area based on the mask image and the search area.
[0053] Specifically, since the mask images are processed based on the search regions, they are corresponding in spatial position. Align the mask images with the search region images at the pixel level to ensure that the coordinate systems of the two are consistent. For example, if the search region is a rectangular area cropped from a panoramic image, the mask images should also correspond to the size and position of this rectangular area. For binary mask images, traverse each pixel point of the mask image. When the pixel value is 1, the pixel point at the same position in the search region is considered to be part of the possible desired region; when the pixel value is 0, the pixel point at the same position in the search region is excluded from the desired region. For grayscale mask images, a threshold can be set (adjusted according to the actual situation). When the pixel grayscale value is greater than or equal to this threshold, the pixel point at the same position in the search region is included in the possible desired region; the part of the search region corresponding to the pixel points with values less than this threshold is excluded.
[0054] In addition, in some scenarios, in order to make the selected desired region more accurate and complete, morphological processing can be performed on the region selected based on the mask images. For example, use the dilation operation to expand the smaller connected parts in the desired region to make it closer to the shape of the real desired region; use the erosion operation to remove some noise points or isolated pixel points to make the region smoother and more accurate.
[0055] In the above multi-instruction image recognition method, by using a pre-trained image processing model and combining voice information as a guiding condition to process the search region. Voice information can provide richer semantic descriptions, enabling the model to more accurately understand the user's intention, and thus more precisely identify and locate the desired region. Introducing voice information as an interaction method, users can interact with the system through voice commands without manually setting parameters tediously or making selections in the image. This interaction method is more natural and convenient, reducing the operation difficulty for users. The voice processing layer converts the voice information into text information and encodes it, and then the image processing layer fuses and processes it with the image encoding information. The whole process reflects the system's comprehensive understanding and processing ability of two different modal information, voice and image. This multi-modal information fusion method enables the system to have a higher level of intelligence, better simulate the human perception and understanding process, and thus more intelligently complete the image analysis task. By delimiting the search region, it avoids the indiscriminate processing of the entire panoramic image, reduces the data processing volume, and improves the processing efficiency. At the same time, the pre-training of the model enables the rapid processing of the input voice and image information and the output of results in practical applications to meet the real-time requirements.
[0056] In one embodiment, as Figure 3As shown, the image processing layer includes: a plurality of sequentially connected encoding layers, a feature fusion layer matching the plurality of encoding layers, and a plurality of sequentially connected decoding layers matching the plurality of encoding layers.
[0057] The encoding layers are used to encode images, the feature fusion layer matching the plurality of encoding layers is used to receive the encoded text information from the speech processing layer, and fuse the encoded image information and the encoded text information; the decoding layers are used to receive the fused information and output the decoded fused information.
[0058] Specifically, the main function of the encoding layer is to encode the input image and extract the features of the image. Generally, a convolutional neural network (CNN) can be used to implement the encoding layer. Each encoding layer usually contains convolutional layers. The convolutional layers perform convolutional operations on the input image through convolutional kernels of different sizes (such as 3x3, 5x5, etc.) to extract the local features of the image. For example, in the first encoding layer, the input image first undergoes convolution through a 3x3 convolutional kernel to generate a set of feature maps, which contain preliminary edge, texture, and other information of the image. To reduce the data volume and extract more advanced semantic features, downsampling operations such as max pooling or average pooling are usually used in the encoding layer. For example, after several convolutional operations, a 2x2 max pooling layer is used to halve the size of the feature map while retaining the most significant features. Multiple encoding layers are connected in sequence. Each layer further extracts more advanced and abstract features based on the features extracted by the previous layer. As the encoding layer deepens, the size of the feature map gradually decreases, while the number of channels gradually increases to capture richer image information. The feature fusion layer is used to receive the encoded text information from the speech processing layer and fuse it with the encoded image information. The encoded text information transmitted from the speech processing layer is usually a vector of a fixed dimension (such as a word embedding vector). First, some transformations may be required on this vector to make its dimension match the channel dimension of the image feature map for the fusion operation. Attention mechanisms and CBN (Conditional Batch Normalization) techniques can be used for the fusion operation. The decoding layer is used to receive the fused information and decode it for output. Generally, operations such as transposed convolution are used to implement the decoding layer. The decoding layer contains multiple transposed convolutional layers. The transposed convolutional layers perform upsampling operations on the input feature map through learned parameters to restore the size of the image. For example, a 2x2 transposed convolutional kernel is used to double the size of the input feature map. The decoding layer matches the encoding layer and is usually the inverse process of the encoding layer in terms of structure. During the decoding process, the fused high-level features are gradually restored to a meaningful image representation. For example, in the last decoding layer, a feature map of the same size as the original input image is output, and this feature map can be further processed (such as using an activation function) to obtain the final output, such as a mask image. During the decoding process, the output of the decoding layer may be connected (skip connection) to the output of the corresponding encoding layer to fuse features at different levels and improve the decoding accuracy and detail restoration ability. This connection method allows the decoding layer to utilize the underlying detail information retained in the encoding layer.
[0059] In some exemplary embodiments, the encoding layer is composed of residual blocks, and the residual blocks encode images using convolutional operations; the decoding layer is composed of inverse residual blocks matching the residual blocks, and the inverse residual blocks decode the fused information through transposed convolution and skip connection splicing.
[0060] Specifically, each residual block typically contains a convolutional layer. First, the input image enters the first convolutional layer, which performs a convolutional operation on the image using a convolutional kernel of a specific size (such as 3x3), extracts local features of the image through the convolutional operation, and obtains a set of feature maps. Then, the feature maps enter the second convolutional layer of the second residual block, and a convolutional operation is performed again to further extract more abstract features. Between these two convolutional layers, batch normalization (BatchNormalization, BN) and an activation function (such as ReLU) are usually used to normalize and non-linearly transform the data to improve the training efficiency and performance of the model. In addition to the convolutional operation, the residual block also introduces a skip connection. The skip connection directly adds the input image or the feature maps after partial convolutional operations to the feature maps processed by multiple convolutional layers. The advantage of this is that it can avoid the problem of gradient disappearance or gradient explosion caused by the increase in the number of layers in a deep neural network, and at the same time helps to retain some underlying information of the original image, enabling the network to better learn features at different levels. The encoding layer is formed by sequentially connecting multiple such residual blocks. As the residual blocks are continuously stacked, the features of the image are gradually extracted and abstracted. After passing through each residual block, the number of channels of the feature maps may increase, while the size may decrease (by setting an appropriate stride in the convolutional layer or adding a downsampling layer, such as a max pooling layer, between the residual blocks).
[0061] The inverse residual block matches the residual block and mainly decodes the fused information through transposed convolution and skip connection splicing. Transposed convolution (also known as deconvolution) is one of the core operations in the decoding layer. Its function is the opposite of convolution and is used to upsample the feature map to restore the image size. For example, using a 2x2 transposed convolution kernel to operate on the input feature map with a specific stride and padding to double its size. In the inverse residual block, in addition to transposed convolution, skip connection splicing is also used. Skip connection splicing is to connect the feature map at the corresponding position in the encoding layer with the feature map after operations such as transposed convolution in the current inverse residual block (usually splicing in the channel dimension). This can introduce the underlying detailed information retained in the encoding layer into the decoding process, helping to improve the decoding accuracy and restore the details of the image. Similar to the residual block, batch normalization and activation functions (such as ReLU) are also used in the inverse residual block to process the data to ensure the stability and performance of the network. Through this step-by-step upsampling and decoding process, the fused information is finally converted into a meaningful image representation, such as an output mask image, for determining the desired region.
[0062] In this embodiment, multiple sequentially connected encoding layers can gradually extract the high-level semantic features of the image through convolution and downsampling operations, effectively capture the key information in the image, reduce the data volume, and improve the processing efficiency of the model and the ability to understand the image. The feature fusion layer fuses the text information encoded by the speech processing layer and the encoded image information, making full use of the information of both speech and image modalities, enabling the model to process by comprehensively considering the speech description and image content, which is more accurate and robust than using only single-modal information and can better understand the specific regions related to speech in the image. The sequentially connected decoding layers decode the fused information through operations such as deconvolution, can restore the abstract features to a specific image representation, accurately output the mask image, and help to accurately determine the position of the desired region in the search region, realizing the precise positioning and segmentation of the target region. This hierarchical architecture has good flexibility and scalability. It can flexibly adjust the number and structure of the encoding layer and decoding layer, as well as the way of feature fusion according to the requirements of specific tasks and data characteristics to adapt to image processing tasks of different complexities. At the same time, it is also convenient to add new modules or improve existing modules to further improve the model performance.
[0063] In one embodiment, as Figure 4 shown, the residual blocks in the multiple encoding layers encode the sequentially received data and transmit it to the next residual block and the feature fusion layer;
[0064] The feature fusion layer fuses the encoded data and then transmits it to the inverse residual block that matches the residual block for encoding;
[0065] In the multiple decoding layers, the inverse residual blocks sequentially decode the received data and transmit it to the previous inverse residual block, and the inverse residual module of the first decoding layer in the multiple decoding layers outputs a mask image.
[0066] In one embodiment, as Figure 5 shown, the speech processing layer includes a text encoding layer, and the encoding layer and the text encoding layer are CLIP encodings; the image processing model further includes: a category recognition layer connected to the encoding layer, and the category recognition layer is used to recognize the category of objects in the image; wherein, CLIP encoding: that is, Contrastive Language-Image Pretraining encoding, is a technology for pre-training on a large-scale text-image pair, which can learn the semantic association between text and image, and encodes text and image into feature vectors with semantic information respectively. The CLIP (Contrastive Language-Image Pretraining) model contains a powerful text encoder, generally based on the Transformer architecture. The pre-processed text digital representation is input into this text encoder, and the self-attention mechanism in the Transformer weights each word or sub-word unit to capture the semantic relationship between them. Through the processing of multiple Transformer modules, the text information is encoded into a feature vector with a fixed dimension, and this vector contains the semantic information of the text. The category recognition layer receives the image feature vector output by the encoding layer. This feature vector already contains the semantic information of the objects in the image and is the basis for object category recognition. The category recognition layer usually consists of one or more fully connected layers, which further transform and map the input feature vector. The number of output nodes of the last fully connected layer is equal to the number of object categories to be recognized.
[0067] As Figure 6 shown, the method further includes:
[0068] S202, obtaining training parameter data.
[0069] S204, loading the pre-training parameters of the CLIP encoding into the image processing model to be trained, freezing the weights of the CLIP encoding, and controlling the CLIP not to participate in the training.
[0070] Among them, the training parameter data can be a data set for training the image processing model to be trained. It contains various information related to the training task, such as the input panoramic image, the corresponding search area annotation, the expected area annotation, the relevant speech information (converted into text description), and the correct results corresponding to these text descriptions (used to calculate the loss value), etc. The pre-trained parameters usually refer to the model parameters obtained after the CLIP encoding is pre-trained on a large-scale data set, including the weights and biases of each layer in the network, etc. These parameters already contain the semantic understanding ability of the CLIP encoding for text and images. The image processing model to be trained is usually an unfinished model, which consists of a CLIP encoding (used to encode speech information and image information), a feature fusion layer (fusing the encoded text information and image information), a decoding layer (decoding the fused information to obtain results such as a mask image), and a class recognition layer (recognizing the classes of objects in the image), etc. Freezing the weights means that during the training process, the weights of some layers in the model (here it is the CLIP encoding) are fixed and not updated. The purpose of doing this is to retain the knowledge learned by the pre-trained model, thereby improving the training speed of the model.
[0071] Specifically, usually, the CLIP encoding does not require additional training and can directly complete tasks such as image classification and retrieval according to the text description. Therefore, the parameters obtained by pre-training the CLIP encoding can be directly used for training processing, and the weights of the CLIP encoding are frozen, and the CLIP does not participate in the training process of the model, thereby improving the training speed of the model. In addition, it should be noted that because this solution wants to directly use the pre-trained parameters in the pre-trained model, however, because the application scenario of this solution is different from the data set during the CLIP training, the image features and text feature spaces obtained directly by the CLIP encoding are theoretically different. In order to match the feature spaces of the image and the text and reduce this difference, a class recognition layer also needs to be added to the image processing module, and the results output by the class recognition layer are used for training.
[0072] S206, use the training parameter data to train the feature fusion layer, decoding layer and class recognition layer in the image processing model to be trained.
[0073] S208, calculate the loss value based on the results input by the class recognition layer and the text description results in the training parameter data.
[0074] Among them, the loss value is usually a numerical value that measures the difference between the model prediction result and the true result.
[0075] Specifically, the image data in the training parameter data is input into the CLIP encoding part of the model for image encoding, and the text description data is input into the CLIP encoding part for text encoding. Then, the encoded image information and text information are passed to the feature fusion layer for fusion. The fused information is decoded through the decoding layer in sequence to output results such as a mask image, and then the object category is recognized through the category recognition layer to output the predicted object category. Based on the prediction result output by the category recognition layer and the text description result (real object category-related description) in the training parameter data, a suitable loss function (such as the cross-entropy loss function) is used to calculate the loss value.
[0076] S210, iteratively adjust the parameters in the feature fusion layer, decoding layer, and category recognition layer of the image processing model to be trained based on the loss value to obtain the image processing model.
[0077] Specifically, according to the calculated loss value, the backpropagation algorithm is used to calculate the gradients of the loss value with respect to the parameters in the feature fusion layer, decoding layer, and category recognition layer. An optimizer (such as stochastic gradient descent, Adam, etc.) is used to update the parameters in the feature fusion layer, decoding layer, and category recognition layer according to the calculated gradients to reduce the loss value. The above processes of data input, forward propagation, loss calculation, backpropagation, and parameter update are repeated for multiple iterative trainings until the loss value converges to a smaller value or reaches a predetermined number of training times, and finally the trained image processing model is obtained.
[0078] In this embodiment, since CLIP encoding has been pre-trained and has learned rich text and image semantic features. Its pre-trained parameters are directly loaded into the image processing model to be trained, avoiding the large amount of computing resources and time costs required for training the CLIP encoding part from scratch. The model can directly utilize the existing feature extraction ability of CLIP encoding and quickly enter the training of the subsequent layers (feature fusion layer, decoding layer, and category recognition layer), significantly shortening the overall training cycle. In addition, focusing the training on the feature fusion layer, decoding layer, and category recognition layer enables the model to specifically learn how to better fuse image and text information, perform accurate decoding, and category recognition according to the specific training tasks and data. This targeted training can make the model perform better on specific tasks.
[0079] In one embodiment, the CBN (Conditional Batch Normalization) technology is used in the feature fusion layer to link the text information encoded by the speech processing layer to the encoding layer through an MLP (Multilayer Perceptron).
[0080] Among them, CBN (Conditional Batch Normalization) is an improvement over traditional Batch Normalization. Traditional batch normalization normalizes the input data on each channel to make the data have zero mean and unit variance, which helps to speed up the model training and improve stability. CBN introduces conditional information and can dynamically adjust the normalization parameters according to additional conditions (such as the speech encoding information here), making the normalization process more flexible and adaptable to specific tasks. MLP (Multilayer Perceptron) is a basic artificial neural network structure composed of multiple fully connected layers, including an input layer, several hidden layers, and an output layer. Each neuron is connected to all neurons in the previous layer and performs calculations through weighted summation and non-linear activation functions. In this scenario, the MLP is used to further process the text information encoded by the speech processing layer and convert it into parameters that can be used by CBN.
[0081] Specifically, the speech processing layer receives speech input, uses speech recognition technology to convert it into text content, encodes the obtained text, and represents it as a feature vector with a fixed dimension. This vector contains the semantic information described by the speech. The encoding layer processes the input image, extracts the features of the image through a series of operations such as convolution and pooling, and outputs a multi-channel feature map. Each channel of this feature map represents different feature information of the image. In the feature fusion layer, a multilayer perceptron (MLP) is constructed. Its input dimension is the dimension of the speech encoding vector, and the output dimension is twice the number of channels of the feature map of the encoding layer. This is because CBN needs to generate two parameters for each channel: a scaling factor and an offset factor. The speech encoding vector is input into the MLP. After being processed by multiple layers of linear transformation and non-linear activation functions, a vector with a length twice the number of channels of the feature map is output. This output vector is split into two vectors with lengths equal to the number of channels of the feature map, which are used as the scaling factor and offset factor required by CBN respectively. For the feature map output by the encoding layer, the mean and variance are calculated for each channel. The feature map is normalized according to the calculated mean and variance to make the data of each channel have zero mean and unit variance. The normalized feature map is scaled and offset using the scaling factor and offset factor output by the MLP. Specifically, for each normalized element, multiply it by the corresponding scaling factor and add the offset factor. The feature map processed by CBN is the result of fusing speech information and image information, which is used as the output of the feature fusion layer. This output can be passed to subsequent network layers for further processing, such as the decoding layer or the classification layer.
[0082] In this embodiment, batch normalization (including CBN) can reduce internal covariate shift, that is, the phenomenon that the distribution of the input data of each layer of the network changes during the training process. This helps to accelerate the training process of the model and enables the model to converge to the optimal solution faster. At the same time, since CBN combines speech information, it can more effectively guide the learning of features and further improve the training efficiency.
[0083] In one embodiment, the method further includes:
[0084] During the inference process of the image processing model, turn off the class recognition layer in the image processing model.
[0085] Specifically, during the operation of the image processing model, since the function of the class recognition layer in the image processing model is to recognize classes and class recognition is not required during the inference process, the class recognition layer can be turned off during the inference process.
[0086] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0087] Based on the same inventive concept, the embodiments of the present disclosure also provide a multi-instruction image recognition device for implementing the multi-instruction image recognition method involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the multi-instruction image recognition device provided below can refer to the limitations on the multi-instruction image recognition method in the above text, and will not be repeated here.
[0088] In one embodiment, as Figure 7 shown, a multi-instruction image recognition device 300 is provided, including: a search area determination module 302, a model processing module 304, and an expected area determination module 306, where:
[0089] The search area determination module 302 is configured to obtain a panoramic image and determine a search area in the panoramic image, where the search area includes an expected area;
[0090] The model processing module 304 is configured to, in response to detecting voice information associated with the expected region, use the voice information as a guiding condition and process the search region by using a pre-trained image processing model to output a mask image. The image processing model includes a voice processing layer and an image processing layer. The voice processing layer is connected to the image processing layer and is configured to recognize the voice information as text information, encode the text information, and transmit the encoded text information to the image processing layer. The image processing layer is configured to encode the search region, receive the encoded text information, fuse the encoded text information with the encoded search region, and decode the fused result to obtain a mask image.
[0091] The expected region determination module 306 is configured to determine an expected region based on the mask image and the search region.
[0092] In an embodiment of the apparatus, the image processing layer includes a plurality of sequentially connected encoding layers, a feature fusion layer matching the plurality of encoding layers, and a plurality of sequentially connected decoding layers matching the plurality of encoding layers.
[0093] The encoding layer is configured to encode an image. The feature fusion layer matching the plurality of encoding layers is configured to receive the text information encoded by the voice processing layer and fuse the encoded image information with the encoded text information. The decoding layer is configured to receive the fused information and decode the fused information for output.
[0094] In an embodiment of the apparatus, the encoding layer is composed of residual blocks, and the residual blocks encode an image by using a convolution operation.
[0095] The decoding layer is composed of inverse residual blocks matching the residual blocks, and the inverse residual blocks decode the fused information through transposed convolution and skip connection splicing.
[0096] In an embodiment of the apparatus, the residual blocks in the plurality of encoding layers encode the sequentially received data and transmit the encoded data to the next residual block and the feature fusion layer.
[0097] The feature fusion layer fuses the encoded data and transmits the fused data to the inverse residual block matching the residual block that performs encoding.
[0098] The inverse residual blocks in the plurality of decoding layers sequentially decode the received data and transmit the decoded data to the previous inverse residual block, and the inverse residual module of the first decoding layer in the plurality of decoding layers outputs a mask image.
[0099] In one embodiment of the device, the speech processing layer includes a text encoding layer, and the encoding layer and the text encoding layer are CLIP encodings; the image processing model further includes: a category recognition layer connected to the encoding layer, and the category recognition layer is used to recognize the category of objects in the image; the device further includes:
[0100] A model training module, configured to obtain training parameter data; load the pre-training parameters of the CLIP encoding into the image processing model to be trained, freeze the weights of the CLIP encoding, and control the CLIP not to participate in the training; use the training parameter data to train the feature fusion layer, decoding layer and category recognition layer in the image processing model to be trained; calculate a loss value based on the result input by the category recognition layer and the text description result in the training parameter data; iteratively adjust the parameters in the feature fusion layer, decoding layer and category recognition layer of the image processing model to be trained based on the loss value to obtain an image processing model.
[0101] In one embodiment of the device, the CBN (Conditional Batch Normalization) technology is used in the feature fusion layer to link the text information encoded by the speech processing layer to the encoding layer through an MLP (Multilayer Perceptron).
[0102] In one embodiment of the device, the device further includes:
[0103] An identification function closing module, configured to close the category recognition layer in the image processing model during the inference process of the image processing model.
[0104] Each module in the above multi-instruction image recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above respective modules.
[0105] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 8As shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a multi-instruction image recognition method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0106] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the computer device to which the solution of the present disclosure is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0107] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0108] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in any of the above method embodiments.
[0109] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in any of the above method embodiments.
[0110] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present disclosure can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present disclosure can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present disclosure can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0111] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0112] The above-described embodiments merely represent several implementation manners of the present disclosure. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present disclosure, several modifications and improvements can still be made, and these all belong to the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the appended claims.
Claims
1. A multi-instruction image recognition method, characterized in that, Applied to a photographing device, the method includes: Obtain a panoramic image, and determine a search area in the panoramic image, where the search area includes a desired area; In response to detecting voice information associated with the desired area, use the voice information as a guiding condition, and use a pre-trained image processing model to process the search area and output a mask image; wherein, the image processing model includes: a voice processing layer and an image processing layer, the voice processing layer is connected to the image processing layer, and is used to recognize the voice information as text information, encode the text information, and then transmit it to the image processing layer; the image processing layer is used to encode the search area, receive the encoded text information, fuse the encoded text information with the encoded search area, and decode the fused result to obtain a mask image; Based on the mask image and the search area, determine the desired area.
2. The method according to claim 1, wherein The image processing layer includes: a plurality of sequentially connected encoding layers, a feature fusion layer matching the plurality of encoding layers, and a plurality of sequentially connected decoding layers matching the plurality of encoding layers; The encoding layer is used to encode an image, and the feature fusion layer matching the plurality of encoding layers is used to receive the text information encoded by the voice processing layer and fuse the encoded image information and the encoded text information; the decoding layer is used to receive the fused information and decode the fused information and output it.
3. The method according to claim 2, wherein The encoding layer is composed of residual blocks, and the residual blocks use convolution operations to encode images; The decoding layer is composed of inverse residual blocks matching the residual blocks, and the inverse residual blocks decode the fused information through transposed convolution and skip connection splicing.
4. The method according to claim 3, wherein The residual blocks in the plurality of encoding layers encode the sequentially received data and then transmit it to the next residual block and the feature fusion layer; The feature fusion layer fuses the encoded data and then transmits it to the inverse residual block matching the residual block that performs encoding; The inverse residual blocks in the plurality of decoding layers sequentially decode the received data and then transmit it to the previous inverse residual block, and the inverse residual module of the first decoding layer in the plurality of decoding layers outputs a mask image.
5. The method according to claim 2, wherein The voice processing layer includes a text encoding layer, and the encoding layer and the text encoding layer are CLIP encodings; The image processing model further includes: a category recognition layer connected to the encoding layer, and the category recognition layer is used to recognize the category of objects in the image; the method further includes: Obtain training parameter data; Load the pre-trained parameters of the CLIP encoding into the image processing model to be trained, freeze the weights of the CLIP encoding, and control the CLIP not to participate in training; Use the training parameter data to train the feature fusion layer, decoding layer and category recognition layer in the image processing model to be trained; Calculate a loss value based on the result input by the category recognition layer and the text description result in the training parameter data; Based on the loss value, iteratively adjust the parameters in the feature fusion layer, decoding layer and category recognition layer of the image processing model to be trained to obtain an image processing model.
6. The method according to any one of claims 2 to 4, characterized in that, In the feature fusion layer, the CBN (Conditional Batch Normalization) technology is used to link the text information encoded by the speech processing layer to the encoding layer through the MLP (Multilayer Perceptron).
7. The method according to claim 4, characterized in that The method further includes: During the inference process of the image processing model, the class recognition layer in the image processing model is turned off.
8. A multi-instruction image recognition device, characterized in that, Applied to a photographing device, the apparatus includes: A search area determination module, configured to obtain a panoramic image and determine a search area in the panoramic image, where the search area includes a desired area; A model processing module, configured to, in response to detecting speech information associated with the desired area, use the speech information as a guiding condition and process the search area by using a pre-trained image processing model to output a mask image; wherein, the image processing model includes: a speech processing layer and an image processing layer, the speech processing layer is connected to the image processing layer and is configured to recognize the speech information as text information, encode the text information, and transmit the encoded text information to the image processing layer; the image processing layer is configured to encode the search area, receive the encoded text information, fuse the encoded text information with the encoded search area, and decode the fused result to obtain a mask image; A desired area determination module, configured to determine a desired area based on the mask image and the search area.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.