Image recognition method, device and equipment

By combining pre-trained intention recognition model and multiple image recognition models, appropriate image recognition model and usage strategies are selected according to image processing requirements, the problem of image processing singularity in the prior art is solved, and the diverse processing requirements of complex images are realized.

CN120220177APending Publication Date: 2025-06-27SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311837090.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, neural network structures can only perform the same processing on images, and cannot meet the diverse and personalized needs of users, and it is difficult to effectively process the diversified processing needs of complex images.

Method used

By using the pre-trained intention recognition model to identify the processing requirements of the image to be recognized, the target intention of image recognition is determined, the target image recognition model corresponding to the target intention and its usage strategy are determined from the pre-trained multiple image recognition models, and the recognition image is recognized according to the usage strategy.

Benefits of technology

It realizes the selection of corresponding image recognition models according to different image processing needs, meets different processing needs, and solves the diverse processing needs of complex images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220177A_ABST
    Figure CN120220177A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition method, device and equipment. The method provided by the invention comprises the following steps: for a to-be-recognized image, performing intention recognition on a processing demand configured for the to-be-recognized image in advance by using a pre-trained intention recognition model, and determining a target intention of image recognition; determining a target image recognition model corresponding to the target intention and a use strategy of the target image recognition model from a plurality of pre-trained image recognition models; and identifying the to-be-identified image by using the target image identification model according to the use strategy to obtain an identification result corresponding to the target intention. The image recognition method, device and equipment provided by the invention are used for solving diversified processing requirements of complex images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an image recognition method, apparatus, and device. Background Art

[0002] With the development of computer vision technology and image processing technology, the demand for processing complex images is increasing day by day. Traditional image processing methods are often limited by the rules planned by staff. With the rise of deep learning, the cutting-edge technology of image processing is to learn a large amount of data through neural networks, and then automatically extract the features of complex images to complete corresponding operations.

[0003] In the prior art, only when staff need to process specific tasks, a corresponding neural network structure is specially configured. Therefore, a neural network structure can only perform the same processing on images and cannot meet the diverse and personalized needs of users. Summary of the Invention

[0004] In view of this, this application provides an image recognition method, apparatus, and device to solve the diverse processing requirements of complex images.

[0005] Specifically, this application is implemented through the following technical solutions:

[0006] In the first aspect of this application, an image recognition method is provided. The method includes:

[0007] For the image to be recognized, use the pre-trained intent recognition model to perform intent recognition on the processing requirements pre-configured for the image to be recognized, and determine the target intent of image recognition;

[0008] From multiple pre-trained image recognition models, determine the target image recognition model corresponding to the target intent and the usage strategy of the target image recognition model;

[0009] According to the usage strategy, use the target image recognition model to recognize the image to be recognized, and obtain the recognition result corresponding to the target intent.

[0010] In the second aspect of this application, an image recognition apparatus is provided. The apparatus includes a recognition module, a determination module, and a processing module; wherein,

[0011] The recognition module is configured to, for the image to be recognized, use the pre-trained intent recognition model to perform intent recognition on the processing requirements pre-configured for the image to be recognized, and determine the target intent of image recognition;

[0012] The determining module is configured to determine, from a plurality of pre-trained image recognition models, a target image recognition model corresponding to the target intent and a usage strategy for the target image recognition model;

[0013] The processing module is configured to recognize the image to be recognized by using the target image recognition model according to the usage strategy, so as to obtain a recognition result corresponding to the target intent.

[0014] A third aspect of the present application provides an image recognition device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any of the methods provided in the first aspect of the present application are implemented.

[0015] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of any of the methods provided in the first aspect of the present application are implemented.

[0016] For the image to be recognized, the image recognition method, device, and equipment provided in the present application first use a pre-trained intent recognition model to perform intent recognition on the processing requirements pre-configured for the image to be recognized, determine the target intent of image recognition, and then determine, from a plurality of pre-trained image recognition models, a target image recognition model corresponding to the target intent and a usage strategy for the target image recognition model. Finally, according to the usage strategy, the target image recognition model is used to recognize the image to be recognized, so as to obtain a recognition result corresponding to the target intent. In this way, the target intent of the processing requirements is determined through the intent recognition model. Subsequently, for different target intents, different image recognition models are selected to process the image to be recognized, which can meet different processing requirements and solve the diverse processing requirements of complex images. Description of the Drawings

[0017] Figure 1 It is a flowchart of the first embodiment of the image recognition method provided by the present application;

[0018] Figure 2 It is a flowchart of the second embodiment of the image recognition method provided by the present application;

[0019] Figure 3 It is a flowchart of the third embodiment of the image recognition method provided by the present application;

[0020] Figure 4 It is a schematic diagram of the implementation principle of the image recognition process shown in an exemplary embodiment of the present application;

[0021] Figure 5 It is a hardware structure diagram of the image recognition device where the image recognition device provided by the present application is located;

[0022] Figure 6 This is a schematic structural diagram of the first embodiment of the image recognition device provided by this application. Specific embodiments

[0023] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0024] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0026] This application provides an image recognition method, device, and equipment to address the diverse processing requirements of complex images.

[0027] For the image recognition method, device, and equipment provided by this application, for the image to be recognized, first, an intention recognition model trained in advance is used to perform intention recognition on the processing requirements configured in advance for the image to be recognized to determine the target intention of image recognition. Then, from multiple image recognition models trained in advance, the target image recognition model corresponding to the target intention and the usage strategy of the target image recognition model are determined. Finally, according to the usage strategy, the target image recognition model is used to recognize the image to be recognized, and the recognition result corresponding to the target intention is obtained. In this way, by using the intention recognition model to determine the target intention of the processing requirements, subsequently, for different target intentions, different image recognition models are selected to process the image to be recognized, which can meet different processing requirements and address the diverse processing requirements of complex images.

[0028] Specific embodiments are given below to introduce the technical solutions of the present application in detail.

[0029] Figure 1 This is a flowchart of the first embodiment of the image recognition method provided by the present application. Please refer to Figure 1 , the method provided in this embodiment may include:

[0030] S101. For the image to be recognized, use a pre-trained intent recognition model to perform intent recognition on the processing requirements pre-configured for the image to be recognized, and determine the target intent of image recognition.

[0031] Specifically, the image to be recognized refers to the image that needs to be recognized currently. It should be noted that the image to be recognized may include text information, pattern and texture information, color information, background information, etc. For example, in one embodiment, when paying water and electricity bills, the image to be recognized is a photo of a water and electricity meter, and the numbers and letters in the water and electricity meter can be recognized based on the image to be recognized.

[0032] Furthermore, the processing requirements are pre-configured by the user for the image to be recognized, and are used to characterize the user's processing requirements, that is, what kind of processing will be performed on the image to be recognized, or what information will be obtained from the image to be recognized, etc. Among them, the processing requirements include at least one modality of data, and at least one modality of data may include text data, voice data, etc.

[0033] In specific implementation, in one embodiment, the processing requirements can be characterized by recognition tasks. For example, in one embodiment, when the image to be recognized is a photo of a water and electricity meter, the processing requirements to be processed can be: recognize the water and electricity bills generated in the image to be recognized. For another example, in another embodiment, when the image to be recognized is a traffic photo, the processing requirements to be processed can be: recognize non-motor vehicle drivers without wearing safety helmets in the image to be recognized. For another example, in yet another embodiment, the processing requirements to be processed can also be: perform human target detection on the image to be recognized.

[0034] In addition, in another embodiment, the processing requirements can also be characterized by the objects of interest in the image to be processed and the operation requirements of the objects of interest. For example, in one embodiment, when the image to be recognized is a photo of a water and electricity meter, the objects of interest are all the numbers and letters in the image to be recognized, and the operation requirements of the objects of interest are to recognize the numbers and letters on the water and electricity meter. For another example, in another embodiment, when the image to be recognized is a traffic photo, the objects of interest are all the drivers in the image to be recognized, and the operation requirements of the objects of interest are to recognize non-motor vehicle drivers without wearing safety helmets.

[0035] Furthermore, the intention recognition model is a pre-trained model. In specific implementation, the processing requirements pre-configured for the image to be recognized are input into the pre-trained intention recognition model, so that the target intention can be output by the intention recognition model. For example, in a possible implementation, regardless of the processing requirements of any modality, the pre-trained intention recognition model can perform intention recognition on it and output the target intention.

[0036] It should be noted that when the processing requirement is in text form, the intention recognition model can be a pre-trained large model, which can classify the processing requirement in text form into different intentions to obtain the target intention. When the processing requirement is in voice form, the intention recognition model can include two parts. The first part is used to convert the processing requirement in voice form into a processing requirement in text form, and the second part is also a pre-trained large model, which is used to classify the processing requirement in text form into different intentions.

[0037] The pre-trained large model can be a natural language processing model. For example, it can be BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), etc. It should be noted that the pre-trained large model can better capture the complex relationships and contexts in language, thereby improving the accuracy of the model.

[0038] It can be understood that the large model refers to a neural network model with a large number of parameters in deep learning. Such models usually have more parameters and layers and can better capture the complex patterns and relationships in the data.

[0039] Optionally, in a possible implementation of this application, the intention recognition model is a pre-trained model based on the Transformer model;

[0040] Among them, the pre-trained model refers to a model pre-trained based on a large number of sample images, mainly composed of an encoder and a decoder. The encoder is used to encode the input processing requirements into a fixed-length vector representation, and the decoder is used to decode this vector representation into the target intention.

[0041] From the perspective of the data processing process, after the encoder obtains the input processing requirements, it encodes the input processing requirements into a vector representation of a fixed length, and at the same time adds a position encoding consistent with the dimension of the word vector to each word vector. The vector representation of the fixed length is input into the encoder to obtain the encoding information. The encoder of the Transformer model includes a multi-head attention layer, a summation and normalization layer, and a fully connected layer. Among them, the attention mechanism directly calculates the attention weights at each position in the encoding process of the sentence through a certain operation; then the hidden vector representation of the entire sentence is calculated in the form of the sum of weights. The multi-head attention mechanism uses independently learned n groups of different linear projections to calculate the query Q, key K, and value V. The outputs of n parallel attention pooling are concatenated together and transformed through another learnable linear projection to produce the final output. The attention value in the multi-head attention layer is:

[0042]

[0043] where dk is the scaling factor and T is the transpose symbol.

[0044] To prevent degradation during the training of the deep neural network and accelerate the training speed and stability, the output matrix of the multi-head attention layer is input into the summation and normalization layer. After using residual connection and normalization adjustment, the matrix is input into the decoder. The Transformer decoder consists of two multi-head self-attention layers based on the multi-head self-attention mechanism connected to a feed-forward neural network. Among them, the first multi-head self-attention layer in the two multi-head self-attention layers uses a masking operation. The K and V matrices of the second multi-head self-attention layer are calculated using the encoding information matrix output by the encoder, and the Q matrix of the second multi-head self-attention layer uses the output of the previous decoder for calculation.

[0045] The intent recognition model provided by the present invention is a pre-trained large model based on the Transformer model structure and pre-trained using a large number of processing requirements. The intent recognition model at least includes an encoder and a decoder; the encoding of the processing requirements is obtained based on the pre-trained intent recognition model; and the intent is recognized based on the encoding.

[0046] S102. Determine the target image recognition model corresponding to the target intent and the usage strategy of the target image recognition model from multiple pre-trained image recognition models.

[0047] Specifically, the image recognition model is pre-trained, and different types of image recognition models can be pre-trained according to actual needs. For example, in one embodiment, the pre-trained image recognition model may include a model for object detection, a model for image classification, a model for image segmentation, a model for face recognition, a model for scene understanding, a model for text recognition, a model for anomaly detection, etc.

[0048] Furthermore, the usage strategy is used to represent the strategy for recognizing the image to be recognized based on the target image recognition model.

[0049] Optionally, the target image recognition model includes at least one model, and the usage strategy of the target image recognition model is used to indicate the input and output of each model in the at least one model, and the determination rule for determining the recognition result corresponding to the target intention by using the output of the at least one model.

[0050] It should be noted that the specific determination rule in the usage strategy is determined according to actual needs, and this is not limited in this embodiment.

[0051] In specific implementation, when the target image recognition model includes one model, the usage strategy is used to indicate the input and output of the model and the determination rule for determining the recognition result by using the output of the one model. At this time, the determination rule may be to determine the output of the model as the recognition result. For example, in one embodiment, the processing requirement is to perform face detection on the image to be recognized. At this time, the target image recognition model includes Model 1 (specifically, Model 1 may be a face detection model), and the usage strategy includes: the input of Model 1 is the image, the output is the face detection result, and the determination rule for the recognition result is to use the output of Model 1 as the recognition result. In other words, this usage strategy is used to indicate that the image to be recognized is input into Model 1, and the output result of Model 1 is directly used as the recognition result corresponding to the target intention.

[0052] Furthermore, when the target image recognition model includes multiple models, the usage strategy is also used to indicate the input and output of each model in the at least one model, and the determination rule for determining the recognition result by using the output of the at least one model.

[0053] At this time, since the target image recognition model includes multiple models, the determination rule for determining the recognition result corresponding to the target intention based on the outputs of multiple models is set according to actual needs, and this is not limited in this embodiment. For example, in one embodiment, Usage Strategy A uses the output result of the last model in the processing process as the recognition result corresponding to the target intention. For another example, in another embodiment, Usage Strategy B uses the output results of all target image recognition models together as the recognition result corresponding to the target intention.

[0054] The image recognition method provided in this embodiment can accurately perform image recognition processing by using a policy to determine the input and output of the target image recognition model and the determination rule for the recognition result corresponding to the target intent.

[0055] Combined with the above introduction, for example, in one embodiment, for Image 1, the processing requirement configured by the user is to recognize the text in Image 1. At this time, after performing intent recognition on Image 1, it is determined that its target intent is to recognize the text in Image 1. Correspondingly, in this step, the target recognition model is determined to be an OCR recognition model, and the usage policy of the target recognition model includes: OCR recognition model (input is the image, output is the recognized text), and the determination rule for the recognition result is: the output of the OCR recognition model is determined as the recognition result corresponding to the target intent.

[0056] Another example, in another embodiment, for Image 2, the processing requirement configured by the user is to recognize the human pose in Image 2. At this time, the target recognition model is determined to include a human target detection model and a human pose estimation model. Correspondingly, the usage policy of the target recognition model is: human detection model (input is the image, output is the detected human target), human pose estimation model (input is the human target, output is the human pose of the human target). Correspondingly, the determination rule is: the output of the human pose estimation model is determined as the recognition result.

[0057] S103. According to the usage policy, use the target image recognition model to recognize the image to be recognized, and obtain the recognition result corresponding to the target intent.

[0058] Specifically, the recognition result is obtained after using the target image recognition model to recognize the image to be recognized according to the usage policy. For example, in one embodiment, combined with the above example, when the image to be recognized is a photo of a water and electricity meter, the recognition result obtained through the target image recognition model can be: 25 cubic meters of water used and 260 degrees of electricity used.

[0059] Furthermore, another example, combined with the example in step S102, for example, for Image 1, according to the usage policy, in this step, Image 1 is input into the OCR recognition model, so that the OCR recognition model outputs the recognized text. Furthermore, finally, the output of the OCR recognition model is determined as the recognition result of the target intent.

[0060] Similarly, for Image 2, according to the usage strategy, first input Image 2 into the human target detection model to output the detected human target by the human target detection model. Further, input the human target into the human pose estimation model to output the human pose by the human pose estimation model. Finally, based on the determination rule in the usage strategy, determine the human pose output by the human pose estimation model as the recognition result of the target intention.

[0061] For the image recognition method provided in this embodiment, for the image to be recognized, first use the pre-trained intention recognition model to perform intention recognition on the processing requirements pre-configured for the image to be recognized, determine the target intention of the image recognition, and then determine the target image recognition model corresponding to the target intention and the usage strategy of the target image recognition model from multiple pre-trained image recognition models. Finally, according to the usage strategy, use the target image recognition model to recognize the image to be recognized to obtain the recognition result corresponding to the target intention. In this way, by using the intention recognition model to determine the target intention of the processing requirements, subsequently, for different target intentions, different image recognition models are selected to process the image to be recognized, which can meet different processing requirements and solve the diverse processing requirements of complex images.

[0062] Figure 2 This is the flowchart of the second embodiment of the image recognition method provided in this application. Please refer to Figure 2 , on the basis of the above embodiment, the processing requirements include at least one type of modal data, and the pre-trained intention recognition model includes sub-intention recognition models corresponding to various modalities; the step of using the pre-trained intention recognition model to determine the target intention of the image to be recognized includes:

[0063] S201. For each type of modal data in the processing requirements, use the sub-intention recognition model corresponding to this modality in the pre-trained intention recognition model to determine the sub-intention of the data of this modality.

[0064] Specifically, the processing requirements include at least one type of modal data. For example, the processing requirements may include at least one of the following types of data: text data, voice data, and video data.

[0065] Furthermore, the pre-trained intention recognition model includes multiple sub-intention recognition models, and the number of multiple sub-intention recognition models is greater than or equal to the number of modalities in the processing requirements, so that all modal processing requirements can be processed. For example, in one embodiment, the sub-intention recognition models include Model S1, Model S2, and Model S3; among them, Model S1 is used to perform intention recognition on text data, Model S2 is used to perform intention recognition on voice data, and Model S3 is used to perform intention recognition on video data.

[0066] In specific implementation, in this step, the data in the processing requirement is input into the corresponding sub-intent recognition model according to its modality, so that the sub-intent of the data of this modality can be output by the sub-intent recognition model.

[0067] For example, in an embodiment, the processing requirement 1 is to recognize a human face in a picture (in text form). At this time, the processing requirement 1 only includes text data. Combining with the above example, at this time, the text data is input into the model S1, and the model S1 performs intent recognition on the text data to obtain the sub-intent 11 of the text data.

[0068] For another example, in another embodiment, the processing requirement 2 is to recognize a human target in an image (in text form) and recognize the human pose of the human target (in voice form). At this time, the processing requirement 2 includes not only text data but also voice data. Combining with the above example, at this time, the text data in the processing requirement 2 is input into the model S1, and the model S1 performs intent recognition on the text data to obtain the sub-intent 21 of the text data; similarly, the voice data in the processing requirement 2 is input into the model S2, and the model S2 performs intent recognition on the voice data to obtain the sub-intent 22 of the voice data.

[0069] Optionally, in a possible implementation manner, the at least one modality of data includes non-text data; the sub-intent recognition model corresponding to the non-text data includes a cascaded text recognition sub-model and a first intent recognition sub-model; wherein, the text recognition sub-model is configured to perform text recognition on the non-text data in the processing requirement to obtain the text information corresponding to the non-text data; the first intent recognition sub-model is configured to perform intent recognition on the text information to obtain the sub-intent corresponding to the non-text data.

[0070] Specifically, the sub-intent recognition model corresponding to the non-text data may include a sub-intent recognition model corresponding to voice data and a sub-intent recognition model corresponding to video data. In addition, the sub-intent recognition model corresponding to the non-text data includes a cascaded text recognition sub-model and a first intent recognition sub-model.

[0071] Specifically, the text recognition sub-model is used to process the non-text data in the processing requirement. In specific implementation, the non-text data in the processing requirement can be input into a pre-trained text recognition sub-model, so that the text information corresponding to the non-text data is output by the text recognition sub-model. For example, in one embodiment, in combination with the above example, the processing requirement includes voice data and text data. At this time, for the voice data, the sub-intent recognition model corresponding to the voice data is used to recognize it. The sub-intent recognition model corresponding to the voice data includes a text recognition sub-model and a first intent recognition sub-model. The sub-intent recognition model corresponding to the voice data first uses the text recognition sub-model to process the voice data to obtain the text information corresponding to the voice data.

[0072] Specifically, the first intent recognition sub-model is pre-trained according to the intent recognition framework. In this embodiment, no limitation is imposed on this. It should be noted that the first intent recognition sub-model is used to process the text information output by the text recognition sub-model for the text recognition sub-model. For example, in one embodiment, in combination with the above example, the text information corresponding to the voice data is input into the first intent recognition sub-model, and the first intent recognition sub-model outputs the corresponding sub-intent.

[0073] For the image recognition method provided in this embodiment, when at least one type of data in the processing requirement includes non-text data, the text information corresponding to the non-text data is obtained based on the text recognition sub-model, and then the text information is processed based on the first intent recognition sub-model to obtain the sub-intent corresponding to the processing requirement. In this way, for a processing requirement with non-text data, its corresponding sub-intent can be obtained quickly and accurately.

[0074] Further, in another possible implementation manner, the at least one type of data includes text data; the sub-model corresponding to the text data includes a second intent recognition sub-model; wherein, the second intent recognition sub-model is used to perform intent recognition on the text data to obtain the sub-intent corresponding to the text data.

[0075] Specifically, the second intent recognition sub-model is similar to the first intent recognition sub-model, and its detailed introduction can refer to the relevant description in the above steps and will not be elaborated here. It should be noted that the second intent recognition sub-model and the first intent recognition sub-model can be the same or different. In this embodiment, no limitation is imposed on this.

[0076] In specific implementation, for example, in one embodiment, the processing requirement only includes text data, and the second intent recognition sub-model is directly used to process the text processing requirement, and the obtained corresponding sub-intent is: recognizing the reading of the water and electricity meter.

[0077] S202. Sub-intents of various modalities of data in the processing requirement are fused to obtain the target intent of the image to be recognized.

[0078] In specific implementation, sub-intents of various modalities of data are combined to obtain the target intent of the image to be recognized. For example, in combination with the second example in step S201, in this step, the sub-intent 21 of text data and the sub-intent 22 of voice data are combined to obtain the target intent of the image to be recognized as: recognizing the human target in the image and the human posture of the human target.

[0079] For the image recognition method provided in this embodiment, for each modality of data in the processing requirement, the sub-intent recognition model corresponding to this modality in the pre-trained intent recognition model is used to determine the sub-intent of the data of this modality, and then the sub-intents of various modalities of data in the processing requirement are fused to obtain the target intent of the image to be recognized. In this way, when the processing requirement includes at least one different modality of data, multiple sub-intent recognition models can be used to process the data of different modalities respectively, and the corresponding sub-intents can be obtained respectively, and finally the target intent can be obtained to process the image to be recognized. In this way, the target intent can be obtained quickly and accurately, and then different image recognition models can be selected based on the target intent to recognize the image to be recognized to meet the processing requirements of users and improve the user experience.

[0080] Figure 3 This is the flowchart of the third embodiment of the image recognition method provided by this application. Please refer to Figure 3 , based on the above embodiment, the step of determining the target image recognition model corresponding to the target intent and the usage strategy of the target image recognition model from the pre-trained multiple image recognition models includes:

[0081] S301. Search for the target corresponding relationship where the target intent is located from the corresponding relationship established in advance among the intent, the image recognition model, and the usage strategy.

[0082] Specifically, an intent corresponds to an image recognition model and the usage strategy of the image recognition model. The usage strategy can be used to indicate the processing order of the image recognition model (which can be indirectly characterized by the input and output of each model) and the determination rule of the recognition result corresponding to the target intent. For example, in one embodiment, the usage strategy a is used to indicate that the image recognition model 1 processes the image to be recognized first, and its processing result is input into the image recognition model 2, and the output of the image recognition model 2 is used as the recognition result.

[0083] It should be noted that the correspondence relationship among the intention, the image recognition model, and the usage strategy can be set based on the principle and function of the image recognition model. The correspondence relationship among the intention, the image recognition model, and the usage strategy can also be set based on historical data. For example, in one embodiment, historical data can be learned, and the correspondence relationship can be established based on the learning result. The historical data can be the historical processing records of the intention, which record the intention, as well as the model used during processing and the usage strategy of the model. By learning the historical data, the correspondence relationship can be accurately established.

[0084] In specific implementation, in this step, after determining the target intention, the image recognition model and the usage strategy corresponding to the target intention can be found through the pre-established correspondence relationship among the intention, the image recognition model, and the usage strategy.

[0085] For example, in one embodiment, the pre-established correspondence relationship among the intention, the image recognition model, and the usage strategy is shown in Table 1.

[0086] Table 1

[0087]

[0088]

[0089] S302. Determine the image recognition model recorded in the target correspondence relationship as the target image recognition model, and determine the usage strategy recorded in the target correspondence relationship as the usage strategy of the target image recognition model.

[0090] In specific implementation, for example, in one embodiment, the target intention is to identify non-motor vehicle drivers not wearing safety helmets. Figure 4 This is the schematic diagram of the implementation of the image recognition process shown in an exemplary embodiment of the present application. Please refer to Figure 4 , through the pre-established correspondence relationship among the intention, the image recognition model, and the usage strategy, the image recognition models corresponding to the target intention are found to include Model 1, Model 2, Model 3, Model 4, and Model 5. The processing process obtained based on the usage strategy is as Figure 4 shown. Finally, the result obtained after processing by Model 5 is the recognition result corresponding to the target intention.

[0091] The image recognition method provided in this embodiment searches for the target correspondence relationship where the target intent is located from the pre-established correspondence relationship among the intent, the image recognition model, and the usage strategy. Furthermore, the image recognition model recorded in the target correspondence relationship is determined as the target image recognition model, and the usage strategy recorded in the target correspondence relationship is determined as the usage strategy of the target image recognition model. In this way, through the pre-established correspondence relationship, the target image recognition model and the usage strategy are obtained through the target intent to process the image to be recognized. In this way, the processing of the image to be recognized can be completed quickly and accurately.

[0092] Corresponding to the foregoing embodiment of an image recognition method, the present application also provides an embodiment of an image recognition device.

[0093] The embodiment of the image recognition device of the present application can be applied to an image recognition device. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the image recognition device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 5 shown, it is a hardware structure diagram of the image recognition device where the image recognition device provided by the present application is located. In addition to Figure 5 the shown processor, memory, network interface, and non-volatile memory, the image recognition device where the device is located in the embodiment usually includes other hardware according to the actual function of the image recognition device, which will not be elaborated here.

[0094] Figure 6 It is a schematic structural diagram of Embodiment 1 of the image recognition device provided by the present application. Please refer to Figure 6 , the device provided in this embodiment includes an identification module 610, a determination module 620, and a processing module 630; where

[0095] The identification module 610 is configured to perform intent recognition on the processing requirements pre-configured for the image to be recognized by using a pre-trained intent recognition model for the image to be recognized, and determine the target intent of the image recognition;

[0096] The determination module 620 is configured to determine the target image recognition model corresponding to the target intent and the usage strategy of the target image recognition model from a plurality of pre-trained image recognition models;

[0097] The processing module 630 is configured to perform recognition on the image to be recognized by using the target image recognition model according to the usage strategy to obtain the recognition result corresponding to the target intent.

[0098] The device of this embodiment can be used to execute Figure 1 the steps of the method embodiment shown above. The specific implementation principle and process are similar and will not be elaborated here.

[0099] Optionally, the processing requirement includes data of at least one modality, and the pre-trained intent recognition model includes sub-intent recognition models corresponding to various modalities;

[0100] The recognition module 610 is specifically configured to, for the data of each modality in the processing requirement, use the sub-intent recognition model corresponding to this modality in the pre-trained intent recognition model to determine the sub-intent of the data of this modality;

[0101] The recognition module 610 is also specifically configured to fuse the sub-intents of the data of various modalities in the processing requirement to obtain the target intent of the image to be recognized.

[0102] Optionally, the determination module 620 is specifically configured to search for the target correspondence relationship where the target intent is located from the correspondence relationship established in advance among the intent, the image recognition model, and the usage strategy;

[0103] The determination module 620 is also specifically configured to determine the image recognition model recorded in the target correspondence relationship as the target image recognition model, and determine the usage strategy recorded in the target correspondence relationship as the usage strategy of the target image recognition model.

[0104] Optionally, the data of the at least one modality includes non-text data; the sub-intent recognition model corresponding to the non-text data includes a cascaded text recognition sub-model and a first intent recognition sub-model; wherein,

[0105] The text recognition sub-model is configured to perform text recognition on the non-text data in the processing requirement to obtain the text information corresponding to the non-text data;

[0106] The first intent recognition sub-model is configured to perform intent recognition on the text information to obtain the sub-intent corresponding to the non-text data.

[0107] Optionally, the data of the at least one modality includes text data; the sub-model corresponding to the text data includes a second intent recognition sub-model; wherein,

[0108] The second intent recognition sub-model is configured to perform intent recognition on the text data to obtain the sub-intent corresponding to the text data.

[0109] Optionally, the target image recognition model includes at least one model, and the usage strategy of the target image recognition model is used to indicate the input and output of each model in the at least one model, and the determination rule for determining the recognition result by using the output of the at least one model.

[0110] Please continue to refer to Figure 5 , this application also provides an image recognition device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any of the methods provided in the first aspect of this application are implemented.

[0111] This application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of any of the methods provided in this application are implemented.

[0112] The implementation process of the functions and roles of each unit in the above device is specifically described in detail in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0113] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0114] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of protection of this application.

Claims

1. An image recognition method, characterized in that, The method includes: For the image to be recognized, use a pre-trained intent recognition model to perform intent recognition on the processing requirements pre-configured for the image to be recognized, and determine the target intent of image recognition; From multiple pre-trained image recognition models, determine the target image recognition model corresponding to the target intent and the usage strategy of the target image recognition model; According to the usage strategy, use the target image recognition model to recognize the image to be recognized, and obtain the recognition result corresponding to the target intent.

2. The method according to claim 1, wherein The processing requirements include data of at least one modality, and the pre-trained intent recognition model includes sub-intent recognition models corresponding to various modalities; using the pre-trained intent recognition model to determine the target intent of the image to be recognized includes: For each modality of data in the processing requirements, use the sub-intent recognition model corresponding to this modality in the pre-trained intent recognition model to determine the sub-intent of the data of this modality; Fuse the sub-intents of the data of various modalities in the processing requirements to obtain the target intent of the image to be recognized.

3. The method according to claim 1, wherein The determining the target image recognition model corresponding to the target intent and the usage strategy of the target image recognition model from multiple pre-trained image recognition models includes: Search for the target correspondence relationship where the target intent is located from the pre-established correspondence relationship among intent, image recognition model, and usage strategy; Determine the image recognition model recorded in the target correspondence relationship as the target image recognition model, and determine the usage strategy recorded in the target correspondence relationship as the usage strategy of the target image recognition model.

4. The method according to claim 2, wherein The data of the at least one modality includes non-text data; the sub-intent recognition model corresponding to the non-text data includes a cascaded text recognition sub-model and a first intent recognition sub-model; where The text recognition sub-model is used to perform text recognition on the non-text data in the processing requirements to obtain the text information corresponding to the non-text data; The first intent recognition sub-model is used to perform intent recognition on the text information to obtain the sub-intent corresponding to the non-text data.

5. The method according to claim 2 or 4, characterized in that, The data of the at least one modality includes text data; the sub-model corresponding to the text data includes a second intent recognition sub-model; where The second intent recognition sub-model is used to perform intent recognition on the text data to obtain the sub-intent corresponding to the text data.

6. The method according to claim 1, characterized in that The target image recognition model includes at least one model, and the usage strategy of the target image recognition model is used to indicate the input and output of each model in the at least one model, and the determination rule for determining the recognition result using the output of the at least one model.

7. An image recognition device, characterized in that, The device includes an identification module, a determination module, and a processing module; where The identification module is used to, for the image to be recognized, use a pre-trained intent recognition model to perform intent recognition on the processing requirements pre-configured for the image to be recognized, and determine the target intent of image recognition; The determining module is configured to determine, from multiple pre-trained image recognition models, the target image recognition model corresponding to the target intent and the usage strategy of the target image recognition model; The processing module is configured to, according to the usage strategy, use the target image recognition model to recognize the image to be recognized, and obtain the recognition result corresponding to the target intent.

8. The device according to claim 7, characterized in that, The processing requirement includes data of at least one modality, and the pre-trained intent recognition model includes sub-intent recognition models corresponding to various modalities; The recognition module is specifically configured to, for the data of each modality in the processing requirement, use the sub-intent recognition model corresponding to the modality in the pre-trained intent recognition model to determine the sub-intent of the data of this modality; The recognition module is further specifically configured to fuse the sub-intents of the data of various modalities in the processing requirement to obtain the target intent of the image to be recognized.

9. An image recognition device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1-6 when executing the program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program implements the steps of the method according to any one of claims 1-6 when executed by the processor.