Image Detection Method, Apparatus, Device, and Storage Medium
By extracting object description information and image features and combining Transformer model for image detection, the problem of insufficient human-computer interaction experience in the prior art is solved, and more efficient and wider application scenarios are achieved.
Patent Information
- Application Number
- CN202210094666.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The existing image detection methods cannot meet the improvement of human-computer interaction experience, especially in terms of efficiency, accessibility and applicability.
By extracting the features of object description information (voice or text) and the features of the image to be processed, and combining the Transformer model to detect the target object, the positioning of the target object in the image is achieved.
It improves the fun and efficiency of human-computer interaction, expands the convenience and applicability of operation, and enhances the application effect in human-computer interaction scenarios.
Smart Images

Figure CN114429187B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to, but not limited to, an image detection method, apparatus, device, and storage medium. Background Art
[0002] As a long-term and challenging problem in computer vision, image detection has become an active research field for decades. The purpose of image detection is to determine whether a target object exists in a given image, and if so, output the spatial position of the target object. Image detection is widely used in many fields such as artificial intelligence, information technology, etc., such as traffic supervision, autonomous driving, human-computer interaction, drones, content-based image retrieval, intelligent video surveillance, and augmented reality.
[0003] With the development of the field of artificial intelligence, people are increasingly pursuing an improved human-computer interaction experience, and existing image detection methods can no longer meet people's needs. Summary of the Invention
[0004] In view of this, embodiments of the present application provide an image detection method, apparatus, device, and storage medium.
[0005] In a first aspect, embodiments of the present application provide an image detection method, the method including: obtaining a first image to be processed and input object description information, where the object description information is a first voice or a first text for describing a first target object to be recognized; extracting features from the object description information to obtain object description features; extracting features from the first image to be processed to obtain first image features; and based on the object description features and the first image features, detecting the first target object in the first image to be processed to obtain a detection result of the first image to be processed.
[0006] In a second aspect, embodiments of the present application provide an image detection model training method, the method including: obtaining an image sample set and an object description sample set, where the object description sample set is a voice sample set or a text sample set for describing an object to be recognized; using the image sample set and the object description sample set to train an image detection model to obtain detection results of each image sample in the image sample set; determining a loss of the image detection model based on the detection results of each image sample; and using the loss to adjust network parameters of the image detection model so that the loss of the detection result output by the adjusted image detection model satisfies a convergence condition.
[0007] In a third aspect, an embodiment of the present application provides an image detection device, the device includes: a first acquisition module, configured to acquire a first image to be processed and input object description information, where the object description information is a first voice or a first text for describing a first target object to be recognized; a first extraction module, configured to extract features from the object description information to obtain object description features; a second extraction module, configured to extract features from the first image to be processed to obtain first image features; a target object detection module, configured to detect the first target object in the first image to be processed based on the object description features and the first image features, and obtain a detection result of the first image to be processed.
[0008] In a fourth aspect, an embodiment of the present application provides a training device for an image detection model, the device includes: a fourth acquisition module, configured to acquire an image sample set and an object description sample set, where the object description sample set is a voice sample set or a text sample set for describing an object to be recognized; a first training module, configured to use the image sample set and the object description sample set to train an image detection model, and obtain detection results of each image sample in the image sample set; a second determination module, configured to determine a loss of the image detection model based on the detection results of each image sample; an adjustment module, configured to use the loss to adjust network parameters of the image detection model, so that the loss of the detection result output by the adjusted image detection model meets a convergence condition.
[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, the device includes: a memory and a processor, the memory stores a computer program that can run on the processor, and when the processor executes the computer program, the steps in the above method are implemented.
[0010] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method are implemented.
[0011] In the embodiments of the present application, by extracting the features of the object description information (the first voice or the first text) and the features of the first image to be processed, the object description features and the first image features are obtained. Then, based on the object description features and the first image features, the first object to be detected is detected in the first image to be processed, and the detection result of the first image to be processed is obtained. That is, by combining the first voice and the first image to be processed, or the first text and the first image to be processed, the detection of the target object in the image is realized. Therefore, it can be widely applied in the scenario of human-computer interaction, using text or voice to issue commands to detect the target object in the image, increasing the interest of human-computer interaction. At the same time, when the object description information is the first voice, due to the advantages of voice in terms of efficiency, accessibility, and applicability, the operation is more convenient and efficient, the user group is wider, the application scenarios are more, and it is more helpful to improve the human-computer interaction experience in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1A is a schematic structural diagram of an image detection system provided by an embodiment of the present application;
[0013] Figure 1B is a schematic flow diagram of an image detection method provided by an embodiment of the present application;
[0014] Figure 2A is a schematic flow diagram of another image detection method provided by an embodiment of the present application;
[0015] Figure 2B is a schematic diagram of different tracking methods provided by an embodiment of the present application;
[0016] Figure 3A and Figure 3B is a schematic flow diagram of a target tracking method provided by an embodiment of the present application;
[0017] Figure 4 is a bar chart and a pie chart of the number and proportion of voice files corresponding to different voice durations in the two major databases of LaSOT and TNL2K provided by an embodiment of the present application;
[0018] Figure 5 is a success score and precision score chart of different models provided by an embodiment of the present application;
[0019] Figure 6 is a schematic structural diagram of an image detection device provided by an embodiment of the present application;
[0020] Figure 7 is a schematic structural diagram of a training device for an image detection model provided by an embodiment of the present application;
[0021] Figure 8 is a schematic diagram of a hardware entity of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0022] The technical solutions of the present application will be further elaborated in detail below in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0023] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0024] In subsequent descriptions, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of description of the present application, and have no specific meaning in themselves. Therefore, "module", "component", or "unit" can be used interchangeably.
[0025] It should be noted that the terms "first", "second", and "third" involved in the embodiments of the present application are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first", "second", and "third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0026] The embodiments of the present application provide an image detection method, which is applied to an electronic device. The electronic device includes, but is not limited to, a mobile phone, a laptop computer, a tablet computer, a handheld Internet device, a multimedia device, a streaming media device, a mobile Internet device, a drone, a robot, or other types of electronic devices. The functions implemented by this method can be realized by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. Obviously, the electronic device includes at least a processor and a storage medium. The processor can be used for image detection, and the memory can be used for storing the data required during the image detection process and the generated data.
[0027] Figure 1A FIG. is an optional schematic architecture diagram of an image detection system 10 according to an embodiment of the present application. Refer to Figure 1A, in some embodiments, the image acquisition device 100 may send the image set to the server 200, which transmits it to the electronic device 300, and the electronic device 300 performs image detection; in some embodiments, the image acquisition device 100 may directly transmit the image set to the electronic device 300 for image detection; in some embodiments, the electronic device 300 may perform image detection using the image set stored locally.
[0028] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The image acquisition device, the electronic device, and the server may be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not make any limitations. The exemplary applications of the electronic device 300 will be described below.
[0029] Embodiments of the present application provide an image detection method, as Figure 1B shown, the method includes:
[0030] Step 102: Obtain a first image to be processed and input object description information, where the object description information is a first voice or a first text that describes a first target object to be recognized.
[0031] Here, the first image to be processed may be a certain image in the image set. The image set may be at least two images in an image folder or an image sequence, that is, images with time series, such as video frames.
[0032] In some implementation manners, the image set may be images collected in real time by an image acquisition device, such as a camera module, provided on the electronic device; in other implementation manners, the image set may be images that need to be subjected to image detection and are transmitted to the electronic device by other devices through instant messaging; in some implementation manners, the image set may also be an image set obtained by the electronic device in response to a task processing instruction by calling the local photo album, and the embodiments of the present application do not make any limitations in this regard.
[0033] Here, the object description information is the input object description information, which reflects that the image detection method is an end-to-end image detection method. The object description information is a description of the first target object to be recognized, such as the category of the first target object (person, vehicle, bird, etc.), attributes (red, long hair, etc.), etc., and is used to recognize the first target object. The description method can be voice or text. For example, if the first target object to be recognized is a bird, the voice can be the first voice of "a bird", and the text can be the first text of "a bird".
[0034] The object description information can be input through artificial voice, or can be input to the electronic device through instant messaging, or can also be that the electronic device responds to a task processing instruction and calls local text or voice.
[0035] In some embodiments, the implementation of obtaining the input object description information in step 102 may include: for example, obtaining the input voice or text, and determining that the input voice or text is the first voice or the first text based on a preset first keyword. Among them, the preset first keyword can be set according to the category of the target object. For example, if the target object to be recognized is an animal, the preset first keyword can be the animal category, such as little bird, little dog, etc. If the input voice or text is a little bird, the input voice or text is the first voice or the first text; the preset first keyword can also be set according to the purpose. For example, if it is necessary to detect the target object, the preset first keyword can be detection. If the input voice or text contains detection, the input voice or text is the first voice or the first text. Also, for example, a trigger button or a trigger word can be set. By first pressing the trigger button or saying the trigger word such as "Xiaoxi Xiaoxi", etc., to trigger the recognition of the first voice or the first text, and then inputting the first voice or the first text, so as to distinguish that the input voice or text is the first voice or the first text.
[0036] Step 104: Extract features from the object description information to obtain object description features;
[0037] Here, step 104 is used to extract the features of the object description information through a network model, and convert the object description information into a vector form containing the features of the object description information, which is convenient for subsequent operations.
[0038] In some embodiments, when the object description information is the first voice, the Wav2Vec2 network model can be used for feature extraction to obtain the first voice features;
[0039] When the object description information is the first text, the BERT network model can be used for feature extraction to obtain the first text features; the network models for extracting the first voice features and the first text features in the embodiments of the present application are not limited.
[0040] Step 106: Extract features from the first image to be processed to obtain first image features;
[0041] Here, step 106 is used to extract the features of the first image to be processed through a network model, and convert the first image to be processed into a vector form containing the features of the first image to be processed, which is convenient for subsequent operations.
[0042] In some embodiments, step 106 can be implemented by using a residual network model such as Resnet34, Resnet50, Resnet101, or Resnet152 to extract features from the first image to be processed to obtain first image features. The network model for extracting the first image features is not limited in the embodiments of the present application.
[0043] Step 108: Based on the object description features and the first image features, detect the first target object in the first image to be processed to obtain the detection result of the first image to be processed.
[0044] Here, the detection result indicates whether the first target object is detected or not. In the case where the detection result indicates that the first target object is detected, the detection result includes the position information of the first target object. The position information can be composed of the coordinates of the upper left corner point and the lower right corner point of the area where the first target object is located. For example, (x1, y1, x2, y2), where (x1, y1) are the coordinates of the upper left corner point and (x2, y2) are the coordinates of the lower right corner point.
[0045] When implementing step 108, the object description features and the first image features can be first fused into one feature, and then input into the Transformer model. The object description features and the first image features are fused through the Transformer model, and finally the detection and positioning of the first target object are realized through the detection head. Among them, the Transformer model is a classic model in NLP (Natural Language Processing) proposed in 2017. The Transformer model uses the self-attention mechanism. The Transformer model includes an encoder and a decoder. The encoder includes a self-attention layer, and the decoder includes a self-attention layer and a cross-attention layer. Through the self-attention mechanism of the Transformer model, the fusion of the object description features and the first image features can be realized, and more attention can be placed on the first target object, and then the position information of the first target object is detected through the detection head.
[0046] In some embodiments, the implementation of step 108 may include:
[0047] Step 1081: Fuse the object description feature and the second image feature to obtain a first fusion feature, where the second image feature includes the first image feature.
[0048] Here, in order to fuse the object description feature and the second image feature to obtain a first fusion feature, the implementation of step 1081 can first convert the object description feature and the second image feature into features with the same number of dimensions and at most one dimension having a different number of dimensions, and then connect the two features with the same number of dimensions and at most one dimension having a different number of dimensions to obtain the first fusion feature.
[0049] Taking the dimension of the second image feature, i.e., the first image feature, as Wg / s×Hg / s×C1 and the dimension of the object description feature (e.g., the first speech feature) as Ns×C2 as an example for illustration. Here, the second image feature is the feature obtained after extracting features from the first image to be processed through the Resnet network model, and the first speech feature is the feature obtained after extracting features through the Wav2Vec2 network model when the object description information is the first speech. The number of dimensions of the second image feature is 3, and the number of dimensions of the first speech feature is 2. In some embodiments, to make the number of dimensions of the second image feature the same as that of the first speech feature, the width dimension and height dimension of the second image feature can be flattened into one dimension, i.e., Wg / s*Hg / s×C1, so as to change the number of dimensions of the second image feature from 3 to 2, making the number of dimensions of the second image feature the same as that of the first speech feature. To make the second image feature and the first speech feature have at most one dimension with a different number of dimensions, in some embodiments, the channel dimension C2 of the first speech feature with the same number of dimensions and the channel dimension C1 of the second image feature can both be converted to Cei, where Cei can be equal to C1, C2, or other numbers. That is, the dimension of the second image feature is Wg / s*Hg / s×Cei, and the dimension of the first speech feature is Ns×Cei, thus obtaining the second image feature and the first speech feature with only one dimension having a different number of dimensions. Then connect the first speech feature and the second image feature with the same number of dimensions and only one dimension having a different number of dimensions to obtain a first fusion feature with a dimension of (Ns + Wg / s*Hg / s)×Cei.
[0050] In some embodiments, the implementation of step 1081 can also first convert the channel dimensions of the object description feature and the second image feature into the same number of dimensions, and then flatten the width dimension and height dimension of the second image feature into one dimension, so as to obtain the object description feature and the second image feature with the same number of dimensions and at most one dimension having a different number of dimensions.
[0051] It should be noted that the method of fusing the object description feature and the second image feature depends on the network model for extracting features from the object description information and the first image to be processed. For example, if the dimensionality of the object description information and the first image to be processed after feature extraction by the network model is 3, the object description feature and the second image feature can both be flattened to obtain two dimensions, and then the dimensionality of one of the dimensions is changed to the same dimensionality, and after connection, the first fusion feature is obtained; for another example, if the dimensionality of the object description information and the first image to be processed after feature extraction by the network model is 2, there is no need to flatten, and directly the dimensionality of one of the dimensions is changed to the same dimensionality, and after connection, the first fusion feature is obtained.
[0052] Step 1082: Based on the first fusion feature, determine the detection result of the first image to be processed.
[0053] In some embodiments, the implementation of step 1082 may include:
[0054] Step 1821: Encode the first fusion feature through a first encoder to obtain a first encoded feature that fuses the object description feature and the second image feature;
[0055] Here, the first encoder is the encoder in the Transformer model, which is used to utilize the self-attention layer in the first encoder to realize the fusion of the object description feature and the second image feature and obtain the first encoded feature.
[0056] Step 1822: Decode the first encoded feature through a first decoder to obtain a first decoded feature that fuses the object description feature and the second image feature;
[0057] Here, the first decoder is the decoder in the Transformer model, which is used to further utilize the self-attention layer and the cross-attention layer in the first decoder to realize the fusion of the object description feature and the second image feature and obtain the first decoded feature.
[0058] Step 1823: Based on the first encoded feature and the first decoded feature, the first detection head realizes the positioning of the first target object to obtain the detection result of the first image to be processed.
[0059] Here, the first detection head fuses the first encoded feature and the first decoded feature, and operates on the fused first encoded feature and first decoded feature to obtain the top-left point probability map and the bottom-right point probability map of the first target object. The maximum values greater than the threshold are determined in the top-left point probability map and the bottom-right point probability map respectively. If there are maximum values greater than the threshold, it means that the first target object exists in the first image to be processed, that is, the detection result indicates that the first target object is detected. At the same time, the positions of the maximum values greater than the threshold are respectively determined as the top-left point coordinates and the bottom-right point coordinates of the first target object. If there are no values greater than the threshold, it means that the first target object does not exist in the first image to be processed, and the detection result indicates that the first target object is not detected.
[0060] In the case where the object description information is the first text, the implementation of step 108 may include: detecting the first target object in the first image to be processed based on the first text feature and the first image feature, and obtaining the detection result of the first image to be processed. Then, in the case where the object description information is the first text, it can be widely applied in the scenario of human-computer interaction. For example, for video classification, the category can be described by text, and it is required to classify the video according to the category. Then, after different categories are recognized in the video, the video can be classified. Another example is to find a target in multiple images. The target can be described by text. Then, after the target is recognized in multiple images, the image where the target is located is output.
[0061] In the case where the object description information is the first voice, the implementation of step 108 may include: detecting the first target object in the first image to be processed based on the first voice feature and the first image feature, and obtaining the detection result of the first image to be processed. Then, in the case where the object description information is the first voice, it can also be widely applied in the scenario of human-computer interaction. For example, for autonomous driving, an instruction can be issued by voice, requiring the vehicle to track the bird in front. Then, after the bird is recognized in the image in front of the vehicle collected by the vehicle, the vehicle can be controlled to track the bird. Another example is for intelligent video surveillance. An instruction can be issued by voice, requiring to find the target person in the surveillance video. Then, after the target person is recognized in the surveillance video, the image where the target person is located is output.
[0062] Compared with text, voice has incomparable advantages in terms of efficiency, accessibility and applicability. In terms of efficiency, humans speak at a speed of about 150 words per minute, while typing is only 40 words per minute. In terms of accessibility, humans can speak from infancy, but recognizing words or typing requires acquired training. There are still about 14% of the illiterate population in the world, and these people cannot type. In terms of applicability, a microphone is smaller and cheaper than a typing device (keyboard or touch screen), and when the hands are occupied, such as when driving, people cannot type. Therefore, when the object description information is the first voice, it can make the operation more convenient and efficient, with a wider user base and more application scenarios, which is more conducive to improving the human-computer interaction experience in actual applications.
[0063] In the embodiment of the present application, the object description features and the first image features are obtained by extracting the features of the object description information (first voice or first text) and the features of the first image to be processed, and then the first target object is detected on the first image to be processed based on the object description features and the first image features to obtain the detection result of the first image to be processed. That is, by combining the first voice and the first image to be processed, or the first text and the first image to be processed, the target object is detected in the image. Therefore, it can be widely used in human-computer interaction scenarios, using text or voice to issue instructions, detecting the target object in the image, and increasing the fun of human-computer interaction; at the same time, when the object description information is the first voice, due to the advantages of voice in efficiency, accessibility and applicability, the operation is more convenient and efficient, the user population is wider, the application scenarios are more, and it is more helpful to improve the human-computer interaction experience in actual applications.
[0064] In some embodiments, after step 108, the following steps may also be included:
[0065] Step 110a: Acquire a second voice;
[0066] Step 112a: When the second speech includes a preset first keyword, extracting features from the second speech to obtain second speech features;
[0067] Here, step 110a and step 112a may refer to step 102 .
[0068] In the case that the second voice includes the preset first keyword, the second voice can be determined to be the voice of the second target object, that is, after completing the detection of the first target object, the detection of the second target object is performed.
[0069] Step 114a: Acquire a second image to be processed;
[0070] Here, step 114a may refer to step 102.
[0071] Step 116a: Extract features from the second image to be processed to obtain third image features;
[0072] Step 118a: Based on the second voice feature and the third image feature, detect the second target object in the second image to be processed to obtain the detection result of the second image to be processed.
[0073] Here, for Step 116a and Step 118a, reference can be made to Step 106 and Step 108.
[0074] In the embodiments of the present application, when obtaining the second voice, by judging the category of the second voice, it is determined whether the second voice is the second target object to be detected, so as to detect the second target object after completing the detection of the first target object.
[0075] In some embodiments, after Step 108, it may further include:
[0076] Step 110b: Obtain the second voice;
[0077] Here, for Step 110b, reference can be made to Step 102.
[0078] Step 112b: When the second voice includes a preset second keyword, determine that the second voice is a correction voice;
[0079] Here, the correction voice is a voice used to correct the size and orientation of the position information of the first target object in the detection result of the first image to be processed, for example: "right down". The preset second keyword can be words representing directions and sizes, such as "up down left right size width narrow length". When the second voice includes the preset second keyword, then the second voice is a correction voice. For example, if the second voice is "left up", which includes "left and up", then the second voice is a correction voice.
[0080] Table 1-1 shows the description of positions in the correction voice, and Table 1-2 shows the description of sizes in the correction voice.
[0081] Table 1-1
[0082]
[0083] Table 1-2
[0084]
[0085] In some embodiments, the correction voice can combine the position description in Table 1-1 and the size description in Table 1-2, such as "move right and become smaller"
[0086] Step 114b: Update the position information of the first target object in the first image to be processed based at least on the corrected speech.
[0087] In some embodiments, the implementation of step 114b may include:
[0088] Step 1141: Determine a corrected convolutional kernel based on the corrected speech.
[0089] In some embodiments, the implementation of step 1141 may include steps 141a to 141c:
[0090] Step 141a: Extract the corrected speech features of the corrected speech.
[0091] Here, step 141a is used to extract the features of the corrected speech through a network model, and convert the corrected speech into a vector form containing the corrected speech features, which is convenient for subsequent operations.
[0092] In some embodiments, the Wav2Vec2 network model may be used for feature extraction to obtain the corrected speech features. The embodiments of the present application do not limit the network model for extracting the corrected speech features.
[0093] Step 141b: Encode the corrected speech features through a second encoder to obtain second encoded features.
[0094] Step 141c: Decode the second encoded features through a second decoder to obtain the corrected speech convolutional kernel.
[0095] Here, steps 141b and 141c may refer to steps 1821 and 1822.
[0096] Among them, the corrected speech convolutional kernel is the convolutional kernel for subsequent convolutional operations. In some embodiments, after decoding the second encoded features through a second decoder to obtain decoded features, the dimensions of the decoded features may be changed through a dimension transformation method to obtain the corrected speech convolutional kernel for subsequent convolutional operations.
[0097] Step 1142: Determine a target mask image based on the position information of the first target object in the first image to be processed and the first image to be processed.
[0098] Here, the target mask image is a mask image containing the position information of the target object.
[0099] In some embodiments, the implementation of step 1142 may include steps 142a and 142b:
[0100] Step 142a: Generate a binary image from the first image to be processed based on the position information of the first target object in the first image to be processed and the first image to be processed.
[0101] Here, in subsequent convolution operations, to reduce the influence of the image outside the first target object in the first image to be processed on the position information of the first target object, a binary image can be generated from the first image to be processed.
[0102] In implementation, based on the position information of the first target object in the first image to be processed, the area where the first target object is located can be set to 1, and other areas in the first image to be processed can be set to 0, so as to generate a binary image from the first image to be processed.
[0103] Step 142b: Copy the binary image M times along the channel dimension direction based on the channel dimension of the first image feature of the first image to be processed to obtain the target mask image.
[0104] Here, since the number of channel dimensions of the binary image is 1, to increase the influence of the position information of the first target object in the first image to be processed in subsequent convolution operations, the binary image can be copied M times, where M can be the number of dimensions of the channel dimension of the first image feature of the first image to be processed; it can also be a number near the number of dimensions of the channel dimension of the first image feature of the first image to be processed, for example, M - 1, M + 1, etc.
[0105] In some embodiments, if a partial image region of the first image to be processed is intercepted for feature extraction, the implementation of step 142b may include: copying the binary image M times along the channel dimension direction based on the channel dimension of the image region feature to obtain the target mask image.
[0106] Step 1143: Update the position information of the first target object in the first image to be processed based on the corrected convolution kernel, the target mask image, and the first image feature of the first image to be processed.
[0107] In some embodiments, if a partial image region of the first image to be processed is intercepted for feature extraction, the implementation of step 1143 may include: updating the position information of the first target object in the first image to be processed based on the corrected convolution kernel, the target mask image, and the image region feature.
[0108] In some embodiments, the implementation of step 1143 may include steps 143a to 143c:
[0109] Step 143a: Fuse the target mask image and the first image feature of the first image to be processed to obtain a second fused feature.
[0110] Here, when the number of dimensions of the target mask image is the same as that of the first image feature, and the dimensions of the width dimension and the height dimension are the same, the implementation of step 143a can be: directly connecting the target mask image and the first image feature of the first image to be processed to obtain a second fusion feature.
[0111] When the number of dimensions of the target mask image is different from that of the first image feature, and the dimensions of the width dimension and the height dimension are different, the implementation of step 143a can be: first, through a dimension transformation method, making the number of dimensions of the target mask image and the first image feature the same, and the dimensions of the width dimension and the height dimension the same, and then obtaining a second fusion feature through connection.
[0112] Step 143b: Performing a convolution operation on the second fusion feature by using the corrected speech convolution kernel to obtain a third fusion feature;
[0113] Here, by performing a convolution operation on the second fusion feature by using the corrected speech convolution kernel, the change information of the position and size in the corrected speech convolution kernel is incorporated into the second fusion feature to obtain a third fusion feature.
[0114] Step 143c: Updating the position information of the first target object in the first image to be processed by a second detection head based on the third fusion feature.
[0115] Here, step 143c can refer to step 1823.
[0116] Since the change information of the position and size in the corrected speech convolution kernel is incorporated into the third fusion feature, after the second detection head performs arithmetic processing on the third fusion feature, the updated position information of the first target object including the change information of the position and size in the corrected speech convolution kernel can be obtained to update the position information of the first target object in the first image to be processed.
[0117] In the embodiment of the present application, after obtaining the detection result of the first image to be processed, by obtaining the corrected speech and based on the corrected convolution kernel, the target mask image and the first image feature of the first image to be processed, the position information of the first target object is updated in the first image to be processed to improve the accuracy of the position information of the first target object.
[0118] When the first object description information is the first speech and the first image to be processed is each image in the image set, the embodiment of the present application further provides an image detection method, and the method includes:
[0119] Step 202: Obtaining each image in the image set and the input first speech, where the first speech is a description of the first target object to be recognized.
[0120] Here, the image set can be one of the following: at least two images in an image folder; an image sequence;
[0121] Step 204: Extract features from the first speech to obtain first speech features;
[0122] Step 206: Extract features from each image in the image set to obtain each image feature;
[0123] Here, steps 202 to 206 can refer to steps 102 to 106;
[0124] Step 208: Based on the first speech features and each image feature, detect the first target object for each image in the image set, obtain and record the detection result of each image in the image set to obtain a recording result; wherein, the detection result indicates whether the first target object is detected or not;
[0125] Here, step 208 can refer to step 108. The difference between step 208 and step 108 is that step 208 needs to obtain the detection result of each image in the image set.
[0126] Step 210: Output at least the recording result.
[0127] In the embodiments of the present application, by obtaining the detection result of each image in the image set, the image with the target object is detected and output in the image set, thereby improving the human-computer interaction experience in practical applications.
[0128] When the object description information is the first speech, the embodiments of the present application further provide an image detection method, and the method includes:
[0129] Step 302: Obtain a first image to be processed and the input first speech, wherein the first speech is a description of the first target object to be recognized;
[0130] Here, step 302 can refer to step 102.
[0131] Step 304: Convert the first speech into a second text;
[0132] Here, step 304 can be implemented by a speech recognition network model such as Transformer, Conformer, etc. to convert the first speech into a second text. The embodiments of the present application do not limit the network model for converting the first speech into a second text.
[0133] Step 306: Extract features from the second text to obtain second text features;
[0134] Step 308: Extract features from the first image to be processed to obtain first image features;
[0135] Step 310: Based on the second text features and the first image features, detect the first target object in the first image to be processed to obtain the detection result of the first image to be processed.
[0136] Here, Steps 306 to 310 can refer to Steps 104 to 108.
[0137] In the embodiments of the present application, by converting the first speech into the second text and then based on the second text and the first image to be processed, the detection result of the first image to be processed is obtained, providing a new method for image detection.
[0138] When the first image to be processed is the i-th frame image in the image sequence and the object description information is the first speech, the embodiments of the present application further provide an image detection method, as Figure 2A shown, the method includes:
[0139] Step 402: Obtain the i-th frame image and the input first speech, where the first speech is a description of the first target object to be recognized;
[0140] Step 404: Extract features from the first speech to obtain first speech features;
[0141] Step 406: Extract features from the i-th frame image to obtain i-th frame image features;
[0142] Step 408: Based on the i-th frame image features and the first speech features, detect the first target object in the i-th frame image to obtain the detection result of the i-th frame image.
[0143] Here, Steps 402 to 408 can refer to Steps 102 to 108.
[0144] After Step 408, the method further includes Step 410a or 410b:
[0145] Step 410a: Extract features from the (i + 1)-th frame image to obtain (i + 1)-th frame image features; Based on the (i + 1)-th frame image features and the first speech features, track the first target object in the (i + 1)-th frame image;
[0146] That is, after obtaining the detection result of the i-th frame image, track the first target object in the (i + 1)-th frame image.
[0147] Figure 2B For the schematic diagrams of different tracking methods, as Figure 2BAs shown, the traditional visual tracking method (i.e., Figure 2B Figure (a) in Figure 2B ) requires the true value box of the target object to be marked in the first frame image and starts tracking from the second frame. Compared with the traditional visual tracking method, the method provided in the embodiments of the present application (i.e.,
[0148] Figures (b) and (c) in
[0149] , where Figure (b) is the method corresponding to the first text of the object description information, and Figure (c) is the method corresponding to the first voice of the object description information) does not depend on the annotation of the true value box of the target object in the first frame image and can directly track the target object. At the same time, in the case of using the combination of voice and image to track the target object, due to the advantages of voice in terms of efficiency, accessibility, and applicability, the operation is more convenient and efficient, the user group is wider, the application scenarios are more, and it is more helpful to improve the human-computer interaction experience in practical applications.
[0150] Step 410b: When the detection result of the i-th frame image indicates that the first target object is detected, based on the position information of the first target object in the detection result of the i-th frame image, track the first target object in the (i + 1)-th frame image.
[0151] Here, the area to be processed is a part of the (i + 1)-th frame image.
[0152] In some embodiments, the implementation of step 410b may include steps 41b1 to 41b3:
[0153] Step 41b1: Based on the position information of the first target object in the detection result of the i-th frame image, determine the (i + 1)-th frame image area in the (i + 1)-th frame image and use the (i + 1)-th frame image area as the area to be processed;
[0154] For example: If the position information of the first target object is (x1, y1, x2, y2), then the central position is ((x1 + x2) / 2, (y1 + y2) / 2).
[0155] In some embodiments, if the position information of the first target object is updated after the position information of the first target object is obtained, the implementation of step 4111 may further include: obtaining the central position of the prediction area where the updated position information of the first target object in the i-th frame image is located, and then, with the central position of the prediction area where the updated position information of the first target object is located as the center, obtaining the (i + 1)-th frame image area.
[0156] Step 4112: With the central position as the center, intercept an N-fold area of the prediction area in the (i + 1)-th frame image to obtain the (i + 1)-th frame image area, and use the (i + 1)-th frame image area as the area to be processed;
[0157] Among them, N can be set according to the computing power of the electronic device. If the computing power of the electronic device is large, N can take a larger value; if the computing power of the electronic device is small, N can take a smaller value. In implementation, with the central position as the center, the long side and the short side of the prediction area can be enlarged by the same multiple to obtain the (i + 1)-th frame image area, or with the central position as the center, the long side and the short side of the prediction area can be enlarged by different multiples to obtain the (i + 1)-th frame image area. In some embodiments, N can be 4.
[0158] Here, the i-th frame can be any frame. If the i-th frame is the first frame, only the first frame image uses the entire image for image detection, and starting from the second frame image, the second frame image area is used for image detection. If the i-th frame is the second frame, only the first two frame images use the entire image for image detection, and starting from the third frame image, the third frame image area is used for image detection. Compared with the method of using the entire image for image detection, using the area to be processed for image detection can reduce the amount of calculation.
[0159] Step 41b2: Extract features from the area to be processed to obtain the features of the area to be processed;
[0160] Here, step 41b2 can refer to step 106.
[0161] Step 41b3: Based on the features of the area to be processed and the first voice feature, perform tracking of the first target object on the (i + 1)-th frame image.
[0162] Here, step 41b3 can refer to step 108 (including step 1081 and step 1082).
[0163] Among them, in step 1081, the second image feature further includes the features of the area to be processed. That is, the implementation of step 1081 can be: fusing the object description feature and the features of the area to be processed to obtain the first fusion feature.
[0164] In some embodiments, when the second image feature is the first image feature and the feature of the area to be processed, the encoder, decoder, and detection head used may be the same or different, and the embodiments of the present application do not limit this.
[0165] In the embodiments of the present application, based on the position information of the first target object in the detection result of the i-th frame image, the i + 1-th frame image area is determined in the i + 1-th frame image, and then the i + 1-th frame image area is used for image detection to implement tracking of the first target object in the i + 1-th frame image. Compared with the method of using the entire image for tracking the first target object, the amount of calculation can be reduced.
[0166] The embodiments of the present application provide a target tracking method, as Figure 3A and Figure 3B shown, the method includes:
[0167] The first part: Obtain the t1-th frame image (i.e., the first image to be processed) and the first voice, where the first voice is a description of the first target object to be recognized; extract features from the first voice using the Wav2Vec2 network model to obtain the first voice feature 101 (i.e., the object description feature); extract features from the t1-th frame image using the Resnet50 network model to obtain the t1-th frame image feature 102 (i.e., the first image feature);
[0168] Here, the first part can be referred to in steps 102 to 106.
[0169] The second part: Fuse the first voice feature and the t1-th frame image feature to obtain the third fusion feature 103 (i.e., the first fusion feature);
[0170] Here, the second part can be referred to in step 1081.
[0171] The third part: Encode the third fusion feature through the first encoder to obtain the third encoded feature 104 (i.e., the first encoded feature) that fuses the first voice feature and the t1-th frame image feature; decode the third encoded feature through the first decoder to obtain the third decoded feature (i.e., the first decoded feature) that fuses the first voice feature and the t1-th frame image feature; based on the third encoded feature and the third decoded feature, the first detection head realizes the positioning of the first target object to obtain the detection result of the t1-th frame image;
[0172] Here, the third part can be referred to in steps 1821 to 1823.
[0173] The fourth part: Obtain the t iThe central position of the prediction region where the position information of the first target object in the detection result of the frame image is located, where i is an integer greater than or equal to 1. When i equals 1, the t i th frame is the t1th frame; with the central position as the center, intercept an N-fold region of the prediction region in the t i+1 th frame image to obtain the t i+1 th frame image region, and use the t i+1 th frame image region as the region to be processed;
[0174] Here, the fourth part can be seen in steps 4111 and 4112.
[0175] Fifth part: Extract features from the region to be processed to obtain the feature 105 of the region to be processed;
[0176] Here, the fifth part can be seen in step 41b2.
[0177] Sixth part: Fuse the first speech feature and the feature of the region to be processed to obtain the fourth fused feature 106 (i.e., the first fused feature); encode the fourth fused feature through a third encoder to obtain the fourth encoded feature 107 (i.e., the first encoded feature) that fuses the first speech feature and the feature of the region to be processed; decode the third encoded feature through a third decoder to obtain the fourth decoded feature (i.e., the first decoded feature) that fuses the first speech feature and the feature of the region to be processed; use a third detection head to locate the first target object based on the fourth encoded feature and the fourth decoded feature to obtain the detection result of the t i+1 th frame image;
[0178] Here, the sixth part can be seen in steps 1081, 1821 to 1823.
[0179] Seventh part: Obtain a second speech; in the case where the second speech includes a preset second keyword, determine the second speech as a corrective speech;
[0180] Here, the seventh part can be seen in steps 110b and 112b.
[0181] Eighth part: Extract the corrective speech feature of the corrective speech; encode the corrective speech feature through a second encoder to obtain a second encoded feature; decode the second encoded feature through a second decoder to obtain the corrective speech convolution kernel.
[0182] Here, the eighth part can be seen in steps 141a to 141c.
[0183] Part IX: Based on the position information of the first target object in the first image to be processed and the first image to be processed, generate a binary image 108 from the first image to be processed; based on the channel dimension of the characteristics of the area to be processed, copy the binary image M times along the channel dimension direction to obtain the target mask image.
[0184] Here, Part IX can refer to Step 142a and Step 142b.
[0185] Part X: Fuse the target mask image and the characteristics of the area to be processed to obtain a second fused feature 109; perform a convolution operation on the second fused feature using the corrected speech convolution kernel to obtain a third fused feature; based on the third fused feature through the second detection head, update the position information of the first target object in the first image to be processed, including the updated upper left corner point coordinates 110 and the updated lower right corner point coordinates 111.
[0186] Here, Part X can refer to Step 143a to Step 143c.
[0187] In the embodiment of the present application, on the first hand, by combining speech and image, the tracking of the first target object is realized. Compared with the traditional visual tracking method, it is not necessary to mark the true value box of the target object in the first frame of image, making the operation more convenient. At the same time, due to the advantages of speech in terms of efficiency, accessibility and applicability, the operation is more convenient and efficient, the user group is wider, the application scenarios are more, and it is more helpful to improve the human-computer interaction experience in practical applications; on the second hand, by using the position information of the first target object in the detection result of the t i th frame image, determine the t i+1 th frame image area, and then use the t i+1 th frame image area to track the first target object in the t i+1 th frame image, which can reduce the calculation amount; on the third hand, by introducing corrected speech to update the position information of the first target object, the result is more accurate and the accuracy of tracking is improved.
[0188] The embodiment of the present application provides an image detection model training method, and the method includes:
[0189] Step 602: Obtain an image sample set and an object description sample set, where the object description sample set is a speech sample set or a text sample set for describing the object to be recognized;
[0190] In some embodiments, the image sample set can be from two major databases, LaSOT and TNL2K. When the object description sample set is a text sample set, the object description sample set can also be from the two major databases, LaSOT and TNL2K; when the object description sample set is a voice sample set, the object description sample set can be obtained by manually voice annotating the two major databases, LaSOT and TNL2K. During implementation, people with different accents can perform various manual voice annotations on the two major databases, LaSOT and TNL2K, to simulate the actual application scenario and reduce the impact of different voice features (such as speed, accent, noise) on the detection results.
[0191] In some embodiments, the object description sample set can be set as six different difficulty levels of voice sample sets as shown in Table 2 to train the model.
[0192] Table 2
[0193] Setting Voice source Proportion of native English speakers Noise MF Machine-converted female voice 100% MM Machine-converted male voice 100% HF Native English female voice 100% ★ HM Native English male voice 100% ★ HC1 Voices of 17 different people 58.8% ★★ HC2 Voices of 17 different people 29.4% ★★
[0194] Among them, the more ★, the greater the noise. The first two voices come from machine conversion and are converted into male voice (MM) and female voice (MF) respectively. Such voices are standard English and have no noise, which rarely exist in reality. The third and fourth voices come from female and male annotators whose native language is English respectively. They provide voice annotations for each video. These two parts of voices will contain a small amount of noise introduced by the recording equipment. To be closer to the actual situation, the fifth and sixth voices come from 17 different people, and only some of them have English as their native language. During implementation, 3,400 videos (among them, 1,400 are from the LaSOT dataset and 2,000 are from TNL2K) are divided into 17 groups, 34 annotators are found, and each annotator annotates 200 segments of voices. The annotated voices are divided into two parts (HC1 and HC2). Since only some of these two parts of voices come from annotators whose native language is English, the annotated voices contain accent variations and the noise level is higher than the first four voice annotations.
[0195] Table 3 is a statistical table of the voice duration and the number of files in the two major databases, LaSOT and TNL2K.
[0196] Table 3
[0197]
[0198] Among them, the voice files are the six different difficulty levels of voice files in Table 2.
[0199] Figure 4Shows the number and proportion of speech files corresponding to different speech durations in the two major databases of LaSOT and TNL2K. Among them, most speech files last for 3 to 4 seconds (s), and the proportion (18%) of files with a duration of 1 to 2 s and greater than 7 s in the TNL2K dataset is greater than the proportion (8%) of files with a duration of 1 to 2 s and greater than 7 s in the LaSOT dataset.
[0200] Step 604: Use the image sample set and the object description sample set to train an image detection model to obtain the detection results of each image sample in the image sample set;
[0201] In some embodiments, the implementation of step 604 may include:
[0202] Step 6041a: Use the image detection model to extract features from the object description sample set to obtain an object description feature set;
[0203] Step 6042a: Use the image detection model to extract features from the image sample set to obtain a first image feature set;
[0204] Step 6043a: Based on the object description feature set and the first image feature set, use the image detection model to detect the object to be recognized in each image sample in the image sample set to obtain the detection results of each image sample in the image sample set.
[0205] Here, steps 6041a to 6043a can refer to steps 104 to 108.
[0206] Step 606: Determine the loss of the image detection model based on the detection results of each image sample;
[0207] Here, the loss function of the image detection model can be:
[0208]
[0209] Among them, L iou is the loss representing the overlap degree between the ground truth box and the predicted box, L1 is the loss representing the difference between the ground truth box and the predicted box, that is, the minimum absolute value error; θ is the network parameter, λ is the weight coefficient of the loss function, b i is the predicted value, is the ground truth.
[0210] The first loss is the loss value when b i is the detection result of each image sample. By substituting the detection results of each image sample and the ground truth into formula (6-1), the first loss is obtained.
[0211] Step 608: Adjust the network parameters of the image detection model using the loss, so that the loss of the detection result output by the adjusted image detection model meets the convergence condition.
[0212] Here, the formula for the network parameters can be:
[0213]
[0214] Where θ is the network parameter and α is the weight coefficient of the loss function.
[0215] By substituting L(θ) into formula (6-2), the adjusted network parameters are obtained.
[0216] In the embodiments of the present application, by obtaining an image sample set and an object description sample set, then using the image sample set and the object description sample set to train an image detection model, obtaining the detection results of each image sample in the image sample set, and then based on the detection results of each image sample, determining the loss of the image detection model, and finally using the loss to adjust the network parameters of the image detection model, the training of the image detection model is realized, so that the loss of the detection result output by the adjusted image detection model can meet the convergence condition.
[0217] In some embodiments, the detection module may include a first detection module and a second detection module, and the implementation of step 604 may include:
[0218] Step 6041b: Use the i-th frame image sample in the image sample set and the object description sample set to train the first detection module to obtain the detection results of each i-th frame image sample in the image sample set;
[0219] Here, step 6041b can refer to steps 104 to 108.
[0220] Step 6042b: When the detection results of each i-th frame image sample in the image sample set indicate that the object to be recognized is detected, use the position information of the object to be recognized in the detection results of each i-th frame image sample in the image sample set, the (i + 1)-th frame image sample and the object description sample set to train the second detection module to obtain the detection results of each (i + 1)-th frame image sample in the image sample set;
[0221] In some embodiments, the implementation of step 6042b may include:
[0222] Step 6421: Based on the position information of the object to be recognized in the detection results of each i-th frame image sample in the image sample set, determine each (i + 1)-th frame image sample region in each of the (i + 1)-th frame image samples, and use each of the (i + 1)-th frame image sample regions as a set of regions to be processed;
[0223] Step 6422: Extract features from the set of regions to be processed to obtain a set of features of the regions to be processed;
[0224] Step 6423: Based on the set of features of the regions to be processed and the set of object description features, train the second detection module to obtain the detection results of each (i + 1)-th frame image sample in the image sample set.
[0225] Here, Steps 6421 to 6423 can refer to Steps 41b1 to 41b3.
[0226] Correspondingly, Step 606, "Determine the loss of the image detection model based on the detection results of each of the image samples", includes:
[0227] Step 6061: Based on the detection results of each i-th frame image sample in the image sample set, determine the third loss of the first detection module;
[0228] Here, the loss function of the first detection module can be:
[0229]
[0230] where g represents the first detection module.
[0231] Step 6062: Based on the detection results of each (i + 1)-th frame image sample in the image sample set, determine the fourth loss of the second detection module;
[0232] Here, the loss function of the second detection module can be:
[0233]
[0234] where l represents the second detection module.
[0235] Step 6063: Determine the loss of the image detection model based on at least the third loss and the fourth loss.
[0236] Here, the loss function of the image detection model can be:
[0237] L(θ) = α g L g (θ) + α l L l (θ) (6 - 5);
[0238] Among them, α is the weight coefficient of the loss function.
[0239] The expression of the corresponding network parameters is:
[0240]
[0241] In the embodiments of the present application, by adding a second detection module and converting the image sample into an image sample region for training, the calculation amount can be reduced.
[0242] The embodiments of the present application also provide an image detection model training method. The image detection model includes a detection module and a correction module. The method includes:
[0243] Step 701: Obtain an image sample set and an object description sample set. Among them, the object description sample set is a voice sample set or a text sample set for describing the object to be recognized;
[0244] Step 702: Use the image sample set and the object description sample set to train the detection module to obtain the detection results of each image sample in the image sample set;
[0245] Here, Step 702 can refer to Steps 104 to 108.
[0246] Step 703: Determine the deviation between the position information of the object to be recognized in the detection results of each image sample and the ground truth;
[0247] Here, the deviation can be the intersection over union (IoU) between the position information of the object to be recognized and the ground truth. Then, the implementation of Step 703 can be to determine the intersection over union (IoU) between the position information of the object to be recognized and the ground truth.
[0248] The embodiments of the present application do not limit the type of the deviation.
[0249] Step 704: When the deviation exceeds the preset range, obtain the correction voice;
[0250] Here, the preset range can be determined according to the type of the deviation and the required accuracy. For example, when the deviation is the intersection over union (IoU) between the position information of the object to be recognized and the ground truth and the accuracy requirement is relatively high, the preset range can be less than 0.4.
[0251] Step 705: At least use the correction voice to train the correction module and update the position information of the object to be recognized in each image sample;
[0252] In some embodiments, the implementation of Step 705 can include:
[0253] Step 7051: Determine a corrected convolution kernel based on the corrected speech;
[0254] Step 7052: Determine a target mask image based on the position information of the object to be recognized in each of the image samples and the image samples;
[0255] Step 7053: Update the position information of the object to be recognized in each of the image samples based on the corrected convolution kernel, the target mask image, and the first image feature set.
[0256] Here, Steps 7051 to 7053 can refer to Steps 141a to 141c.
[0257] Step 706: Determine a first loss of the correction module based on the updated position information of the object to be recognized;
[0258] Here, the loss function of the correction module can be:
[0259]
[0260] where t represents the correction module.
[0261] Step 707: Determine a second loss of the detection module based on the detection results of the image samples;
[0262] Here, the loss function of the detection module can be:
[0263]
[0264] In the case where the detection module includes a first detection module and a second detection module:
[0265] The loss function of the detection module can be:
[0266] L r (θ) = α g L g (θ) + α l L l (θ) (7 - 3);
[0267] where g represents the first detection module and l represents the second detection module.
[0268] Step 709: Determine the loss of the image detection model based on the first loss and the second loss.
[0269] Here, the loss function of the image detection model can be:
[0270] L(θ) = α r L r(θ) + α t L t (θ)(7 - 4);
[0271] Step 710: Adjust the network parameters of the image detection model using the loss, so that the loss of the detection result output by the adjusted image detection model meets the convergence condition.
[0272] Here, the formula for the network parameters can be:
[0273]
[0274] In the embodiments of the present application, by adding a correction module, the accuracy of the detection result is improved.
[0275] Based on the above method, in order to compare the image detection model provided by the embodiments of the present application with other models, all models are trained using the LaSOT dataset and the speech from the MM settings in Table 2. Among them, the size of the image input to the first detection module of the image detection model provided by the embodiments of the present application is 640×640 pixels, and the size of the image input to the second detection module is 320×320 pixels. The entire model is trained in an end-to-end manner, with a total of 250 epochs. When the I ou is less than 0.4, the correction module is triggered.
[0276] During implementation, a precision map and a success map are used to evaluate the performance of different models. Among them, the success map represents the ratio of the frames in which the I ou value between the predicted box and the ground truth box is higher than the predefined overlap threshold. The precision map represents the percentage of frames in which the position error between the predicted box and the ground truth box is less than the predefined threshold. According to the success map and the precision map, a success score and a precision score can be obtained. Figure 5 Shows the success scores and precision scores of different models. It can be seen that the success score of 0.70 and the precision score of 0.743 of the image detection model provided by the embodiments of the present application are higher than the success scores and precision scores of other models, indicating the feasibility and high accuracy of the model provided by the embodiments of the present application.
[0277] Based on the foregoing embodiments, the embodiments of the present application provide an image detection device. The device includes each module included, as well as each sub-module included in each module, each unit included in each sub-module, and each sub-unit included in each unit, all of which can be implemented by an electronic device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0278] Figure 6 This is a schematic diagram of the composition structure of an image detection device provided by an embodiment of the present application. As Figure 6 shown, the image detection device 600 includes a first acquisition module 601, a first extraction module 602, a second extraction module 603, and a first target object detection module 604, where:
[0279] The first acquisition module 601 is configured to acquire a first image to be processed and input object description information, where the object description information is a first voice or a first text for describing a first target object to be recognized;
[0280] The first extraction module 602 is configured to extract features from the object description information to obtain object description features;
[0281] The second extraction module 603 is configured to extract features from the first image to be processed to obtain first image features;
[0282] The target object detection module 604 is configured to detect the first target object in the first image to be processed based on the object description features and the first image features, and obtain a detection result of the first image to be processed.
[0283] In some embodiments, the object description information is the first voice, and the first image to be processed is each image in an image set; the image detection device 600 further includes: a second acquisition module, configured to acquire and record the detection result of each image in the image set to obtain a recording result; the detection result indicates whether the first target object is detected or not; an output module, configured to output at least the recording result.
[0284] In some embodiments, after acquiring the first voice, the image detection device 600 further includes: a conversion module, configured to convert the first voice into a second text; the first extraction module 602 is further configured to extract features from the second text to obtain second text features; the target object detection module 604 is further configured to detect the first target object in the first image to be processed based on the second text features and the first image features, and obtain a detection result of the first image to be processed.
[0285] In some embodiments, the first image to be processed is the i-th frame image in an image sequence, and the image detection device 600 further includes at least one of the following: a first tracking module, configured to extract features from the (i + 1)-th frame image to obtain the (i + 1)-th frame image features; and based on the (i + 1)-th frame image features and the first voice features, perform tracking of the first target object on the (i + 1)-th frame image; a second tracking module, configured to, when the detection result of the i-th frame image indicates that the first target object is detected, perform tracking of the first target object on the (i + 1)-th frame image based on the position information of the first target object in the detection result of the i-th frame image.
[0286] In some embodiments, the second tracking module includes: a first determination sub-module, configured to determine a (i + 1)-th frame image region in the (i + 1)-th frame image based on the position information of the first target object in the detection result of the i-th frame image, and use the (i + 1)-th frame image region as a region to be processed; an extraction sub-module, configured to extract features from the region to be processed to obtain region-to-be-processed features; and a tracking sub-module, configured to perform tracking of the first target object on the (i + 1)-th frame image based on the region-to-be-processed features and the first voice features.
[0287] In some embodiments, the first determination sub-module includes: an acquisition unit, configured to acquire the central position of a prediction region where the position information of the first target object in the detection result of the i-th frame image is located; and a cropping unit, configured to crop an N-fold region of the prediction region in the (i + 1)-th frame image with the central position as the center to obtain the (i + 1)-th frame image region.
[0288] In some embodiments, the target object detection module 604 includes: a fusion sub-module, configured to fuse the object description features and the second image features to obtain first fusion features, where the second image features include the first image features or the region-to-be-processed features; and a second determination sub-module, configured to determine the detection result of the first image to be processed based on the first fusion features.
[0289] In some embodiments, the second determination sub-module includes: an encoding unit, configured to perform encoding processing on the first fusion features through a first encoder to obtain first encoded features that fuse the object description features and the second image features; a decoding unit, configured to perform decoding processing on the first encoded features through a first decoder to obtain first decoded features that fuse the object description features and the second image features; and a positioning unit, configured to implement positioning of the first target object based on the first encoded features and the first decoded features through a first detection head to obtain the detection result of the first image to be processed.
[0290] In some embodiments, after obtaining the detection result of the first image to be processed, the image detection device 600 further includes: a third acquisition module, configured to acquire a second voice; a first extraction module 602, further configured to perform feature extraction on the second voice to obtain a second voice feature when the second voice includes a preset first keyword; a first acquisition module 601, further configured to acquire a second image to be processed; a second extraction module 603, further configured to perform feature extraction on the second image to be processed to obtain a third image feature; and a target object detection module 604, further configured to detect a second target object in the second image to be processed based on the second voice feature and the third image feature, so as to obtain the detection result of the second image to be processed.
[0291] In some embodiments, after obtaining the detection result of the first image to be processed, the image detection device 600 further includes: a third acquisition module, configured to acquire a second voice; a first determination module, configured to determine that the second voice is a correction voice when the second voice includes a preset second keyword; and an update module, configured to update the position information of the first target object in the first image to be processed at least based on the correction voice.
[0292] In some embodiments, the update module includes: a third determination sub-module, configured to determine a correction convolution kernel based on the correction voice; a fourth determination sub-module, configured to determine a target mask image based on the position information of the first target object in the first image to be processed and the first image to be processed; and a first update sub-module, configured to update the position information of the first target object in the first image to be processed based on the correction convolution kernel, the target mask image, and the first image feature of the first image to be processed.
[0293] In some embodiments, the fourth determination sub-module includes: a generation unit, configured to generate a binary image from the first image to be processed based on the position information of the first target object in the first image to be processed and the first image to be processed; and a replication unit, configured to replicate the binary image M times along the channel dimension direction based on the channel dimension of the first image feature of the first image to be processed, so as to obtain the target mask image.
[0294] In some embodiments, the third determination sub-module includes: a first extraction unit, configured to extract a correction voice feature of the correction voice; an encoding unit, configured to encode the correction voice feature through a second encoder to obtain a second encoded feature; and a decoding unit, configured to decode the second encoded feature through a second decoder to obtain the correction voice convolution kernel.
[0295] In some embodiments, the first update sub-module includes: a fusion unit configured to fuse the first image features of the target mask map and the first image to be processed to obtain second fused features; a convolution unit configured to perform a convolution operation on the second fused features by using the corrected speech convolution kernel to obtain third fused features; and an update unit configured to update the position information of the first target object in the first image to be processed by a second detection head based on the third fused features.
[0296] Figure 7 The following is a schematic structural diagram of a training device for an image detection model provided by an embodiment of the present application. As Figure 7 shown, the training device 700 for the image detection model includes:
[0297] A fourth acquisition module 701, configured to acquire an image sample set and an object description sample set, where the object description sample set is a speech sample set or a text sample set for describing an object to be recognized;
[0298] A first training module 702, configured to use the image sample set and the object description sample set to train an image detection model to obtain detection results of each image sample in the image sample set;
[0299] A second determination module 703, configured to determine the loss of the image detection model based on the detection results of each image sample;
[0300] An adjustment module 704, configured to use the loss to adjust network parameters of the image detection model so that the loss of the detection results output by the adjusted image detection model meets a convergence condition.
[0301] In some embodiments, the first training module includes: a second extraction sub-module, configured to use the image detection model to extract features from the object description sample set to obtain an object description feature set; a third extraction sub-module, configured to use the image detection model to extract features from the image sample set to obtain a first image feature set; and a target object detection sub-module, configured to detect the object to be recognized in each image sample in the image sample set by using the image detection model based on the object description feature set and the first image feature set to obtain detection results of each image sample in the image sample set.
[0302] In some embodiments, the image detection model includes a detection module and a correction module, and the first training module includes: a training sub-module, configured to use the image sample set and the object description sample set to train the detection module to obtain detection results of each image sample in the image sample set;
[0303] After obtaining the detection results of each image sample in the image sample set, the training device 700 of the image detection model further includes: a third determination module, which determines the deviation between the position information of the object to be recognized in the detection results of each image sample and the ground truth; a fifth acquisition module, which is used to acquire a correction voice when the deviation exceeds a preset range; a second training module, which is used to train the correction module at least using the correction voice, and update the position information of the object to be recognized in each image sample;
[0304] A fourth determination module, which is used to determine the first loss of the correction module based on the updated position information of the object to be recognized; correspondingly, the second determination module includes: a fifth determination sub-module, which is used to determine the second loss of the detection module based on the detection results of each image sample; a sixth determination sub-module, which is used to determine the loss of the image detection model based on the first loss and the second loss.
[0305] In some embodiments, the second training module includes: a seventh determination sub-module, which is used to determine a correction convolution kernel based on the correction voice; an eighth determination sub-module, which is used to determine a target mask map based on the position information of the object to be recognized in each image sample and each image sample; a second update sub-module, which is used to update the position information of the object to be recognized in each image sample based on the correction convolution kernel, the target mask map, and the first image feature set.
[0306] In some embodiments, the detection module includes a first detection module and a second detection module; the first training module includes: a first training sub-module, which is used to train the first detection module using the i-th frame image sample in the image sample set and the object description sample set to obtain the detection results of each i-th frame image sample in the image sample set; a second training sub-module, which is used to train the second detection module using the position information of the object to be recognized in the detection results of each i-th frame image sample in the image sample set, the (i + 1)-th frame image sample, and the object description sample set when the detection results of each i-th frame image sample in the image sample set indicate that the object to be recognized is detected, to obtain the detection results of each (i + 1)-th frame image sample in the image sample set;
[0307] Correspondingly, the second determination module includes: a ninth determination sub-module, configured to determine a third loss of the first detection module based on the detection results of each i-th frame image sample in the image sample set; a tenth determination sub-module, configured to determine a fourth loss of the second detection module based on the detection results of each (i + 1)-th frame image sample in the image sample set; an eleventh determination sub-module, configured to determine the loss of the image detection model based on at least the third loss and the fourth loss.
[0308] In some embodiments, the second training sub-module includes: a determination unit, configured to determine, based on the position information of the object to be recognized in the detection results of each i-th frame image sample in the image sample set, each (i + 1)-th frame image sample region in each (i + 1)-th frame image sample, and use each (i + 1)-th frame image sample region as a set of regions to be processed; a second extraction unit, configured to perform feature extraction on the set of regions to be processed to obtain a set of region features to be processed; a training unit, configured to train the second detection module based on the set of region features to be processed and the set of object description features to obtain the detection results of each (i + 1)-th frame image sample in the image sample set.
[0309] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0310] It should be noted that in the embodiments of the present application, if the above image detection method and related training method are implemented in the form of software function modules and sold or used as an independent product, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a ROM (Read Only Memory), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0311] Correspondingly, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements the steps in the image detection method and related training method provided in the above embodiments.
[0312] Correspondingly, an embodiment of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above image detection method and related training method are implemented.
[0313] It should be noted here that the descriptions of the above storage medium and platform embodiments are similar to those of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and platform embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0314] It should be noted that Figure 8 is a schematic diagram of a hardware entity of the electronic device according to an embodiment of the present application. As Figure 8 shown, the hardware entity of the electronic device 800 includes: a processor 801, a communication interface 802, and a memory 803, where
[0315] The processor 801 generally controls the overall operation of the electronic device 800.
[0316] The communication interface 802 can enable the electronic device 800 to communicate with other platforms, electronic devices, or servers through a network.
[0317] The memory 803 is configured to store instructions and applications executable by the processor 801, and can also cache data to be processed or already processed by the processor 801 and each module in the electronic device 800 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by FLASH (flash memory) or RAM (Random Access Memory).
[0318] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical, or other forms.
[0319] The units described as separate components above may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0320] In addition, each functional unit in the embodiments of the present application can all be integrated into one processing module, or each unit can be separately used as one unit, or two or more units can be integrated into one unit; the above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0321] The methods disclosed in the several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0322] The features disclosed in the several product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0323] The features disclosed in the several method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0324] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An image detection method, characterized in that, Including: Obtain a first image to be processed and input object description information, where the object description information is the first speech or the first text for describing a first target object to be recognized; Extract features from the object description information to obtain object description features; Extract features from the first image to be processed to obtain first image features; Based on the object description features and the first image features, detect the first target object in the first image to be processed to obtain a detection result of the first image to be processed; the detection result indicates whether the first target object is detected or not; in the case where the detection result indicates that the first target object is detected, the detection result includes the position information of the first target object; Obtain a second speech; In the case where the second speech includes a preset second keyword, determine that the second speech is a correction speech; the correction speech is used to correct the size and orientation of the position information of the first target object in the detection result of the first image to be processed; Update the position information of the first target object in the first image to be processed at least based on the correction speech.
2. The method according to claim 1, wherein The object description information is the first speech, and the first image to be processed is each image in an image set; The method further includes: Obtain and record the detection results of each image in the image set to obtain a recording result; Output at least the recording result.
3. The method according to claim 1, characterized in that After obtaining the first speech, it further includes: Convert the first speech into a second text; Extract features from the second text to obtain second text features; The detecting the first target object in the first image to be processed based on the object description features and the first image features to obtain a detection result of the first image to be processed includes: Detect the first target object in the first image to be processed based on the second text features and the first image features to obtain a detection result of the first image to be processed.
4. The method according to claim 2, wherein The first image to be processed is the i-th frame image in an image sequence, and the method further includes at least one of the following: Extract features from the (i + 1)-th frame image to obtain (i + 1)-th frame image features; Track the first target object in the (i + 1)-th frame image based on the (i + 1)-th frame image features and first speech features; In the case where the detection result of the i-th frame image indicates that the first target object is detected, track the first target object in the (i + 1)-th frame image based on the position information of the first target object in the detection result of the i-th frame image.
5. The method according to claim 4, characterized in that, The tracking the first target object in the (i + 1)-th frame image based on the position information of the first target object in the detection result of the i-th frame image includes: Based on the position information of the first target object in the detection result of the i-th frame image, determine an (i + 1)-th frame image region in the (i + 1)-th frame image, and use the (i + 1)-th frame image region as a region to be processed; Extract features from the region to be processed to obtain region-to-be-processed features; Based on the characteristics of the area to be processed and the first speech feature, track the first target object in the (i + 1)-th frame image.
6. The method according to claim 5, characterized in that, Determining the (i + 1)-th frame image area in the (i + 1)-th frame image based on the position information of the first target object in the detection result of the i-th frame image includes: Obtaining the central position of the prediction area where the position information of the first target object in the detection result of the i-th frame image is located; Taking the central position as the center, intercepting an N-fold area of the prediction area in the (i + 1)-th frame image to obtain the (i + 1)-th frame image area.
7. The method according to any one of claims 1 to 6, characterized in that, Detecting the first target object in the first image to be processed based on the object description feature and the first image feature to obtain the detection result of the first image to be processed, including: Fusing the object description feature and the second image feature to obtain a first fusion feature, where the second image feature includes the first image feature or the characteristics of the area to be processed; Determining the detection result of the first image to be processed based on the first fusion feature.
8. The method according to claim 7, wherein Determining the detection result of the first image to be processed based on the first fusion feature includes: Encoding the first fusion feature through a first encoder to obtain a first encoded feature that fuses the object description feature and the second image feature; Decoding the first encoded feature through a first decoder to obtain a first decoded feature that fuses the object description feature and the second image feature; Implementing the positioning of the first target object based on the first encoded feature and the first decoded feature through a first detection head to obtain the detection result of the first image to be processed.
9. The method according to claim 1, characterized in that, After obtaining the detection result of the first image to be processed, it further includes: Obtaining a second speech; When the second speech includes a preset first keyword, extracting the feature of the second speech to obtain a second speech feature; Obtaining a second image to be processed; Extracting the feature of the second image to be processed to obtain a third image feature; Detecting the second target object in the second image to be processed based on the second speech feature and the third image feature to obtain the detection result of the second image to be processed.
10. The method according to claim 1, characterized in that, Updating the position information of the first target object in the first image to be processed at least based on the corrected speech includes: Determining a correction convolution kernel based on the corrected speech; Determining a target mask image based on the position information of the first target object in the first image to be processed and the first image to be processed; Updating the position information of the first target object in the first image to be processed based on the correction convolution kernel, the target mask image, and the first image feature of the first image to be processed.
11. The method according to claim 10, wherein Determining the target mask image based on the position information of the first target object in the first image to be processed and the first image to be processed includes: Generating a binary image from the first image to be processed based on the position information of the first target object in the first image to be processed and the first image to be processed; Based on the channel dimension of the first image feature of the first image to be processed, copy the binary map M times along the channel dimension direction to obtain the target mask map.
12. The method according to claim 10, characterized in that, The determining the correction convolution kernel based on the corrected speech includes: Extracting the corrected speech feature of the corrected speech; Encoding the corrected speech feature through a second encoder to obtain a second encoded feature; Decoding the second encoded feature through a second decoder to obtain the corrected speech convolution kernel.
13. The method according to claim 10, wherein The updating the position information of the first target object in the first image to be processed based on the second speech convolution kernel, the target mask map, and the first image feature of the first image to be processed includes: Fusing the target mask map and the first image feature of the first image to be processed to obtain a second fused feature; Performing a convolution operation on the second fused feature by using the corrected speech convolution kernel to obtain a third fused feature; Updating the position information of the first target object in the first image to be processed by a second detection head based on the third fused feature.
14. A method for training an image detection model, characterized in that, The method includes: Obtaining an image sample set and an object description sample set, where the object description sample set is a speech sample set or a text sample set for describing the object to be recognized; Training an image detection model by using the image sample set and the object description sample set to obtain detection results of each image sample in the image sample set; Determining the loss of the image detection model based on the detection results of each image sample; Adjusting the network parameters of the image detection model by using the loss so that the loss of the detection results output by the adjusted image detection model satisfies a convergence condition; The image detection model includes a correction module, and the detection result indicates whether the object to be recognized is detected or not. When the detection result indicates that the object to be recognized is detected, the detection result includes the position information of the object to be recognized; Determining the deviation between the position information of the object to be recognized in the detection results of each image sample and the ground truth; when the deviation exceeds a preset range, obtaining corrected speech; the corrected speech is used to correct the size and orientation of the position information of the object to be recognized in each image sample; Training the correction module at least by using the corrected speech to update the position information of the object to be recognized in each image sample.
15. The method according to claim 14, wherein The training the image detection model by using the image sample set and the object description sample set to obtain the detection results of each image sample in the image sample set includes: Extracting object description features by using the image detection model for the object description sample set to obtain an object description feature set; Extracting first image features by using the image detection model for the image sample set to obtain a first image feature set; Based on the set of object description features and the set of first image features, use the image detection model to detect the object to be recognized in each image sample in the image sample set, and obtain the detection results of each image sample in the image sample set.
16. The method according to claim 14, characterized in that, The image detection model further includes a detection module. The training of the image detection model using the image sample set and the object description sample set to obtain the detection results of each image sample in the image sample set includes: Training the detection module using the image sample set and the object description sample set to obtain the detection results of each image sample in the image sample set. Determining the loss of the image detection model based on the detection results of each image sample includes: Determining the first loss of the correction module based on the updated position information of the object to be recognized. Determining the second loss of the detection module based on the detection results of each image sample. Determining the loss of the image detection model based on the first loss and the second loss.
17. The method according to claim 14, wherein The training of the correction module using at least the correction speech to update the position information of the object to be recognized in each image sample includes: Determining a correction convolution kernel based on the correction speech. Determining a target mask image based on the position information of the object to be recognized in each image sample and each image sample. Updating the position information of the object to be recognized in each image sample based on the correction convolution kernel, the target mask image, and the set of first image features.
18. The method according to claim 16, wherein The detection module includes a first detection module and a second detection module. The training of the image detection model using the image sample set and the object description sample set to obtain the detection results of each image sample in the image sample set includes: Training the first detection module using the i-th frame image sample in the image sample set and the object description sample set to obtain the detection results of each i-th frame image sample in the image sample set. When the detection results of each i-th frame image sample in the image sample set indicate that the object to be recognized is detected, training the second detection module using the position information of the object to be recognized in the detection results of each i-th frame image sample in the image sample set, the (i + 1)-th frame image sample, and the object description sample set to obtain the detection results of each (i + 1)-th frame image sample in the image sample set. Correspondingly, determining the loss of the image detection model based on the detection results of each image sample includes: Determining the third loss of the first detection module based on the detection results of each i-th frame image sample in the image sample set. Determining the fourth loss of the second detection module based on the detection results of each (i + 1)-th frame image sample in the image sample set. Determining the loss of the image detection model based on at least the third loss and the fourth loss.
19. The method according to claim 18, characterized in that, Training the second detection module by using the position information of the object to be recognized in the detection results of each i-th frame image sample in the image sample set, the (i + 1)-th frame image sample, and the object description sample set to obtain the detection results of each (i + 1)-th frame image sample in the image sample set, including: Based on the position information of the object to be recognized in the detection results of each i-th frame image sample in the image sample set, determining each (i + 1)-th frame image sample region in each (i + 1)-th frame image sample, and using each (i + 1)-th frame image sample region as a set of regions to be processed; Performing feature extraction on the set of regions to be processed to obtain a set of features of regions to be processed; Training the second detection module based on the set of features of regions to be processed and the set of object description features to obtain the detection results of each (i + 1)-th frame image sample in the image sample set.
20. An image detection device, characterized in that, Including: A first acquisition module, configured to acquire a first image to be processed and input object description information, where the object description information is a first voice or a first text for describing a first target object to be recognized; A first extraction module, configured to perform feature extraction on the object description information to obtain object description features; A second extraction module, configured to perform feature extraction on the first image to be processed to obtain first image features; A target object detection module, configured to detect the first target object in the first image to be processed based on the object description features and the first image features to obtain a detection result of the first image to be processed; the detection result indicates whether the first target object is detected or not; in the case where the detection result indicates that the first target object is detected, the detection result includes the position information of the first target object; A third acquisition module, configured to acquire a second voice; A first determination module, configured to determine the second voice as a correction voice in the case where the second voice includes a preset second keyword; the correction voice is used to correct the size and orientation of the position information of the first target object in the detection result of the first image to be processed; An update module, configured to update the position information of the first target object in the first image to be processed at least based on the correction voice.
21. A training device for an image detection model, characterized in that, Including: A fourth acquisition module, configured to acquire an image sample set and an object description sample set, where the object description sample set is a set of voice samples or a set of text samples for describing an object to be recognized; A first training module, configured to train an image detection model by using the image sample set and the object description sample set to obtain the detection results of each image sample in the image sample set; A second determination module, configured to determine the loss of the image detection model based on the detection results of each image sample; An adjustment module, configured to adjust the network parameters of the image detection model by using the loss so that the loss of the detection results output by the adjusted image detection model satisfies a convergence condition; The image detection model includes a correction module, and the detection result represents whether the object to be recognized is detected or not. When the detection result represents that the object to be recognized is detected, the detection result includes the position information of the object to be recognized. A third determination module determines the deviation between the position information of the object to be recognized in the detection results of the respective image samples and the ground truth. A fifth acquisition module is configured to acquire a correction voice when the deviation exceeds a preset range; the correction voice is used to correct the size and orientation of the position information of the object to be recognized in the respective image samples. A second training module is configured to train the correction module at least using the correction voice, and update the position information of the object to be recognized in the respective image samples.
22. An electronic device, comprising a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, the steps in the method according to any one of claims 1 to 19 are implemented.
23. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 19 are implemented.
Citation Information
Patent Citations
Target retrieval method and device
CN111782921A
Image segmentation method and device, computer equipment and storage medium
CN112818955A
Target detection method and device
CN113837257A
Apparatus for generating annotated image information using multimodal input data, apparatus for training an artificial intelligence model using annotated image information, and methods thereof
US20210326643A1