Training method of image detection model, and image detection method and device
The neural network model determines the parameters and prediction similar information of the sample sub-image, which solves the problem of reconstructing the model when adding new candidate object categories in the prior art, and realizes efficient image detection and expanded application scenarios.
Patent Information
- Application Number
- CN202311581725.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
When adding new candidate object categories in the prior art, it is necessary to reconstruct neural network models and train them, resulting in the impact of image detection efficiency.
The parameters of the sample sub-image are determined through the neural network model, and the predicted similar information is determined based on the sample category text and the parameters of the sample sub-image, so as to realize image detection of objects of any category, avoiding the need for model iteration.
It improves image detection efficiency, expands application scenarios, and can detect objects in other categories other than limited categories without model iteration, saving iteration time.
Smart Images

Figure CN120032197A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a training method for an image detection model, an image detection method and a device. Background Art
[0002] With the continuous development of artificial intelligence technology, image detection technology is applied in many industries such as autonomous driving, medicine, and industry. Image detection technology is a technology that detects a given image to determine whether a specific object exists in the image.
[0003] In the related art, image detection can be implemented through an image detection model. Generally, a given image can be framed by an image detection model to obtain multiple image regions. Then, each image region is classified by the image detection model to obtain the probability that each image region belongs to multiple candidate object categories. After that, for any image region, the candidate object category corresponding to the largest probability is used as the object category to which the image region belongs. If there is an image region of a specific object category, it is determined that a specific object exists in the image.
[0004] However, when a new candidate object category needs to be added, the above technology needs to reconstruct the neural network model and train the neural network model to obtain an image detection model. Since model training requires a certain amount of iteration time, the efficiency of image detection is affected. Summary of the invention
[0005] The present application provides a training method, an image detection method and a device for an image detection model, which can be trained to obtain an image detection model for image detection of objects of any category, thereby improving image detection efficiency. The technical solution includes the following contents.
[0006] In a first aspect, a method for training an image detection model is provided, the method comprising:
[0007] Acquire a sample image and a sample category text, wherein the sample image includes at least one sample object, and the sample category text is text about the category to which a specified object in the at least one sample object belongs;
[0008] Determining parameters of at least one sample sub-image based on the sample image by a neural network model, wherein the sample sub-image is an image region in the sample image that contains the sample object;
[0009] For any sample sub-image, determining predicted similarity information of the any sample sub-image based on the sample category text and the parameters of the any sample sub-image by the neural network model, wherein the predicted similarity information is used to characterize the degree of similarity between the category to which the sample object contained in the any sample sub-image belongs and the category to which the specified object belongs;
[0010] Based on the predicted similarity information of each sample sub-image, the neural network model is trained to obtain an image detection model, and the image detection model is used to detect whether there is an object of the target category in the reference image.
[0011] In a second aspect, an image detection method is provided, the method comprising:
[0012] Acquire a reference image and a target category text, wherein the reference image includes at least one reference object;
[0013] Determine, by means of an image detection model, parameters of at least one reference sub-image based on the reference image, wherein the reference sub-image is an image region in the reference image that contains the reference object, and the image detection model is trained according to the method shown in the first aspect;
[0014] For any reference sub-image, determining target similarity information of any reference sub-image based on the target category text and parameters of any reference sub-image by the image detection model, wherein the target similarity information is used to characterize the degree of similarity between the category to which the reference object included in any reference sub-image belongs and the target category involved in the target category text;
[0015] If there is target similarity information satisfying the similarity information condition, it is determined that there is an object of the target category in the reference image.
[0016] In a third aspect, a training device for an image detection model is provided, the device comprising:
[0017] An acquisition module, used for acquiring a sample image and a sample category text, wherein the sample image includes at least one sample object, and the sample category text is a text about a category to which a specified object in the at least one sample object belongs;
[0018] A determination module, configured to determine parameters of at least one sample sub-image based on the sample image through a neural network model, wherein the sample sub-image is an image region in the sample image that contains the sample object;
[0019] The determination module is further used to determine, for any sample sub-image, predicted similarity information of the any sample sub-image based on the sample category text and the parameters of the any sample sub-image through the neural network model, wherein the predicted similarity information is used to characterize the similarity between the category to which the sample object contained in the any sample sub-image belongs and the category to which the specified object belongs;
[0020] The training module is used to train the neural network model based on the predicted similarity information of each sample sub-image to obtain an image detection model, and the image detection model is used to detect whether there is an object of the target category in the reference image.
[0021] In a possible implementation, the determination module is used to extract features of the sample image through the neural network model to obtain features of the sample image; based on the features of the sample image, determine parameters and indicators of multiple sample candidate boxes, the indicators of the sample candidate boxes are used to characterize the possibility that the sample object is included in the selection area of the sample candidate box in the sample image; for any sample candidate box, if the indicators of any sample candidate box meet the indicator conditions, determine the parameters of any sample candidate box as parameters of a sample sub-image.
[0022] In a possible implementation, the determination module is used to crop the sample image based on the parameters of any one of the sample sub-images through the neural network model to obtain any one of the sample sub-images; and determine the predicted similarity information based on the sample category text and any one of the sample sub-images through the neural network model.
[0023] In a possible implementation, the determination module is used to perform feature extraction on the sample category text through the neural network model to obtain sample category features; perform feature extraction on any one of the sample sub-images through the neural network model to obtain a first feature of any one of the sample sub-images; and determine the predicted similarity information based on the sample category features and the first features of any one of the sample sub-images.
[0024] In a possible implementation, the determination module is used to crop the features of the sample image based on the parameters of any one of the sample sub-images through the neural network model to obtain the second feature of any one of the sample sub-images; and determine the predicted similarity information based on the sample category features and the second feature of any one of the sample sub-images through the neural network model.
[0025] In a possible implementation, the training module is used to obtain the labeled category information corresponding to each sample sub-image, and the labeled category information corresponding to the sample sub-image is used to characterize whether the sample object contained in the sample sub-image belongs to the same category as the specified object; determine the first loss between the predicted similarity information and the labeled category information corresponding to each sample sub-image; based on the first loss, train the neural network model to obtain an image detection model.
[0026] In a possible implementation, the training module is used to obtain annotation parameters, where the annotation parameters are used to characterize the size and position of the image area where the specified object is located in the sample image; for any sample sub-image, an intersection-and-union ratio is determined based on the annotation parameters and the parameters of any sample sub-image, and the annotation category information of any sample sub-image is determined based on the intersection-and-union ratio.
[0027] In a possible implementation manner, the predicted similarity information includes first similarity information and second similarity information determined in different ways;
[0028] The training module is used to determine the first loss based on the first sub-loss and the second sub-loss;
[0029] The first sub-loss is the loss between the first similarity information and the labeled category information corresponding to each sample sub-image, and the second sub-loss is the loss between the second similarity information and the labeled category information corresponding to each sample sub-image.
[0030] In a possible implementation, the training module is used to determine a second loss between a first feature and a second feature of each sample sub-image; based on the first loss and the second loss, the neural network model is trained to obtain an image detection model.
[0031] In a possible implementation, the training module is used to determine a third loss between the annotation parameters and the parameters of each sample sub-image; based on the first loss and the third loss, the neural network model is trained to obtain an image detection model.
[0032] In a fourth aspect, an image detection device is provided, the device comprising:
[0033] An acquisition module, used to acquire a reference image and a target category text, wherein the reference image includes at least one reference object;
[0034] a determination module, configured to determine parameters of at least one reference sub-image based on the reference image through an image detection model, wherein the reference sub-image is an image region in the reference image that contains the reference object, and the image detection model is trained according to the method shown in the first aspect;
[0035] The determination module is further configured to determine, for any reference sub-image, target similarity information of the reference sub-image based on the target category text and parameters of the reference sub-image through the image detection model, wherein the target similarity information is used to characterize the degree of similarity between the category to which the reference object included in the reference sub-image belongs and the target category involved in the target category text;
[0036] The determination module is further configured to determine whether an object of the target category exists in the reference image if target similarity information that satisfies the similarity information condition exists.
[0037] In a possible implementation, the determination module is used to perform feature extraction on the reference image through the image detection model to obtain the features of the reference image; based on the features of the reference image, determine parameters and indicators of multiple reference candidate boxes, the indicators of the reference candidate boxes are used to characterize the possibility that the reference object is included in the selection area of the reference candidate box in the reference image; for any reference candidate box, if the indicators of any reference candidate box meet the indicator conditions, determine the parameters of any reference candidate box as the parameters of a reference sub-image.
[0038] In a possible implementation, the determination module is used to crop the reference image based on the parameters of any one of the reference sub-images through the image detection model to obtain any one of the reference sub-images; and determine the target similarity information based on the target category text and any one of the reference sub-images through the image detection model.
[0039] In one possible implementation, the determination module is used to crop the features of the reference image based on the parameters of any one of the reference sub-images through the image detection model to obtain the second features of any one of the reference sub-images; and determine the target similarity information based on the target category features and the second features of any one of the reference sub-images through the image detection model.
[0040] In a fifth aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the electronic device implements the method shown in the first aspect or the second aspect above.
[0041] In the sixth aspect, a computer-readable storage medium is also provided, wherein at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to enable the electronic device to implement the method shown in the first aspect or the second aspect above.
[0042] In the seventh aspect, a computer program is also provided, wherein the computer program is at least one, and the at least one computer program is loaded and executed by a processor to enable the electronic device to implement the method shown in the first aspect or the second aspect above.
[0043] In an eighth aspect, a computer program product is also provided, wherein at least one computer program is stored in the computer program product, and the at least one computer program is loaded and executed by a processor to enable an electronic device to implement the method shown in the first aspect or the second aspect above.
[0044] The technical solution provided by this application brings at least the following beneficial effects:
[0045] In the technical solution provided by the present application, the parameters of the sample sub-image are determined by a neural network model, and the predicted similarity information is determined based on the sample category text and the parameters of the sample sub-image, so as to achieve the determination of the similarity between the category to which the sample object contained in the sample sub-image belongs and the category to which the specified object corresponding to the sample category text belongs. When the neural network model is subsequently trained based on the predicted similarity information to obtain the image detection model, the image detection model and the neural network model have the same functions. In other words, the image detection model can determine the similarity between the category to which the object contained in the image belongs and the category to which the object corresponding to the category text belongs, thereby determining whether the image contains the object corresponding to the category text, and achieving object detection of any category on the image. Among them, the sample category text corresponds to a limited number of categories, and the image detection model can detect objects of any category on the image, which expands the application scenarios, and even if the image is detected for objects of other categories outside the limited categories, there is no need to iterate the model, which saves iteration time and improves image detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 It is a schematic diagram of an implementation environment of a training method for an image detection model or an image detection method provided in an embodiment of the present application;
[0048] Figure 2It is a flowchart of a training method of an image detection model provided in an embodiment of the present application;
[0049] Figure 3 is a schematic diagram for determining first similar information provided in an embodiment of the present application;
[0050] Figure 4 is a schematic diagram for determining a second loss provided in an embodiment of the present application;
[0051] Figure 5 is a flow chart of an image detection method provided in an embodiment of the present application;
[0052] Figure 6 is a schematic diagram of object detection provided by an embodiment of the present application;
[0053] Figure 7 It is a training framework diagram of an image detection model provided in an embodiment of the present application;
[0054] Figure 8 It is a structural schematic diagram of a training device for an image detection model provided in an embodiment of the present application;
[0055] Fig. 9 is a structural schematic diagram of an image detection device provided in an embodiment of the present application;
[0056] Fig.10 It is a structural diagram of a terminal device provided in an embodiment of the present application;
[0057] Fig.11 It is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0058] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0059] Image detection technology is a technology that detects a given image to determine whether a specific object exists in the image. It is widely used in many industries such as autonomous driving, medicine, and industry. In some cases, image detection can be achieved through image detection models.
[0060] In the related art, a given image can be framed by an image detection model to obtain multiple image regions. Then, each image region is classified by the image detection model to obtain the probability that each image region belongs to multiple candidate object categories. Afterwards, for any image region, the candidate object category corresponding to the largest probability is used as the object category to which the image region belongs. If a specific object category exists in the object category to which each image region belongs, it is determined that a specific object exists in the image, and the image region of the specific object category is output. Among them, the image region of the specific object category is the image region where the specific object is located.
[0061] For scenes where objects change frequently, such as recommendation scenes and shopping scenes, it is necessary to continuously add candidate object categories. For the above technology, it is necessary to reconstruct the neural network model and train the neural network model to obtain the image detection model. Since model training requires a certain amount of iteration time, the image detection efficiency is affected.
[0062] Based on the above reasons, the embodiments of the present application provide an image detection model training method and an image detection method, which can be used to train an image detection model and implement image detection of any category through the image detection model, thereby saving the iteration time required for model training and improving image detection efficiency.
[0063] Figure 1 is a schematic diagram of an implementation environment of an image detection model training method or an image detection method provided in an embodiment of the present application, such as Figure 1 As shown, the implementation environment includes a terminal device 101 and a server 102. The training method of the image detection model or the image detection method in the embodiment of the present application can be executed by the terminal device 101, or by the server 102, or by the terminal device 101 and the server 102 together.
[0064] The terminal device 101 may be a smart phone, a game console, a desktop computer, a tablet computer, a laptop computer, a smart TV, a smart car device, a smart voice interaction device, a smart home appliance, etc.
[0065] The server 102 may be a single server, or a server cluster consisting of multiple servers, or any one of a cloud computing platform and a virtualization center, which is not limited in the embodiments of the present application. The server 102 may be connected to the terminal device 101 via a communication network, which may be a wired network or a wireless network. The server 102 may have functions such as data processing, data storage, and data transmission and reception, which are not limited in the embodiments of the present application. The number of the terminal device 101 and the server 102 is not limited, and may be one or more.
[0066] In an exemplary embodiment, the terminal device 101 or the server 102 pre-trains an image detection model. When a reference image and target category text are obtained, the image detection model determines the sub-image where the target category object is located based on the reference image and the target category text.
[0067] In an exemplary embodiment, the server 102 pre-trains an image detection model. When the terminal device 101 obtains a reference image and a target category text, the reference image and the target category text are sent to the server 102 via a communication network. After receiving the reference image and the target category text, the server 102 determines the sub-image where the target category object is located based on the reference image and the target category text through the image detection model, and sends the sub-image where the target category object is located to the terminal device 101 via the communication network, so that the terminal device 101 displays the sub-image where the target category object is located. Optionally, the aforementioned process includes steps S1 to S8 as shown below.
[0068] Step S1, extracting features from a sample image to obtain features of the sample image, and determining parameters of a sample sub-image based on the features of the sample image.
[0069] Step S2: cropping the sample image based on the parameters of the sample sub-image to obtain a sample sub-image.
[0070] Step S3: extracting features from the sample sub-image to obtain a first feature of the sample sub-image.
[0071] Step S4: determining first similarity information based on the sample category text and the first feature of the sample sub-image.
[0072] Step S5: cropping the features of the sample image based on the parameters of the sample sub-image to obtain a second feature of the sample sub-image.
[0073] Step S6, determining second similarity information based on the sample category text and the second feature of the sample sub-image.
[0074] Step S7, training an image detection model based on the first similarity information and the second similarity information. It should be noted that the implementation of steps S1 to S7 can be found in the following relevant Figure 2 The description of is not repeated here.
[0075] Step S8, using the image detection model to determine the sub-image where the object of the target category is located based on the reference image and the target category text. It should be noted that the implementation content of step S8 can be seen below. Figure 5 The description of is not repeated here.
[0076] The optional embodiments of the present application can be implemented based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0077] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0078] Computer vision (CV) technology is a science that studies how to make machines "see". To put it more specifically, computer vision technology refers to the use of cameras and computers to replace human eyes to identify, detect and measure targets, and other machine vision. Through further graphics processing, the computer processing becomes an image that is more suitable for human eyes to see or transmit to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the field of vision, such as swing transformers (Swin-Transformer), vision transformers (ViT), vision MoE (VisionMoE, V-MoE), and masked auto-encoders (MAE), can be quickly and widely applied to downstream specific tasks after fine tuning (Fine Tune). Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, map construction and other technologies.
[0079] See also Figure 2 , Figure 2 Figure 2 is a method for training an image detection model provided by an embodiment of the present application. This method can be applied to the above-mentioned implementation environment, and an image detection model can be trained. This image detection model is suitable for performing image detection of any category. For ease of description, the terminal device 101 or the server 102 that executes the method shown in the embodiment of the present application can be referred to as an electronic device. That is to say, the method shown in the embodiment of the present application can be executed by an electronic device. As Figure 2 shown, the method includes the following steps.
[0080] Step 201, obtain a sample image and a sample category text. The sample image includes at least one sample object, and the sample category text is text related to the category to which a specified object among the at least one sample objects belongs.
[0081] In the embodiment of the present application, the electronic device can obtain multiple sample images. The embodiment of the present application does not limit the acquisition method of any sample image. Exemplarily, the electronic device can obtain the input sample image, or the electronic device can read the sample image from the Internet, or the electronic device can collect the sample image, or the electronic device can read the sample image from other devices, and so on.
[0082] The sample image includes at least one sample object. Any sample object can be an actual object, for example, living things such as cats and dogs; buildings such as houses and bridges; vehicles; trees; etc. Any sample object can also be a fictional object, for example, dragons, monsters, etc. In the embodiment of the present application, at least one of the sample objects included in the sample image includes a specified object, and there is at least one specified object. For example, the sample image includes two sample objects, a cat and a dog, where the dog is the specified object.
[0083] In the embodiment of the present application, the electronic device can obtain the sample category text. The embodiment of the present application does not limit the acquisition method of the sample category text. Exemplarily, the electronic device can obtain the input sample category text, or the electronic device can read the sample category text from the Internet, or the electronic device can read the sample category text from other devices, or the electronic device can generate the sample category text based on the category to which the specified object belongs, and so on. The embodiment of the present application does not limit the generation method of the sample category text. Exemplarily, the electronic device includes a candidate category template, and combines the category to which the specified object belongs with the candidate category template to obtain the sample category text. For example, the candidate category template is "This is an image of XX". When the specified object is a dog, the sample category text is "This is an image of a dog". Or, the electronic device can generate the sample category text based on the category to which the specified object belongs through a large language model. The generation method will not be elaborated here.
[0084] In an exemplary embodiment, the electronic device can obtain one or several object detectors that have been applied in actual scenes. The object detector can implement fixed-category object detection on the image and obtain the image area where the fixed-category object is located. Among them, the detection principles of each object detector are similar. Simply put, the object detector can regard the image area where the object is located in the image as the foreground area, and regard the image area where the non-object is located in the image as the background area, based on which, the foreground area is detected from the image. Then, the object detector classifies the foreground area. Generally, a normalized exponential function (for example, a Softmax function) is used for classification to obtain the probability that the object in the foreground area belongs to each candidate category, and the candidate category corresponding to the maximum probability is used as the category to which the object in the foreground area belongs. Optionally, the candidate category is a pre-set category, and each candidate category may or may not include a background category.
[0085] It is understandable that the object detector is a trained model. The training data set of the object detector includes at least one training data, each training data includes a training image and a label obtained by manually annotating the training image, and the label includes the annotation category and the annotation parameters of the image area where the object of the annotated category in the training image is located. For example, the training data includes the training image, the annotation category "dog", and the size parameters and position parameters of the image area where the "dog" is located. It is understandable that the number of annotated categories is limited and is confined to a fixed set. In general, the training image can be input into the initial model, and the prediction parameters of the foreground area in the training image are determined by the initial model, such as the size parameters and position parameters of the foreground area, and the prediction category to which the object in the foreground area belongs is determined. Afterwards, based on the prediction category, the annotation category, the prediction parameters of the foreground area, and the annotation parameters of the image area where the object of the annotated category is located, the initial model is trained to obtain the object detector.
[0086] In exemplary embodiment A1, the electronic device may obtain a training data set of an object detector, use the training image as a sample image, use the annotated category as the category to which the specified object belongs, and generate sample category text based on the annotated category.
[0087] In exemplary embodiment A2, the electronic device may acquire a sample image, input the sample image into an object detector, detect a foreground area from the sample image through the object detector, and determine the category to which the object in the foreground area belongs. The category to which the object in the foreground area belongs may be used as the category to which the specified object belongs, and a sample category text may be generated based on the category to which the object in the foreground area belongs.
[0088] It can be understood that there are multiple sample images, each of which corresponds to at least one sample category text. A sample image and each sample category text corresponding to the sample image can be used as a sample data set. Generally speaking, the number of sample data sets is greater than a set number. The set number can be determined based on manual experience or randomly. For example, the set number is 200,000, that is, the number of sample data sets is greater than 200,000. The sample data set can be obtained based on Embodiment A1 or A2. When the number of sample data sets is not greater than the set number, the sample data set can also be obtained based on another embodiment in Embodiments A1 and A2. It is even possible to manually determine the category to which a specified object belongs from a sample image, and generate a sample category text based on the category to which the specified object belongs.
[0089] Step 202: Determine parameters of at least one sample sub-image based on the sample image through a neural network model, where the sample sub-image is an image region in the sample image that contains a sample object.
[0090] In the embodiment of the present application, the sample image includes at least one sample object, and the image region where the sample object is located in the sample image may be referred to as a sample sub-image. For example, if the sample image includes a cat and a dog, then the image region where the cat is located in the sample image is a sample sub-image, and the image region where the dog is located in the sample image is also a sample sub-image. It can be understood that the sample sub-image includes at least one sample object.
[0091] The sample image can be input into the neural network model, and the parameters of at least one sample sub-image can be determined by the neural network model. The embodiments of the present application do not limit the parameters of the sample sub-image. Exemplarily, the parameters of the sample sub-image can include the size parameters of the sample sub-image. For example, the size parameters of the sample sub-image include at least one item such as the area of the sample sub-image, the side length of the sample sub-image, the perimeter of the sample sub-image, and the vertex coordinates of the sample sub-image. The size of the sample sub-image is characterized by the size parameters of the sample sub-image. For another example, the parameters of the sample sub-image can also include the position parameters of the sample sub-image. For example, the position parameters of the sample sub-image include the vertex coordinates of the sample sub-image, the centroid coordinates of the sample sub-image, the center coordinates of the sample sub-image, etc. The position parameters of the sample sub-image are used to characterize the position of the sample sub-image relative to the sample image.
[0092] The embodiments of the present application do not limit the model structure, model parameters, etc. of the neural network model. Exemplarily, the neural network model includes at least one module including a parameter determination network, a first encoder, a second encoder, etc. The structure and function of each module are described below and will not be repeated here. It is understandable that the structure of the neural network model is different, and the method of determining the parameters of the sample sub-image through the neural network model is also different.
[0093] Optionally, the parameter determination network is used to determine the parameters of at least one sample sub-image based on the sample image. The embodiment of the present application does not limit the network structure, network parameters, etc. of the parameter determination network. Exemplarily, the parameter determination network includes a region proposal network (RPN), a region convolutional neural network (R-CNN) based on a selective search algorithm, a fast region convolutional neural network (Fast R-CNN), etc.
[0094] It is understandable that the parameter determination network has different structures and the method of determining the parameters of the sample sub-image also differs. The following takes the parameter determination network including RPN as an example to describe the content of determining the parameters of the sample sub-image through RPN. The method of determining the parameters of the sample sub-image through R-CNN, Fast R-CNN, etc. is not repeated here.
[0095] In an exemplary embodiment, step 202 includes steps 2021 to 2023 (not shown in the figure).
[0096] Step 2021, extract features of the sample image through a neural network model to obtain features of the sample image.
[0097] In an exemplary embodiment, the RPN includes a convolutional neural network, and the convolutional neural network includes at least one convolutional layer, any of which may be any of a standard convolutional layer, a hole convolutional layer, a separable convolutional layer, a grouped convolutional layer, a deconvolutional layer, etc. The sample image is convolved at least once by the convolutional neural network to obtain the features of the sample image, and the pixel value of each pixel point in the sample image is fused with the pixel values of the surrounding pixels to obtain the feature value of a feature point in the features of the sample image, so that the feature value of a feature point can represent the pixel value of a pixel block composed of multiple pixels in the sample image, that is, the feature value of the feature point has a strong representation ability.
[0098] In another exemplary implementation, the RPN may include a Vision Encoder. Optionally, the Vision Encoder includes at least one Transformer layer, and the Transformer layers are connected in series. That is, the input of the first Transformer layer includes the sample image, the input of the next Transformer layer includes the output of the previous Transformer layer, and the output of the last Transformer layer includes the features of the sample image. The structures, functions, etc. of each Transformer layer are the same. Hereinafter, an example in which the Vision Encoder includes one Transformer layer will be described.
[0099] Optionally, the sample image is evenly divided into multiple sample image patches, the pixel values of each sample image patch are linearly mapped to obtain the input vector of each sample image patch. The input vectors of each sample image patch are input into the Transformer layer. The Transformer layer includes a first normalization layer, a multi-head attention layer, a second normalization layer, and a multi-layer perceptron. First, the input vectors of each sample image patch are normalized by the first normalization layer to obtain the first processing result of each sample image patch. Then, the first processing results of each sample image patch are subjected to attention processing based on the attention mechanism by the multi-head attention layer to obtain the attention processing results of each sample image patch. Next, the input vectors and the attention processing results of each sample image patch are concatenated to obtain the first concatenation result of each sample image patch, and the first concatenation result of each sample image patch is normalized by the second normalization layer to obtain the second processing result of each sample image patch. After that, the second processing results of each sample image patch are subjected to at least one of linear mapping and non-linear mapping by the multi-layer perceptron to obtain the mapping results of each sample image patch. Then, the first concatenation result and the mapping results of each sample image patch are concatenated to obtain the second concatenation result of each sample image patch. The second concatenation results of each sample image patch are concatenated or fused to obtain the features of the sample image.
[0100] Alternatively, a set input vector may be concatenated before the input vectors of each sample image patch. For example, the set input vector is a zero matrix or a one matrix or an identity matrix, etc. The set input vector can be processed in a similar manner as the input vectors of the sample image patches described above to obtain a target output vector. The target output vector can be used as the features of the sample image.
[0101] It should be noted that the above only describes two possible feature extraction methods. When applied, there may be other feature extraction methods, which will not be described here. Among them, the features of the sample image can represent the image semantics of the sample image. As the name implies, image semantics is to use language to describe the visual content of the image. Since the sample image includes at least one sample object, the features of the sample image can reflect the position, color, texture, appearance, posture and other information of each sample object, the relative position of each sample object, the color, texture, appearance and other information of the environment in which the sample object is located, and so on. For example, the features of the sample image can reflect the position, color, texture, appearance, posture and other information of the dog in the sample image, the position, color, texture, appearance, posture and other information of the cat in the sample image, the relative position of the cat and the dog, and the color, texture, appearance and other information of the environment in which the cat and the dog are located.
[0102] Step 2022: Determine parameters and indices of multiple sample candidate boxes based on the features of the sample image. The indices of the sample candidate boxes are used to characterize the possibility that the sample object is included in the box selection area of the sample candidate box in the sample image.
[0103] In the embodiment of the present application, the features of the sample image may be convolved to obtain image features of a set dimension. For example, if the features of the sample image are N×16×16, then the features of the sample image may be convolved to obtain image features of 256×16×16.
[0104] On the one hand, by performing convolution processing on the image features of the set dimension, a first convolution result is obtained, and the first convolution result is the parameter of multiple sample candidate frames. For example, the image features of 256×16×16 are convolved to obtain a first convolution result of 36×16×16 (i.e., 16×16×9×4). Optionally, in the first convolution result, 16×16 represents the number of feature points, each feature point corresponds to 9 sample candidate frames, each sample candidate frame corresponds to 4 coordinates, and the size and position of the sample candidate frame are represented by the 4 coordinates. For example, the 4 coordinates corresponding to a sample candidate frame are: the horizontal coordinate of the center point of the sample candidate frame; the vertical coordinate of the center point of the sample candidate frame; the length of the sample candidate frame; the height of the sample candidate frame.
[0105] On the other hand, by performing convolution processing on the image features of the set dimension, a second convolution result is obtained, and the second convolution result is an index of multiple sample candidate frames. For example, the image features of 256×16×16 are convolved to obtain a second convolution result of 18×16×16 (i.e., 16×16×9×2). Optionally, in the second convolution result, 16×16 represents the number of feature points, each feature point corresponds to 9 sample candidate frames, each sample candidate frame corresponds to 2 scores, and the index of the sample candidate frame can be determined based on at least one of the 2 scores. Among them, the two scores corresponding to a sample candidate frame respectively represent: the possibility that the frame selection area of the sample candidate frame in the sample image belongs to the foreground area; the possibility that the frame selection area of the sample candidate frame in the sample image belongs to the background area. As described above, the foreground area includes the sample object, and the background area does not include the sample object. Therefore, the index of the sample candidate frame can represent the possibility that the frame selection area of the sample candidate frame in the sample image includes the sample object.
[0106] Step 2023: for any sample candidate box, if the index of any sample candidate box meets the index condition, the parameters of any sample candidate box are determined as the parameters of a sample sub-image.
[0107] In the embodiment of the present application, the index of any sample candidate box may be a foreground score, which represents the possibility that the selected area of the sample candidate box in the sample image belongs to the foreground area, and the higher the foreground score, the higher the possibility. If the foreground score is greater than the first score threshold, the index of the sample candidate box meets the index condition, and the parameters of the sample candidate box can be determined as the parameters of a sample sub-image.
[0108] Alternatively, the index of any sample candidate box may be a background score, which represents the possibility that the selected area of the sample candidate box in the sample image belongs to the background area, and the higher the background score, the higher the possibility. If the background score is less than the second score threshold, the index of the sample candidate box meets the index condition, and the parameters of the sample candidate box can be determined as the parameters of a sample sub-image. The first score threshold is greater than, equal to, or less than the second score threshold.
[0109] Alternatively, the index of any sample candidate frame may include a foreground score and a background score. If the foreground score is greater than a first score threshold and the background score is less than a second score threshold, the index of the sample candidate frame satisfies the index condition, and the parameters of the sample candidate frame may be determined as parameters of a sample sub-image.
[0110] Step 203, for any sample sub-image, the predicted similarity information of any sample sub-image is determined based on the sample category text and the parameters of any sample sub-image through a neural network model, and the predicted similarity information is used to characterize the similarity between the category to which the sample object contained in any sample sub-image belongs and the category to which the specified object belongs.
[0111] In the embodiment of the present application, on the one hand, the sample category text is related to the category to which the specified object belongs, that is, the sample category text can reflect the category to which the specified object belongs. On the other hand, the sample sub-image contains the sample object and can reflect the color, texture, appearance and other information of the sample object, that is, the sample sub-image can reflect the category to which the sample object belongs, and the sample sub-image can be located based on the parameters of the sample sub-image. According to the above content, the neural network model can determine the predicted similarity information of the sample sub-image based on the sample category text and the parameters of the sample sub-image, and characterize the similarity between the category to which the sample object belongs and the category to which the specified object belongs through the predicted similarity information of the sample sub-image.
[0112] Optionally, the predicted similarity information is a probability value, and the larger the probability value, the higher the similarity between the category to which the sample object belongs and the category to which the specified object belongs. Alternatively, the predicted similarity information is a positive number, and the larger the value, the higher the similarity. Alternatively, the predicted similarity information is a matrix, and the larger the norm or rank of the matrix, the higher the similarity.
[0113] It is understandable that the structures of the neural network models are different, and the methods for determining the predicted similarity information of the sample sub-images are also different. Two possible implementation methods are described below by way of example, see implementation method B1 and implementation method B2 respectively.
[0114] First, implementation B1 is described. In implementation B1, step 203 includes steps B11 to B12 (not shown in the figure).
[0115] Step B11, cropping the sample image based on the parameters of any sample sub-image through a neural network model to obtain any sample sub-image.
[0116] As mentioned above, the parameters of the sample sub-image include the position parameter and size parameter of the sample sub-image, wherein the size parameter represents the size of the sample sub-image, and the position parameter represents the position of the sample sub-image in the sample image. Based on this, an image region of the size represented by the size parameter can be cropped at the position represented by the position parameter in the sample image through a neural network model, and the image region is the sample sub-image.
[0117] Since the parameters of the sample sub-image represent the size of the sample sub-image and the position of the sample sub-image in the sample image, when the sample image is cropped based on the parameters of the sample sub-image, the cropping error is small and the accuracy of the sample sub-image is high. Further, when the predicted similarity information is subsequently determined based on the sample sub-image and the image detection model is trained based on the predicted similarity information, the accuracy of the image detection model is also high due to the high accuracy of the sample sub-image, thereby improving the accuracy of the image detection result.
[0118] Step B12, determining predicted similarity information based on the sample category text and any sample sub-image through a neural network model.
[0119] In an embodiment of the present application, the neural network model includes a text-image multimodal network. The text-image multimodal network is a network model implemented based on text-image multimodal technology, which is used to align the semantics of text and images. A variety of downstream tasks can be implemented based on the text-image multimodal network, such as visual question answering (VQA) tasks, image retrieval tasks, image classification tasks, etc. Among them, in an embodiment of the present application, the image classification task is mainly implemented based on the text-image multimodal network.
[0120] The embodiment of the present application does not limit the network structure, network parameters, etc. of the image-text multimodal network. Exemplarily, the image-text multimodal network is a contrastive language-image pretraining (CLIP) network, a ViT network, a vision-language matching (VLM) network, a vision-language contrastive learning (VLC) network, etc.
[0121] Optionally, the sample category text and any sample sub-image are input into a text-image multimodal network, and the predicted similarity information of the sample sub-image is determined based on the text-image semantic alignment of the sample category text and the sample sub-image through the text-image multimodal network. The predicted similarity information is used to characterize the degree of similarity between the category to which the sample object contained in the sample sub-image belongs and the category to which the specified object belongs.
[0122] Optionally, the predicted similarity information is a probability value, and the larger the probability value, the higher the similarity between the category to which the sample object belongs and the category to which the specified object belongs. Alternatively, the predicted similarity information is a positive number, and the larger the value, the higher the similarity. Alternatively, the predicted similarity information is a matrix, and the larger the norm or rank of the matrix, the higher the similarity.
[0123] In an exemplary embodiment, step B12 includes steps B121 to B123 (not shown in the figure).
[0124] Step B121, extracting features of the sample category text through a neural network model to obtain sample category features.
[0125] In the embodiment of the present application, the graphic-text multimodal network includes a first encoder, the sample category text is input into the first encoder, and the first encoder performs feature extraction on the sample category text to obtain the sample category feature. The embodiment of the present application does not limit the model structure, model parameters, etc. of the first encoder. It is understandable that the structure of the first encoder is different, and the feature extraction method is also different.
[0126] Exemplarily, the first encoder is a text encoder (Text Encoder). Optionally, the text encoder includes at least one transformer (Transformer) layer, and each transformer layer is connected in series. That is, the input of the first transformer layer includes sample category text, the input of the next transformer layer includes the output of the previous transformer layer, and the output of the last transformer layer includes sample category features. The structure and function of each transformer layer are the same, and the following is explained by taking the example of a text encoder including one transformer layer.
[0127] Optionally, the sample category text includes multiple sample words, and each sample word is mapped to a corresponding word vector to obtain a word vector of each sample word. The word vector of each sample word is input into the Transformer layer, and the Transformer layer includes a multi-head attention layer, a first normalization layer, a forward feedback layer, and a second normalization layer. First, the word vector of each sample word is subjected to attention processing based on the attention mechanism by the multi-head attention layer to obtain the attention processing result of each sample word. Then, the attention processing result of each sample word is normalized by the first normalization layer to obtain the first processing result of each sample word. Next, the word vector of each sample word and the first processing result are spliced to obtain the first splicing result of each sample word, and the first splicing result of each sample word is fed forward by the forward feedback layer to obtain the forward feedback result of each sample word. After that, the forward feedback result of each sample word is normalized by the second normalization layer to obtain the second processing result of each sample word. The first splicing result and the second processing result of each sample word are spliced to obtain the second splicing result of each sample word. The second concatenation results of each sample word are concatenated or fused to obtain the sample category feature.
[0128] Alternatively, the word vector of the start character (for example, the start character is a CLS character) may be concatenated before the word vector of each sample word. The word vector of the start character may be processed similarly according to the processing process of the word vector of the sample word to obtain a second concatenation result of the start character. The second concatenation result of the start character may be used as a sample category feature.
[0129] For another example, the first encoder is a residual network (Residual Network, Resnet). Optionally, the residual network includes at least one residual block, and each residual block is connected in series. That is, the input of the first residual block includes the sample category text, the input of the next residual block includes the output of the previous residual block, and the output of the last residual block includes the sample category feature. The structure and function of each residual block are the same, and the following is an example of a residual network including a residual block.
[0130] Optionally, after mapping each sample word in the sample category text to a corresponding word vector, the residual block is input, and the residual block includes a plurality of convolution layers of different convolution scales connected in series. The word vectors of each sample word can be convolved at different convolution scales through each convolution layer to obtain the convolution results of each sample word. The word vectors of each sample word and the convolution results are spliced to obtain the splicing results of each sample word. The splicing results of each sample word are spliced or fused to obtain the sample category features.
[0131] It should be noted that the above only describes two possible feature extraction methods. When applied, there may be other feature extraction methods, which will not be described here. Among them, the sample category feature can represent the semantics of the sample category text.
[0132] Step B122, extracting features from any sample sub-image through a neural network model to obtain a first feature of any sample sub-image.
[0133] In the embodiment of the present application, the image-text multimodal network includes a second encoder, the sample sub-image is input into the second encoder, and the second encoder extracts features of the sample sub-image to obtain a first feature of the sample sub-image. The embodiment of the present application does not limit the model structure, model parameters, etc. of the second encoder. It can be understood that the structure of the second encoder is different, and the feature extraction method is also different.
[0134] Optionally, the second encoder is a convolutional neural network or a visual encoder, etc. Among them, step 2021 has described the content of feature extraction of the sample image through the convolutional neural network and the visual encoder. Based on the feature extraction principle described in step 2021, feature extraction of the sample sub-image can be achieved through the convolutional neural network and the visual encoder, and its implementation process will not be repeated.
[0135] The first feature of the sample sub-image can represent the image semantics of the sample sub-image. Since the sample sub-image includes the sample object, the first feature of the sample sub-image can reflect the position, color, texture, appearance, posture and other information of the sample object and the color, texture, appearance and other information of the environment where the sample object is located, etc. For example, the feature of the sample sub-image can reflect the position, color, texture, appearance, posture and other information of the dog in the sample sub-image and the color, texture, appearance and other information of the environment where the dog is located.
[0136] Step B123, determining predicted similarity information based on the sample category feature and the first feature of any sample sub-image.
[0137] In the embodiment of the present application, the first feature distance between the sample category feature and the first feature of the sample sub-image can be calculated. The embodiment of the present application does not limit the calculation method of the first feature distance. For example, any distance formula such as the Euclidean distance formula, the cosine distance formula, the Manhattan distance formula, and the Chebyshev distance formula can be used to calculate the first feature distance.
[0138] Next, the predicted similarity information is determined based on the first feature distance. It can be understood that the smaller the first feature distance is, the more similar the sample category feature is to the first feature of the sample sub-image, indicating that the category to which the specified object belongs is closer to the category to which the sample object contained in the sample sub-image belongs, that is, the higher the similarity between the category to which the sample object belongs represented by the predicted similarity information and the category to which the specified object belongs is.
[0139] In an exemplary implementation, the predicted similarity information determined by implementation B1 is the first similarity information or the second similarity information mentioned below. Figure 3 , Figure 3 It is a schematic diagram for determining first similar information provided in an embodiment of the present application. The first similar information is determined by implementation method B1.
[0140] like Figure 3 As shown, on the one hand, the text encoder extracts features from the sample category text "an image containing a dog" to obtain text features, which are the sample category features mentioned above. On the other hand, a sample sub-image is cropped from the sample image, and the visual encoder extracts features from the sample sub-image to obtain a first visual feature, which is the first feature of the sample sub-image mentioned above. Next, the first similarity information is calculated based on the text features and the first visual features.
[0141] Next, implementation B2 is described. In implementation B2, step 203 includes steps B21 and B22 (not shown in the figure).
[0142] Step B21, cutting the features of the sample image based on the parameters of any sample sub-image through a neural network model to obtain a second feature of any sample sub-image.
[0143] In the embodiment of the present application, the parameters of the sample sub-image include the position parameter and the size parameter of the sample sub-image, wherein the size parameter represents the size of the sample sub-image, and the position parameter represents the position of the sample sub-image in the sample image. In addition, the features of the sample image include the feature values of multiple feature points, and one feature point corresponds to a pixel block composed of multiple pixel points in the sample image. Based on this, the features of the sample image can be cropped based on the correspondence between each feature point in the features of the sample image and each pixel block in the sample image through a neural network model, that is, a feature block of a size represented by a corresponding size parameter is cropped at a position represented by the position parameter in the features of the sample image, and the feature block is the second feature of the sample sub-image.
[0144] Optionally, the features of the sample image may be pooled, or the feature blocks corresponding to the parameters of the sample sub-image in the features of the sample image may be pooled to obtain the features of the sample image after the pooling process. Then, the feature blocks are cropped from the features of the sample image after the pooling process. Alternatively, the feature blocks are first cropped from the features of the sample image, and then the feature blocks are pooled to obtain the second features of the sample sub-image. Through the pooling process, the dimension of the features is reduced, and while reducing the feature quantity, the effective information of the features is retained, thereby improving the operation speed and ensuring the accuracy of the operation results.
[0145] Since there is a correspondence between the feature points in the features of the sample image and the pixel blocks in the sample image, the second feature of the sample sub-image can be directly obtained by cropping the features of the sample image, which improves the feature determination efficiency of the sample sub-image, helps to speed up the reasoning speed of the model, and improves the image detection efficiency.
[0146] Step B22, determining predicted similarity information based on the sample category feature and the second feature of any sample sub-image through a neural network model.
[0147] In the embodiment of the present application, the second feature distance between the sample category feature and the second feature of the sample sub-image can be calculated. The embodiment of the present application does not limit the calculation method of the second feature distance. For example, any distance formula such as the Euclidean distance formula, the cosine distance formula, the Manhattan distance formula, the Chebyshev distance formula, etc. can be used to calculate the second feature distance.
[0148] Next, the predicted similarity information of the sample sub-image is determined based on the second feature distance. It can be understood that the smaller the second feature distance is, the more similar the sample category feature is to the second feature of the sample sub-image, indicating that the category to which the specified object belongs is closer to the category to which the sample object contained in the sample sub-image belongs, that is, the higher the similarity between the category to which the sample object belongs represented by the predicted similarity information and the category to which the specified object belongs.
[0149] In an exemplary implementation, the predicted similarity information determined by implementation B2 is the first similarity information or the second similarity information mentioned below.
[0150] In actual application, the predicted similarity information can be determined based on both implementation method B1 and implementation method B2. It can be understood that the predicted similarity information with the same characterization content but different determination methods can be obtained through implementation method B1 and implementation method B2. For the sake of distinction, the predicted similarity information determined by implementation method B1 can be referred to as the first similarity information mentioned below, and the predicted similarity information determined by implementation method B2 can be referred to as the second similarity information mentioned below. Alternatively, the predicted similarity information determined by implementation method B1 can be referred to as the second similarity information mentioned below, and the predicted similarity information determined by implementation method B2 can be referred to as the first similarity information mentioned below.
[0151] That is, the predicted similarity information may be only the first similarity information or the second similarity information, or the predicted similarity information may include the first similarity information and the second similarity information, or the predicted similarity information may be obtained by any calculation such as summing or averaging the first similarity information and the second similarity information.
[0152] Step 204: Based on the predicted similarity information of each sample sub-image, the neural network model is trained to obtain an image detection model, and the image detection model is used to detect whether there is an object of the target category in the reference image.
[0153] In the embodiment of the present application, the loss of the neural network model can be determined based on the predicted similarity information corresponding to each sample sub-image. The neural network model is trained once by the loss of the neural network model to obtain a trained neural network model.
[0154] If the trained neural network model meets the training end condition, the trained neural network model is used as the image detection model. If the trained neural network model does not meet the training end condition, the trained neural network model is used as the neural network model for the next training, and the neural network model is trained next time according to the implementation content of step 202 to step 204, until the trained neural network model meets the training end condition, and the trained neural network model is used as the image detection model.
[0155] The embodiments of the present application do not limit the content of the trained neural network model satisfying the training end conditions. Exemplarily, the trained neural network model satisfying the training end conditions include: the number of training times corresponding to the trained neural network model reaches a set number of times, or the model error of the trained neural network model is less than the set error, or the model index of the trained neural network model is greater than the set index. Among them, the model index of the trained neural network model is used to characterize the image detection effect of the trained neural network model.
[0156] In a possible implementation, step 204 includes steps 2041 to 2043 (not shown in the figure).
[0157] Step 2041 : obtaining the annotated category information corresponding to each sample sub-image. The annotated category information corresponding to the sample sub-image is used to indicate whether the sample object contained in the sample sub-image and the specified object belong to the same category.
[0158] In the embodiment of the present application, if the sample object included in the sample sub-image belongs to the same category as the specified object, the mark category information corresponding to the sample sub-image is a first value. If the sample object included in the sample sub-image belongs to a different category than the specified object, the mark category information corresponding to the sample sub-image is a second value. The first value and the second value are different values.
[0159] For example, if the sample sub-image includes a cat and the designated object is a dog, the flag category information corresponding to the sample sub-image is 0; if the sample sub-image includes a dog and the designated object is a dog, the flag category information corresponding to the sample sub-image is 1.
[0160] The embodiment of the present application does not limit the method for obtaining the annotated category information corresponding to the sample sub-image. For example, the sample sub-image can be annotated manually so that the electronic device can obtain the annotated category information corresponding to the sample sub-image.
[0161] For example, in an exemplary embodiment, step 2041 includes: obtaining annotation parameters, the annotation parameters are used to characterize the size and position of the image area where the specified object is located in the sample image; for any sample sub-image, determining the intersection-and-union ratio based on the annotation parameters and the parameters of any sample sub-image, and determining the annotation category information of any sample sub-image based on the intersection-and-union ratio.
[0162] As mentioned above, the electronic device can obtain a training data set for the object detector, and the training data set includes labels obtained by manually annotating the training images, and the labels include annotation parameters of the image area where the object of the annotated category in the training image is located. Wherein, when the sample image is a training image, the annotation parameters of the sample image are the annotation parameters of the image area where the object of the annotated category in the training image is located.
[0163] Alternatively, the electronic device inputs the sample image into the object detector, determines the prediction parameters of the foreground area in the sample image through the object detector, and determines the prediction category to which the object in the foreground area belongs. If the prediction category to which the object in the foreground area belongs is the category to which the specified object belongs, the prediction parameters of the foreground area in the sample image are used as the annotation parameters of the sample image.
[0164] Alternatively, the sample image may be annotated manually so that the electronic device can obtain the annotation parameters of the sample image.
[0165] It can be understood that the annotation parameters of the sample image include an annotation size parameter and an annotation position parameter, wherein the annotation size parameter represents the size of the image area where the specified object is located in the sample image, and the annotation position parameter represents the position of the image area where the specified object is located relative to the sample image.
[0166] The sample image includes at least one sample sub-image. For any sample sub-image, the parameters of the sample sub-image include a size parameter and a position parameter of the sample sub-image, wherein the size parameter of the sample sub-image represents the size of the sample sub-image, and the position parameter of the sample sub-image represents the position of the sample sub-image relative to the sample image.
[0167] Since the annotation parameters can reflect the position and size of the image area where the specified object is located, and the parameters of the sample sub-image can reflect the position and size of the sample sub-image, based on the annotation parameters and the parameters of the sample sub-image, the parameters of the intersection area formed by the image area where the specified object is located and the sample sub-image, and the parameters of the union area formed by the image area where the specified object is located and the sample sub-image can be determined.
[0168] The parameters of the intersection region include the position parameter and size parameter of the intersection region, the position parameter of the intersection region represents the position of the intersection region, and the size parameter of the intersection region represents the size of the intersection region. Similarly, the parameters of the union region include the position parameter and size parameter of the union region, the position parameter of the union region represents the position of the union region, and the size parameter of the union region represents the size of the union region.
[0169] The intersection-and-union ratio can be determined based on the parameters of the intersection area and the parameters of the union area, and the intersection-and-union ratio is used to characterize the degree of overlap between the image area where the specified object is located and the sample sub-image. Optionally, the larger the intersection-and-union ratio, the higher the degree of overlap between the image area where the specified object is located and the sample sub-image.
[0170] If the intersection-and-union ratio is greater than the set ratio, it means that the image region where the specified object is located has a high degree of overlap with the sample sub-image, and the labeled category information of the sample sub-image can be determined to be the first value. In this case, the labeled category information of the sample sub-image indicates that the sample object contained in the sample sub-image belongs to the same category as the specified object. On the contrary, if the intersection-and-union ratio is not greater than the set ratio, it means that the image region where the specified object is located has a low degree of overlap with the sample sub-image, and the labeled category information of the sample sub-image can be determined to be the second value. In this case, the labeled category information of the sample sub-image indicates that the sample object contained in the sample sub-image belongs to a different category than the specified object.
[0171] The embodiment of the present application does not limit the method for determining the set ratio. For example, the set ratio is a value set based on manual experience, for example, the set ratio is 0.7. Alternatively, the set ratio is a value obtained through experimental verification, and the verification method is not repeated here.
[0172] Step 2042: determine a first loss between the predicted similarity information and the labeled category information corresponding to each sample sub-image.
[0173] In the embodiment of the present application, for any sample sub-image, the loss between the predicted similarity information and the labeled category information corresponding to the sample sub-image can be determined, and the information error between the predicted similarity information and the labeled category information corresponding to the sample sub-image can be represented by the loss. The losses between the predicted similarity information and the labeled category information corresponding to each sample sub-image are summed, averaged, etc. to obtain a first loss, and the statistical error between the predicted similarity information and the labeled category information corresponding to each sample sub-image is represented by the first loss.
[0174] As described above, the predicted similarity information can be determined based on only one implementation, for example, by implementation B1 or B2. If the predicted similarity information determined by this implementation is the first similarity information mentioned below, the first loss is the first sub-loss mentioned below. If the predicted similarity information determined by this implementation is the second similarity information mentioned below, the first loss is the second sub-loss mentioned below.
[0175] Of course, in actual application, the predicted similarity information can be determined based on different implementations, for example, by implementations B1 and B2. In this case, the predicted similarity information includes the first similarity information and the second similarity information determined by different methods. Step 2042 includes: determining the first loss based on the first sub-loss and the second sub-loss.
[0176] Among them, the first sub-loss is the loss between the first similarity information corresponding to each sample sub-image and the labeled category information, and the second sub-loss is the loss between the second similarity information corresponding to each sample sub-image and the labeled category information.
[0177] In one possible implementation, the first similar information is determined by implementation B1, but this is only one possible implementation, and in actual application, there may be other implementations. For example, the first similar information is determined by implementation B2 or other implementations, and the contents of other implementations are not repeated here.
[0178] The loss between the first similarity information and the labeled category information corresponding to any sample sub-image can be determined, and the information error between the first similarity information and the labeled category information corresponding to the sample sub-image can be represented by the loss. The losses between the first similarity information and the labeled category information corresponding to each sample sub-image are summed, averaged, etc. to obtain a first sub-loss, and the statistical error between the first similarity information and the labeled category information corresponding to each sample sub-image is represented by the first sub-loss. Optionally, the first loss is a first sub-loss.
[0179] Optionally, the first similarity information of each sample sub-image is determined based on the sample category text and the first feature of each sample sub-image by a neural network model, and the first similarity information of each sample sub-image is determined by s R Characterize the first similarity information of each sample sub-image. In addition, the labeled category information of each sample sub-image is characterized by t. Optionally, the first sub-loss is calculated based on the cross entropy loss formula by calculating the first similarity information and labeled category information corresponding to each sample sub-image. The calculation method is shown in the following formula (1).
[0180] L CR1 =CrossEntropy(s R ,t) Formula (1)
[0181] Among them, L CE1 Represents the first sub-loss, and CrossEntropy represents the cross entropy loss formula.
[0182] In a possible implementation manner, the second similarity information is determined by implementation manner B2, but this is only a possible implementation manner. In actual application, there may be other implementation manners.
[0183] The loss between the second similarity information and the labeled category information corresponding to any sample sub-image can be determined, and the information error between the second similarity information and the labeled category information corresponding to the sample sub-image can be represented by the loss. The losses between the second similarity information and the labeled category information corresponding to each sample sub-image are summed, averaged, etc. to obtain a second sub-loss, and the statistical error between the second similarity information and the labeled category information corresponding to each sample sub-image is represented by the second sub-loss. Optionally, the first loss is the second sub-loss.
[0184] Optionally, the second similarity information of each sample sub-image is determined based on the sample category text and the second feature of each sample sub-image by a neural network model, and the second similarity information of each sample sub-image is determined by s ROI Characterize the second similarity information of each sample sub-image. In addition, the labeled category information of each sample sub-image is characterized by t. Optionally, the second sub-loss is calculated based on the cross entropy loss formula by calculating the second similarity information and labeled category information corresponding to each sample sub-image. The calculation method is shown in formula (2) shown below.
[0185] L CE2 =CrossEntropy(s ROI ,t) Formula (2)
[0186] Among them, L CE2 Represents the second sub-loss, and CrossEntropy represents the cross entropy loss formula.
[0187] In a possible implementation, the first sub-loss and the second sub-loss are weighted summed, weighted averaged, etc. to obtain the first loss. By determining the first sub-loss between the first similarity information corresponding to each sample sub-image and the labeled category information, and the second sub-loss between the second similarity information corresponding to each sample sub-image and the labeled category information, and training the neural network model based on at least one of the first sub-loss and the second sub-loss, the first similarity information and the second similarity information determined by the model can be continuously approximated to the labeled category information, thereby improving the accuracy of the first similarity information and the second similarity information, thereby improving the accuracy of the image detection model.
[0188] Step 2043: Based on the first loss, the neural network model is trained to obtain an image detection model.
[0189] In the embodiment of the present application, the first loss can be used as the loss of the neural network model, and the neural network model is trained by the loss of the neural network model to obtain the image detection model. The training method has been described above and will not be repeated here.
[0190] It should be noted that in actual application, other losses can also be designed to determine the loss of the neural network model based on the first loss and other losses to train the neural network model. The embodiments of the present application do not limit other losses. The following exemplarily describes two possible other losses and corresponding training methods, such as the description of implementation method C1 and implementation method C2.
[0191] In implementation C1, step 2043 includes: determining a second loss between a first feature and a second feature of each sample sub-image; and training a neural network model based on the first loss and the second loss to obtain an image detection model.
[0192] In the embodiment of the present application, for any sample sub-image, the loss between the first feature and the second feature of any sample sub-image can be determined, and the feature error between the first feature and the second feature of the sample sub-image can be characterized by the loss. The losses between the first feature and the second feature of each sample sub-image are summed, averaged, etc. to obtain a second loss, and the statistical error between the first feature and the second feature of each sample sub-image can be characterized by the second loss.
[0193] Optionally, by Characterize the first feature of the i-th sample sub-image, through Characterize the second feature of the i-th sample sub-image. Optionally, the second loss is calculated based on the Manhattan distance formula by calculating the first feature and the second feature of each sample sub-image, and the calculation method is shown in the following formula (3).
[0194]
[0195] Among them, L distillation represents the second loss, ∑ represents the summation symbol, The norm between the first feature and the second feature of the i-th sample sub-image is characterized. The norm is also called the Manhattan distance, which is the sum of the absolute values of each element in the feature obtained by subtracting the first feature from the second feature.
[0196] See also Figure 4 , Figure 4 FIG. 1 is a schematic diagram for determining a second loss provided in an embodiment of the present application. Figure 4 As shown, on the one hand, a sample sub-image is cropped from a sample image, and features are extracted from the sample sub-image through a visual encoder to obtain a first feature of the sample sub-image. On the other hand, features are extracted from the sample image through a visual encoder to obtain features of the sample image, and the sample image is processed through a parameter determination network to obtain parameters of the sample sub-image. The second feature of the sample sub-image is obtained by performing region of interest pooling (ROI Pooling) and cropping on feature blocks corresponding to the parameters of the sample sub-image in the features of the sample image. Afterwards, a second loss is determined based on the first feature and the second feature of the sample sub-image.
[0197] The first loss and the second loss can be weighted summed, weighted averaged, etc. to obtain the loss of the neural network model. The neural network model is trained by the loss of the neural network model to obtain the image detection model. The training method will not be described in detail here.
[0198] It can be understood that the first feature of the sample sub-image is obtained by extracting features from the sample sub-image, and the sample sub-image is obtained by cropping the sample image based on the parameters of the sample sub-image. Since the cropping error is small, the accuracy of the sample sub-image is high, so that the accuracy of the first feature of the sample sub-image is high. By determining the second loss between the first feature and the second feature of each sample sub-image, and training the neural network model based on the second loss, it is possible to implement knowledge distillation of the second feature of the sample sub-image based on the first feature of the sample sub-image, and implement model distillation of the determination network of the second feature based on the determination network of the first feature, so that the model can continuously extract the second feature with higher accuracy, thereby improving the accuracy of the image detection model.
[0199] In implementation C2, step 2043 includes: determining a third loss between the annotation parameters and the parameters of each sample sub-image; and training the neural network model based on the first loss and the third loss to obtain an image detection model.
[0200] In the embodiment of the present application, for any sample sub-image, the loss between the annotation parameter and the parameter of any sample sub-image can be determined, and the parameter error between the annotation parameter and the parameter of the sample sub-image can be characterized by the loss. The loss between the annotation parameter and the parameters of each sample sub-image is summed, averaged, etc. to obtain a third loss, and the statistical error between the annotation parameter and the parameters of each sample sub-image is characterized by the third loss.
[0201] The first loss and the third loss can be weighted summed, weighted averaged, etc. to obtain the loss of the neural network model. The neural network model is trained by the loss of the neural network model to obtain the image detection model. The training method will not be described in detail here.
[0202] In an embodiment of the present application, by determining a third loss between the annotation parameters and the parameters of each sample sub-image, and training a neural network model based on the third loss, the parameters of the sub-image determined by the model can be continuously approached to the annotation parameters, thereby improving the accuracy of the parameters, thereby improving the accuracy of the sample sub-image, and the first feature and the second feature of the sample sub-image, and further improving the accuracy of the image detection model.
[0203] It can be seen from the comprehensive step 204 that the loss of the neural network model can be determined based on at least one of the first sub-loss and the second sub-loss, or based on at least one of the first sub-loss and the second sub-loss and at least one of the second loss and the third loss. Exemplarily, the first sub-loss, the second sub-loss, the second loss and the third loss can be weighted summed, weighted averaged, etc. to obtain the loss of the neural network model.
[0204] Optionally, the loss L of the neural network model is: L = LCE1 +L distillation +L reg +L CE2 Among them, L CE1 Characterize the first sub-loss, L CE2 Characterize the second sub-loss, L distillation Characterize the second loss, L reg Characterize the third loss.
[0205] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions. For example, the sample images, sample category texts, etc. involved in this application are all obtained with full authorization.
[0206] In the above method, the parameters of the sample sub-image are determined by the neural network model, and the predicted similarity information is determined based on the sample category text and the parameters of the sample sub-image, so as to achieve the determination of the similarity between the category to which the sample object contained in the sample sub-image belongs and the category to which the specified object corresponding to the sample category text belongs. When the neural network model is subsequently trained based on the predicted similarity information to obtain the image detection model, the image detection model and the neural network model have the same functions. In other words, the image detection model can determine the similarity between the category to which the object contained in the image belongs and the category to which the object corresponding to the category text belongs, thereby determining whether the image contains the object corresponding to the category text, and achieving object detection of any category on the image. Among them, the sample category text corresponds to a limited number of categories, and the image detection model can detect objects of any category on the image, which expands the application scenarios, and even if the image is detected for objects of other categories outside the limited categories, there is no need to iterate the model, which saves iteration time and improves image detection efficiency.
[0207] See also Figure 5 , Figure 5 This is an image detection method provided by an embodiment of the present application. The method can be applied to the above implementation environment. Figure 2 The image detection model shown in the figure performs image detection of any category. For the convenience of description, the terminal device 101 or the server 102 that executes the method shown in the embodiment of the present application can be referred to as an electronic device, that is, the method shown in the embodiment of the present application can be executed by an electronic device. Figure 5 As shown, the method includes the following steps.
[0208] Step 501: Acquire a reference image and target category text, where the reference image includes at least one reference object.
[0209] It is understandable that the implementation method of step 501 is similar to the implementation principle of step 201, and the description of step 201 can be seen, which will not be repeated here.
[0210] Step 502: Determine parameters of at least one reference sub-image based on a reference image using an image detection model, where the reference sub-image is an image region in the reference image that contains a reference object.
[0211] Among them, the image detection model is based on Figure 2 The training method of the relevant image detection model is used for training. It is understandable that the implementation method of step 502 is similar to the implementation principle of step 202, which can be seen in the description of step 202 and will not be repeated here.
[0212] In an exemplary embodiment, step 502 includes: extracting features of a reference image through an image detection model to obtain features of the reference image; determining parameters and indicators of multiple reference candidate frames based on the features of the reference image, wherein the indicators of the reference candidate frames are used to characterize the possibility that a reference object is included in a selection area of the reference candidate frame in the reference image; for any reference candidate frame, if the indicators of any reference candidate frame meet the indicator conditions, determining the parameters of any reference candidate frame as the parameters of a reference sub-image.
[0213] It can be understood that the implementation method of the above steps is similar to the implementation principle of steps 2021 to 2023. Please refer to the description of steps 2021 to 2023 and will not be repeated here.
[0214] Step 503, for any reference sub-image, determine the target similarity information of any reference sub-image based on the target category text and the parameters of any reference sub-image through the image detection model, and the target similarity information is used to characterize the degree of similarity between the category to which the reference object included in any reference sub-image belongs and the target category involved in the target category text.
[0215] It is understandable that the implementation method of step 503 is similar to the implementation principle of step 203, and the description of step 203 can be seen, which will not be repeated here.
[0216] In an exemplary implementation, step 503 includes: cropping the reference image based on the parameters of any reference sub-image by the image detection model to obtain any reference sub-image; determining target similarity information based on the target category text and any reference sub-image by the image detection model. Optionally, the target similarity information is referred to as third similarity information.
[0217] It can be understood that the implementation method of the above steps is similar to the implementation principle of implementation method B1. Please refer to the description of implementation method B1 and will not be repeated here.
[0218] In another exemplary implementation, step 503 includes: using an image detection model to crop the features of the reference image based on the parameters of any reference sub-image to obtain a second feature of any reference sub-image; and using the image detection model to determine target similarity information based on the target category feature and the second feature of any reference sub-image. Optionally, the target similarity information is referred to as fourth similarity information.
[0219] It can be understood that the implementation method of the above steps is similar to the implementation principle of implementation method B2. Please refer to the description of implementation method B2 and will not be repeated here.
[0220] Step 504: If there is target similarity information satisfying the similarity information condition, it is determined that there is an object of the target category in the reference image.
[0221] In the embodiment of the present application, the target similarity information of any reference sub-image includes the third similarity information of the reference sub-image, and the third similarity information is used to characterize the similarity between the category to which the reference object contained in the reference sub-image belongs and the target category.
[0222] Optionally, the third similarity information is a probability value, and the larger the probability value is, the higher the similarity between the category to which the reference object belongs and the target category is. Alternatively, the third similarity information is a positive number, and the larger the value is, the higher the similarity is. Alternatively, the third similarity information is a matrix, and the larger the norm or rank of the matrix is, the higher the similarity is.
[0223] In an exemplary embodiment, if the third similarity information is a probability value or a positive number, for any reference sub-image, if the third similarity information of the reference sub-image is greater than a threshold value (which can be obtained based on artificial experience or experimental verification, and the verification method is not limited), then the target similarity information of the reference sub-image satisfies the similarity information condition. If the third similarity information is a matrix, for any reference sub-image, if the norm or rank of the third similarity information of the reference sub-image is greater than a threshold value, then the target similarity information of the reference sub-image satisfies the similarity information condition. In these two cases, there is an object of the target category in the reference image, and there is an object of the target category in the reference sub-image, and the reference sub-image is the sub-image where the object of the target category is located.
[0224] Alternatively, the target similarity information of any reference sub-image includes fourth similarity information of the reference sub-image, and the fourth similarity information is used to represent the degree of similarity between the category to which the reference object contained in the reference sub-image belongs and the target category.
[0225] Optionally, the fourth similarity information is a probability value, and the larger the probability value is, the higher the similarity between the category to which the reference object belongs and the target category is. Alternatively, the fourth similarity information is a positive number, and the larger the value is, the higher the similarity is. Alternatively, the fourth similarity information is a matrix, and the larger the norm or rank of the matrix is, the higher the similarity is. For any reference sub-image, if the third similarity information of the reference sub-image or the norm, rank, etc. of the third similarity information is greater than a threshold value, then the target similarity information of the reference sub-image satisfies the similarity information condition, there is an object of the target category in the reference image, and there is an object of the target category in the reference sub-image, and the reference sub-image is the sub-image where the object of the target category is located.
[0226] Alternatively, the target similarity information of any reference sub-image includes the third similarity information and the fourth similarity information of the reference sub-image. Optionally, for any reference sub-image, if the third similarity information of the reference sub-image or the norm, rank, etc. of the third similarity information is greater than a threshold, and the fourth similarity information or the norm, rank, etc. of the fourth similarity information is greater than a threshold, then the target similarity information of the reference sub-image satisfies the similarity information condition, there is an object of the target category in the reference image, and there is an object of the target category in the reference sub-image, and the reference sub-image is the sub-image where the object of the target category is located.
[0227] See also Figure 6 , Figure 6 Schematic diagram of an object detection provided by an embodiment of the present application. Figure 6 As shown, on the one hand, the text encoder is used to extract features of the target category text "an image containing a dog" to obtain text features, which are the target category features mentioned above. On the other hand, the visual encoder is used to extract features of the reference image to obtain features of the reference image, and the reference image is processed by the parameter determination network to obtain parameters of the reference sub-image. The visual features of the reference sub-image are obtained by pooling and cropping the feature blocks corresponding to the parameters of the reference sub-image in the features of the reference image, and the visual features correspond to the second features of the reference sub-image mentioned above. Next, the target similarity information is calculated based on the text features and the visual features, and the object detection result is determined based on the target similarity information.
[0228] Optionally, if the object detection result is that the reference sub-image includes an object of the target type, the parameters of the reference sub-image determined by the parameter determination network can be obtained to obtain the parameters of the sub-image where the object of the target type is located, and based on the parameters of the sub-image where the object of the target type is located, the reference image is cropped to obtain the sub-image where the object of the target type is located.
[0229] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions. For example, the reference images and target category texts involved in this application are all obtained with full authorization.
[0230] In the above method, the parameters of the reference sub-image are determined by the image detection model, and the target similarity information is determined based on the target category text and the parameters of the reference sub-image, thereby determining the similarity between the category of the reference object contained in the reference sub-image and the target category corresponding to the target category text. Subsequently, it is possible to determine whether the reference image contains an object of the target category based on the target similarity information, thereby realizing object detection of any category in the image, expanding the application scenarios, and eliminating the need for model iteration, thus saving iteration time and improving image detection efficiency.
[0231] The above describes the training method of the image detection model and the image detection method provided by the embodiment of the present application from the perspective of method steps. The following is a systematic and comprehensive description. Figure 7 , Figure 7 It is a training framework diagram of an image detection model provided in an embodiment of the present application. The image detection model is obtained by training a neural network model, and the neural network model includes a text encoder, a visual encoder and a parameter determination network.
[0232] First, if Figure 7 As shown. On the one hand, the text encoder is used to extract features from the sample category text "an image containing a dog" to obtain sample category features. On the other hand, a sample sub-image is cropped from the sample image, and the visual encoder is used to extract features from the sample sub-image to obtain the first feature of the sample sub-image. On the other hand, the visual encoder is used to extract features from the sample image to obtain the features of the sample image, and the sample image is processed by the parameter determination network to obtain the parameters of the sample sub-image. The second feature of the sample sub-image is obtained by pooling and cropping the feature blocks corresponding to the parameters of the sample sub-image in the features of the sample image.
[0233] Then, if Figure 7As shown. In the first aspect, the first similarity information is calculated based on the sample category feature and the first feature of the sample sub-image, and the first sub-loss is determined based on the first similarity information. In the second aspect, the second similarity information is calculated based on the sample category feature and the second feature of the sample sub-image, and the second sub-loss is determined based on the second similarity information. In the third aspect, the second loss is determined based on the first feature of the sample sub-image and the second feature of the sample sub-image. In the fourth aspect, the third loss is determined based on the parameters of the sample sub-image.
[0234] Then, based on the first sub-loss, the second sub-loss, the second loss and the third loss, the loss of the neural network model is calculated, and the neural network model is trained through the loss of the neural network model to obtain an image detection model.
[0235] Afterwards, when it is necessary to perform object detection on the reference image, target category text related to the target category can be generated. The reference image and target category text are input into the image detection model, and the target similarity information of the reference sub-image is determined by the image detection model. The determination method can be seen in the relevant Figure 5 If there is target similarity information of the reference sub-image that satisfies the similarity information condition, it is determined that the reference image includes an object of the target category, thereby realizing object detection of the target category on the reference image.
[0236] It is understandable that the target category and the sample category text may involve the same or different categories. That is to say, in actual application scenarios, when it is necessary to detect objects of a newly added category, it is only necessary to input the image and the category text related to the newly added category into the image detection model, and the image detection model can be used to detect objects of the newly added category, thereby improving the application scenarios of image detection and eliminating the need to iteratively update the image detection model, saving the time required for iterative updates and improving image detection efficiency.
[0237] Figure 8 FIG. 1 is a schematic diagram of a structure of a training device for an image detection model provided in an embodiment of the present application. Figure 8 As shown, the device comprises:
[0238] An acquisition module 801 is used to acquire a sample image and a sample category text, wherein the sample image includes at least one sample object, and the sample category text is a text about the category to which a specified object in the at least one sample object belongs;
[0239] A determination module 802 is used to determine parameters of at least one sample sub-image based on the sample image through a neural network model, where the sample sub-image is an image region in the sample image that contains a sample object;
[0240] The determination module 802 is further used to determine, for any sample sub-image, predicted similarity information of any sample sub-image based on the sample category text and the parameters of any sample sub-image through a neural network model, where the predicted similarity information is used to characterize the similarity between the category to which the sample object contained in any sample sub-image belongs and the category to which the specified object belongs;
[0241] The training module 803 is used to train the neural network model based on the predicted similarity information of each sample sub-image to obtain an image detection model, and the image detection model is used to detect whether there is an object of the target category in the reference image.
[0242] In one possible implementation, a determination module 802 is used to extract features of a sample image through a neural network model to obtain features of the sample image; based on the features of the sample image, parameters and indicators of multiple sample candidate boxes are determined, and the indicators of the sample candidate boxes are used to characterize the possibility that the sample object is included in the selection area of the sample candidate box in the sample image; for any sample candidate box, if the indicators of any sample candidate box meet the indicator conditions, the parameters of any sample candidate box are determined as the parameters of a sample sub-image.
[0243] In a possible implementation, the determination module 802 is used to crop the sample image based on the parameters of any sample sub-image through a neural network model to obtain any sample sub-image; and determine the predicted similarity information based on the sample category text and any sample sub-image through the neural network model.
[0244] In one possible implementation, the determination module 802 is used to extract features of the sample category text through a neural network model to obtain sample category features; extract features of any sample sub-image through a neural network model to obtain a first feature of any sample sub-image; and determine the first similarity information based on the sample category features and the first feature of any sample sub-image.
[0245] In one possible implementation, the determination module 802 is used to crop the features of the sample image based on the parameters of any sample sub-image through a neural network model to obtain a second feature of any sample sub-image; and determine the predicted similarity information based on the sample category features and the second feature of any sample sub-image through a neural network model.
[0246] In one possible implementation, the training module 803 is used to obtain the labeled category information corresponding to each sample sub-image, where the labeled category information corresponding to the sample sub-image is used to characterize whether the sample object contained in the sample sub-image belongs to the same category as the specified object; determine the first loss between the predicted similarity information and the labeled category information corresponding to each sample sub-image; and train the neural network model based on the first loss to obtain an image detection model.
[0247] In one possible implementation, the training module 803 is used to obtain annotation parameters, which are used to characterize the size and position of the image area where the specified object is located in the sample image; for any sample sub-image, the intersection-and-union ratio is determined based on the annotation parameters and the parameters of any sample sub-image, and the annotation category information of any sample sub-image is determined based on the intersection-and-union ratio.
[0248] In a possible implementation, the predicted similarity information includes first similarity information and second similarity information determined in different ways;
[0249] A training module 803, configured to determine a first loss based on the first sub-loss and the second sub-loss;
[0250] Among them, the first sub-loss is the loss between the first similarity information corresponding to each sample sub-image and the labeled category information, and the second sub-loss is the loss between the second similarity information corresponding to each sample sub-image and the labeled category information.
[0251] In a possible implementation, the training module 803 is used to determine the second loss between the first feature and the second feature of each sample sub-image; based on the first loss and the second loss, the neural network model is trained to obtain the image detection model.
[0252] In a possible implementation, the training module 803 is used to determine a third loss between the annotation parameters and the parameters of each sample sub-image; based on the first loss and the third loss, the neural network model is trained to obtain an image detection model.
[0253] In the above device, the parameters of the sample sub-image are determined through a neural network model, and based on the sample category text and the parameters of the sample sub-image, the predicted similarity information is determined, realizing the determination of the similarity degree between the category to which the sample object included in the sample sub-image belongs and the category to which the specified object corresponding to the sample category text belongs. When training an image detection model based on the predicted similarity information on the neural network model subsequently, the functions of the image detection model and the neural network model are the same. That is to say, the image detection model can determine the similarity degree between the category to which the object included in the image belongs and the category to which the object corresponding to the category text belongs, so as to determine whether the image contains the object corresponding to the category text, realizing object detection for any category of images. Among them, the sample category text corresponds to a limited number of categories, while the image detection model can perform object detection for any category of images, expanding the application scenarios. Moreover, even when performing object detection for categories other than the limited categories on the image, there is no need to perform model iteration, saving the iteration time and improving the image detection efficiency.
[0254] It should be understood that Figure 8 When the above-provided device realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.
[0255] Fig. 9 The following shows a schematic structural diagram of an image detection device provided by an embodiment of the present application. As Fig. 9 shown, the device includes:
[0256] An acquisition module 901, configured to acquire a reference image and a target category text, where the reference image includes at least one reference object;
[0257] A determination module 902, configured to determine the parameters of at least one reference sub-image based on the reference image through an image detection model, where the reference sub-image is an image area in the reference image that contains a reference object, and the image detection model is trained according to the training method of the image detection model related to Figure 2 ;
[0258] The determination module 902 is further configured to, for any one reference sub-image, determine the target similarity information of any one reference sub-image based on the target category text and the parameters of any one reference sub-image through the image detection model, where the target similarity information is used to characterize the similarity degree between the category to which the reference object included in any one reference sub-image belongs and the target category involved in the target category text;
[0259] The determination module 902 is further configured to determine whether an object of the target category exists in the reference image if target similarity information that satisfies the similarity information condition exists.
[0260] In one possible implementation, a determination module 902 is used to extract features of a reference image through an image detection model to obtain features of the reference image; based on the features of the reference image, parameters and indicators of multiple reference candidate boxes are determined, and the indicators of the reference candidate boxes are used to characterize the possibility that the reference object is included in the selection area of the reference candidate box in the reference image; for any reference candidate box, if the indicators of any reference candidate box meet the indicator conditions, the parameters of any reference candidate box are determined as the parameters of a reference sub-image.
[0261] In one possible implementation, the determination module 902 is used to crop the reference image based on the parameters of any reference sub-image through an image detection model to obtain any reference sub-image; and determine the target similarity information based on the target category text and any reference sub-image through the image detection model.
[0262] In one possible implementation, the determination module 902 is used to crop the features of the reference image based on the parameters of any reference sub-image through an image detection model to obtain a second feature of any reference sub-image; and determine the target similarity information based on the target category features and the second feature of any reference sub-image through the image detection model.
[0263] In the above device, the parameters of the reference sub-image are determined by the image detection model, and the target similarity information is determined based on the target category text and the parameters of the reference sub-image, thereby realizing the similarity between the category to which the reference object contained in the reference sub-image belongs and the target category corresponding to the target category text. Subsequently, it is possible to determine whether the reference image contains an object of the target category based on the target similarity information, thereby realizing object detection of any category in the image, expanding the application scenarios, and eliminating the need for model iteration, thus saving iteration time and improving image detection efficiency.
[0264] It should be understood that the above Fig. 9 When the device provided realizes its functions, only the division of the above-mentioned functional modules is used as an example for illustration. In practical applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0265] Fig.10The structure block diagram of a terminal device 1000 provided by an exemplary embodiment of the present application is shown. The terminal device 1000 includes: a processor 1001 and a memory 1002 .
[0266] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0267] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one computer program, which is used to be executed by the processor 1001 to implement the training method of the image detection model or the image detection method provided in the method embodiment of the present application.
[0268] In some embodiments, the terminal device 1000 may further optionally include: a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002 and the peripheral device interface 1003 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 1003 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007 and a power supply 1008.
[0269] The peripheral device interface 1003 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, and the peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, and the peripheral device interface 1003 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.
[0270] The radio frequency circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1004 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1004 converts the electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 1004 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1004 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.
[0271] The display screen 1005 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1005 is a touch display screen, the display screen 1005 also has the ability to collect touch signals on the surface or above the surface of the display screen 1005. The touch signal can be input to the processor 1001 as a control signal for processing. At this time, the display screen 1005 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 1005 can be one, arranged on the front panel of the terminal device 1000; in other embodiments, the display screen 1005 can be at least two, respectively arranged on different surfaces of the terminal device 1000 or in a folding design; in other embodiments, the display screen 1005 can be a flexible display screen, arranged on a curved surface or a folding surface of the terminal device 1000. Even, the display screen 1005 can also be arranged as a non-rectangular irregular figure, that is, a special-shaped screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0272] The camera assembly 1006 is used to capture images or videos. Optionally, the camera assembly 1006 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0273] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 1001 for processing, or input them into the radio frequency circuit 1004 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal device 1000. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1007 may also include a headphone jack.
[0274] The power supply 1008 is used to power various components in the terminal device 1000. The power supply 1008 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0275] In some embodiments, the terminal device 1000 further includes one or more sensors 1009 . The one or more sensors 1009 include, but are not limited to: an acceleration sensor 1011 , a gyroscope sensor 1012 , a pressure sensor 1013 , an optical sensor 1014 , and a proximity sensor 1015 .
[0276] The acceleration sensor 1011 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal device 1000. For example, the acceleration sensor 1011 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 1001 can control the display screen 1005 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor 1011. The acceleration sensor 1011 can also be used to collect game or user motion data.
[0277] The gyro sensor 1012 can detect the body direction and rotation angle of the terminal device 1000, and the gyro sensor 1012 can cooperate with the acceleration sensor 1011 to collect the user's 3D actions on the terminal device 1000. The processor 1001 can implement the following functions based on the data collected by the gyro sensor 1012: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0278] The pressure sensor 1013 can be set on the side frame of the terminal device 1000 and / or the lower layer of the display screen 1005. When the pressure sensor 1013 is set on the side frame of the terminal device 1000, the user's holding signal of the terminal device 1000 can be detected, and the processor 1001 performs left and right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 1013. When the pressure sensor 1013 is set on the lower layer of the display screen 1005, the processor 1001 controls the operability controls on the UI interface according to the user's pressure operation on the display screen 1005. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0279] The optical sensor 1014 is used to collect the ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 according to the ambient light intensity collected by the optical sensor 1014. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is reduced. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera assembly 1006 according to the ambient light intensity collected by the optical sensor 1014.
[0280] The proximity sensor 1015, also called a distance sensor, is usually arranged on the front panel of the terminal device 1000. The proximity sensor 1015 is used to collect the distance between the user and the front of the terminal device 1000. In one embodiment, when the proximity sensor 1015 detects that the distance between the user and the front of the terminal device 1000 is gradually decreasing, the processor 1001 controls the display screen 1005 to switch from the screen-on state to the screen-off state; when the proximity sensor 1015 detects that the distance between the user and the front of the terminal device 1000 is gradually increasing, the processor 1001 controls the display screen 1005 to switch from the screen-off state to the screen-on state.
[0281] Those skilled in the art will understand that Fig.10 The structure shown in the figure does not constitute a limitation on the terminal device 1000, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0282] Fig.11This is a schematic diagram of the structure of the server provided in the embodiment of the present application. The server 1100 may have relatively large differences due to different configurations or performances, and may include one or more processors 1101 and one or more memories 1102, wherein the one or more memories 1102 store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors 1101 to implement the training method of the image detection model or the image detection method provided in the above-mentioned various method embodiments. Exemplarily, the processor 1101 is a CPU. Of course, the server 1100 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 1100 may also include other components for implementing device functions, which will not be repeated here.
[0283] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above-mentioned image detection model training methods or image detection methods.
[0284] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0285] In an exemplary embodiment, a computer program is also provided. The computer program is at least one, and the at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above-mentioned image detection model training methods or image detection methods.
[0286] In an exemplary embodiment, a computer program product is also provided, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above-mentioned image detection model training methods or image detection methods.
[0287] It should be understood that the "plurality" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0288] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0289] The above description is only an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for an image detection model, It is characterized in that The method comprises: Acquire a sample image and a sample category text, wherein the sample image includes at least one sample object, and the sample category text is text about the category to which a specified object in the at least one sample object belongs; Determining parameters of at least one sample sub-image based on the sample image by a neural network model, wherein the sample sub-image is an image region in the sample image that contains the sample object; For any sample sub-image, determining predicted similarity information of the any sample sub-image based on the sample category text and the parameters of the any sample sub-image by the neural network model, wherein the predicted similarity information is used to characterize the degree of similarity between the category to which the sample object contained in the any sample sub-image belongs and the category to which the specified object belongs; Based on the predicted similarity information of each sample sub-image, the neural network model is trained to obtain an image detection model, and the image detection model is used to detect whether there is an object of the target category in the reference image.
2. The method according to claim 1, It is characterized in that The determining the parameters of at least one sample sub-image based on the sample image by using a neural network model comprises: Extracting features of the sample image by using the neural network model to obtain features of the sample image; Based on the features of the sample image, determining parameters and indices of a plurality of sample candidate boxes, wherein the indices of the sample candidate boxes are used to characterize the possibility that the sample object is included in the box selection area of the sample candidate boxes in the sample image; For any sample candidate box, if the index of any sample candidate box meets the index condition, the parameters of any sample candidate box are determined as the parameters of a sample sub-image.
3. The method according to claim 1, It is characterized in that The step of determining predicted similarity information of any one of the sample sub-images based on the sample category text and the parameters of any one of the sample sub-images by using the neural network model includes: Cropping the sample image based on the parameters of any one of the sample sub-images using the neural network model to obtain any one of the sample sub-images; The predicted similarity information is determined by the neural network model based on the sample category text and any one of the sample sub-images.
4. The method according to claim 3, It is characterized in that The step of determining the predicted similarity information based on the sample category text and any one of the sample sub-images by using the neural network model includes: Extracting features of the sample category text by using the neural network model to obtain sample category features; Extracting features from any one of the sample sub-images using the neural network model to obtain a first feature of any one of the sample sub-images; The predicted similarity information is determined based on the sample category feature and the first feature of any one of the sample sub-images.
5. The method according to claim 1, It is characterized in that The step of determining predicted similarity information of any one of the sample sub-images based on the sample category text and the parameters of any one of the sample sub-images by using the neural network model includes: Cutting the features of the sample image based on the parameters of any one of the sample sub-images by using the neural network model to obtain a second feature of any one of the sample sub-images; The predicted similarity information is determined by the neural network model based on the sample category feature and the second feature of any one of the sample sub-images.
6. The method according to claim 1, It is characterized in that The method of training the neural network model based on the predicted similarity information of each sample sub-image to obtain an image detection model includes: Acquire the annotated category information corresponding to each sample sub-image, where the annotated category information corresponding to the sample sub-image is used to indicate whether the sample object contained in the sample sub-image and the specified object belong to the same category; Determine a first loss between the predicted similarity information and the labeled category information corresponding to each sample sub-image; Based on the first loss, the neural network model is trained to obtain an image detection model.
7. The method according to claim 6, It is characterized in that The obtaining of the labeling category information corresponding to each sample sub-image includes: Acquire annotation parameters, where the annotation parameters are used to characterize the size and position of the image region where the designated object is located in the sample image; For any sample sub-image, an intersection-over-union ratio is determined based on the annotation parameter and the parameter of any sample sub-image, and annotation category information of any sample sub-image is determined based on the intersection-over-union ratio.
8. The method according to claim 6, It is characterized in that The predicted similarity information includes first similarity information and second similarity information determined in different ways; The determining of the first loss between the predicted similarity information and the labeled category information corresponding to each sample sub-image includes: Determining a first loss based on the first sub-loss and the second sub-loss; The first sub-loss is the loss between the first similarity information and the labeled category information corresponding to each sample sub-image, and the second sub-loss is the loss between the second similarity information and the labeled category information corresponding to each sample sub-image.
9. The method according to claim 6, It is characterized in that The step of training the neural network model based on the first loss to obtain an image detection model includes: determining a second loss between the first feature and the second feature of each sample sub-image; Based on the first loss and the second loss, the neural network model is trained to obtain an image detection model.
10. The method according to claim 6, It is characterized in that The step of training the neural network model based on the first loss to obtain an image detection model includes: Determining a third loss between the annotation parameters and the parameters of each sample sub-image; Based on the first loss and the third loss, the neural network model is trained to obtain an image detection model.
11. An image detection method, It is characterized in that The method comprises: Acquire a reference image and a target category text, wherein the reference image includes at least one reference object; Determining parameters of at least one reference sub-image based on the reference image by an image detection model, wherein the reference sub-image is an image region in the reference image that contains the reference object, and the image detection model is trained according to the method according to any one of claims 1 to 10; For any reference sub-image, determining target similarity information of any reference sub-image based on the target category text and parameters of any reference sub-image by the image detection model, wherein the target similarity information is used to characterize the degree of similarity between the category to which the reference object included in any reference sub-image belongs and the target category involved in the target category text; If there is target similarity information satisfying the similarity information condition, it is determined that there is an object of the target category in the reference image.
12. The method according to claim 11, It is characterized in that The determining, based on the reference image, parameters of at least one reference sub-image by using an image detection model comprises: Extracting features from the reference image using the image detection model to obtain features of the reference image; Determining parameters and indices of a plurality of reference candidate frames based on features of the reference image, wherein the indices of the reference candidate frames are used to characterize the possibility that a frame selection area of the reference candidate frames in the reference image includes the reference object; For any reference candidate box, if an index of the any reference candidate box meets an index condition, the parameters of the any reference candidate box are determined as parameters of a reference sub-image.
13. The method according to claim 11, It is characterized in that The determining the target similarity information of any one of the reference sub-images based on the target category text and the parameters of any one of the reference sub-images by the image detection model includes: Cropping the reference image based on the parameters of any one of the reference sub-images using the image detection model to obtain any one of the reference sub-images; The target similarity information is determined by the image detection model based on the target category text and any one of the reference sub-images.
14. The method according to claim 11, It is characterized in that The determining the target similarity information of any one of the reference sub-images based on the target category text and the parameters of any one of the reference sub-images by the image detection model includes: The image detection model is used to crop the features of the reference image based on the parameters of any one of the reference sub-images to obtain a second feature of any one of the reference sub-images; The target similarity information is determined by the image detection model based on the target category feature and the second feature of any one of the reference sub-images.
15. A training device for an image detection model, It is characterized in that The device comprises: An acquisition module, used for acquiring a sample image and a sample category text, wherein the sample image includes at least one sample object, and the sample category text is a text about a category to which a specified object in the at least one sample object belongs; A determination module, configured to determine parameters of at least one sample sub-image based on the sample image through a neural network model, wherein the sample sub-image is an image region in the sample image that contains the sample object; The determination module is further used to determine, for any sample sub-image, predicted similarity information of the any sample sub-image based on the sample category text and the parameters of the any sample sub-image through the neural network model, wherein the predicted similarity information is used to characterize the similarity between the category to which the sample object contained in the any sample sub-image belongs and the category to which the specified object belongs; The training module is used to train the neural network model based on the predicted similarity information of each sample sub-image to obtain an image detection model, and the image detection model is used to detect whether there is an object of the target category in the reference image.
16. An image detection device, It is characterized in that The device comprises: An acquisition module, used to acquire a reference image and a target category text, wherein the reference image includes at least one reference object; a determination module, configured to determine parameters of at least one reference sub-image based on the reference image through an image detection model, wherein the reference sub-image is an image region in the reference image that contains the reference object, and the image detection model is trained according to the method according to any one of claims 1 to 10; The determination module is further configured to determine, for any reference sub-image, target similarity information of the reference sub-image based on the target category text and parameters of the reference sub-image through the image detection model, wherein the target similarity information is used to characterize the degree of similarity between the category to which the reference object included in the reference sub-image belongs and the target category involved in the target category text; The determination module is further configured to determine whether an object of the target category exists in the reference image if target similarity information that satisfies the similarity information condition exists.
17. An electronic device, It is characterized in that The electronic device comprises a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor so that the electronic device implements the method according to any one of claims 1 to 14.
18. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor to enable the electronic device to implement any one of the methods described in claims 1 to 14.
19. A computer program product, It is characterized in that The computer program product stores at least one computer program, and the at least one computer program is loaded and executed by a processor so that the electronic device implements the method according to any one of claims 1 to 14.
Citation Information
Cited By
Zero sample learning-based dangerous event detection method and system in driving scene
CN120431550A