Target detection method and related equipment
By combining image and text features in the object detection method for two-way update processing, the problem of low target detection accuracy in the prior art is solved, and more efficient target recognition is achieved.
Patent Information
- Application Number
- CN202311726849.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the detection accuracy of object detection is low, especially when object detection is performed through deep neural network models, it is difficult to effectively identify the target in the image.
A target detection method is proposed, by determining the image to be detected and the corresponding text description information, performing feature extraction and semantic feature encoding at least one scale, combining the image and text features to perform bidirectional update processing, and finally determining the target object detection area based on similarity matching.
Through bidirectional feature interaction and similarity matching, the characteristic power of the feature is enhanced, the accuracy of object detection is improved, and the target in the image can be more effectively identified.
Smart Images

Figure CN120147602A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an object detection method and related devices. Background Art
[0002] With the development of computer technology, the application of artificial intelligence is becoming more and more extensive. The technology of object detection by relying on machine learning means of artificial intelligence has gradually become a mainstream research direction of object detection. The task of object detection is to find the objects of interest in an image and determine their categories and positions, such as detecting faces, vehicles or buildings from the image.
[0003] In the current related technologies, generally, a deep neural network model is used for object detection. Specifically, the deep neural network model extracts features from the image to be detected, and directly performs object recognition on the extracted image feature map. The detection accuracy of such an object detection method is relatively low. Summary of the Invention
[0004] Embodiments of this application provide an object detection method and related devices. The related devices may include an object detection device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of object detection.
[0005] Embodiments of this application provide an object detection method, including:
[0006] Determine the image to be detected and the text description information for the image to be detected;
[0007] Extract features of the image to be detected at at least one scale to obtain initial image feature maps of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword;
[0008] Based on the text feature information, perform update processing on the initial image feature maps at each scale to obtain target image feature maps at each scale; and based on the target image feature maps at each scale, perform update processing on the text feature information to obtain target text feature information, where the target text feature information includes target text features corresponding to each keyword;
[0009] According to a preset object detection frame, perform object recognition on the target image feature map to obtain a plurality of candidate object detection regions;
[0010] Perform object coding prediction processing on each candidate object detection region to obtain region object coding information of each candidate object detection region;
[0011] Determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the region object encoding information of each candidate object detection region and each target text feature.
[0012] Correspondingly, an embodiment of the present application provides a target detection device, including:
[0013] A determination unit, configured to determine an image to be detected and text description information for the image to be detected;
[0014] A feature extraction unit, configured to perform feature extraction on the image to be detected at at least one scale to obtain an initial image feature map of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword;
[0015] An update unit, configured to perform update processing on the initial image feature maps of each scale based on the text feature information to obtain target image feature maps of each scale; and perform update processing on the text feature information based on the target image feature maps of each scale to obtain target text feature information, where the target text feature information includes target text features corresponding to each keyword;
[0016] An identification unit, configured to perform object identification on the target image feature map according to a preset object detection frame to obtain a plurality of candidate object detection regions;
[0017] An encoding unit, configured to perform object encoding prediction processing on each candidate object detection region to obtain region object encoding information of each candidate object detection region;
[0018] A region determination unit, configured to determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the region object encoding information of each candidate object detection region and each target text feature.
[0019] Optionally, in some embodiments of the present application, the update unit may include a first selection subunit, a first fusion subunit, a sampling subunit, and a second fusion subunit, as follows:
[0020] The first selection subunit is configured to select an initial image feature map of a target scale from the initial image feature maps of each scale, and perform sampling processing on the initial image feature map of the target scale to obtain a sampling image feature map corresponding to the target scale;
[0021] The first fusion subunit is configured to perform fusion processing on the text feature information and the initial image feature map for each initial image feature map of the reference scale, to obtain the image feature map of the reference scale, where the reference scale is other scales except the target scale;
[0022] The sampling subunit is configured to perform sampling processing on the image feature map corresponding to the reference scale, to obtain the sampled image feature map of the reference scale;
[0023] The second fusion subunit is configured to fuse the image feature map of the reference scale with the sampled image feature map of the adjacent scale, to obtain the target image feature map of the reference scale; and determine the target image feature map of the target scale based on the initial image feature map of the target scale.
[0024] Optionally, in some embodiments of the present application, the first fusion subunit may specifically be configured to, for each initial image feature map of the reference scale, fuse the initial image feature map with the text features of each keyword respectively, to obtain the fused image feature maps corresponding to each keyword; select the target fused image feature map from the fused image feature maps corresponding to each keyword; perform regression processing on the target fused image feature map, to obtain the processed image feature map; and fuse the processed image feature map with the initial image feature map, to obtain the image feature map corresponding to the reference scale.
[0025] Optionally, in some embodiments of the present application, the updating unit may further include a determination subunit, a second selection subunit, and a return subunit, as follows:
[0026] The determination subunit is configured to, for the target image feature map of each scale, use the target image feature map of the scale as the new initial image feature map of the scale;
[0027] The second selection subunit is configured to select the initial image feature map of the new target scale from the new initial image feature maps of each scale;
[0028] The return subunit is configured to return and execute the step of performing sampling processing on the initial image feature map of the target scale, to obtain the sampled image feature map corresponding to the target scale, to obtain the new target image feature maps of each scale.
[0029] Optionally, in some embodiments of the present application, the updating unit may further include a pooling subunit, an aggregation subunit, an attention processing subunit, and an updating subunit, as follows:
[0030] The pooling subunit is configured to perform pooling processing on the target image feature maps of each scale, to obtain the pooled image feature maps corresponding to each scale;
[0031] An aggregation subunit, configured to aggregate the pooled image feature maps corresponding to each scale to obtain an aggregated image feature map;
[0032] An attention processing subunit, configured to perform attention processing on the text feature information based on the aggregated image feature map to obtain processed text feature information;
[0033] An update subunit, configured to perform update processing on the text feature information based on the processed text feature information to obtain target text feature information.
[0034] Optionally, in some embodiments of the present application, the feature extraction unit may specifically be configured to perform feature extraction on the image to be detected at at least one scale through a target detection model to obtain initial image feature maps of the image to be detected at at least one scale.
[0035] Optionally, in some embodiments of the present application, the target detection device may further include a training unit, and the training unit is configured to train a preset target detection model.
[0036] Optionally, in some embodiments of the present application, the training unit may specifically be configured to obtain training data, where the training data includes a plurality of sample images and sample texts; perform feature extraction on the sample images at at least one scale through a preset target detection model to obtain initial image feature maps of the sample images at at least one scale; perform semantic feature encoding on the sample texts to obtain text feature information corresponding to the sample texts; perform update processing on the initial image feature maps at each scale based on the text feature information to obtain target image feature maps at each scale; perform update processing on the text feature information based on the target image feature maps at each scale to obtain target text feature information; perform object recognition on the target image feature maps according to a preset object detection frame to obtain a plurality of object detection regions; perform object coding prediction processing on each object detection region to obtain region object coding information of each object detection region; determine positive sample texts and negative sample texts corresponding to each object detection region from the sample texts; and adjust parameters of the preset target detection model based on a first similarity between the region object coding information of the object detection region and the target text feature information of its positive sample text, and a second similarity between the region object coding information of the object detection region and the target text feature information of its negative sample text to obtain a target detection model.
[0037] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.
[0038] An electronic device provided by the embodiments of the present application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the object detection method provided by the embodiments of the present application.
[0039] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the object detection method provided by the embodiments of the present application are implemented.
[0040] In addition, the embodiments of the present application also provide a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps in the object detection method provided by the embodiments of the present application are implemented.
[0041] The embodiments of the present application provide an object detection method and related devices; an image to be detected and text description information for the image to be detected can be determined; feature extraction is performed on the image to be detected at at least one scale to obtain initial image feature maps of the image to be detected at at least one scale; semantic feature encoding is performed on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword; based on the text feature information, the initial image feature maps at each scale are updated to obtain target image feature maps at each scale; and based on the target image feature maps at each scale, the text feature information is updated to obtain target text feature information, where the target text feature information includes target text features corresponding to each keyword; according to a preset object detection frame, object recognition is performed on the target image feature maps to obtain multiple candidate object detection regions; object coding prediction processing is performed on each candidate object detection region to obtain region object coding information for each candidate object detection region; based on the similarity between the region object coding information of each candidate object detection region and each target text feature, at least one target object detection region of the object corresponding to each keyword is determined from each candidate object detection region.
[0042] This application can perform two-way feature interaction between image features and text, and then use the interacted features to perform similarity matching between regions and text, so as to detect the desired target from the image to be detected. This enhances the representational power of the features and can improve the accuracy of target detection. Description of the Drawings
[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0044] Figure 1a is a schematic diagram of the scenario of the target detection method provided by the embodiment of this application;
[0045] Figure 1b is a flowchart of the target detection method provided by the embodiment of this application;
[0046] Figure 1c is a model structure diagram of the target detection method provided by the embodiment of this application;
[0047] Figure 1d is another model structure diagram of the target detection method provided by the embodiment of this application;
[0048] Figure 2 is another flowchart of the target detection method provided by the embodiment of this application;
[0049] Figure 3 is a schematic diagram of the structure of the target detection device provided by the embodiment of this application;
[0050] Figure 4 is a schematic diagram of the structure of the electronic device provided by the embodiment of this application. Detailed Embodiments
[0051] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of this application.
[0052] The embodiments of this application provide a target detection method and related devices. The related devices may include a target detection device, an electronic device, a computer-readable storage medium, and a computer program product. The target detection device may be specifically integrated in the electronic device, and the electronic device may be a device such as a terminal or a server.
[0053] It can be understood that the object detection method in this embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as a limitation to this application.
[0054] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0055] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0056] As Figure 1a shown, taking the example of a terminal and a server jointly executing the object detection method. The object detection system provided in the embodiment of this application includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected through a network, for example, through a wired or wireless network connection, etc., where the object detection device can be integrated in the server.
[0057] Among them, the server 11 can be used to: receive the image to be detected and the text description information for the image to be detected sent by the terminal 10; extract features of the image to be detected at at least one scale to obtain the initial image feature maps of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes the text features corresponding to each keyword; based on the text feature information, perform update processing on the initial image feature maps at each scale to obtain the target image feature maps at each scale; and based on the target image feature maps at each scale, perform update processing on the text feature information to obtain target text feature information, where the target text feature information includes the target text features corresponding to each keyword; according to a preset object detection frame, perform object recognition on the target image feature map to obtain multiple candidate object detection regions; perform object coding prediction processing on each candidate object detection region to obtain the region object coding information of each candidate object detection region; based on the similarity between the region object coding information of each candidate object detection region and each target text feature, determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region. Among them, the server 11 can be a single server, or a server cluster or cloud server composed of multiple servers.
[0058] Among them, the terminal 10 can be used to: obtain the image to be detected and the text description information for the image to be detected, and send the image to be detected and the text description information to the server 11, so that the server 11 performs target detection on the image to be detected. The server 11 can also send the recognized detection result to the terminal 10, that is, send at least one target object detection region of the object corresponding to each keyword to the terminal 10, and the terminal 10 can receive the target object detection region sent by the server 11. Among them, the terminal 10 can include a mobile phone, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a tablet computer, a laptop computer, or a personal computer (PC, Personal Computer), etc. A client can also be set on the terminal 10, and the client can be an application client or a browser client, etc.
[0059] The steps such as target detection in the above-mentioned server 11 can also be executed by the terminal 10.
[0060] The target detection method provided by the embodiments of this application relates to computer vision technology, natural language processing, and machine learning in the field of artificial intelligence.
[0061] Among them, Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. Among them, artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, intelligent transportation, and several other major directions.
[0062] Among them, Computer Vision (CV) is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing image processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0063] Among them, Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.
[0064] Among them, Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0065] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0066] This embodiment will be described from the perspective of the target detection device, which can be specifically integrated in an electronic device, and the electronic device can be a device such as a server or a terminal.
[0067] It can be understood that in the specific implementation of this application, data related to user information and the like are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0068] As Figure 1b shown, the specific process of the target detection method can be as follows:
[0069] 101. Determine the image to be detected and the text description information for the image to be detected.
[0070] Among them, the image to be detected may include one or more target objects to be detected, which is specifically an image that needs to identify the specific location of the target object. The image type of the image to be detected is not limited. For the image to be detected containing multiple target objects, the multiple target objects to be detected may be of the same type or different types.
[0071] Among them, the target object to be detected can be various types of objects to be recognized. For example, it can be various types such as people, animals, text, buildings, icons, etc., and this embodiment does not limit this.
[0072] Among them, the text description information can be various types of descriptive texts about the target objects in the image to be detected. For example, the text description information can be noun phrases such as the category and attributes of the target object to be detected, or descriptive statements about the target object, etc. The text description information contains the semantic information of the target object.
[0073] For example, the text description information can be "Aman and a woman are skiing with a dog", which includes the target objects to be detected, namely "man", "woman", and "dog".
[0074] 102. Extract features of the image to be detected at at least one scale to obtain initial image feature maps of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword.
[0075] Among them, the size of the scale can be set according to the actual situation. In one embodiment, image features of the image to be detected at three different scales can be extracted, and these three scales can be 1 / 8, 1 / 16, and 1 / 32 of the image resolution of the image to be detected. For example, when the image resolution of the image to be detected is 640x640, the three initial image feature maps extracted can be 80x80, 40x40, and 20x20.
[0076] Among them, a neural network model can be used to extract features of the image to be detected at at least one scale. Here, the neural network model can be a Visual Geometry Group Network (VGGNet), a Residual Network (ResNet), a Dense Convolutional Network (DenseNet), etc. However, it should be understood that the neural network in this embodiment is not limited to the several types listed above.
[0077] Among them, for the text description information, keyword extraction can be performed first, and then semantic feature encoding can be performed on each extracted keyword through a text encoder. Specifically, keyword extraction can be to extract noun entities in the text description information.
[0078] In a specific embodiment, the text encoder can be based on a 12-layer Transformer structure. When each keyword is input into the text encoder, a 512-dimensional text vector can be output. For example, Figure 1cAs shown, during the training phase of the text encoder, the input to the text encoder can be a series of nouns (including categories). The corresponding lexical encodings, i.e., text feature information, are extracted using the text encoder. During the testing phase of the text encoder, predefined vocabulary can be used for input, or vocabulary defined by the user can be used for input. These vocabulary are encoded offline and stored in the parameters of the network model, so that during the deployment process, it is no longer necessary to repeatedly calculate the encodings of predefined or user-defined vocabulary (i.e., text features).
[0079] 103. Based on the text feature information, update the initial image feature maps at each scale to obtain the target image feature maps at each scale; and based on the target image feature maps at each scale, update the text feature information to obtain the target text feature information, where the target text feature information includes the target text features corresponding to each keyword.
[0080] Among them, bidirectional interaction and fusion can be performed on the image features and text features to obtain updated lexical encodings with image information (i.e., target text feature information) and image features with text information (i.e., target image feature maps).
[0081] Optionally, in this embodiment, the step of "based on the text feature information, update the initial image feature maps at each scale to obtain the target image feature maps at each scale" may include:
[0082] Select the initial image feature map of the target scale from the initial image feature maps at each scale, and perform sampling processing on the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale;
[0083] For each initial image feature map of the reference scale, fuse the text feature information and the initial image feature map to obtain the image feature map of the reference scale, where the reference scale is other scales except the target scale;
[0084] Perform sampling processing on the image feature map corresponding to the reference scale to obtain the sampled image feature map of the reference scale;
[0085] Fuse the image feature map of the reference scale with the sampled image feature map of the adjacent scale to obtain the target image feature map of the reference scale; and based on the initial image feature map of the target scale, determine the target image feature map of the target scale.
[0086] Among them, the target scale can specifically be the smallest scale among all scales. Sampling the initial image feature map of the target scale can specifically be performing upsampling on the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale. Among them, the essence of upsampling is to enlarge the image and image interpolation, and the interpolation method can be the nearest neighbor method, bilinear interpolation method, cubic convolution interpolation method, etc.
[0087] For example, when performing image feature extraction on the image to be detected at three scales, the three initial image feature maps obtained can be 80x80, 40x40, and 20x20, which are respectively denoted as C 3 、C 4 and C 5 . Specifically, C 5 can be the initial image feature map of the target scale, while C 3 、C 4 are the initial image feature maps of the reference scales.
[0088] Among them, there are various fusion methods between the text feature information and the initial image feature map, and this embodiment does not limit this. For example, the fusion method can be splicing processing or weighted fusion, etc.
[0089] Among them, sampling the image feature map of the reference scale can specifically be performing upsampling on the image feature map of the reference scale to obtain the sampled image feature map of the reference scale.
[0090] Among them, the adjacent scale of a certain scale can refer to the largest scale among the scales smaller than this scale. Specifically, it can also refer to the scale that is half of this scale. For example, there are scales 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32. Among them, the adjacent scale of 1 / 8 is 1 / 16.
[0091] Optionally, in this embodiment, the step of "for each initial image feature map of the reference scale, fusing the text feature information and the initial image feature map to obtain the image feature map of the reference scale" may include:
[0092] For each initial image feature map of the reference scale, fusing the initial image feature map with the text features of each keyword respectively to obtain the fused image feature maps corresponding to each keyword;
[0093] Selecting the target fused image feature map from the fused image feature maps corresponding to each keyword;
[0094] Performing regression processing on the target fused image feature map to obtain the processed image feature map;
[0095] Fuse the processed image feature map with the initial image feature map to obtain the image feature map corresponding to the reference scale.
[0096] There are various ways to fuse the initial image feature map with the text features of each keyword respectively. For example, the fusion method can be weighted summation or the like.
[0097] Among them, to select the target fused image feature map from the fused image feature maps corresponding to each keyword, specifically, the maximum value can be selected as the target fused image feature map through the max function.
[0098] There are various ways to fuse the processed image feature map with the initial image feature map, and this embodiment does not limit this. For example, the fusion method can be addition.
[0099] Optionally, in this embodiment, after the step of "determining the target image feature map of the target scale based on the initial image feature map of the target scale", the following steps may further be included:
[0100] For the target image feature map of each scale, use the target image feature map of this scale as the new initial image feature map of this scale;
[0101] Select the initial image feature map of the new target scale from the new initial image feature maps of each scale;
[0102] Return to execute the step of sampling the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale, and obtain the new target image feature maps of each scale.
[0103] Among them, in this embodiment, the fusion of multi-scale image features and text features may include two rounds. In the first round, the sampling process of the feature map is specifically upsampling, and the target scale is the smallest scale. In the second round, the sampling process of the feature map is specifically downsampling, and the new target scale is the largest scale.
[0104] Optionally, in this embodiment, the step of "updating the text feature information based on the target image feature maps of each scale to obtain the target text feature information" may include:
[0105] Perform pooling processing on the target image feature maps of each scale to obtain the pooled image feature maps corresponding to each scale;
[0106] Aggregate the pooled image feature maps corresponding to each scale to obtain the aggregated image feature map;
[0107] Based on the aggregated image feature map, perform attention processing on the text feature information to obtain processed text feature information;
[0108] Based on the processed text feature information, perform update processing on the text feature information to obtain target text feature information.
[0109] Among them, the target image feature maps of each scale can be pooled to obtain pooled image feature maps of the same size.
[0110] Among them, aggregating the aggregated image scale feature maps corresponding to each scale can specifically be splicing processing, etc.
[0111] In some embodiments, based on the processed text feature information and the updated text feature information, specifically, the processed text feature information can be fused with the text feature information, and the fusion method of the two can be weighted fusion, etc.
[0112] In a specific embodiment, a multi-scale feature fusion network can be used to increase the fusion of text-to-image features and image-to-text features, and realize updating image features based on text feature information and updating text features based on image features. The network architecture of this multi-scale feature fusion network is as Figure 1d shown. The multi-scale feature fusion network is specifically a Vision-Language Path Aggregation Network (Vision-Language PAN), which supports the input of text features and multi-scale image features.
[0113] Among them, for the fusion of text to image (Text to Image), that is, the process of updating image features based on text features, this embodiment can be implemented through the T-CSPLayer (Cross-Stage Layer) module in the multi-scale feature fusion network. The T-CSPLayer includes a text-to-image Max-Sigmoid fusion module based on the attention mechanism. For the initial image feature map X l of a certain scale as input, and the text feature information is W, its update method is shown in formula (1):
[0114]
[0115] Among them, W j represents the text features of each keyword, X l ′ represents the image feature map of this scale, and δ is the Sigmoid function. Since PAN contains 4 layers of T-CSPLayer, a total of 4 text-to-image fusions will be carried out.
[0116] Among them, refer toFigure 1d , the T-CSPLayer can first process the initial image feature map through the Dark Bottleneck bottleneck network module, then fuse the processed initial image feature map and text features through the Max-Sigmoid fusion module, and finally fuse (concat) the processed image feature map output by the Max-Sigmoid fusion module with the initial image feature map to obtain the image feature map.
[0117] Specifically, PAN contains two rounds of multi-scale fusion. The first round of multi-scale fusion is the process from top to bottom from C 5 to C 3 . C 3 to C 5 respectively represent the initial image feature maps of each scale. C 3 and C 4 will fuse the text through the T-CSPLayer. Specifically, C 3 and C 4 are processed through formula (1) to obtain C 3 ′ and C 4 ′ (i.e., the image feature maps of the corresponding scales). Then, upsampling is performed on C 3 ′, C 4 ′ and the unprocessed C 5 to obtain the sampled image feature maps of each scale; then, C 3 ′ and C 4 ′ are respectively fused with the sampled image feature maps of their adjacent scales to obtain the target image feature maps at these two scales. For C 5 , it can be directly used as the target image feature map at this scale, thus completing the first round of multi-scale fusion.
[0118] For the second round of multi-scale fusion, it is the process from bottom to top from C 3 to C 5 . The target image feature maps of each scale obtained in the first round can be used as the new C 3 , C 4 and C 5 . Among them, C 4 and C 5 will fuse the text through the T-CSPLayer. Specifically, C 4 and C 5 are processed through formula (1) to obtain C 4 ′ and C 5 ′ (i.e., the image feature maps of the corresponding scales). Then, downsampling is performed on C 4 ′, C 5 ′ and the new C 3 to obtain the sampled image feature maps of each scale; then, C4 ′, C 5 ′ are respectively fused with the sampling image feature maps of their adjacent scales to obtain the target image feature maps at these two scales. For the new C 3 , it can be directly used as the target image feature map at this scale, thus completing the second round of multi-scale fusion.
[0119] Among them, for the fusion of image to text, that is, the process of updating text features based on image features, this embodiment adopts a pooling image attention mechanism for updating. After the first round of multi-scale fusion of image features, this embodiment can use the Max Pooling method to process the target image feature maps at 3 scales to obtain 3 image features of size 3x3 (i.e., the pooled image feature maps), and flatten them to obtain 27 image region encodings X (i.e., the aggregated image feature maps). Then, the multi-head attention mechanism (Multi-Head Convolutional Attention, MHCA) is used to update the text features, as shown in formula (2):
[0120] W′ = W + MultiHead-Attention(W, X, X) (2)
[0121] Among them, W represents the text feature information, and W′ represents the target text feature information.
[0122] Optionally, it is also possible not to update the image features and text features through the Vision-Language PAN (Visual-Language Path Aggregation Network) in the above embodiment, and this embodiment does not limit this.
[0123] 104. According to the preset object detection frame, perform object recognition on the target image feature map to obtain multiple candidate object detection regions.
[0124] Among them, the preset object detection frame can be slid on the target image feature map to perform object region recognition on the target image feature map to obtain multiple candidate object detection regions. Among them, the preset object detection frame can be a detection frame with multiple scales.
[0125] Specifically, for the target image feature map at each scale, object recognition can be performed according to the preset object detection frame on the target image feature map at this scale to obtain multiple candidate object detection regions corresponding to this scale.
[0126] 105. Perform object encoding prediction processing on each candidate object detection region to obtain the region object encoding information of each candidate object detection region.
[0127] Among them, the obtained regional object encoding information can specifically be regarded as the lexical encoding corresponding to the objects included in the candidate object detection regions. Here, the candidate object detection regions include the candidate object detection regions corresponding to each scale.
[0128] 106. Determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the regional object encoding information of each candidate object detection region and each target text feature.
[0129] Among them, the target object detection region represents the position information of each object in the image to be detected.
[0130] Among them, for each candidate object detection region, the regional object encoding information of the candidate object detection region can be respectively calculated for similarity with the target text features of each keyword. According to the calculation results, the candidate object detection region is determined as the target object detection region of the object corresponding to a certain keyword. That is to say, it can be determined that the candidate object detection region contains the object corresponding to the keyword. Specifically, the candidate object detection region with a similarity greater than the preset similarity to the target text feature of a certain keyword can be selected as the target object detection region of the object corresponding to the keyword.
[0131] Optionally, in this embodiment, the step of "performing feature extraction on the image to be detected at at least one scale to obtain the initial image feature map of the image to be detected at at least one scale" may include:
[0132] Performing feature extraction on the image to be detected at at least one scale through a target detection model to obtain the initial image feature map of the image to be detected at at least one scale.
[0133] Among them, the target detection model can be a neural network model, and the type of the neural network model is not limited in this embodiment. Specifically, the target detection model can be an object detection model based on vision-language modeling, which can include the multi-scale feature fusion network mentioned in the above embodiments, etc.
[0134] Optionally, in this embodiment, before the step of "performing feature extraction on the image to be detected at at least one scale through a target detection model to obtain the initial image feature map of the image to be detected at at least one scale", it may further include:
[0135] Obtain training data, where the training data includes a plurality of sample images and sample texts;
[0136] By presetting a target detection model, feature extraction is performed on the sample image at at least one scale to obtain initial image feature maps of the sample image at at least one scale; and semantic feature encoding is performed on the sample text to obtain text feature information corresponding to the sample text;
[0137] Based on the text feature information, update processing is performed on the initial image feature maps at each scale to obtain target image feature maps at each scale; and based on the target image feature maps at each scale, update processing is performed on the text feature information to obtain target text feature information;
[0138] According to a preset object detection frame, object recognition is performed on the target image feature map to obtain multiple object detection regions;
[0139] Object coding prediction processing is performed on each object detection region to obtain region object coding information of each object detection region;
[0140] Positive sample texts and negative sample texts corresponding to each object detection region are determined from the sample text;
[0141] Based on the first similarity between the region object coding information of the object detection region and the target text feature information of its positive sample text, and the second similarity between the region object coding information of the object detection region and the target text feature information of its negative sample text, the parameters of the preset target detection model are adjusted to obtain a target detection model.
[0142] Among them, the sample text here is specifically various types of text description information. For example, the sample text can be a noun phrase such as a category or an attribute, or a descriptive statement, etc.
[0143] Among them, based on the first similarity and the second similarity, a contrast loss value can be calculated, and then according to the contrast loss value, the parameters of the preset target detection model are adjusted.
[0144] Among them, the training process can be to first calculate the contrast loss value, and then use the backpropagation algorithm to adjust the parameters of the preset target detection model, optimize the parameters of the preset target detection model based on the contrast loss value, so that the contrast loss value is less than the preset loss value.
[0145] For open scenario detection for large-scale vocabulary, training on traditional datasets with limited vocabulary cannot endow the model with satisfactory generalization ability and open zero-shot detection ability. This embodiment proposes region-text contrast learning for open vocabulary detection. Compared with traditional object detection training that uses bounding box and class annotations, this embodiment further abstracts them into bounding box and text annotations, i.e., using a text to describe the semantic information of the corresponding region, and the text can be a class, a noun phrase, or even a text description sentence. Compared with traditional bounding box and class annotations, region-text pairs contain a large amount of semantic information, which can help the model achieve stronger semantic understanding ability.
[0146] Specifically, this embodiment proposes a contrast learning method for region-text pairs. During the pre-training process, for each sample image, this embodiment randomly selects K texts (K can be set to 80), which include texts that match the bounding boxes in the image. Each bounding box has a positive sample text, and the rest are negative samples. A contrast learning loss function is constructed by combining the similarity between the object and the text. Among them, the contrast learning loss function can be based on the cross-entropy loss.
[0147] The detection method based on vision-language pre-training in this application is no longer restricted to the traditional detection idea of training on a dataset with limited categories. Instead, it makes full use of pre-training with large-scale data to improve the open recognition ability of the detection model. The pre-training data includes large-scale detection data and image-text data. Its core idea is to construct matching pairs of regions and texts. The text is no longer restricted to categories, but is extended to noun phrases and text descriptions. Each bounding box corresponds to a piece of text. By constructing the matching relationship between regions and texts, the object detection model is directly trained, and a two-way image-visual feature interaction is constructed. Finally, the model is trained using the matching similarity between regions and texts. Its training is no longer restricted by categories and can support any text input. After the pre-training of this type of method, it has extremely strong zero-shot detection and recognition abilities, and can easily achieve good performance in the detection tasks of some specific scenarios through fine-tuning.
[0148] Currently, the related technologies generally have low accuracy in object detection problems for large-scale vocabulary in open scenarios. Moreover, due to the use of complex models, the overall operation efficiency is low, the inference speed is slow, and it is difficult to be directly deployed on the edge device. This application proposes a real-time and open vocabulary image object detection framework, which has achieved leading accuracy, high inference speed, and a simple and easy-to-deploy model.
[0149] Here, the vocabulary can specifically represent a set of object categories or other noun phrases. Open vocabulary means that the number of vocabulary is not restricted and can include any noun phrase, and the set can be dynamically adjusted.
[0150] The present application provides an efficient object detection model structure based on vision-language modeling. Different from traditional object detection models that only support image input, the present application supports both image and text input, and can detect specific target objects according to the user's text. Specifically, as Figure 1c shown, the overall structure of the model is presented, which mainly includes a text encoder and an object detection model for images. The text encoder encodes the user input text (such as "A man and a woman are skiing with a dog") or the user-defined object categories to obtain corresponding vocabulary embedding features (i.e., the text feature information in the above embodiments). The input image will pass through a fully convolutional YOLO backbone network to extract three different-scale image features (i.e., the initial image feature maps in the above embodiments).
[0151] Among them, the backbone network is an image encoder for extracting image features, specifically multi-scale features.
[0152] After obtaining the multi-scale image features and vocabulary embedding features, the multi-scale image features and vocabulary embedding features will interact and fuse bidirectionally in the Vision-Language PAN module to update the image features and text features. The specific processing process of the Vision-Language PAN (Vision-Language Path Aggregation Network) module can refer to the above embodiments and Figure 1d , which will not be elaborated here.
[0153] Among them, the updated image features will further pass through two separate networks, namely the Text Contrastive Head network and the Box Head network. The Box Head network can predict the vocabulary encoding (i.e., the above regional object encoding information) and the object box (i.e., the candidate object detection region) of each object pixel by pixel. In the Text Contrastive Head network, the vocabulary encoding of the object will be calculated for similarity with the vocabulary embedding features obtained by the above update (the target text features of each keyword, such as the target text features corresponding to "man", "woman", "dog"), and the category corresponding to each object in the image to be detected will be determined according to the similarity.
[0154] The object detection method provided by this application can achieve leading accuracy in the zero-shot detection problem of large-scale vocabulary at a real-time inference speed. On the publicly available large-scale vocabulary instance segmentation dataset (Large Vocabulary Instance Segmentation, LVIS), the detection accuracy of this application can reach up to 35.4 AP (AP, average precision, an object detection evaluation metric), and the inference speed of 52.0 FPS (frame rate) can be achieved on an NVIDIA V100 GPU. The data comparison with current related detection methods (such as GLIP-T, GLIPv2-T, GroundingDINO-T, and DetCLIP-T) is shown in Table 1 as follows:
[0155] Table 1
[0156]
[0157] Among them, Table 1 shows the comparison of zero-shot evaluation results between this application and current related technologies on the LVIS dataset. Among them, rare, common, and frequent respectively represent the small, medium, and large numbers of objects in the corresponding categories. Generally, it is difficult to accurately detect the categories with a small number of objects.
[0158] As can be seen from Table 1, the object detection method provided by this application can greatly improve the accuracy of object detection.
[0159] Specifically, this application can be developed on the open-source image object detection framework YOLOv8. YOLO (You Only Look Once) is a series of efficient image object detection models. This application can be applied to browsers. When the user opens the camera, all objects captured by the camera can be detected in real time through the method provided by this application.
[0160] As can be seen from the above, in this embodiment, the image to be detected and the text description information for the image to be detected can be determined; feature extraction is performed on the image to be detected at at least one scale to obtain the initial image feature maps of the image to be detected at at least one scale; and semantic feature encoding is performed on each keyword in the text description information to obtain text feature information, where the text feature information includes the text features corresponding to each keyword; based on the text feature information, the initial image feature maps at each scale are updated to obtain the target image feature maps at each scale; and based on the target image feature maps at each scale, the text feature information is updated to obtain target text feature information, where the target text feature information includes the target text features corresponding to each keyword; according to a preset object detection frame, object recognition is performed on the target image feature map to obtain a plurality of candidate object detection regions; object encoding prediction processing is performed on each candidate object detection region to obtain the region object encoding information of each candidate object detection region; based on the similarity between the region object encoding information of each candidate object detection region and each target text feature, at least one target object detection region of the object corresponding to each keyword is determined from each candidate object detection region.
[0161] This application can perform two-way feature interaction between image features and text, and then use the features after interaction to perform similarity matching between regions and text to detect the target to be detected from the image to be detected, which enhances the representativeness of the features and can improve the accuracy of target detection.
[0162] According to the method described in the previous embodiment, the following will take the specific integration of the target detection device in the server as an example for further detailed description.
[0163] An embodiment of this application provides a target detection method, as Figure 2 shown, the specific process of this target detection method can be as follows:
[0164] 201. The server determines the image to be detected and the text description information for the image to be detected.
[0165] Among them, the image to be detected may include one or more target objects to be detected, and specifically, it is an image that needs to identify the specific location of the target object. The image type of the image to be detected is not limited.
[0166] Among them, the target object to be detected can be various types of targets to be recognized. For example, it can be various types such as people, animals, text, buildings, icons, etc., and this embodiment does not limit this.
[0167] Among them, the text description information can be various types of description texts about the target object in the image to be detected. For example, the text description information can be noun phrases such as the category and attributes of the target object to be detected, or descriptive statements about the target object. The text description information contains the semantic information of the target object.
[0168] 202. The server extracts features of the image to be detected at at least one scale, and obtains initial image feature maps of the image to be detected at at least one scale.
[0169] Among them, the size of the scale can be set according to the actual situation. In one embodiment, image features of three different scales of the image to be detected can be extracted, and these three scales can be 1 / 8, 1 / 16, and 1 / 32 of the image resolution of the image to be detected. For example, when the image resolution of the image to be detected is 640x640, the three initial image feature maps extracted can be 80x80, 40x40, and 20x20.
[0170] 203. The server performs semantic feature encoding on each keyword in the text description information to obtain text feature information, and the text feature information includes text features corresponding to each keyword.
[0171] Among them, for the text description information, keyword extraction can be performed first, and then semantic feature encoding is performed on each extracted keyword through a text encoder. Specifically, keyword extraction can be to extract noun entities in the text description information.
[0172] 204. The server performs update processing on the initial image feature maps of each scale based on the text feature information to obtain target image feature maps of each scale.
[0173] Optionally, in this embodiment, the step of "performing update processing on the initial image feature maps of each scale based on the text feature information to obtain target image feature maps of each scale" may include:
[0174] Select the initial image feature map of the target scale from the initial image feature maps of each scale, and perform sampling processing on the initial image feature map of the target scale to obtain a sampled image feature map corresponding to the target scale;
[0175] For each initial image feature map of the reference scale, the text feature information and the initial image feature map are fused to obtain an image feature map of the reference scale, and the reference scale is other scales except the target scale;
[0176] Perform sampling processing on the image feature map corresponding to the reference scale to obtain a sampled image feature map of the reference scale;
[0177] Fuse the image feature map of the reference scale with the sampled image feature maps of adjacent scales to obtain the target image feature map of the reference scale; and determine the target image feature map of the target scale based on the initial image feature map of the target scale.
[0178] Among them, the target scale can specifically be the smallest scale among all scales. Sampling the initial image feature map of the target scale can specifically be performing upsampling on the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale.
[0179] Among them, sampling the image feature map of the reference scale can specifically be performing upsampling on the image feature map of the reference scale to obtain the sampled image feature map of the reference scale.
[0180] Optionally, in this embodiment, the step of "for each initial image feature map of the reference scale, fuse the text feature information and the initial image feature map to obtain the image feature map of the reference scale" may include:
[0181] For each initial image feature map of the reference scale, fuse the initial image feature map with the text features of each keyword respectively to obtain the fused image feature maps corresponding to each keyword;
[0182] Select the target fused image feature map from the fused image feature maps corresponding to each keyword;
[0183] Perform regression processing on the target fused image feature map to obtain the processed image feature map;
[0184] Fuse the processed image feature map with the initial image feature map to obtain the image feature map corresponding to the reference scale.
[0185] Optionally, in this embodiment, after the step of "determine the target image feature map of the target scale based on the initial image feature map of the target scale", it may further include:
[0186] For the target image feature map of each scale, use the target image feature map of the scale as the new initial image feature map of the scale;
[0187] Select the new initial image feature map of the new target scale from the new initial image feature maps of each scale;
[0188] Return to execute the step of sampling the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale, and obtain the new target image feature maps of each scale.
[0189] Among them, in this embodiment, the fusion of multi-scale image features and text features may include two rounds. In the first round, the sampling process of the feature map may specifically be an upsampling process, and the target scale is the smallest scale. In the second round, the sampling process of the feature map may specifically be a downsampling process, and the new target scale is the largest scale.
[0190] 205. The server updates the text feature information based on the target image feature maps of each scale to obtain target text feature information, where the target text feature information includes target text features corresponding to each keyword.
[0191] Optionally, in this embodiment, the step of "updating the text feature information based on the target image feature maps of each scale to obtain target text feature information" may include:
[0192] Performing pooling processing on the target image feature maps of each scale to obtain pooled image feature maps corresponding to each scale;
[0193] Aggregating the pooled image feature maps corresponding to each scale to obtain an aggregated image feature map;
[0194] Performing attention processing on the text feature information based on the aggregated image feature map to obtain processed text feature information;
[0195] Updating the text feature information based on the processed text feature information to obtain target text feature information.
[0196] Among them, pooling processing may be performed on the target image feature maps of each scale, and the obtained pooled image feature maps may be of the same size.
[0197] Among them, aggregating the aggregated image scale feature maps corresponding to each scale may specifically be splicing processing or the like.
[0198] 206. The server performs object recognition on the target image feature map according to a preset object detection frame to obtain multiple candidate object detection regions.
[0199] Among them, the preset object detection frame may be slid on the target image feature map to perform object region recognition on the target image feature map to obtain multiple candidate object detection regions. Among them, the preset object detection frame may be a detection frame with multiple scales.
[0200] 207. The server performs object coding prediction processing on each candidate object detection region to obtain region object coding information of each candidate object detection region.
[0201] 208. The server determines at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the region object encoding information of each candidate object detection region and each target text feature.
[0202] Among them, a candidate object detection region with a similarity greater than a preset similarity to the target text feature of a certain keyword can be selected as the target object detection region of the object corresponding to the keyword.
[0203] As can be seen from the above, in this embodiment, the server can determine the image to be detected and the text description information for the image to be detected; extract features of the image to be detected at at least one scale to obtain initial image feature maps of the image to be detected at at least one scale; perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes the text features corresponding to each keyword; based on the text feature information, perform update processing on the initial image feature maps at each scale to obtain target image feature maps at each scale; and based on the target image feature maps at each scale, perform update processing on the text feature information to obtain target text feature information, where the target text feature information includes the target text features corresponding to each keyword; perform object recognition on the target image feature maps according to a preset object detection frame to obtain a plurality of candidate object detection regions; perform object encoding prediction processing on each candidate object detection region to obtain the region object encoding information of each candidate object detection region; and determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the region object encoding information of each candidate object detection region and each target text feature.
[0204] This application can perform two-way feature interaction between image features and text, and then use the interacted features to perform similarity matching between regions and text to detect the target to be detected from the image to be detected, thereby enhancing the representational power of the features and improving the accuracy of target detection.
[0205] To better implement the above method, an embodiment of this application further provides a target detection device, as Figure 3 shown. The target detection device may include a determination unit 301, a feature extraction unit 302, an update unit 303, an identification unit 304, an encoding unit 305, and a region determination unit 306, as follows:
[0206] (1) The determination unit 301;
[0207] The determination unit is used to determine the image to be detected and the text description information for the image to be detected.
[0208] (2) Feature extraction unit 302;
[0209] The feature extraction unit is used to perform feature extraction on the image to be detected at at least one scale to obtain an initial image feature map of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword.
[0210] Optionally, in some embodiments of the present application, the feature extraction unit may specifically be used to perform feature extraction on the image to be detected at at least one scale through a target detection model to obtain an initial image feature map of the image to be detected at at least one scale.
[0211] Optionally, in some embodiments of the present application, the target detection device may further include a training unit, and the training unit is used to train a preset target detection model.
[0212] Optionally, in some embodiments of the present application, the training unit may specifically be used to obtain training data, where the training data includes a plurality of sample images and sample texts; perform feature extraction on the sample images at at least one scale through a preset target detection model to obtain initial image feature maps of the sample images at at least one scale; perform semantic feature encoding on the sample texts to obtain text feature information corresponding to the sample texts; based on the text feature information, perform update processing on the initial image feature maps at each scale to obtain target image feature maps at each scale; and based on the target image feature maps at each scale, perform update processing on the text feature information to obtain target text feature information; perform object recognition on the target image feature maps according to a preset object detection frame to obtain a plurality of object detection regions; perform object encoding prediction processing on each object detection region to obtain region object encoding information of each object detection region; determine positive sample texts and negative sample texts corresponding to each object detection region from the sample texts; and adjust the parameters of the preset target detection model based on a first similarity between the region object encoding information of the object detection region and the target text feature information of its positive sample text, and a second similarity between the region object encoding information of the object detection region and the target text feature information of its negative sample text to obtain a target detection model.
[0213] (3) Update unit 303;
[0214] An update unit, configured to update the initial image feature maps of each scale based on the text feature information to obtain the target image feature maps of each scale; and update the text feature information based on the target image feature maps of each scale to obtain target text feature information, where the target text feature information includes the target text features corresponding to each keyword.
[0215] Optionally, in some embodiments of the present application, the update unit may include a first selection subunit, a first fusion subunit, a sampling subunit, and a second fusion subunit, as follows:
[0216] The first selection subunit is configured to select the initial image feature map of the target scale from the initial image feature maps of each scale, and perform sampling processing on the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale;
[0217] The first fusion subunit is configured to, for the initial image feature map of each reference scale, fuse the text feature information and the initial image feature map to obtain the image feature map of the reference scale, where the reference scale is other scales except the target scale;
[0218] The sampling subunit is configured to perform sampling processing on the image feature map corresponding to the reference scale to obtain the sampled image feature map of the reference scale;
[0219] The second fusion subunit is configured to fuse the image feature map of the reference scale with the sampled image feature map of the adjacent scale to obtain the target image feature map of the reference scale; and determine the target image feature map of the target scale based on the initial image feature map of the target scale.
[0220] Optionally, in some embodiments of the present application, the first fusion subunit may specifically be configured to, for the initial image feature map of each reference scale, fuse the initial image feature map with the text features of each keyword respectively to obtain the fused image feature maps corresponding to each keyword; select the target fused image feature map from the fused image feature maps corresponding to each keyword; perform regression processing on the target fused image feature map to obtain the processed image feature map; and fuse the processed image feature map with the initial image feature map to obtain the image feature map corresponding to the reference scale.
[0221] Optionally, in some embodiments of the present application, the update unit may further include a determination subunit, a second selection subunit, and a return subunit, as follows:
[0222] The determining subunit is configured to use the target image feature map of each scale as the new initial image feature map of the scale.
[0223] A second selection subunit, configured to select the initial image feature map of the new target scale from the new initial image feature maps of each scale.
[0224] A return subunit, configured to return the step of performing sampling processing on the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale, so as to obtain the new target image feature maps of each scale.
[0225] Optionally, in some embodiments of the present application, the updating unit may further include a pooling subunit, an aggregation subunit, an attention processing subunit, and an updating subunit, as follows:
[0226] The pooling subunit is configured to perform pooling processing on the target image feature maps of each scale to obtain the pooled image feature maps corresponding to each scale.
[0227] The aggregation subunit is configured to perform aggregation processing on the pooled image feature maps corresponding to each scale to obtain the aggregated image feature map.
[0228] The attention processing subunit is configured to perform attention processing on the text feature information based on the aggregated image feature map to obtain the processed text feature information.
[0229] The updating subunit is configured to perform updating processing on the text feature information based on the processed text feature information to obtain the target text feature information.
[0230] (4) Recognition unit 304;
[0231] The recognition unit is configured to perform object recognition on the target image feature map according to a preset object detection frame to obtain a plurality of candidate object detection regions.
[0232] (5) Encoding unit 305;
[0233] The encoding unit is configured to perform object encoding prediction processing on each candidate object detection region to obtain the region object encoding information of each candidate object detection region.
[0234] (6) Region determination unit 306;
[0235] The region determination unit is configured to determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the region object encoding information of each candidate object detection region and each target text feature.
[0236] As can be seen from the above, in this embodiment, the determination unit 301 can determine the image to be detected and the text description information for the image to be detected; the feature extraction unit 302 can extract features of the image to be detected at at least one scale to obtain initial image feature maps of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword; the update unit 303 can perform update processing on the initial image feature maps at each scale based on the text feature information to obtain target image feature maps at each scale; and perform update processing on the text feature information based on the target image feature maps at each scale to obtain target text feature information, where the target text feature information includes target text features corresponding to each keyword; the recognition unit 304 can perform object recognition on the target image feature maps according to a preset object detection frame to obtain a plurality of candidate object detection regions; the encoding unit 305 can perform object encoding prediction processing on each candidate object detection region to obtain region object encoding information of each candidate object detection region; and the region determination unit 306 can determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the region object encoding information of each candidate object detection region and each target text feature.
[0237] This application can perform two-way feature interaction between image features and text, and then use the interacted features to perform similarity matching between regions and text, so as to detect the target to be detected from the image to be detected, which enhances the representativeness of the features and can improve the accuracy of target detection.
[0238] The embodiment of this application also provides an electronic device, as Figure 4 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of this application. The electronic device can be a terminal or a server, etc. Specifically:
[0239] The electronic device can include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404, and other components. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in
[0240] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the electronic device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.
[0241] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the electronic device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0242] The electronic device further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0243] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0244] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to achieve various functions as follows:
[0245] Determine the image to be detected and the text description information for the image to be detected; perform feature extraction on the image to be detected at at least one scale to obtain the initial image feature maps of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes the text features corresponding to each keyword; based on the text feature information, perform update processing on the initial image feature maps at each scale to obtain the target image feature maps at each scale; and based on the target image feature maps at each scale, perform update processing on the text feature information to obtain target text feature information, where the target text feature information includes the target text features corresponding to each keyword; according to a preset object detection frame, perform object recognition on the target image feature maps to obtain multiple candidate object detection regions; perform object coding prediction processing on each candidate object detection region to obtain the region object coding information of each candidate object detection region; based on the similarity between the region object coding information of each candidate object detection region and each target text feature, determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region.
[0246] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.
[0247] As can be seen from the above, in this embodiment, it is possible to determine the image to be detected and the text description information for the image to be detected; perform feature extraction on the image to be detected at at least one scale to obtain the initial image feature maps of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes the text features corresponding to each keyword; based on the text feature information, perform update processing on the initial image feature maps at each scale to obtain the target image feature maps at each scale; and based on the target image feature maps at each scale, perform update processing on the text feature information to obtain target text feature information, where the target text feature information includes the target text features corresponding to each keyword; according to a preset object detection frame, perform object recognition on the target image feature maps to obtain multiple candidate object detection regions; perform object coding prediction processing on each candidate object detection region to obtain the region object coding information of each candidate object detection region; based on the similarity between the region object coding information of each candidate object detection region and each target text feature, determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region.
[0248] This application can perform two-way feature interaction between image features and text, and then use the features after interaction to perform similarity matching between regions and text, so as to detect the desired target from the image to be detected, which enhances the representational power of the features and can improve the accuracy of target detection.
[0249] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0250] Therefore, an embodiment of this application provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any of the object detection methods provided by the embodiments of this application. For example, the instructions can perform the following steps:
[0251] Determine the image to be detected and the text description information for the image to be detected; perform feature extraction on the image to be detected at at least one scale to obtain the initial image feature maps of the image to be detected at at least one scale; and perform semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes the text features corresponding to each keyword; based on the text feature information, perform update processing on the initial image feature maps at each scale to obtain the target image feature maps at each scale; and based on the target image feature maps at each scale, perform update processing on the text feature information to obtain target text feature information, where the target text feature information includes the target text features corresponding to each keyword; according to a preset object detection frame, perform object recognition on the target image feature maps to obtain multiple candidate object detection regions; perform object encoding prediction processing on each candidate object detection region to obtain the region object encoding information of each candidate object detection region; based on the similarity between the region object encoding information of each candidate object detection region and each target text feature, determine at least one target object detection region of the object corresponding to each keyword from each candidate object detection region.
[0252] For the specific implementation of the above operations, reference can be made to the previous embodiments and will not be elaborated here.
[0253] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0254] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the object detection methods provided by the embodiments of the present application, the beneficial effects achievable by any of the object detection methods provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated herein.
[0255] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the above object detection aspect.
[0256] The above has introduced in detail an object detection method and related devices provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A target detection method, characterized in that, comprising: determining an image to be detected and text description information for the image to be detected; performing feature extraction on the image to be detected at at least one scale to obtain initial image feature maps of the image to be detected at at least one scale; and performing semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword; based on the text feature information, performing update processing on the initial image feature maps at each scale to obtain target image feature maps at each scale; and based on the target image feature maps at each scale, performing update processing on the text feature information to obtain target text feature information, where the target text feature information includes target text features corresponding to each keyword; according to a preset object detection frame, performing object recognition on the target image feature map to obtain a plurality of candidate object detection regions; performing object coding prediction processing on each candidate object detection region to obtain region object coding information of each candidate object detection region; based on the similarity between the region object coding information of each candidate object detection region and each target text feature, determining at least one target object detection region of the object corresponding to each keyword from each candidate object detection region.
2. The method according to claim 1, characterized in that, the performing update processing on the initial image feature maps at each scale based on the text feature information to obtain target image feature maps at each scale includes: selecting an initial image feature map of a target scale from the initial image feature maps at each scale, and performing sampling processing on the initial image feature map of the target scale to obtain a sampled image feature map corresponding to the target scale; for each initial image feature map of a reference scale, fusing the text feature information and the initial image feature map to obtain an image feature map of the reference scale, where the reference scale is other scales except the target scale; performing sampling processing on the image feature map corresponding to the reference scale to obtain a sampled image feature map of the reference scale; fusing the image feature map of the reference scale with the sampled image feature map of an adjacent scale to obtain a target image feature map of the reference scale; and determining a target image feature map of the target scale based on the initial image feature map of the target scale.
3. The method according to claim 2, characterized in that, the fusing the text feature information and the initial image feature map for each initial image feature map of a reference scale to obtain an image feature map of the reference scale includes: for each initial image feature map of a reference scale, fusing the initial image feature map with the text features of each keyword respectively to obtain fused image feature maps corresponding to each keyword; selecting a target fused image feature map from the fused image feature maps corresponding to each keyword; performing regression processing on the target fused image feature map to obtain a processed image feature map; Fuse the processed image feature map with the initial image feature map to obtain the image feature map corresponding to the reference scale.
4. The method according to claim 2, wherein, after determining the target image feature map of the target scale based on the initial image feature map of the target scale, it further includes: For the target image feature map of each scale, use the target image feature map of this scale as the new initial image feature map of this scale; Select the initial image feature map of the new target scale from the new initial image feature maps of each scale; Return to execute the step of sampling the initial image feature map of the target scale to obtain the sampled image feature map corresponding to the target scale, and obtain the new target image feature maps of each scale.
5. The method according to claim 1, wherein, the updating process of the text feature information based on the target image feature maps of each scale to obtain the target text feature information includes: Perform pooling processing on the target image feature maps of each scale to obtain the pooled image feature maps corresponding to each scale; Aggregate the pooled image feature maps corresponding to each scale to obtain the aggregated image feature map; Based on the aggregated image feature map, perform attention processing on the text feature information to obtain the processed text feature information; Based on the processed text feature information, perform updating processing on the text feature information to obtain the target text feature information.
6. The method according to claim 1, wherein, the extracting the initial image feature map of the to-be-detected image at at least one scale to obtain the initial image feature map of the to-be-detected image at at least one scale includes: Through a target detection model, extract features of the to-be-detected image at at least one scale to obtain the initial image feature map of the to-be-detected image at at least one scale.
7. The method according to claim 6, wherein, before extracting the features of the to-be-detected image at at least one scale through the target detection model to obtain the initial image feature map of the to-be-detected image at at least one scale, it further includes: Obtain training data, where the training data includes a plurality of sample images and sample texts; Through a preset target detection model, extract features of the sample images at at least one scale to obtain the initial image feature maps of the sample images at at least one scale; and perform semantic feature encoding on the sample texts to obtain the text feature information corresponding to the sample texts; Based on the text feature information, perform updating processing on the initial image feature maps of each scale to obtain the target image feature maps of each scale; and based on the target image feature maps of each scale, perform updating processing on the text feature information to obtain the target text feature information; According to a preset object detection frame, perform object recognition on the target image feature map to obtain a plurality of object detection regions; Perform object coding prediction processing on each object detection region to obtain the region object coding information of each object detection region. Determine the positive sample text and negative sample text corresponding to each object detection region from the sample text; Based on the first similarity between the region object coding information of the object detection region and the target text feature information of its positive sample text, and the second similarity between the region object coding information of the object detection region and the target text feature information of its negative sample text, adjust the parameters of the preset target detection model to obtain the target detection model.
8. An object detection device Characterized in that It includes: A determination unit for determining an image to be detected and text description information for the image to be detected; A feature extraction unit for performing feature extraction on the image to be detected at at least one scale to obtain an initial image feature map of the image to be detected at at least one scale; and performing semantic feature encoding on each keyword in the text description information to obtain text feature information, where the text feature information includes text features corresponding to each keyword; An update unit for updating the initial image feature map of each scale based on the text feature information to obtain the target image feature map of each scale; and updating the text feature information based on the target image feature map of each scale to obtain target text feature information, where the target text feature information includes target text features corresponding to each keyword; An identification unit for performing object identification on the target image feature map according to a preset object detection frame to obtain a plurality of candidate object detection regions; An encoding unit for performing object coding prediction processing on each candidate object detection region to obtain region object coding information of each candidate object detection region; A region determination unit for determining at least one target object detection region of the object corresponding to each keyword from each candidate object detection region based on the similarity between the region object coding information of each candidate object detection region and each target text feature.
9. An electronic device Characterized in that It includes a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to execute the operations in the object detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium Characterized in that The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the object detection method according to any one of claims 1 to 7.
11. A computer program product, including a computer program or instruction Characterized in that When the computer program or instruction is executed by a processor, the steps in the object detection method according to any one of claims 1 to 7 are implemented.