Object detection method and device, electronic equipment and readable storage medium

By using surround-view image fusion technology in the inspection of three-dimensional structural products and employing an attention weight matrix to fuse feature images, the problem of inaccurate detection results is solved, achieving more efficient and accurate detection.

CN115311652BActive Publication Date: 2026-01-30THUNDERSOFT (NANJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210820212.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2026-01-30
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

Existing technologies for inspecting three-dimensional structural products are prone to inaccurate results due to changes in defect morphology caused by displacement, rotation, and environmental influences, which increases the probability of over-detection and under-detection.

Method used

By acquiring multiple images from a surround view, and fusing the feature images using an attention weight matrix, the feature information from different perspectives is fully utilized to improve detection accuracy and efficiency.

Benefits of technology

It improves the accuracy and efficiency of object detection, reduces the probability of over-detection and under-detection, and significantly enhances the reliability of detection results, especially in the detection of three-dimensional structural products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311652B_ABST
    Figure CN115311652B_ABST
Patent Text Reader

Abstract

This application discloses an object detection method, apparatus, electronic device, and readable storage medium. The method includes: acquiring a view to be detected, the view to be detected including at least a first image and a second image, wherein the first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of a target object; extracting features from the view to be detected to obtain a feature image for each image; fusing the feature images of at least the two adjacent images based on an attention weight matrix to obtain a target feature image; the attention weight matrix is ​​used to characterize the degree of attention of the left image to the right image in the two images; and detecting object information of the target object based on the target feature image, the object information including: category information of the target object, and / or, position information of the target object in the view to be detected. According to the embodiments of this application, the target object can be accurately detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information processing technology, and in particular relates to an object detection method, apparatus, electronic device and readable storage medium. Background Technology

[0002] Currently, with the continuous development of artificial intelligence, neural network models are also widely used in product inspection, such as inspecting products produced on the production line to detect defective products, which can reduce manual input and improve inspection efficiency.

[0003] The product is usually a three-dimensional structure, and displacement and rotation are inevitable during the shooting process. As well as the influence of the shooting environment, the defect shape will change, resulting in inaccurate final inspection results. Summary of the Invention

[0004] This application provides an object detection method, apparatus, electronic device, and readable storage medium, which can solve the problem of low accuracy in current object detection.

[0005] In a first aspect, embodiments of this application provide an object detection method, the method comprising:

[0006] Obtain the view to be detected, which includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object.

[0007] Feature extraction is performed on the view to be detected to obtain the feature image of each image;

[0008] Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain the target feature image; the attention weight matrix is ​​used to characterize the degree of attention of the left image to the right image in the two images;

[0009] Based on the target feature image, detect the object information of the target object, which includes: the category information of the target object, and / or, the position information of the target object in the view to be detected.

[0010] Secondly, embodiments of this application provide an object detection device, the device comprising:

[0011] The acquisition module is used to acquire the view to be detected. The view to be detected includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object.

[0012] The extraction module is used to extract features from the view to be detected, and obtain the feature image of each image;

[0013] The fusion module is used to fuse the feature images of at least two adjacent images based on the attention weight matrix to obtain a target feature image; the attention weight matrix is ​​used to characterize the degree of attention of the left image to the right image in the two images;

[0014] The detection module is used to detect object information of the target object based on the target feature image. The object information includes: the category information of the target object, and / or the position information of the target object in the view to be detected.

[0015] Thirdly, embodiments of this application provide an electronic device, the device including: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the method as described in the first aspect or any possible implementation of the first aspect.

[0016] Fourthly, embodiments of this application provide a readable storage medium storing computer program instructions that, when executed by a processor, implement the method as described in the first aspect or any possible implementation thereof.

[0017] In the object detection method provided in this application, a view to be detected is obtained, which includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object. Feature extraction is performed on the view to be detected to obtain a feature image for each image. Based on an attention weight matrix that characterizes the degree of attention of the left image to the right image in the two images, the feature images of at least the two adjacent images are fused to obtain a target feature image. Here, feature images obtained from different perspectives can be fully utilized for fusion, and the feature image from one perspective can effectively supplement the feature image from the adjacent perspective, making the information included in the fused target feature image richer. Detecting the object information of the target object based on the target feature image can quickly and accurately infer the object information of the target object, improving detection efficiency and detection accuracy. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram illustrating the training and application processes of an object detection model provided in an embodiment of this application;

[0020] Figure 2This is a flowchart of a training method for an object detection model provided in an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of a model structure provided in an embodiment of this application;

[0022] Figure 4 This is a flowchart of an object detection method provided in an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of the structure of an object detection device provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] The features and exemplary embodiments of various aspects of this application will now be described in detail. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only configured to explain this application and are not configured to limit this application. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples of this application.

[0026] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0027] First, the technical terms used in the embodiments of this application will be introduced.

[0028] Attention mechanisms are a method that automatically assigns attention to different parts or regions of an image based on image features, allocating higher attention to key regions.

[0029] Feature extraction networks, which can be convolutional neural networks, take the original image as input and output a 3D feature matrix of size M x N x C, where M x N is the number of feature vectors and C is the dimension of the feature vectors. Typically, M and N decrease with increasing network depth, while C increases with increasing network depth.

[0030] The detection head is a type of convolutional neural network. Its input is the feature matrix output by the feature extraction network, which has a shape of MxNxC. Its output is a matrix of size MxNxD, where D = (1 + 4 + NUM_C), and NUM_C represents the number of defect categories.

[0031] Four of the D-dimensional vectors are used to regress the defect location, which, when represented by a rectangle, are the x-coordinate of the rectangle's center, y-coordinate of the center, the width of the rectangle, and the height of the rectangle. One of the D-dimensional vectors is used to indicate whether a defect exists, while the remaining NUM_C vectors are used to indicate the category to which the defect belongs.

[0032] A detection model typically consists of two parts: a feature extraction network and a detection head. If N feature matrices are extracted from the feature extraction network, then there are corresponding N detection heads.

[0033] Downsampling, also known as extraction, involves taking samples from a sequence of sample values ​​at intervals of several samples to obtain a new sequence, which is called downsampling of the original sequence.

[0034] The inspection process identifies normal samples as defective products.

[0035] Missed inspections result in defective products being identified as normal samples.

[0036] A view is an image taken by a camera at a fixed angle; images taken by cameras at different angles are considered different views.

[0037] A valve stem, also known as a tire valve, is used for inflating and deflating tires. It contains a copper core and is wrapped in rubber. It is usually cylindrical in shape.

[0038] Valve nozzle defects refer to visible defects on the valve nozzle surface caused by external pressure or manufacturing processes, with a width and height greater than 1mm.

[0039] The object detection method provided in this application embodiment can be applied to at least the following application scenarios, which will be described below.

[0040] With the continuous development of artificial intelligence, neural network models are also widely used in product inspection, such as inspecting products produced on the production line to detect defective products, which can reduce manual input and improve inspection efficiency.

[0041] The product is typically a three-dimensional structure. Taking a valve stem as an example, its cylindrical shape causes light scattering. Multiple cameras are needed to capture images from different angles, allowing overlapping of images from adjacent cameras to detect scattering areas. However, this increases camera costs. Furthermore, the images from multiple cameras are not fused, leading to significant changes in defect morphology due to angle or lighting disturbances. This greatly affects the detection results, increasing the probability of both over-detection and under-detection.

[0042] Because valve stems inevitably undergo displacement and rotation during the shooting process, their defect morphology will change. Furthermore, environmental factors such as surface dust can make defects difficult to identify. The same applies to other products; unavoidable displacement and rotation during shooting, along with environmental factors, will alter their defect morphology, leading to inaccurate final inspection results.

[0043] Based on the above application scenarios, the object detection method provided in the embodiments of this application will be described in detail below.

[0044] The object detection model provided in the embodiments of this application will be described in general below.

[0045] Figure 1 This is a schematic diagram illustrating the training and application process of an object detection model provided in an embodiment of this application, as shown below. Figure 1 As shown, it is divided into training process 110 and application process 120.

[0046] In the training process 110, multiple sample data are acquired. Each sample data includes a sample image 111 and preset sample information 115. The sample image 111 includes at least a first sample image and a second sample image, which are two adjacent sample images obtained from multiple sample images captured around the sample object. The sample image 111 is input into the preset model 112 to extract features from the sample image 111, resulting in a sample feature image for each sample image 111. Based on the sample attention weight matrix determined according to the sample feature image of each sample image, the two sample feature images are fused to obtain the target sample feature image 113. Here, feature images obtained from different perspectives are fully utilized for fusion, and the sample feature images of adjacent views are used to supplement the sample feature images of the view.

[0047] Then, based on the target sample feature image, the sample object information 114 of the sample object is detected. The sample object information 114 includes: the category information of the sample object, and / or, the position information of the sample object in the sample image. Based on the sample object information 114 and the preset sample information 115, the preset model 112 is trained to continuously improve the detection capability of the model until the preset model meets the preset training conditions, and the object detection model 122 is obtained.

[0048] In application process 120, a view to be detected 121 is acquired. The view to be detected 121 includes at least a first image and a second image, which are two adjacent images from multiple images captured around the target object. The view to be detected 121 is input into a trained object detection model 122, and feature extraction is performed on the view to be detected 121 to obtain a feature image for each image. Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain a target feature image 123. The attention weight matrix is ​​used to characterize the degree of attention the left-hand image pays to the right-hand image. Here, the fusion of feature images obtained from different perspectives is fully utilized, effectively supplementing the feature image of the adjacent perspective with the feature image of one side. Finally, based on the target feature image 123, the object information 124 of the target object is detected. Since the trained object detection model 122 also has good inference ability for features that are difficult to identify or determine, it can effectively improve detection efficiency and accuracy.

[0049] The training method and object determination method of the object detection model provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0050] Figure 2 This is a flowchart illustrating a training method for an object detection model provided in an embodiment of this application.

[0051] like Figure 2 As shown, the training method for this object detection model may include steps 210-240, as detailed below:

[0052] Step 210: Acquire multiple sample data. Each sample data includes a sample image and preset sample information. The sample image includes at least a first sample image and a second sample image. The sample image is two adjacent sample images among the multiple sample images obtained by taking a surround shot of the sample object.

[0053] Step 220: Input the sample image into the preset model to detect the sample object information of the sample object. The sample object information includes: the category information of the sample object, and / or, the position information of the sample object in the sample image.

[0054] Step 230: Train the preset model based on the sample object information and preset sample information until the preset model meets the preset training conditions to obtain the object detection model.

[0055] The contents of steps 210-230 are described below:

[0056] Step 210 is involved.

[0057] Acquire multiple sample data, each sample data includes a sample image and preset sample information. The sample image includes at least: a first sample image and a second sample image. The sample image is two adjacent sample images among the multiple sample images obtained by taking a surround shot of the sample object.

[0058] A sample image is two adjacent sample images among multiple sample images obtained by taking a panoramic photo of the sample object. It means that the first sample image and the second sample image included in the sample image may have at least some overlapping image regions.

[0059] The preset sample information may include: category information of pre-labeled sample objects, and / or, position information of pre-labeled sample objects in the sample image.

[0060] Step 220 is involved.

[0061] The sample image is input into the preset model to detect the sample object information, which includes: the category information of the sample object, and / or the position information of the sample object in the sample image.

[0062] In one possible embodiment, the preset model includes: an initial feature extraction network, an initial attention network, and an initial detection network; step 220 may specifically include the following steps:

[0063] The sample images are input into the preset model, and the initial feature extraction network is used to extract features from the sample images to obtain the sample feature images of each sample image;

[0064] The sample feature image of each sample image is input into the initial attention network. Based on the sample attention weight matrix, the two sample images are fused to obtain the target sample feature image. The sample attention weight matrix is ​​used to characterize the degree of attention of the sample image on the left to the sample image on the right.

[0065] The target sample feature image is input into the initial detection network. Based on the target sample feature image, the sample object information of the sample object is detected. The sample object information includes: the category information of the sample object, and / or, the position information of the sample object in the sample image.

[0066] Specifically, it can be as follows: Figure 3 As shown, firstly, the sample images are input into a preset model. Through an initial feature extraction network, features are extracted from the sample images to obtain the sample feature image for each sample image. Figure 3 Feature map in.

[0067] The initial feature extraction network can be ResNet18, which uses a convolutional neural network to extract image features to obtain the sample feature image for each sample image.

[0068] Among them, the feature extraction network ResNet18 is a convolutional neural network containing multiple sets of convolutional blocks. After each convolutional block, the obtained feature map is downsampled by 2 times to obtain a new sample feature image. Overall, it can be downsampled by 32 times.

[0069] Each sample image, after feature extraction by ResNet18, yields a sample feature image of size (M, N, C).

[0070] The sample feature image of each sample image is input into the initial attention network. Based on the sample attention weight matrix, the two sample images are fused to obtain the target sample feature image, that is, the feature map after fusion. The sample attention weight matrix is ​​used to characterize the degree of attention of the sample image on the left to the sample image on the right.

[0071] In the initial attention network, a sample attention weight matrix is ​​determined based on two sample images to characterize the degree of attention of the sample image on the left to the sample image on the right. Then, based on the sample attention weight matrix, the two sample images are fused to obtain the target sample feature image.

[0072] Then, the target sample feature image is input into the initial detection network, i.e., the detection head. Based on the target sample feature image, the network detects the sample object information, which includes: the category information of the sample object, and / or, the position information of the sample object in the sample image. Specifically, the detection of the sample object information can be implemented through a fully connected module.

[0073] The fully connected module is a neural network consisting of two fully connected layers and an activation layer in between. Its input and output shapes are identical. The fused feature map is input into the fully connected module, which outputs a final feature map whose shape is restored from (MxN, C) to its original shape (M, N, C). This restored feature map is then input into the detection head for defect location regression and defect category classification.

[0074] Step 230 is involved.

[0075] Step 230: Train the preset model based on the sample object information and preset sample information until the preset model meets the preset training conditions to obtain the object detection model.

[0076] Specifically, the loss value can be determined based on the sample object information and preset sample information. The preset model is then trained based on the loss value until the preset model meets the preset training conditions, such as the loss value of the preset model meeting the preset convergence condition, or the loss value being less than the preset threshold, thus obtaining a trained object detection model.

[0077] The object detection model training method provided in this application acquires multiple sample data, each including a sample image and preset sample information. The sample image consists of two adjacent sample images from multiple images captured around the sample object. The sample images are input into a preset model, where an initial feature extraction network extracts features from each image to obtain a sample feature image. The sample feature images are then input into an initial attention network, where the two sample images are fused based on a sample attention weight matrix to obtain a target sample feature image. Here, the fusion of sample feature images obtained from different perspectives is fully utilized, supplementing the information of the adjacent sample feature image with the sample feature image from one side view. The target sample feature image is then input into an initial detection network, where the sample object information of the detected object is determined. This sample object information includes the object's category information and / or its position information within the sample image. Finally, the preset model is trained based on the sample object information and the preset sample information, continuously improving the model's detection capability until the preset model meets the preset training conditions, thus obtaining the object detection model.

[0078] based on Figure 2 The present application also provides a model construction method for the object detection model training method shown, which may specifically include:

[0079] A pre-defined network structure with multiple layers is constructed, including a feature extraction network, an attention network, and a detection network.

[0080] Each layer of the preset network structure includes a feature extraction network for feature extraction; an attention network for fusing the output of the feature extraction network; and a detection network for calculating sample object information based on the output of the attention network, and calculating the loss value based on the sample object information and preset object information.

[0081] A pre-defined network structure is trained based on the loss value to determine the parameters of multiple trained neural network layers at each of the multiple layers.

[0082] The trained preset network structure is determined as the detection model.

[0083] Figure 4 This is a flowchart of an object detection method provided in an embodiment of this application.

[0084] like Figure 4 As shown, the object detection method may include steps 410-440. This method is applied to an object detection device, as detailed below:

[0085] Step 410: Obtain the view to be detected. The view to be detected includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object.

[0086] Step 420: Perform feature extraction on the view to be detected to obtain the feature image of each image.

[0087] Step 430: Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain the target feature image; the attention weight matrix is ​​used to characterize the degree of attention of the left image to the right image in the two images.

[0088] Step 440: Based on the target feature image, detect the object information of the target object. The object information includes: the category information of the target object, and / or, the position information of the target object in the view to be detected.

[0089] In the object detection method provided in this application, a view to be detected is obtained, which includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object. Feature extraction is performed on the view to be detected to obtain a feature image for each image. Based on an attention weight matrix that characterizes the degree of attention of the left image to the right image in the two images, the feature images of at least the two adjacent images are fused to obtain a target feature image. Here, feature images obtained from different perspectives can be fully utilized for fusion, and the feature image from one perspective can effectively supplement the feature image from the adjacent perspective, making the information included in the fused target feature image richer. Detecting the object information of the target object based on the target feature image can quickly and accurately infer the object information of the target object, improving detection efficiency and detection accuracy.

[0090] The contents of steps 410-440 are described below:

[0091] Step 410 is involved.

[0092] Obtain the view to be detected, which includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object.

[0093] The first image and the second image are two adjacent images among multiple images obtained by taking a panoramic shot of the target object, meaning that at least some image regions may overlap between the first image and the second image.

[0094] Step 420 is involved.

[0095] Feature extraction is performed on the view to be detected to obtain the feature image of each image.

[0096] In one possible embodiment, the view to be detected corresponds to view identification information. Step 420 may specifically include the following steps:

[0097] Feature extraction is performed on the view to be detected to obtain the view feature image;

[0098] The first projected image is determined based on the view identification information;

[0099] Determine the second projected image based on the view feature image;

[0100] For each view to be detected, the view feature image, the first projection image, and the second projection image are added together to obtain the feature image of each image.

[0101] Among them, the view identification information is used to identify the view. For example, each view is numbered according to the event sequence of the surround shooting, such as: first image, second image, Nth view, etc., where N is an integer greater than 1.

[0102] In the steps of determining the first projected image based on the view identification information and determining the second projected image based on the view feature image, the following position projection formula can be used:

[0103] concat(sin(2πXB T ), cos(2πXB) T ))

[0104] Where X represents the original coordinate vector. If it is a one-dimensional coordinate, the vector dimension is 1; if it is a two-dimensional coordinate, the vector dimension is 2, and so on. X is the feature map of (M, N, C).

[0105] B is the projection matrix, whose elements are derived from a standard Gaussian distribution and have a shape of (Z / 2, the dimension of the original coordinate vector), where Z represents the dimension of the projected position vector. Concat represents the concatenation operation, which concatenates two vectors of shape (1, Z / 2) together, resulting in a final projected position vector with dimension Z.

[0106] The Gaussian distribution, also known as the normal distribution, is the most important continuous probability distribution in statistics. Studies have shown that in physical sciences and economics, the distribution of a large amount of data usually follows a Gaussian distribution. Therefore, when we are unclear about the potential distribution pattern of data, we can preferentially use the Gaussian distribution to approximate or accurately describe it.

[0107] A projection matrix is ​​a matrix A that is both symmetric and idempotent. A matrix is ​​a symmetric matrix (a matrix whose elements are equal about the main diagonal).

[0108] For each view to be detected, the view feature image, the first projection image and the second projection image are added together to obtain the feature image corresponding to each view to be detected. This results in a new feature image containing positional information, which is the feature image corresponding to each view to be detected. The shape of the feature image corresponding to each view to be detected is still (M, N, C).

[0109] Specifically, the step of determining the first projected image based on the view identification information mentioned above may include the following steps:

[0110] Based on the preset projection relationship, the view identification information is transformed into a first position vector;

[0111] The first projected image is determined based on the first position vector;

[0112] Specifically, the step of determining the second projected image based on the view feature image mentioned above may include the following steps:

[0113] Based on a preset projection relationship, the view feature image is transformed into a second position vector;

[0114] The second projected image is determined based on the second position vector.

[0115] Among them, based on the preset projection relationship, the view identification information is transformed into a first position vector;

[0116] The first projected image is determined based on the first position vector.

[0117] Specifically, the view identification information to which the image belongs, i.e. which view it is, is converted into a first position vector using the position projection formula. The dimension of the first position vector is the same as the vector dimension of the view feature image.

[0118] Each view feature image has a view position projection of shape (1, C), which is expanded to (1, 1, C) and copied to size (M, N, C).

[0119] Based on a preset projection relationship, the view feature image is transformed into a second position vector;

[0120] The second projected image is determined based on the second position vector.

[0121] The position of the feature map vector, i.e. the 2D coordinates in the feature map, is transformed into a second position vector using the position projection formula. The vector dimension of the second position vector is the same as the dimension of the view feature image.

[0122] Each view feature image has a second projection image, the shape of which is (M, N, C).

[0123] In one possible embodiment, feature extraction is performed on the view to be detected to obtain a feature image for each image, including:

[0124] The image to be detected is input into a pre-trained object detection model, which includes a feature extraction network, an attention network, and a detection network. The feature extraction network extracts features from the view to be detected to obtain a feature image for each image.

[0125] Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain the target feature image, including:

[0126] The feature images of two contiguous images are input into the attention network to determine the attention weight matrix;

[0127] Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain the target feature image;

[0128] Based on the target feature image, detect the object information of the target object, including:

[0129] The target feature image is input into the detection network to obtain the object information of the target object.

[0130] Pre-trained object detection models can quickly and accurately extract feature images from each image, including the feature images of the first and second images. They can also quickly and accurately determine the attention weight matrix based on the feature images of two adjacent images.

[0131] Based on the attention weight matrix determined by the feature images of two adjacent images, the feature images of at least two adjacent images are fused to obtain the target feature image. This method can make full use of feature images obtained from different perspectives for fusion, and can effectively supplement the feature image of the adjacent perspective with the feature image of one side, making the details included in the fused target feature image richer. Finally, the object information of the target object is detected based on the target feature image of the trained detection network, which can quickly and accurately infer the object information of the target object, improving detection efficiency and detection accuracy.

[0132] Step 430 is involved.

[0133] Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain the target feature image; the attention weight matrix is ​​used to characterize the degree of attention of the left image to the right image in the two images.

[0134] In one possible embodiment, the following steps may be included prior to step 430:

[0135] The attention weight matrix is ​​determined based on the feature images of the two images.

[0136] Specifically, the attention weight matrix can be determined based on the following formulas for Softmax and attention weights:

[0137]

[0138] Y is a scalar, which is 1. In the formula, Z represents a matrix consisting of C vectors. d Let d represent the d-th vector in the Z matrix.

[0139] The output of Softmax is the probability of each class being selected. It maps some inputs to real numbers between 0 and 1, and normalization ensures that the sum is 1, so the sum of the probabilities of multiple classes is also exactly 1.

[0140] The outputs of the Softmax function are interrelated, and their probabilities always sum to 1; for example, 0.04 + 0.21 + 0.05 + 0.70 = 1.00. Therefore, in the softmax function, if the probability of one class increases, the probability of other classes will necessarily decrease. The Softmax function is used if the model's outputs are mutually exclusive and only one class can be selected.

[0141] Based on the feature images of the two images, the attention weight matrix can be determined using the following attention weight calculation formula:

[0142]

[0143] Where Q and K represent matrices consisting of query vectors and key vectors, with shape (N, d), where d is the dimension of the vectors and N represents the number of vectors. After softmax, an attention weight matrix A of shape (NxN) is obtained, where the sum of the values ​​in each row of the matrix is ​​1. The element Aij in the matrix represents the attention assignment value of the i-th Qi to the j-th Kj.

[0144] Where Q is the feature image of the first image and K is the feature image of the second image.

[0145] The query vector Q is used to match the key vector K of all elements in the sequence. The attention weights are obtained by multiplying the current element's own Q by the K of all elements. Then, these attention weights are cross-multiplied by the value vectors at their corresponding positions, and the summation yields the self-attention result for the current element. The query vector Q is mainly used to calculate the attention for the current element itself, while the key vector K is mainly used to calculate the attention for other elements.

[0146] The step of determining the attention weight matrix based on the feature images of the two images mentioned above can be implemented according to the following formula:

[0147]

[0148] This step represents the extraction and fusion of information from K based on the attention of Q to each K. The fused information... The shape is the same as the original Q.

[0149] in, K is the target feature image, K is the second feature image of the second image, and A is the first attention weight matrix.

[0150] The feature map containing location information is flattened, i.e. its shape becomes (M x N, C). For each feature map, the attention weight matrix between it and its neighboring views is calculated. Then, the feature map is updated by information fusion based on the attention weights, thus obtaining the target feature map.

[0151] In one possible embodiment, the view to be detected mentioned above also includes a third image that is adjacent to the other side of the first image or the second image. The step of determining the attention weight matrix based on the feature images of the two images may specifically include the following steps:

[0152] Based on the feature images of the first image and the second image, a first attention weight matrix is ​​determined; and,

[0153] The second attention weight matrix is ​​determined based on the feature images of the third image and the feature images of the images adjacent to the third image;

[0154] Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain the target feature image, including:

[0155] Based on the first attention weight matrix, the feature images of the first image and the second image are fused to obtain the fused feature image;

[0156] Based on the second attention weight matrix, the feature image of the third image and the fused feature image are fused to obtain the target feature image.

[0157] The view to be detected is obtained by taking sequential images around the target object. The view to be detected also includes a third image that is adjacent to the other side of the first or second image. That is, the third image is adjacent to the first image, and the shooting order may include: third image, first image, second image; or, second image, first image, third image; or, first image, second image, third image; or, third image, second image, first image.

[0158] The view to be detected also includes a third image that is connected to the other side of the first image or the second image. The step of determining the attention weight matrix based on the feature images of the two images may specifically include the following steps: determining a first attention weight matrix based on the feature images of the first image and the feature images of the second image; and determining a second attention weight matrix based on the feature images of the third image and the feature images of the image connected to the third image.

[0159] Taking the shooting order as second image, first image, and third image as an example. The first and second images are contiguous, and the first and third images are also contiguous. A first attention weight matrix is ​​determined based on the first and second images respectively; and a second attention weight matrix is ​​determined based on the first and third images.

[0160] Accordingly, the step of fusing the feature images of at least two adjacent images based on the attention weight matrix to obtain the target feature image may specifically include the following steps:

[0161] Based on the first attention weight matrix, the feature images of the first image and the second image are fused to obtain the fused feature image;

[0162] Based on the second attention weight matrix, the feature image of the third image and the fused feature image are fused to obtain the target feature image.

[0163] Specifically, for each feature image, information from the adjacent right-hand feature image can be fused first, and then information from the adjacent left-hand feature image can be fused to achieve information fusion of multiple feature images.

[0164] Therefore, by fully utilizing feature images obtained from different perspectives for fusion, the feature image from the right-hand perspective can effectively supplement the feature image from the adjacent perspective, resulting in a fused feature image. Then, the feature image from the left-hand perspective is used to further supplement the fused feature image, yielding a target feature image. This results in a richer information content in the fused target feature image, facilitating subsequent detection of target object information based on the target feature image.

[0165] Step 440: Based on the target feature image, detect the object information of the target object. The object information includes: the object's category information, and / or, the object's position information in the view to be detected.

[0166] In the object detection method provided in this application, a view to be detected is obtained, which includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object. Feature extraction is performed on the view to be detected to obtain a feature image for each image. Based on an attention weight matrix that characterizes the degree of attention of the left image to the right image in the two images, the feature images of at least the two adjacent images are fused to obtain a target feature image. Here, feature images obtained from different perspectives can be fully utilized for fusion, and the feature image from one perspective can effectively supplement the feature image from the adjacent perspective, making the information included in the fused target feature image richer. Detecting the object information of the target object based on the target feature image can quickly and accurately infer the object information of the target object, improving detection efficiency and detection accuracy.

[0167] Based on the above Figure 4 The object detection method shown in this application also provides an object detection device, such as... Figure 5 As shown, the device 500 may include:

[0168] The acquisition module 510 is used to acquire the view to be detected, which includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object.

[0169] The extraction module 520 is used to extract features from the view to be detected, and obtain the feature image of each image.

[0170] The fusion module 530 is used to fuse the feature images of at least two adjacent images based on the attention weight matrix to obtain the target feature image; the attention weight matrix is ​​used to characterize the degree of attention of the left image to the right image in the two images.

[0171] The detection module 540 is used to detect object information of the target object based on the target feature image. The object information includes: category information of the target object, and / or, position information of the target object in the view to be detected.

[0172] In one possible implementation, the device 500 may further include:

[0173] The first determining module is used to determine the attention weight matrix based on the feature images of the two images.

[0174] In one possible implementation, the view to be detected further includes a third image adjoining the other side of the first or second image, and the first determining module is specifically used for:

[0175] Based on the feature images of the first image and the second image, a first attention weight matrix is ​​determined; and,

[0176] The second attention weight matrix is ​​determined based on the feature images of the third image and the feature images of the images adjacent to the third image.

[0177] Fusion module 530 is specifically used for:

[0178] Based on the first attention weight matrix, the feature images of the first image and the second image are fused to obtain the fused feature image;

[0179] Based on the second attention weight matrix, the feature image of the third image and the fused feature image are fused to obtain the target feature image.

[0180] In one possible implementation, the view identifier information corresponding to the view to be detected is extracted by module 520, which is specifically used for:

[0181] Feature extraction is performed on the view to be detected to obtain the view feature image.

[0182] Extraction module 520 may specifically include:

[0183] The second determining module is used to determine the first projected image based on the view identification information.

[0184] The third determining module is used to determine the second projected image based on the view feature image.

[0185] The addition module is used to add the view feature image, the first projection image and the second projection image for each view to be detected, so as to obtain the feature image of each image.

[0186] In one possible implementation, the second determining module is specifically used for:

[0187] Based on the preset projection relationship, the view identification information is transformed into a first position vector;

[0188] The first projected image is determined based on the first position vector.

[0189] The third determining module is specifically used for:

[0190] Based on a preset projection relationship, the view feature image is transformed into a second position vector;

[0191] The second projected image is determined based on the second position vector.

[0192] In one possible implementation, the extraction module 520 is specifically used for:

[0193] The image to be detected is input into a pre-trained object detection model, which includes a feature extraction network, an attention network, and a detection network. The feature extraction network extracts features from the view to be detected to obtain a feature image for each image.

[0194] Fusion module 530 is specifically used for:

[0195] The feature images of two contiguous images are input into the attention network to determine the attention weight matrix;

[0196] Based on the attention weight matrix, the feature images of at least two adjacent images are fused to obtain the target feature image.

[0197] The detection module 540 is specifically used to input the target feature image into the detection network to obtain the object information of the target object.

[0198] In one possible implementation, the device 500 may further include:

[0199] The first acquisition module is used to acquire multiple sample data. Each sample data includes a sample image and preset sample information. The sample image includes at least a first sample image and a second sample image. The sample image is two adjacent sample images among multiple sample images obtained by taking a surround shot of the sample object.

[0200] The input module is used to input sample images into a preset model to detect sample object information, which includes: sample object category information, and / or, sample object location information in the sample image.

[0201] The training module is used to train the preset model based on the sample object information and preset sample information until the preset model meets the preset training conditions, thus obtaining the object detection model.

[0202] In one possible implementation, the pre-defined model includes: an initial feature extraction network, an initial attention network, and an initial detection network; the input module is specifically used for:

[0203] The sample images are input into the preset model, and the initial feature extraction network is used to extract features from the sample images to obtain the sample feature images of each sample image;

[0204] The sample feature image of each sample image is input into the initial attention network. Based on the sample attention weight matrix, the two sample images are fused to obtain the target sample feature image. The sample attention weight matrix is ​​used to characterize the degree of attention of the sample image on the left to the sample image on the right.

[0205] The target sample feature image is input into the initial detection network. Based on the target sample feature image, the sample object information of the sample object is detected. The sample object information includes: the category information of the sample object, and / or, the position information of the sample object in the sample image.

[0206] The object detection apparatus provided in this application acquires a view to be detected, which includes at least a first image and a second image. The first image and the second image are two adjacent images among multiple images obtained by taking a surround shot of the target object. Feature extraction is performed on the view to be detected to obtain a feature image for each image. Based on an attention weight matrix that characterizes the degree of attention the left image pays to the right image in the two images, the feature images of at least the two adjacent images are fused to obtain a target feature image. Here, feature images obtained from different perspectives can be fully utilized for fusion, and the feature image from one perspective can effectively supplement the feature image from the adjacent perspective, making the information included in the fused target feature image richer. Detecting the object information of the target object based on the target feature image enables rapid and accurate inference of the object information of the target object, improving detection efficiency and accuracy.

[0207] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown.

[0208] An electronic device may include a processor 601 and a memory 602 storing computer program instructions.

[0209] Specifically, the processor 601 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0210] Memory 602 may include a large-capacity memory for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 602 is a non-volatile solid-state memory. In a particular embodiment, memory 602 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0211] The processor 601 implements any of the object detection methods in the embodiments shown in the figure by reading and executing computer program instructions stored in the memory 602.

[0212] In one example, the electronic device may also include a communication interface 603 and a bus 610. For example, Figure 6 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 610 and complete communication with each other.

[0213] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0214] Bus 610 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 610 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0215] The electronic device can execute the object detection method in the embodiments of this application, thereby achieving the combination Figures 1 to 4 The described object detection method.

[0216] Furthermore, in conjunction with the object detection method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; these computer program instructions are implemented when executed by a processor. Figures 1 to 4 Object detection methods in [the context of the text].

[0217] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0218] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0219] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0220] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A method of object detection, characterized by, The method comprises: obtaining a to-be-detected view, the to-be-detected view comprising at least a first image and a second image, the first image and the second image being two images adjacent to each other in a plurality of images obtained by surrounding shooting of a target object; performing feature extraction on the to-be-detected view to obtain a feature image of each of the images; fusing, based on an attention weight matrix, at least the feature images of the two adjacent images to obtain a target feature image, the attention weight matrix being used to represent a degree of attention of an image on the left to an image on the right in the two images; detecting, according to the target feature image, object information of the target object, the object information comprising category information of the target object and / or position information of the target object in the to-be-detected view; the feature extraction on the to-be-detected view to obtain the feature image of each of the images comprises: inputting the to-be-detected image into a pre-trained object detection model, the object detection model comprising a feature extraction network, an attention network and a detection network, and performing feature extraction on the to-be-detected view by the feature extraction network to obtain the feature image of each of the images; the fusing, based on the attention weight matrix, at least the feature images of the two adjacent images to obtain the target feature image comprises: inputting the feature images of the two adjacent images into the attention network to determine the attention weight matrix; fusing, based on the attention weight matrix, at least the feature images of the two adjacent images to obtain the target feature image; the detection, according to the target feature image, of the object information of the target object comprises: inputting the target feature image into the detection network to obtain the object information of the target object.

2. The method of claim 1, wherein, Before the fusing, based on the attention weight matrix, at least the feature images of the two adjacent images to obtain the target feature image, the method further comprises: determining the attention weight matrix according to the feature images of the two images.

3. The method of claim 2, wherein, The to-be-detected view further comprises a third image adjacent to the other side of the first image or the second image, and the determination of the attention weight matrix according to the feature images of the two images comprises: determining a first attention weight matrix according to the feature image of the first image and the feature image of the second image; and determining a second attention weight matrix according to the feature image of the third image and the feature image of an image adjacent to the third image; the fusing, based on the attention weight matrix, at least the feature images of the two adjacent images to obtain the target feature image comprises: fusing, based on the first attention weight matrix, the feature image of the first image and the feature image of the second image to obtain a fused feature image; fusing, based on the second attention weight matrix, the feature image of the third image and the fused feature image to obtain the target feature image.

4. The method of claim 1, wherein, The to-be-detected view corresponds to view identification information, and the feature extraction on the to-be-detected view to obtain the feature image of each of the images comprises: perform feature extraction on the to-be-detected view to obtain a view feature image; determine a first projection image according to the view identification information; determine a second projection image according to the view feature image; add the view feature image, the first projection image and the second projection image to obtain a feature image of each of the images.

5. The method of claim 4, wherein, The method further comprises the following steps before the step of performing feature extraction on the to-be-detected view to obtain a feature image of each of the images: obtain a plurality of sample data, each of the sample data comprising a sample image and preset sample information, the sample image comprising at least a first sample image and a second sample image, the sample image being two adjacent sample images in a plurality of sample images obtained by surrounding shooting of a sample object; input the sample image into a preset model to detect sample object information of the sample object, the sample object information comprising category information of the sample object and / or position information of the sample object in the sample image; train the preset model according to the sample object information and the preset sample information until the preset model satisfies a preset training condition to obtain the object detection model. The preset model comprises an initial feature extraction network, an initial attention network and an initial detection network, and the step of inputting the sample image into the preset model to detect sample object information of the sample object comprises the following steps: input the sample image into the preset model to perform feature extraction on the sample image by using the initial feature extraction network to obtain a sample feature image of each of the sample images; 6. The method of claim 1, wherein, input the sample feature image of each of the sample images into the initial attention network to fuse the two sample images based on a sample attention weight matrix to obtain a target sample feature image, the sample attention weight matrix being used to represent the attention degree of a sample image located on the left side to a sample image located on the right side in the two sample images; input the target sample feature image into the initial detection network to detect sample object information of the sample object according to the target sample feature image, the sample object information comprising category information of the sample object and / or position information of the sample object in the sample image. The device comprises: an acquisition module configured to acquire to-be-detected views, the to-be-detected views comprising at least a first image and a second image, the first image and the second image being two adjacent images in a plurality of images obtained by surrounding shooting of a target object; 7. The method of claim 6, wherein, ​ ​ ​ ​ 8. An object detection device, characterized by, ​ ​ An extraction module is configured to perform feature extraction on the to-be-detected view to obtain a feature image of each of the images; A fusion module is configured to fuse at least the feature images of the two images adjacent to each other based on an attention weight matrix to obtain a target feature image, wherein the attention weight matrix is used to represent the attention degree of the image on the left to the image on the right among the two images; A detection module is configured to detect object information of the target object according to the target feature image, wherein the object information includes category information of the target object and / or position information of the target object in the to-be-detected view. The extraction module is specifically configured to: input the to-be-detected image into a pre-trained object detection model, wherein the object detection model includes a feature extraction network, an attention network and a detection network, perform feature extraction on the to-be-detected view by using the feature extraction network to obtain a feature image of each of the images; The fusion module is specifically configured to: input the feature images of the two images adjacent to each other into the attention network to determine the attention weight matrix; fuse at least the feature images of the two images adjacent to each other based on the attention weight matrix to obtain the target feature image; The detection module is specifically configured to: input the target feature image into the detection network to obtain the object information of the target object.

9. An electronic device, comprising: The device includes a processor and a memory having computer program instructions stored therein; the processor implements the object detection method according to any one of claims 1-7 when executing the computer program instructions.

10. A readable storage medium, characterized by, The computer readable storage medium has computer program instructions stored thereon, and the computer program instructions are executed by the processor to implement the object detection method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image detection method and device, electronic equipment and storage medium

    CN114663670A

  • Blurred image correction method and apparatus, computer device, and storage medium

    WO2022142009A1