Multi-view Object Detection or Model Training Method, Apparatus, Electronic Device and Medium

Through the multi-head cross-attention network, the characteristics of each perspective are integrated into the multi-view image, the problem of limited accuracy of object detection of multi-view image in the prior art is solved, and high-accurate object detection is achieved.

CN119850937BActive Publication Date: 2025-07-01ZHEJIANG PECKERAI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510329993.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-01
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing object detection algorithm has limited improvement in detection accuracy in multi-view images.

Method used

By obtaining the acquired images from multiple perspectives of the same scene, feature extraction is performed, visual features of interest and position of interest are obtained, and attention results between each perspective are calculated using a multi-head cross-attention network, and the attention results are fused to generate the multi-head attention results, and the target characteristics are finally updated and the target detection results are determined.

Benefits of technology

It has achieved a significant improvement in object detection accuracy in multi-view images, and enrich image features and improve representation through multi-view features fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850937B_ABST
    Figure CN119850937B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-view object detection or model training method, device, electronic device and medium. The method includes: performing feature extraction processing on acquisition images of at least two views for the same scene to obtain the interested visual features and interested position features corresponding to each of the views; performing linear processing on the interested visual features corresponding to the view according to the interested position features corresponding to the view to obtain an attention key matrix corresponding to the view under at least one attention head; calculating the attention results between every two views; performing fusion according to the attention results between the two views under each of the attention heads to obtain a multi-head attention result of the two views; determining the target features of each of the views; and determining the object detection results of the acquisition images of each of the views for the target features of each of the views. Embodiments of the present invention can make full use of image features of multiple views and greatly improve the accuracy of object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a multi-view object detection or model training method, device, electronic device and medium. Background Art

[0002] Object detection is usually a detection method for detecting whether there is an object of interest in an image.

[0003] Currently, object detection algorithms have high detection accuracy in single-view images, but for two-view or even more-view images, the improvement in detection accuracy is limited. Summary of the Invention

[0004] The present invention provides a multi-view object detection or model training method, device, electronic device and medium, which can make full use of multi-view image features and greatly improve object detection accuracy.

[0005] According to one aspect of the present invention, there is provided a multi-view object detection method, the method comprising:

[0006] Performing feature extraction processing on at least two perspectives of captured images for the same scene to obtain the visual features of interest and the position features of interest corresponding to each of the perspectives;

[0007] For each of the perspectives, determining an attention key matrix corresponding to the perspective under at least one attention head according to the visual features of interest and the position features of interest corresponding to the perspective;

[0008] For each of the attention heads, calculating the attention results between every two perspectives according to the attention key matrices corresponding to each of the perspectives;

[0009] For every two perspectives, fusing the attention results between the two perspectives under each of the attention heads to obtain the multi-head attention results of the two perspectives;

[0010] Determining the target features of each of the perspectives according to the multi-head attention results of each of the two perspectives and the visual features of interest of each of the perspectives;

[0011] Determining the object detection results of the captured images of each of the perspectives according to the target features of each of the perspectives.

[0012] According to one aspect of the present invention, there is provided a method for training a multi-view object detection model, the method comprising:

[0013] Performing feature extraction processing on at least two perspectives of sample images for the same scene through a siamese network feature extraction network in the multi-view object detection model to obtain the visual features of interest and the position features of interest corresponding to each of the perspectives;

[0014] Through the multi-head cross-attention network in the multi-view object detection model, for each of the views, according to the visual features of interest and the location features of interest corresponding to the view, determine the attention key matrix corresponding to the view under at least one attention head;

[0015] Through the multi-head cross-attention network, for each of the attention heads, according to the attention key matrices corresponding to each of the views, calculate the attention results between every two views;

[0016] Through the multi-head cross-attention network, for every two views, fuse the attention results between the two views under each of the attention heads to obtain the multi-head attention results of the two views;

[0017] Through the regression and classification network in the multi-view object detection model, according to the multi-head attention results of each of the two views and the visual features of interest of each of the views, determine the target features of each of the views;

[0018] Through the regression and classification network, for the target features of each of the views, determine the predicted detection results of the sample images of each of the views;

[0019] According to the differences between the standard detection results and the predicted detection results of the sample images of each of the views, adjust the parameters of the siamese network feature extraction network and the multi-head cross-attention network.

[0020] According to another aspect of the present invention, there is provided a multi-view object detection device, the device includes:

[0021] The siamese network feature extraction network in the multi-view object detection model, which is used to perform feature extraction processing on the captured images of at least two views of the same scene to obtain the visual features of interest and the location features of interest corresponding to each of the views;

[0022] The multi-head cross-attention network in the multi-view object detection model, which is used to, for each of the views, according to the visual features of interest and the location features of interest corresponding to the view, determine the attention key matrix corresponding to the view under at least one attention head;

[0023] The multi-head cross-attention network, which is used to, for each of the attention heads, according to the attention key matrices corresponding to each of the views, calculate the attention results between every two views;

[0024] The multi-head cross-attention network, which is used to, for every two views, fuse the attention results between the two views under each of the attention heads to obtain the multi-head attention results of the two views;

[0025] In the regression and classification network of the multi-view object detection model, it is used to determine the object features of each view according to the multi-head attention results of each of the two views and the visual features of interest of each view.

[0026] The regression and classification network is used to determine the object detection results of the acquired images of each view for the object features of each view.

[0027] According to another aspect of the present invention, there is provided a training device for a multi-view object detection model, and the device includes:

[0028] The feature extraction network of the twin network in the multi-view object detection model is used to perform feature extraction processing on sample images of at least two views of the same scene, and obtain the visual features of interest and the location features of interest corresponding to each view.

[0029] The multi-head cross-attention network in the multi-view object detection model is used to determine the attention key matrix corresponding to each view under at least one attention head according to the visual features of interest and the location features of interest corresponding to the view.

[0030] The multi-head cross-attention network is used to calculate the attention results between every two views according to the attention key matrices corresponding to each view for each attention head.

[0031] The multi-head cross-attention network is used to fuse the attention results between the two views under each attention head for every two views, and obtain the multi-head attention results of the two views.

[0032] In the regression and classification network of the multi-view object detection model, it is used to determine the object features of each view according to the multi-head attention results of each of the two views and the visual features of interest of each view.

[0033] The regression and classification network is used to determine the predicted detection results of the sample images of each view for the object features of each view.

[0034] The model training module is used to adjust the parameters of the twin network feature extraction network and the multi-head cross-attention network according to the difference between the standard detection results and the predicted detection results of the sample images of each view.

[0035] According to another aspect of the present invention, there is provided an electronic device, and the electronic device includes:

[0036] At least one processor; and

[0037] A memory communicatively connected to the at least one processor; wherein,

[0038] The memory stores a computer program executable by the at least one processor. When executed by the at least one processor, the computer program enables the at least one processor to execute the multi-view object detection method or the training method of the multi-view object detection model according to any embodiment of the present invention.

[0039] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the multi-view object detection method or the training method of the multi-view object detection model according to any embodiment of the present invention when executed.

[0040] In the technical solution of the embodiment of the present invention, by acquiring acquisition images of multiple perspectives of the same scene, performing feature extraction to obtain the visual features of interest and the position features of interest, and for the visual features of interest and the position features of interest corresponding to each perspective, determining the attention key matrix corresponding to each perspective under each attention head, for each attention head, calculating the attention results of each perspective for other perspectives, and fusing the attention results of different attention heads to obtain the multi-head attention results of each perspective for other perspectives. Finally, based on the multi-head attention results of the perspective for other perspectives, updating the visual features of interest of this perspective, determining the target features of this perspective, and based on the target features, determining the object detection result of this perspective. It is possible to add the scene information of other perspectives to the visual features of interest of each perspective, enrich the image features of each perspective, improve the representativeness of the visual features of interest, and realize the fusion of image data of multiple perspectives, thereby improving the object detection accuracy.

[0041] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 is a flowchart of a multi-view object detection method provided according to an embodiment of the present invention;

[0044] Figure 2 is a flowchart of a multi-view object detection method provided according to an embodiment of the present invention;

[0045] Figure 3 It is a flowchart of a method for training a multi-view object detection model provided according to an embodiment of the present invention;

[0046] Figure 4 It is a scenario diagram of a method for training a multi-view object detection model provided according to an embodiment of the present invention;

[0047] Figure 5 It is a schematic structural diagram of a multi-view object detection device provided according to an embodiment of the present invention;

[0048] Figure 6 It is a schematic structural diagram of a training device for a multi-view object detection model provided according to an embodiment of the present invention;

[0049] Figure 7 It is a schematic structural diagram of an electronic device for implementing the multi-view object detection method or the method for training a multi-view object detection model according to an embodiment of the present invention. Detailed implementation manners

[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0052] Figure 1The flowchart of a multi-view object detection method provided by an embodiment of the present invention. The embodiment of the present invention is applicable to the situation of object detection based on multi-view acquired images. This method can be executed by a multi-view object detection device, which can be implemented in the form of hardware and / or software. The multi-view object detection device can be configured in an electronic device with multi-view object detection function, such as a client or a server. The client can include: mobile phones, tablet computers, laptop computers, desktop computers, wearable devices, etc.

[0053] See Figure 1 the multi-view object detection method shown in

[0054] S101. Perform feature extraction processing on the acquired images of at least two views for the same scene to obtain the interested visual features and interested position features corresponding to each view.

[0055] Among them, the same scene can refer to the same space or the same product. Image acquisition devices with different positions and / or angles can be set, and at least two views of acquired images are collected through at least one image acquisition device. At least two views of acquired images can be collected by using image acquisition devices at different positions, or at least two views of acquired images can be collected by using image acquisition devices at the same position at different angles. One acquired image corresponds to one view. The interested visual feature is used to describe the interested image features in one view. The interested position feature is used to describe the interested position features in one view. Feature extraction processing is performed on the acquired images to determine at least one interested region, and to determine the visual feature and position feature of each interested region. Among them, the visual feature can be a feature vector obtained by extracting the image features of the interested region. The position feature can be the positioning information of the interested region. For example, the position feature includes key point coordinates, width, height, etc. where the interested region is a rectangle, and the key point coordinates can include the upper left vertex coordinates, upper right vertex coordinates, lower left vertex coordinates, or lower right vertex coordinates of the interested region.

[0056] In some embodiments, a siamese network feature extraction module can be used to process the acquired images to obtain the visual features of interest and the position features of interest of at least one region of interest. For example, the siamese network feature extraction module can be a Faster R-CNN (Faster Region-based Convolutional Neural Networks) network. Faster R-CNN includes a parameter-shared Backbone (backbone network), a parameter-independent FPN (Feature Pyramid Networks), and an RPN (Region Proposal Network). In one example, for a certain perspective, the acquired image of this perspective is input into the Backbone for processing, and a feature map of this perspective is output. One perspective corresponds to one Backbone, and the parameters between the Backbones corresponding to different perspectives are shared. Parameter sharing not only reduces the video memory occupancy but also avoids redundant calculations, further improving the computational efficiency of the network. The feature map of this perspective is input into the FPN for processing, and a set F of feature vectors of the regions of interest in the vision is output and used as the visual features of interest of this perspective; and the feature map is input into the RPN for processing, and a set P of the positions of the regions of interest is output and used as the position features of interest of this perspective. Among them, F includes the visual feature vectors of m regions of interest, and P includes the positions of m regions of interest. The visual feature of interest , the position feature of interest , , is the upper left coordinate of this region of interest, is the width of this region of interest, is the height of this region of interest. For example, the visual features of interest F1 and the position features of interest P1 of the first perspective, and the visual features of interest F2 and the position features of interest P2 of the second perspective are obtained.

[0057] S102. For each of the perspectives, according to the visual features of interest and the position features of interest corresponding to the perspective, determine the attention key matrix corresponding to the perspective under at least one attention head.

[0058] Among them, the attention head is used to describe the attention dimension of the model to the key information in the image. The attention key matrix is used to calculate the attention weight to assist the model in focusing on the key information. Based on the Transform algorithm, processing the visual features of interest and the location features of interest can obtain the attention key matrix of at least one attention head. One attention head corresponds to one attention key matrix, and the attention key matrix includes matrices such as Query, Key, and Value.

[0059] In some embodiments, linear transformation can be performed on the visual features of interest and the location features of interest to obtain a set of QKV matrices, which are used as the attention key matrix. When performing different linear transformations, multiple sets of QKV matrices, that is, multiple attention key matrices, can be obtained. One linear transformation corresponds to obtaining the attention key matrix of one attention head.

[0060] In some embodiments, linear transformation can be performed on the visual features of interest to obtain the QKV matrix, and the QKV matrix can be adjusted according to the location features of interest, and the adjusted QKV matrix is used as the attention key matrix.

[0061] S103. For each of the attention heads, according to the attention key matrix corresponding to each of the perspectives, calculate the attention results between every two perspectives.

[0062] Among them, the attention result can refer to the degree of attention to the features of another perspective from one perspective. The attention results between every two perspectives can refer to the attention result of the first perspective to the second perspective and the attention result of the second perspective to the first perspective. Multiple perspectives can be combined in pairs to obtain at least one set of two perspectives.

[0063] S104. For every two perspectives, fuse the attention results between the two perspectives under each of the attention heads to obtain the multi-head attention result of the two perspectives.

[0064] Among them, the multi-head attention result can refer to the degree of attention to the features of another perspective from one perspective in multiple dimensions. Fusing the attention results between two perspectives under each attention head actually means fusing the attention results of multiple attention heads for a set of two perspectives. In some embodiments, a set of two perspectives includes perspective A and perspective B, and there are 3 attention heads for this set of perspectives. The attention results of perspective A to perspective B under the 3 attention heads can be fused to obtain the multi-head attention result of perspective A to perspective B; and the attention results of perspective B to perspective A under the 3 attention heads can be fused to obtain the multi-head attention result of perspective B to perspective A.

[0065] S105. Determine the target feature of each perspective according to the multi-head attention results of the two perspectives and the visual features of interest of each perspective.

[0066] Among them, the target feature can refer to the feature obtained by enhancing the visual feature of interest through the multi-head attention result. The key information of other perspectives is focused on in the target feature, and the key image information of other perspectives that is more representative can be effectively added to the visual feature of this perspective, enriching the visual feature of this perspective. Obtain the fusion result of all the multi-head attention results of a perspective and the visual feature of interest of this perspective as the target feature of this perspective.

[0067] S106. Determine the target detection result of the captured image of each perspective for each target feature of each perspective.

[0068] Among them, the target feature of a perspective can be non-linearly transformed, expressed, and predicted to obtain the target detection result of the captured image of this perspective. The target detection result includes at least one detection box and the classification label of each detection box. In some embodiments, the item being transported can be photographed to obtain the captured image, and the target detection is performed on the captured image. For example, the object of target detection is the item transported on the conveyor belt, and the position of the item is detected for sorting the item, etc. Another example is that the object of target detection is a dangerous object, and it is detected whether there is a dangerous object on the item for security inspection of the item, etc.

[0069] The technical solution of the embodiment of the present invention obtains the captured images of multiple perspectives of the same scene, performs feature extraction to obtain the visual features of interest and the position features of interest, and for the visual features of interest and the position features of interest corresponding to each perspective, determines the attention key matrix corresponding to each perspective under each attention head. For each attention head, calculates the attention results of each perspective for other perspectives, and fuses the attention results of different attention heads to obtain the multi-head attention results of each perspective for other perspectives. Finally, based on the multi-head attention results of the perspective for other perspectives, updates the visual features of interest of this perspective, determines the target feature of this perspective, and based on this target feature, determines the target detection result of this perspective. It can add the scene information of other perspectives to the visual features of interest of each perspective, enrich the image features of each perspective, improve the representativeness of the visual features of interest, realize the fusion of image data of multiple perspectives, and thus improve the accuracy of target detection.

[0070] Figure 2The flowchart of a multi - perspective object detection method provided by an embodiment of the present invention. Based on the above - mentioned embodiment, in the embodiment of the present invention, "calculating the attention result between every two perspectives according to the attention key matrix corresponding to each perspective" is specifically implemented as follows: for every two perspectives, fusing the query matrix in the attention key matrix of the first perspective and the key matrix in the attention key matrix of the second perspective to obtain the weight matrix of the first perspective with respect to the second perspective; according to the weight matrix of the first perspective with respect to the second perspective and the value matrix in the attention key matrix of the second perspective, calculating the attention matrix of the first perspective with respect to the second perspective as the attention result between the two perspectives. It should be noted that for the parts not detailed in the embodiment of the present invention, reference can be made to the descriptions of other embodiments.

[0071] See Figure 2 The multi - perspective object detection method shown, includes:

[0072] S201. Perform feature extraction processing on the collected images of at least two perspectives for the same scene to obtain the corresponding interesting visual features and interesting position features for each perspective.

[0073] S202. For each perspective, determine the attention key matrix corresponding to the perspective under at least one attention head according to the corresponding interesting visual features and interesting position features of the perspective.

[0074] In an optional embodiment, the method of linearly processing the interesting visual features corresponding to the perspective according to the interesting position features corresponding to the perspective to obtain the attention key matrix corresponding to the perspective under at least one attention head includes: linearly processing the interesting visual features corresponding to the perspective to obtain the initial attention matrix of the perspective under at least one attention head; calculating the central coordinates of at least one interesting position feature corresponding to the perspective; splicing the central coordinates of each interesting position feature to obtain the interesting center vector corresponding to the perspective; encoding the interesting center vector corresponding to the perspective to obtain the position encoding corresponding to the perspective; for each attention head, fusing the initial attention matrix of the perspective and the corresponding position encoding to obtain the attention key matrix corresponding to the perspective.

[0075] Among them, the initial attention matrix includes the QKV matrix. The initial attention matrix represents the QKV matrix obtained by processing the interesting visual features of the perspective through the attention mechanism. The center of interest vector is used to describe the position of the region of interest. The center coordinates can refer to the coordinates of the center point of the region of interest, the region of interest is a rectangle, and the center point is the intersection of the diagonals of the rectangle. According to the coordinates, width, and height of the key points in each feature of the interesting position, the center coordinates of the feature of the interesting position are calculated. The center coordinates of different features of the interesting positions of the same perspective are concatenated to obtain the center of interest vector of this perspective. The center of interest vector is encoded to extract features from the center of interest vector. The initial attention matrix is fused with the position encoding to fuse the initial attention matrix of the perspective with the position information to enrich the feature information of the initial attention matrix. In some embodiments, the position encoding is added to the query matrix of the initial attention matrix to update the query matrix, the position encoding is added to the key matrix of the initial attention matrix to update the key matrix, and the value matrix of the initial attention matrix can remain unchanged. The updated initial attention matrix is the attention key matrix.

[0076] In some embodiments, the visual features F1 of the region of interest of the first perspective are used to generate the initial attention matrices of multiple attention heads through multiple linear transformations. An initial attention matrix includes a query, a key, and a value matrix:

[0077] , ,

[0078] where W 1,q , W 1,k and W 1,v are parameters that can be learned during model training and are the weight matrices of the query matrix, the key matrix, and the value matrix respectively.

[0079] Correspondingly, the visual features F2 of the region of interest of the second perspective are used to generate the initial attention matrices of multiple attention heads through multiple linear transformations:

[0080] , ,

[0081] where the position encoding is usually a vector with the same dimension as the feature map (interesting visual features), and it is added element-wise to the original features so that the model can maintain sensitivity to the position while paying attention to the feature content. The center point coordinates of each region of interest are encoded through sine-cosine position encoding:

[0082]

[0083]

[0084]

[0085] Among them, is the central point coordinate of the nth region of interest in the first perspective, is the region of interest center vector, which is a vector formed by concatenating the central point coordinates of all regions of interest in the first perspective, is the position encoding of all regions of interest in the first perspective obtained through linear transformation. d is the number of regions of interest, and d is also the dimension of the region of interest center vector, is the mapping matrix, is the bias term.

[0086] Among them, after the query matrix and the key matrix in the initial attention matrix of the first perspective are respectively accumulated with the position encoding, the updated query matrix and the key matrix are obtained:

[0087] ,

[0088] Similarly, after the query matrix and the key matrix in the initial attention matrix of the second perspective are respectively accumulated with the position encoding, the updated query matrix and the key matrix are obtained:

[0089] ,

[0090] It can be seen that by adding the central coordinate encoding of the region of interest position feature to the initial attention matrix obtained by the linear transformation of the visual feature in the perspective, the position encoding is obtained, realizing the fusion of the visual feature and the position feature. While paying attention to the image feature content, it can maintain the sensitivity to the position, and by encoding the central coordinates to obtain the position encoding, the acquisition operation of the position feature can be simplified, and the position feature that conforms to the geometric feature of the region of interest can be obtained as the fusion object, enabling the natural fusion of the region of interest position feature and the region of interest visual feature. Finally, based on the fused attention key matrix, richer feature information can be extracted, improving the accuracy of object detection.

[0091] S203. For each pair of perspectives, fuse the query matrix in the attention key matrix of the first perspective and the key matrix in the attention key matrix of the second perspective to obtain the weight matrix of the first perspective with respect to the second perspective.

[0092] Among them, in each group of two perspectives, one perspective is determined as the first perspective and the other perspective is determined as the second perspective. The fusion method can be to calculate the dot product of the query matrix and the key matrix. The weight matrix is used to describe the degree of attention of each region of interest in the first perspective to each region of interest in the second perspective.

[0093] In some embodiments, the following formula can be used to calculate the weight matrix of the first perspective on the second perspective :

[0094]

[0095] Wherein, represents the dimension of the key vector, is a scaling factor to avoid the dot product being too large.

[0096] In some embodiments, the weight matrix Attention of the second perspective on the first perspective 2→1 .

[0097] S204. According to the weight matrix of the first perspective on the second perspective and the median matrix in the attention key matrix of the second perspective, calculate the attention matrix of the first perspective on the second perspective as the attention result between the two perspectives.

[0098] Among them, the attention matrix is used to describe the degree of attention to the features of the second perspective starting from the first perspective. The weight matrix of the first perspective on the second perspective and the median matrix in the attention key matrix of the second perspective are fused to obtain the attention matrix.

[0099] In some embodiments, the following formula can be used to perform weighted summation on the value matrix V2 of the second perspective based on the weight matrix Attention 1→2 to obtain the attention matrix :

[0100]

[0101] In some embodiments, the weight matrix of the second perspective on the first perspective .

[0102] S205. For each pair of perspectives, fuse the attention results between the two perspectives under each attention head to obtain the multi-head attention result of the two perspectives.

[0103] In an alternative embodiment, the step of fusing the attention results between the two perspectives under each attention head to obtain the multi-head attention result of the two perspectives includes: fusing the attention matrices of the first perspective to the second perspective under each attention head to obtain the fused attention matrix of the first perspective to the second perspective; wherein, the multi-head attention result of the two perspectives includes the fused attention matrix of the first perspective to the second perspective.

[0104] Specifically, the attention matrices of the first perspective to the second perspective under each attention head are concatenated to obtain the fused attention matrix of the first perspective to the second perspective. In some embodiments, when expanding from a single attention head to multiple single-head attention heads, the attention matrices are concatenated and then through a linear transformation, the final multi-head cross-attention result is obtained. That is, the following formula can be used to calculate the fused attention matrix of the first perspective to the second perspective :

[0105]

[0106] wherein, the multi-head attention result may further include the fused attention matrix of the second perspective to the first perspective .

[0107] It can be seen that by concatenating the attention matrices of the first perspective to the second perspective under each attention head to obtain the fused attention matrix of the first perspective to the second perspective, key information with multiple dimensions can be added to the fused attention matrix, improving the representativeness of the fused attention matrix.

[0108] S206. Determine the target features of each perspective according to the multi-head attention results of the two perspectives and the visually interesting features of each perspective.

[0109] In an alternative embodiment, the step of determining the target features of each perspective according to the multi-head attention results of the two perspectives and the visually interesting features of each perspective includes: for each perspective, updating the visually interesting feature of the perspective according to the fused attention matrix of the two perspectives to which the perspective belongs; for each perspective, performing feature extraction processing on the updated visually interesting feature of the perspective to obtain the target feature of the perspective.

[0110] Among them, for each perspective, the fusion attention matrix of the perspective with respect to each other perspective is used to adjust and update the visually interesting features of the perspective, so as to obtain the updated visually interesting features of the perspective. In some embodiments, the fusion attention matrix of the perspective with respect to each other perspective is superimposed on the visually interesting features to obtain the updated visually interesting features. Non-linear processing is performed on the updated visually interesting features to obtain the target features. For example, the fusion attention matrix of the i-th perspective with respect to the j-th perspective is superimposed on the visually interesting features of the i-th perspective to obtain the updated visually interesting features of the i-th perspective. Wherein, i is not equal to j, i ∈ R, j ∈ R, and R is the set of perspectives.

[0111] In one example, the number of perspectives is two. The visually interesting features of the first perspective and the fusion attention matrix of the first perspective with respect to the second perspective are subjected to residual connection and layer normalization processing to obtain the updated visually interesting features. :

[0112]

[0113] The updated visually interesting features of the first perspective Through a feedforward neural network (FFM), non-linearity and representational ability are added to obtain the target feature Fffm,1.

[0114] It can be seen that through the fusion attention matrices of the two perspectives to which each perspective belongs, the visually interesting features of the perspective are updated, and the feature extraction process is continued to obtain the target feature of the perspective, so as to fully fuse the multi-perspective feature information and improve the accuracy of target detection.

[0115] S207. For the target features of each of the perspectives, determine the target detection results of the captured images of each of the perspectives.

[0116] In the technical solution of the embodiment of the present invention, for every two perspectives, the query matrix of one perspective is subjected to an attention mechanism calculation with the key matrix and value matrix of the other perspective to generate a cross-perspective attention matrix, so as to pay attention to multiple relevant features of the other perspective in one perspective, realizing the full mutual fusion of the features of the two perspectives from a global perspective, enriching the image features of each perspective, increasing the representativeness of the image features of each perspective, and thus performing target detection based on the fused target features, which can improve the accuracy of target detection.

[0117] In a scenario, the number of perspectives is two. Feature extraction processing is performed on the captured images of the two perspectives of the same scenario to obtain the visual features of interest and the location features of interest corresponding to each perspective. For each perspective, linear processing is performed on the visual features of interest corresponding to that perspective to obtain the initial attention matrix of that perspective. For each perspective, the central coordinates of at least one location feature of interest corresponding to that perspective are calculated and concatenated to obtain the center vector of interest corresponding to the perspective. The center vector of interest corresponding to that perspective is encoded to obtain the position encoding corresponding to that perspective. For each attention head, the initial attention matrix corresponding to that perspective and the corresponding position encoding are fused to obtain the attention key matrix corresponding to the perspective. For each attention head, the query matrix in the attention key matrix of the first perspective and the key matrix in the attention key matrix of the second perspective are fused to obtain the weight matrix of the first perspective with respect to the second perspective. According to the weight matrix of the first perspective with respect to the second perspective and the value matrix in the attention key matrix of the second perspective, the attention matrix of the first perspective with respect to the second perspective is calculated. The attention matrices of the first perspective with respect to the second perspective under each attention head are fused to obtain the fused attention matrix of the first perspective with respect to the second perspective. According to the visual features of interest of the first perspective and the fused attention matrix of the first perspective with respect to the second perspective, the visual features of interest of the first perspective are updated. The updated visual features of interest of the first perspective are processed to obtain the target features of the first perspective. According to the target features of the first perspective, the target detection result of the captured image of the first perspective is determined. For each attention head, the query matrix in the attention key matrix of the second perspective and the key matrix in the attention key matrix of the first perspective are fused to obtain the weight matrix of the second perspective with respect to the first perspective. According to the weight matrix of the second perspective with respect to the first perspective and the value matrix in the attention key matrix of the first perspective, the attention matrix of the second perspective with respect to the first perspective is calculated. The attention matrices of the second perspective with respect to the first perspective under each attention head are fused to obtain the fused attention matrix of the second perspective with respect to the first perspective. According to the visual features of interest of the second perspective and the fused attention matrix of the second perspective with respect to the first perspective, the visual features of interest of the second perspective are updated. The updated visual features of interest of the second perspective are processed to obtain the target features of the second perspective. According to the target features of the second perspective, the target detection result of the captured image of the second perspective is determined.

[0118] Figure 3The flowchart of a training method for a multi - perspective object detection model provided by an embodiment of the present invention. The embodiment of the present invention is applicable to the situation of training a multi - perspective object detection model. This method can be executed by a training device for a multi - perspective object detection model. The training device for a multi - perspective object detection model can be implemented in the form of hardware and / or software, and can be configured in an electronic device that undertakes the training function of the multi - perspective object detection model, such as a client or a server. The client can include: mobile phones, tablet computers, laptop computers, desktop computers, wearable devices, etc.

[0119] See Figure 3 The training method of the multi - perspective object detection model shown, includes:

[0120] S301. Through the siamese network feature extraction network in the multi - perspective object detection model, perform feature extraction processing on sample images of at least two perspectives for the same scene, and obtain the interested visual features and interested position features corresponding to each of the perspectives.

[0121] Among them, multi - perspective image acquisition is performed on the same scene to obtain sample images. In some embodiments, images of multiple items in multiple perspectives can be collected first as sample images, and images of the same item in multiple perspectives are labeled with the same identification information. At the same time, items are labeled in each sample image, and the target box of the labeled item is used as the standard detection result. The sample images are divided into a training set and a test set according to a ratio of 4:1. For example, the item is a logistics parcel. In the embodiment of the present invention, the sample images are images of the same item in multiple perspectives.

[0122] Among them, the siamese network feature extraction network is used to extract the interested visual features and interested position features.

[0123] S302. Through the multi - head cross - attention network in the multi - perspective object detection model, for each of the perspectives, determine the attention key matrix corresponding to the perspective under at least one attention head according to the interested visual features and interested position features corresponding to the perspective.

[0124] Among them, the multi - head cross - attention network is used to generate an initial attention matrix, superimpose position encoding to obtain the attention key matrix, and realize the fusion of the attention key matrices across perspectives to obtain the target features of each perspective, and detect the target detection result based on the target features.

[0125] S303. Through the multi - head cross - attention network, for each of the attention heads, calculate the attention results between every two perspectives according to the attention key matrices corresponding to each of the perspectives.

[0126] S304. Through the multi-head cross-attention network, for each pair of perspectives, fuse according to the attention results between the two perspectives under each attention head to obtain the multi-head attention results of the two perspectives.

[0127] S305. Through the multi-head cross-attention network, determine the target features of each perspective according to the multi-head attention results of each pair of perspectives and the visually interesting features of each perspective.

[0128] S306. Through the multi-head cross-attention network, for the target features of each perspective, determine the predicted detection results of the sample images of each perspective.

[0129] Among them, the predicted detection result is the target detection result output by the multi-head cross-attention network.

[0130] S307. According to the differences between the standard detection results and the predicted detection results of the sample images of each perspective, adjust the parameters of the siamese network feature extraction network and the multi-head cross-attention network.

[0131] Among them, the standard detection result is the correct target detection result in the sample image, which can be the target detection result obtained by manually annotating the sample image. For each perspective, calculate the difference between the standard detection result and the predicted detection result, calculate the loss value of this perspective according to the difference of this perspective, accumulate the loss values of all perspectives, and adjust the parameters of the siamese network feature extraction network and the multi-head cross-attention network. It can be determined that the training of the siamese network feature extraction network and the multi-head cross-attention network is completed when the loss value is the smallest, the loss value converges, or the number of iterations is greater than or equal to the number threshold, that is, the training of the multi-perspective target detection model is completed. The multi-perspective target detection model can be released for application.

[0132] In some embodiments, the loss value L can be calculated based on the following formula:

[0133]

[0134] where L i,cls is the classification loss value of the i-th perspective, and L i,reg is the regression loss value of the i-th perspective. The above formula is used to calculate the loss value L when the number of perspectives is two. Among them, the classification loss value is calculated using cross-entropy loss. The regression loss value is calculated using DIoU (Distance Intersection over Union). is the weight coefficient.

[0135] The technical solution of the embodiment of the present invention obtains sample images of multiple perspectives of the same scene, extracts features, obtains visual features of interest and position features of interest, and for the visual features of interest and position features of interest corresponding to each perspective, determines the attention key matrix corresponding to each perspective under each attention head. For each attention head, calculates the attention results of each perspective for other perspectives, and fuses the attention results of different attention heads to obtain the multi-head attention results of each perspective for other perspectives. Finally, based on the multi-head attention results of the perspective for other perspectives, updates the visual features of interest of this perspective, determines the target features of this perspective, and based on these target features, determines the target detection result of this perspective. It can add the scene information of other perspectives to the visual features of interest of each perspective, enrich the image features of each perspective, improve the representativeness of the visual features of interest, and achieve effective fusion of image data of multiple perspectives. Through cross-perspective feature fusion, the model can learn the common feature information between different perspectives, has stronger generalization ability, and improves the target detection accuracy of the model.

[0136] In a scene, such as Figure 4 shown, two perspectives are used to collect images of the same scene, and a sample image corresponding to the perspective is obtained.

[0137] The first stage (blue line) is the Faster R-CNN Learnable Proposals: Through the Faster R-CNN network, the collected images are processed to obtain visual features of the region of interest and position features of the region of interest. Based on the Region of Interest Pool method, through two Backbones with shared parameters, the features in the sample images of the two perspectives are respectively extracted to obtain the feature maps of each perspective. For each perspective, the feature map of this perspective is input into the FPN with independent parameters for processing, and a set of feature vectors F of the visual features of the region of interest is output and used as the visual features of interest of this perspective; and the feature map is input into the RPN for processing, and a set of positions of the region of interest is output and used as the position features of interest of this perspective.

[0138] The second stage (red line): Through the multi-head cross-attention network, the interested visual features and interested position features are processed to obtain a fused attention matrix, which is used as the multi-head attention result. Specifically, for each perspective, the interested visual features corresponding to this perspective are linearly processed to obtain the Q matrix, K matrix, and V matrix, which serve as the initial attention matrices for this perspective. For each perspective, the central coordinates of at least one interested position feature corresponding to this perspective are calculated, and the obtained interested center vectors corresponding to the perspective are concatenated; the interested center vectors corresponding to this perspective are encoded to obtain the position encoding corresponding to this perspective; for each attention head, the initial attention matrix corresponding to this perspective and the corresponding position encoding are fused to obtain the attention key matrix corresponding to this perspective, namely Figure 4 Q1, K1, V1, Q2, K2, and V2 among them. Using the CrossAttention network, calculate the weight matrix of the first perspective on the second perspective , based on the weight matrix weighted sum the value matrix V2 of the second perspective to obtain the attention matrix . Concatenate the attention matrices, and then through a linear transformation, obtain the final multi-head cross-attention result, that is, calculate the fused attention matrix of the first perspective on the second perspective , and process the interested visual features of the first perspective and the fused attention matrix of the first perspective on the second perspective through residual connection and layer normalization (Add Norm) to obtain the interested visual features of the first perspective , thereby realizing the acquisition of cross-view features. Similarly, calculate the interested visual features of the second perspective .

[0139] The third stage (green line): Through the regression classification network, the interested visual features are processed to obtain the object detection results, that is, the object bounding boxes and the classification labels of each object bounding box. Through a feedforward neural network (Feedforward NeuralNetwork), perform a non-linear transformation on the interested visual features of the first perspective to obtain the object feature F of the first perspective ffm,1 . Through residual connection and layer normalization (Add Norm), based on the interested visual features of the first perspective and the object feature F ffm,1 , determine the object detection results, that is, the object bounding boxes (Boxes) and classification labels (Labels) in the captured image of the first perspective. Similarly, detect the object bounding boxes and classification labels in the captured image of the second perspective

[0140] Among them, Figure 4In the collected images, the target boxes of different colors represent different classification labels. For example, green boxes, blue boxes, and red boxes represent objects of different categories. Through experiments, algorithm comparison experiments of single-view and dual-view are carried out on the dataset. When the false positive ratio on each picture is controlled to be 0.5 and 1 respectively, the recall rates are calculated as shown in Table 1 below:

[0141] Table 1

[0142] Viewpoint R@0.5 R@1 Single Viewpoint 1 0.9465 0.9551 Single Viewpoint 2 0.9186 0.9306 Dual Viewpoint 0.9568 0.9646

[0143] Judging from the experimental results, the multi-view object detection method proposed in the embodiment of the present invention improves the detection accuracy and reduces false positives. Among them, false positive (False Positive, FP) refers to the situation where the model predicts the target box as a positive class (positive), but the classification label of the actual target box is a negative class.

[0144] The embodiment of the present invention adopts a twin network feature extraction network to control the sharing of some parameters, reduces the video memory occupation and the amount of calculation. By using the high-dimensional feature after linear transformation of the central coordinates of the region of interest as the position encoding, compared with the complex spatial information decomposed by sine-cosine encoding, the central point coordinates are more in line with the geometric features of the region of interest, making the position encoding more naturally fused with the features of the region of interest; through the multi-head cross-attention mechanism, multi-view feature information is fully fused to achieve efficient cross-view feature fusion and improve the accuracy of object detection; through cross-view feature fusion and position encoding, the model can learn the common feature information between different views, has stronger generalization ability, realizes flexible multi-view adaptability, and improves the object detection accuracy of the model.

[0145] Figure 5 It is a schematic structural diagram of a multi-view object detection device provided by an embodiment of the present invention. The embodiment of the present invention is applicable to the situation of object detection based on multi-view collected images. The device can execute the multi-view object detection method. The device can be implemented in the form of hardware and / or software, and the device can be configured in an electronic device carrying the multi-view object detection function.

[0146] See Figure 5 The multi-view object detection device shown in the figure, including:

[0147] The twin network feature extraction network 501 in the multi-view object detection model is used to perform feature extraction processing on the collected images of at least two views of the same scene, and obtain the corresponding visual features of interest and position features of interest for each of the views;

[0148] In the multi-view object detection model, the multi-head cross-attention network 502 is used to determine, for each of the views, an attention key matrix corresponding to the view under at least one attention head according to the visual features of interest and the location features of interest corresponding to the view.

[0149] The multi-head cross-attention network 502 is used to calculate, for each of the attention heads, the attention result between every two views according to the attention key matrices corresponding to the views.

[0150] The multi-head cross-attention network 502 is used to fuse, for every two views, the attention results between the two views under each of the attention heads to obtain the multi-head attention result of the two views.

[0151] In the multi-view object detection model, the regression and classification network 503 is used to determine the target features of each view according to the multi-head attention results of each two views and the visual features of interest of each view.

[0152] The regression and classification network 503 is used to determine the object detection result of the captured image of each view for the target features of each view.

[0153] The technical solution of the embodiment of the present invention can add the scene information of other views to the visual features of interest of each view, enrich the image features of each view, improve the representativeness of the visual features of interest, realize the fusion of image data of multiple views, and thus improve the object detection accuracy by acquiring the captured images of multiple views of the same scene, performing feature extraction to obtain the visual features of interest and the location features of interest, determining, for the visual features of interest and the location features of interest corresponding to each view, the attention key matrix corresponding to each view under each attention head, calculating, for each attention head, the attention result of each view for other views, fusing the attention results of different attention heads to obtain the multi-head attention result of each view for other views, finally updating the visual features of interest of the view based on the multi-head attention result of the view for other views, determining the target features of the view, and determining the object detection result of the view based on the target features.

[0154] Optionally, the multi-head cross-attention network 502 is specifically used for:

[0155] For every two views, fuse the query matrix in the attention key matrix of the first view and the key matrix in the attention key matrix of the second view to obtain the weight matrix of the first view for the second view.

[0156] Calculate the attention matrix of the first perspective to the second perspective based on the weight matrix of the second perspective with respect to the first perspective and the median matrix in the attention key matrix of the second perspective, as the attention result between the two perspectives.

[0157] Optionally, the multi-head cross-attention network 502 is specifically configured to:

[0158] Fuse the attention matrices of the first perspective to the second perspective under each attention head to obtain the fused attention matrix of the first perspective to the second perspective; wherein, the multi-head attention result between the two perspectives includes the fused attention matrix of the first perspective to the second perspective.

[0159] Optionally, the regression classification network 503 is specifically configured to:

[0160] For each of the perspectives, update the interesting visual features of the perspective according to the fused attention matrix of the two perspectives to which the perspective belongs;

[0161] For each of the perspectives, perform feature extraction processing on the updated interesting visual features of the perspective to obtain the target features of the perspective.

[0162] Optionally, the multi-head cross-attention network 502 is specifically configured to:

[0163] Perform linear processing on the interesting visual features corresponding to the perspective to obtain the initial attention matrix of the perspective under at least one attention head;

[0164] Calculate the central coordinates of at least one interesting position feature corresponding to the perspective;

[0165] Concatenate the central coordinates of each interesting position feature to obtain the interesting center vector corresponding to the perspective;

[0166] Encode the interesting center vector corresponding to the perspective to obtain the position encoding corresponding to the perspective;

[0167] For each attention head, fuse the initial attention matrix of the perspective and the corresponding position encoding to obtain the attention key matrix corresponding to the perspective.

[0168] The multi-perspective object detection device provided by the embodiments of the present invention can execute the multi-perspective object detection method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0169] Figure 6The figure is a schematic structural diagram of a training device for a multi-view object detection model provided by an embodiment of the present invention. The embodiment of the present invention is applicable to the situation of training a multi-view object detection model. The device can execute the training method of the multi-view object detection model. The device can be implemented in the form of hardware and / or software, and can be configured in an electronic device carrying the training function of the multi-view object detection model.

[0170] Referring to Figure 6 the training device for the multi-view object detection model shown in the figure, includes:

[0171] The siamese network feature extraction network 601 in the multi-view object detection model is used to perform feature extraction processing on sample images of at least two views for the same scene, and obtain the interested visual features and interested position features corresponding to each of the views;

[0172] The multi-head cross-attention network 602 in the multi-view object detection model is used to, for each of the views, determine an attention key matrix corresponding to the view under at least one attention head according to the interested visual features and interested position features corresponding to the view;

[0173] The multi-head cross-attention network 602 is used to, for each of the attention heads, calculate the attention results between every two views according to the attention key matrices corresponding to each of the views;

[0174] The multi-head cross-attention network 602 is used to, for every two views, fuse the attention results between the two views under each of the attention heads to obtain the multi-head attention results of the two views;

[0175] The regression and classification network 603 in the multi-view object detection model is used to determine the target features of each of the views according to the multi-head attention results of each of the two views and the interested visual features of each of the views;

[0176] The regression and classification network 603 is used to, for the target features of each of the views, determine the predicted detection results of the sample images of each of the views;

[0177] The model training module 604 is used to adjust the parameters of the siamese network feature extraction network and the multi-head cross-attention network according to the difference between the standard detection results and the predicted detection results of the sample images of each of the views.

[0178] The technical solution of the embodiment of the present invention obtains sample images from multiple perspectives of the same scene, extracts features, and obtains visual features of interest and position features of interest. For the visual features of interest and position features of interest corresponding to each perspective, an attention key matrix corresponding to each perspective under each attention head is determined. For each attention head, the attention results of each perspective on other perspectives are calculated, and the attention results of different attention heads are fused to obtain the multi-head attention results of each perspective on other perspectives. Finally, based on the multi-head attention results of the perspective on other perspectives, the visual features of interest of this perspective are updated, the target features of this perspective are determined, and based on the target features, the target detection result of this perspective is determined. It is possible to add the scene information of other perspectives to the visual features of interest of each perspective, enrich the image features of each perspective, improve the representativeness of the visual features of interest, and achieve effective fusion of image data from multiple perspectives. Through cross-perspective feature fusion, the model learns the common feature information between different perspectives, has stronger generalization ability, and improves the target detection accuracy of the model.

[0179] The training device of the multi-perspective target detection model provided by the embodiment of the present invention can execute the training method of the multi-perspective target detection model provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0180] In the technical solution of the embodiment of the present invention, the acquisition, storage, and application of the involved data all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0181] Figure 7 FIG. shows a schematic structural diagram of an electronic device 700 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0182] As Figure 7As shown, the electronic device 700 includes at least one processor 701 and a memory communicatively connected to the at least one processor 701, such as read-only memory (ROM) 702, random access memory (RAM) 703, etc. The memory stores a computer program executable by the at least one processor. The processor 701 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0183] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0184] The processor 701 can be various general and / or special processing components with processing and computing capabilities. Some examples of the processor 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 701 executes the various methods and processes described above, such as a multi-view object detection method or a training method for a multi-view object detection model.

[0185] In some embodiments, the multi-view object detection method or the training method for the multi-view object detection model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the processor 701, one or more steps of the multi-view object detection method or the training method for the multi-view object detection model described above can be executed. Alternatively, in other embodiments, the processor 701 can be configured to execute the multi-view object detection method or the training method for the multi-view object detection model in any other appropriate way (e.g., by means of firmware).

[0186] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0187] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0188] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0190] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0191] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS (Virtual Private Server) services.

[0192] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0193] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-view target detection method, characterized in that: The method comprises: Through the twin network feature extraction network, feature extraction processing is performed on the collected images of at least two perspectives of the same scene to obtain the visual features of interest and the position features of interest corresponding to each of the perspectives; the twin network feature extraction network includes a backbone network with shared parameters, a feature pyramid network with independent parameters, and a region candidate network, and different perspectives correspond to different backbone networks and share parameters; the visual features of interest include a feature vector of interest for at least one region of interest, and the position features of interest include the positions of each of the regions of interest; For each of the perspectives, linearly process the visual features of interest corresponding to the perspective to obtain an initial attention matrix of the perspective under at least one attention head; Calculate the center coordinates of at least one interesting position feature corresponding to the viewing angle; The center coordinates of each of the features at the position of interest are spliced ​​to obtain a center vector of interest corresponding to the viewing angle; Encoding a center of interest vector corresponding to the viewing angle to obtain a position code corresponding to the viewing angle; For each of the attention heads, add the query matrix and the key matrix in the initial attention matrix of the perspective to the corresponding position code element by element to obtain the attention key matrix corresponding to the perspective; For each attention head, the query matrix in the attention key matrix of the first perspective and the key matrix in the attention key matrix of the second perspective are fused to obtain the weight matrix of the first perspective to the second perspective; Calculate the attention matrix of the first perspective to the second perspective according to the weight matrix of the first perspective to the second perspective and the median matrix of the attention key matrix of the second perspective; The attention matrices of the first perspective to the second perspective under each attention head are concatenated and linearly transformed to obtain the fused attention matrix of the first perspective to the second perspective; The visual features of interest from the first perspective and the fusion attention matrix of the first perspective to the second perspective are processed by residual connection and layer normalization to obtain updated visual features of interest from the first perspective; Processing the updated visual features of interest from the first perspective to obtain target features from the first perspective; Determine the target detection result of the captured image of the first perspective according to the target feature of the first perspective; For each attention head, the query matrix in the attention key matrix of the second perspective and the key matrix in the attention key matrix of the first perspective are fused to obtain the weight matrix of the second perspective to the first perspective; Calculate the attention matrix of the second perspective to the first perspective according to the weight matrix of the second perspective to the first perspective and the median matrix of the attention key matrix of the first perspective; The attention matrices of the second perspective to the first perspective under each attention head are concatenated and linearly transformed to obtain the fused attention matrix of the second perspective to the first perspective; The visual features of interest of the second perspective and the fusion attention matrix of the second perspective to the first perspective are processed by residual connection and layer normalization to obtain the updated visual features of interest of the second perspective; Processing the updated visual features of interest from the second perspective to obtain target features from the second perspective; According to the target feature of the second viewing angle, the target detection result of the captured image of the second viewing angle is determined.

2. A training method for a multi-view target detection model, characterized in that: The method comprises: Through the twin network feature extraction network in the multi-view target detection model, feature extraction processing is performed on sample images of at least two viewpoints of the same scene to obtain the visual features of interest and the position features of interest corresponding to each of the viewpoints; the twin network feature extraction network includes a backbone network with shared parameters, a feature pyramid network with independent parameters, and a region candidate network. Different viewpoints correspond to different backbone networks and share parameters; the visual features of interest include a feature vector of interest for at least one region of interest, and the position features of interest include the positions of each of the regions of interest; Through the multi-head cross-attention network in the multi-view target detection model, for each of the viewpoints, linear processing is performed on the visual features of interest corresponding to the viewpoint to obtain the initial attention matrix of the viewpoint under at least one attention head; the center coordinates of at least one position feature of interest corresponding to the viewpoint are calculated; the center coordinates of each position feature of interest are spliced ​​to obtain the center vector of interest corresponding to the viewpoint; the center vector of interest corresponding to the viewpoint is encoded to obtain the position encoding corresponding to the viewpoint; for each of the attention heads, the query matrix and the key matrix in the initial attention matrix of the viewpoint are added element by element to the corresponding position encoding to obtain the attention key matrix corresponding to the viewpoint; Through the multi-head cross attention network, for each of the attention heads, the query matrix in the attention key matrix of the first perspective and the key matrix in the attention key matrix of the second perspective are fused to obtain a weight matrix of the first perspective to the second perspective; according to the weight matrix of the first perspective to the second perspective and the median matrix of the attention key matrix of the second perspective, the attention matrix of the first perspective to the second perspective is calculated; the query matrix in the attention key matrix of the second perspective and the key matrix in the attention key matrix of the first perspective are fused to obtain a weight matrix of the second perspective to the first perspective; according to the weight matrix of the second perspective to the first perspective and the median matrix of the attention key matrix of the first perspective, the attention matrix of the second perspective to the first perspective is calculated; Through the multi-head cross attention network, the attention matrices of the first perspective to the second perspective under each attention head are spliced ​​and linearly transformed to obtain a fused attention matrix of the first perspective to the second perspective; the attention matrices of the second perspective to the first perspective under each attention head are spliced ​​and linearly transformed to obtain a fused attention matrix of the second perspective to the first perspective; Through the regression classification network in the multi-view target detection model, the visual features of interest of the first view and the fusion attention matrix of the first view to the second view are subjected to residual connection and layer normalization processing to obtain updated visual features of interest of the first view; the updated visual features of interest of the first view are processed to obtain target features of the first view; the visual features of interest of the second view and the fusion attention matrix of the second view to the first view are subjected to residual connection and layer normalization processing to obtain updated visual features of interest of the second view; the updated visual features of interest of the second view are processed to obtain target features of the second view; Determine the predicted detection results of the sample images of each of the viewing angles according to the target features of each of the viewing angles through the regression classification network; According to the difference between the standard detection results and the predicted detection results of the sample images of each perspective, the parameters of the twin network feature extraction network and the multi-head cross attention network are adjusted.

3. A multi-view target detection device, characterized in that: The device comprises: The twin network feature extraction network in the multi-view target detection model is used to perform feature extraction processing on the collected images of at least two viewpoints of the same scene to obtain the visual features of interest and the position features of interest corresponding to each of the viewpoints; the twin network feature extraction network includes a backbone network with shared parameters, a feature pyramid network with independent parameters, and a region candidate network. Different viewpoints correspond to different backbone networks and share parameters; the visual features of interest include a feature vector of interest for at least one region of interest, and the position features of interest include the positions of each of the regions of interest; The multi-head cross-attention network in the multi-view target detection model is used to perform linear processing on the visual features of interest corresponding to each of the view angles to obtain the initial attention matrix of the view angle under at least one attention head; calculate the center coordinates of at least one position feature of interest corresponding to the view angle; splice the center coordinates of each position feature of interest to obtain the center vector of interest corresponding to the view angle; encode the center vector of interest corresponding to the view angle to obtain the position code corresponding to the view angle; for each of the attention heads, add the query matrix and the key matrix in the initial attention matrix of the view angle to the corresponding position code element by element to obtain the attention key matrix corresponding to the view angle; The multi-head cross attention network is used to fuse the query matrix in the attention key matrix of the first perspective and the key matrix in the attention key matrix of the second perspective for each of the attention heads to obtain a weight matrix of the first perspective to the second perspective; calculate the attention matrix of the first perspective to the second perspective according to the weight matrix of the first perspective to the second perspective and the median matrix of the attention key matrix of the second perspective; fuse the query matrix in the attention key matrix of the second perspective and the key matrix in the attention key matrix of the first perspective to obtain a weight matrix of the second perspective to the first perspective; calculate the attention matrix of the second perspective to the first perspective according to the weight matrix of the second perspective to the first perspective and the median matrix of the attention key matrix of the first perspective; The multi-head cross attention network is used to splice and linearly transform the attention matrices of the first perspective to the second perspective under each attention head to obtain a fused attention matrix of the first perspective to the second perspective; splice and linearly transform the attention matrices of the second perspective to the first perspective under each attention head to obtain a fused attention matrix of the second perspective to the first perspective; The regression classification network in the multi-view target detection model is used to perform residual connection and layer normalization processing on the visual features of interest of the first view and the fusion attention matrix of the first view to the second view to obtain the updated visual features of interest of the first view; process the updated visual features of interest of the first view to obtain the target features of the first view; perform residual connection and layer normalization processing on the visual features of interest of the second view and the fusion attention matrix of the second view to the first view to obtain the updated visual features of interest of the second view; process the updated visual features of interest of the second view to obtain the target features of the second view; The regression classification network is used to determine the target detection result of the captured image of the first perspective according to the target features of the first perspective; and to determine the target detection result of the captured image of the second perspective according to the target features of the second perspective.

4. A training device for a multi-view target detection model, characterized in that: The device comprises: The twin network feature extraction network in the multi-view target detection model is used to perform feature extraction processing on sample images of at least two viewpoints of the same scene to obtain the visual features of interest and the position features of interest corresponding to each of the viewpoints; the twin network feature extraction network includes a backbone network with shared parameters, a feature pyramid network with independent parameters, and a region candidate network, and different viewpoints correspond to different backbone networks and share parameters; the visual features of interest include a feature vector of interest of at least one region of interest, and the position features of interest include the positions of each of the regions of interest; The multi-head cross-attention network in the multi-view target detection model is used to perform linear processing on the visual features of interest corresponding to each of the view angles to obtain the initial attention matrix of the view angle under at least one attention head; calculate the center coordinates of at least one position feature of interest corresponding to the view angle; splice the center coordinates of each position feature of interest to obtain the center vector of interest corresponding to the view angle; encode the center vector of interest corresponding to the view angle to obtain the position code corresponding to the view angle; for each of the attention heads, add the query matrix and the key matrix in the initial attention matrix of the view angle to the corresponding position code element by element to obtain the attention key matrix corresponding to the view angle; The multi-head cross attention network is used to fuse the query matrix in the attention key matrix of the first perspective and the key matrix in the attention key matrix of the second perspective for each of the attention heads to obtain a weight matrix of the first perspective to the second perspective; calculate the attention matrix of the first perspective to the second perspective according to the weight matrix of the first perspective to the second perspective and the median matrix of the attention key matrix of the second perspective; fuse the query matrix in the attention key matrix of the second perspective and the key matrix in the attention key matrix of the first perspective to obtain a weight matrix of the second perspective to the first perspective; calculate the attention matrix of the second perspective to the first perspective according to the weight matrix of the second perspective to the first perspective and the median matrix of the attention key matrix of the first perspective; The multi-head cross attention network is used to splice and linearly transform the attention matrices of the first perspective to the second perspective under each attention head to obtain a fused attention matrix of the first perspective to the second perspective; splice and linearly transform the attention matrices of the second perspective to the first perspective under each attention head to obtain a fused attention matrix of the second perspective to the first perspective; The regression classification network in the multi-view target detection model is used to perform residual connection and layer normalization processing on the visual features of interest of the first view and the fusion attention matrix of the first view to the second view to obtain the updated visual features of interest of the first view; process the updated visual features of interest of the first view to obtain the target features of the first view; perform residual connection and layer normalization processing on the visual features of interest of the second view and the fusion attention matrix of the second view to the first view to obtain the updated visual features of interest of the second view; process the updated visual features of interest of the second view to obtain the target features of the second view; The regression classification network is used to determine the predicted detection results of the sample images of each of the viewing angles according to the target features of each of the viewing angles; The model training module is used to adjust the parameters of the twin network feature extraction network and the multi-head cross attention network according to the difference between the standard detection results and the predicted detection results of the sample images of each perspective.

5. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multi-view target detection method described in claim 1, or the training method of the multi-view target detection model described in claim 2.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the processor to implement the multi-view target detection method described in claim 1, or the training method of the multi-view target detection model described in claim 2 when executed.

Citation Information

Patent Citations

  • Cross-view-angle associated double-unmanned-aerial-vehicle cooperative target detection method

    CN118429620A