Detection model training method and detection method based on heterogeneous modal data fusion

By integrating video image data and camera pose parameters in mobile object detection, the problem of detecting false alarms and missed alarms in complex backgrounds in the prior art is solved, and higher detection accuracy and robustness are achieved.

CN116844082BActive Publication Date: 2025-06-06HANGZHOU EBOYLAMP ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310662042.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-06
Publication Date
2025-06-06
Estimated Expiration
2043-06-06

AI Technical Summary

Technical Problem

Existing mobile object detection methods are prone to false alarms and missed alarms in complex backgrounds, and are difficult to adapt to environmental impacts such as lighting changes, occlusion, and camera shake.

Method used

The detection model training method based on heterogeneous mode data fusion is adopted to fuse video image data and camera pose parameters through deep learning to realize mobile object detection, improving the robustness and accuracy of the model.

Benefits of technology

It effectively solves the false alarm and missed alarm problems of mobile object detection in complex backgrounds, improves the accuracy of detection results and the robustness of the model, and is suitable for detection scenarios in static and dynamic backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844082B_ABST
    Figure CN116844082B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and specifically to a detection model training method and a detection method based on heterogeneous modal data fusion. The method comprises the following steps: A, obtaining a training video set and storing it as a label file; B, extracting image features in the training video through an image feature extractor to obtain a first fusion feature, and extracting camera posture parameter features in the training video through a text feature extractor to obtain a second fusion feature; C, realizing the fusion of the first fusion feature and the second fusion feature through a heterogeneous modal data fusion device to obtain an output vector; D, training the image feature extractor, the text feature extractor and the heterogeneous modal data fusion device according to the obtained training video set, the label file and the output vector, and repeatedly iterating and executing steps B to C until the loss function converges to a preset value to obtain a trained detection model, so as to realize mobile target detection and improve the robustness and accuracy of the mobile target detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a detection model training method and a detection method based on heterogeneous modal data fusion. Background Art

[0002] Moving target detection is one of the basic research directions of computer vision. Its purpose is to remove the long-term unchanged background information from a large amount of video sequence information and retain the information of moving objects. Although target detection has been widely used in the field of computer vision and has achieved certain results. However, target detection can only detect targets of preset categories. Once the usage scenario is changed, retraining is required to obtain better detection results. In addition, the existing target detection algorithm can only distinguish the position of a specific target in the image, but cannot determine the motion state of the target, that is, whether it is stationary or moving, making the target detection algorithm unusable in some special application scenarios, such as smart security scenarios such as regional intrusion detection.

[0003] There are some existing moving target detection methods, but there are still some problems. When the video background is complex, it will make it difficult to detect the moving target. The classic moving target detection method is difficult to cope with complex monitoring environments, such as lighting changes, occlusions, camera shake, shadows, etc., which will cause a large number of false positives and missed positives. Summary of the invention

[0004] In view of the above-mentioned problems existing in the prior art, the present invention provides a detection model training method and a detection method based on heterogeneous modal data fusion, which can abandon the traditional background modeling method, take deep learning as the technical basis, fuse heterogeneous modal data such as video and camera pose parameters, realize mobile target detection, and improve the robustness and accuracy of the mobile target detection model.

[0005] A detection model training method based on heterogeneous modal data fusion, comprising the steps of:

[0006] A, obtain the training video set and store it as a label file;

[0007] B, extracting image features in the training video through an image feature extractor to obtain a first fusion feature, and extracting camera pose parameter features in the training video through a text feature extractor to obtain a second fusion feature;

[0008] C, through the heterogeneous modality data fuser, the first fusion feature and the second fusion feature are fused to obtain an output vector;

[0009] D. According to the acquired training video set, label file and output vector, the image feature extractor, text feature extractor and heterogeneous modal data fusion device are trained, the loss function is calculated, and steps B to C are iterated repeatedly until the loss function converges to a preset value, so as to obtain a detection model including the trained image feature extractor, text feature extractor and heterogeneous modal data fusion device.

[0010] As a preferred implementation, in step A, storing as a label file specifically includes: processing the acquired training video, extracting two consecutive frames of images in the training video as image pairs, marking multiple key feature points in each frame, and calculating the homography transformation matrix between the image pairs based on the feature points to obtain a label file, wherein the label file includes multiple key feature points and a homography transformation matrix.

[0011] As a preferred implementation, in step B, the image feature extractor includes a first convolution block, a second convolution block, a third convolution block and a feature fusion block, and the feature fusion block is connected to the first convolution block, the second convolution block and the third convolution block respectively;

[0012] The first convolution block: used to perform convolution processing on the first frame of two consecutive frames in the training video;

[0013] The second convolution block is used to perform convolution processing on the second frame of two consecutive frames of images in the training video. The third convolution block is used to perform convolution processing on the combined image of two consecutive frames of images in the training video.

[0014] As a preferred implementation, in step B, the extraction of the camera pose parameter features is achieved by sequentially connecting an encoder module and a first multi-layer perceptron.

[0015] As a preferred implementation, in step B, before extracting the high-level features of camera pose parameters in the training video by the text feature extractor, the high-level features of camera pose parameters in the training video are also preprocessed, and the preprocessing is specifically as follows:

[0016] The positions are encoded according to the position order of the parameters to obtain the corresponding position encoding feature vector. The encoding formula is:

[0017]

[0018]

[0019]

[0020] Among them, PE represents position encoding, pos represents the parameter position number, d_index represents the position of the encoded feature vector, and d represents the dimension of the encoded feature vector.

[0021] As a preferred implementation, in step C, the heterogeneous modality data fuser includes a preprocessing block, a multiplication module, a weighted addition module and a second multilayer perceptron connected in sequence, and the preprocessing block is also connected to the weighted addition module.

[0022] As a preferred implementation, step C specifically includes:

[0023] C1, performing normalization preprocessing on the first fusion feature and the second fusion feature to obtain the preprocessed first fusion feature and the second fusion feature;

[0024] C2, element-wise multiplication of the preprocessed first fusion feature and the second fusion feature to obtain a first heterogeneous modality data fusion feature vector;

[0025] C3, performing weighted processing on the preprocessed first fusion feature and the second fusion feature, and performing addition processing on the first heterogeneous modality data fusion feature vector to obtain a second heterogeneous modality data fusion feature vector;

[0026] C4, feature mapping is performed on the second heterogeneous modality data fusion feature vector through a second multi-layer perceptron to obtain an output vector.

[0027] As a preferred implementation, in step C, the calculation formula for the second heterogeneous modality data fusion feature vector is:

[0028] F 2 =F 1 +w 1 f 1 +w 2 f 2

[0029] Among them, F 1 represents the first heterogeneous modality data fusion feature vector, F 2 represents the second heterogeneous modality data fusion feature vector, f 1 represents the first fusion feature after preprocessing, w 1 represents the weight of the first fusion feature after preprocessing, f 2 represents the second fusion feature after preprocessing, w 2 Represents the weight of the second fusion feature after preprocessing.

[0030] As a preferred implementation, in step D, the loss function is:

[0031] L=λ 1 L 1 +λ 2 L 2

[0032]

[0033]

[0034] Among them, h i represents the true value of the position element in the homography matrix, H i Represents the predicted value of the position element in the homography matrix, i represents the position element, x represents the true value of the horizontal coordinate of the center point of the moving target, X represents the predicted value of the horizontal coordinate of the center point of the moving target, y represents the true value of the vertical coordinate of the center point of the moving target, Y represents the predicted value of the horizontal coordinate of the center point of the moving target, w represents the true value of the width of the moving target, W represents the predicted value of the width of the moving target, h represents the true value of the height of the moving target, and H represents the predicted value of the height of the moving target. 1 represents the weight of the pose loss, λ 2 Represents the weight of the position loss.

[0035] A method for detecting a moving target, using the detection model trained by the above training method to detect a moving target, comprises the following steps:

[0036] S1, obtain a video to be detected, select adjacent frame images of the video, record the first frame of the adjacent frames as the first image, record the second frame of the adjacent frames as the second image, and obtain corresponding camera pose parameters;

[0037] S2, combining the first image and the second image to obtain a third image;

[0038] S3, inputting the first image, the second image, the third image and corresponding camera pose parameters into the detection model to obtain the position information of the moving target.

[0039] The beneficial technical effects of the present invention include:

[0040] By fusing two heterogeneous modal data, video image data and camera pose parameter information, mobile target detection is achieved, which can effectively solve the problem that existing mobile target detection methods are easily affected by the environment, resulting in a large number of false alarms and missed alarms. It can be applied to mobile target detection scenarios under different backgrounds, i.e. static backgrounds and dynamic backgrounds. Compared with the mobile target detection methods in the prior art, the mobile target detection method of heterogeneous modal data fusion proposed in the present invention has a low false alarm rate, a high detection rate, a more accurate detection result, a high model effectiveness, and good robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 It is a flow chart of a detection model training method based on heterogeneous modal data fusion in the present invention;

[0043] Figure 2 A detection model block diagram based on heterogeneous modal data fusion in an embodiment;

[0044] Figure 3 The figure is a flow chart of a moving target detection method in the present invention. DETAILED DESCRIPTION

[0045] The following describes the embodiments of the present invention through specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0046] Embodiment 1:

[0047] Reference Figure 1 , a detection model training method based on heterogeneous modal data fusion, comprising the steps of:

[0048] Step A, obtaining a training video set and storing it as a label file, wherein the training video set includes a simulation video set and a real video set. Step A specifically includes:

[0049] Step A1, establish a simulation video set, which includes different scene videos and corresponding virtual camera pose information. Each scene is rendered to generate a simulation video by adjusting the virtual camera position and shooting angle. The number of different scenes is 100, and the number of videos for each scene is 100. A total of 10,000 videos are collected in the simulation video set. Each video is 5 minutes long, and the resolution of the video is 1920*1080 pixels. The virtual camera pose information includes the camera position and camera pose. Among them, the camera position is a three-dimensional coordinate, expressed as (x, y, z), and the camera pose is the rotation of the camera around the three coordinate axes when generating the scene video, expressed as (dx, dy, dz).

[0050] A real video set is established, and the application scene is shot with a real camera to obtain real videos and corresponding real camera pose information. The real camera pose information includes camera position and camera attitude.

[0051] Step A2, in the simulation video, five different geometric figures are randomly generated, and the geometric figures are uniformly changed in position in the simulation video, so as to simulate the movement of the moving target in the video. The geometric figures include triangles, rectangles, pentagons and random irregular quadrilaterals, the size of the area of ​​a single geometric figure is 25*25 pixels to 270*150 pixels, and the ratio of the area of ​​a single geometric figure to the area of ​​the image satisfies 0.0002≤s≤0.02, where s represents the ratio of the area of ​​a single geometric figure to the area of ​​the image.

[0052] After obtaining the simulation video set, it is also preprocessed to extract continuous frames in the simulation video, and take two adjacent frames as an image pair. The first key feature points of the two adjacent frames are marked respectively, which are recorded as (x1, y1, w1, h1), where x1 represents the x-axis horizontal coordinate of the center point of the moving target in the image coordinate system, y1 represents the y-axis vertical coordinate of the center point of the moving target in the image coordinate system, w1 represents the width of the moving target, and h1 represents the height of the moving target. The first homography transformation matrix between two consecutive frames is calculated based on the first key feature point.

[0053] In the real video, after obtaining the real video set, it is also preprocessed to extract continuous frames in the real video, and two adjacent frames are taken as an image pair. The second key feature points of the two adjacent frames are marked respectively, which are recorded as (x2, y2, w2, h2), wherein x2 represents the x-axis horizontal coordinate of the center point of the moving target in the image coordinate system, y2 represents the y-axis vertical coordinate of the center point of the moving target in the image coordinate system, w1 represents the width of the moving target, and h1 represents the height of the moving target. The second homography transformation matrix between two consecutive frames is calculated based on the second key feature points.

[0054] Step A3: In the simulation video, the first key feature point and the first homography transformation matrix are stored as a label file; in the real video, the second key feature point and the second homography transformation matrix are stored as a label file.

[0055] Step B: extracting image features in the training video through an image feature extractor to obtain a first fusion feature, and extracting high-level features of camera posture parameters in the training video through a text feature extractor to obtain a second fusion feature.

[0056] The image feature extractor includes a first convolution block, a second convolution block, a third convolution block and a feature fusion block, and the feature fusion block is connected to the first convolution block, the second convolution block and the third convolution block respectively.

[0057] The first convolution block: used to perform convolution processing on the first frame of two consecutive frames in the training video;

[0058] The second convolution block: used to perform convolution processing on the second frame of two consecutive frames in the training video;

[0059] The third convolution block is used to perform convolution processing on the combined image of two consecutive frames in the training video.

[0060] The first convolution block and the second convolution block are ResNet101 network structures with 3-channel inputs. The last layer of the network is connected to a fully connected layer with 128-dimensional output. The first convolution block and the second convolution block are used to extract features from the training video to obtain the first feature output and the second feature output. The third convolution block is a ResNet50 network structure with 6-channel inputs. The last layer of the network is connected to a fully connected layer with 256-dimensional output. The third convolution block is used to extract features from the training video to obtain the third feature output. The first feature output, the second feature output, and the third feature output are fused by feature vector concatenation to obtain the first fused feature.

[0061] The text feature extractor includes six encoder modules with the same structure connected in sequence and a first multi-layer perceptron, the encoder module includes a multi-head self-attention unit, a first residual unit, a forward reasoning unit and a second residual unit connected in sequence, and the first residual unit and the second residual unit both include vector addition and normalization. The first multi-layer perceptron is a fully connected neural network including one hidden layer, and the number of neurons in the hidden layer is 1024.

[0062] The input of the text feature extractor is the camera pose parameters, including the three-dimensional coordinates of the camera position (x, y, z) and the rotation of the camera around the three coordinate axes (dx, dy, dz), and the data types are all floating point. Before entering the text feature extractor, the six parameters are also preprocessed, and the six parameters are represented as encoding vectors using 32-bit binary codes. Then, the positions are encoded according to the position order of the six parameters to obtain the corresponding position encoding feature vectors. The position encoding feature vector of each parameter is added to the encoding vector as the input of the text feature extractor. After being processed by the encoder module in the text feature extractor, six parameter features are obtained. The six parameter features are spliced ​​and forward calculated through the first multi-layer perceptron to obtain the second fusion feature.

[0063] The position is encoded according to the position order of the six parameters to obtain the corresponding position encoding feature vector. The encoding formula is:

[0064]

[0065]

[0066]

[0067] Among them, PE represents position encoding, pos represents the parameter position number, d_index represents the position of the encoded feature vector, and d represents the dimension of the encoded feature vector.

[0068] Step C, through the heterogeneous modality data fuser, realizes the fusion of the first fusion feature and the second fusion feature to obtain an output vector. Step C specifically includes:

[0069] Step C1, performing normalization preprocessing on the first fusion feature and the second fusion feature to obtain the preprocessed first fusion feature and the second fusion feature, and the function used for normalization is the softmax function.

[0070] Step C2, performing element-wise multiplication on the preprocessed first fusion feature and the second fusion feature, that is, multiplying the elements at corresponding positions, to obtain a first heterogeneous modality data fusion feature vector.

[0071] Step C3, weighting the preprocessed first fusion feature and the second fusion feature, and adding them to the first heterogeneous modality data fusion feature vector to obtain the second heterogeneous modality data fusion feature vector, and the calculation formula is:

[0072] F 2 =F 1 +w 1 f 1 +w 2 f 2

[0073] Among them, F 1 represents the first heterogeneous modality data fusion feature vector, F 2 represents the second heterogeneous modality data fusion feature vector, f 1 represents the first fusion feature after preprocessing, w 1 represents the weight of the first fusion feature after preprocessing, f 2 represents the second fusion feature after preprocessing, w 2 represents the weight of the second fusion feature after preprocessing, w 1 and w 2 Satisfy w 1 +w 2 =1.0,w 1 =0.75, w 2 =0.25

[0074] Furthermore, in step C, the weight of the first fusion feature after preprocessing satisfies: 0.7≤w 1≤0.9, indicating that the importance of image features is higher than that of text features.

[0075] Step C4, feature mapping the second heterogeneous modal data fusion feature vector through a second multi-layer perceptron to obtain an output vector. The output vector includes the homography matrix of adjacent frames of the training video, the confidence of the moving target, the horizontal coordinate of the center point of the moving target, the vertical coordinate of the center point of the moving target, the width of the moving target, and the height of the moving target. The dimension of the output vector is 59, where the first 9 elements of the vector represent the homography matrix of the image transformation relationship of adjacent frames of the training video, and the last 50 elements are grouped into 5 elements, for a total of 10 groups, each group of elements represents the confidence of the moving target, the horizontal coordinate of the center point of the moving target, the vertical coordinate of the center point of the moving target, the width of the moving target, and the height of the moving target.

[0076] Step D, refer to Figure 2 , according to the obtained training video set, label file and output vector, the image feature extractor, text feature extractor and heterogeneous modal data fusion device are trained, the loss function is calculated, and steps B to C are iteratively executed repeatedly until the loss function converges to a preset value, so as to obtain a detection model including the trained image feature extractor, text feature extractor and heterogeneous modal data fusion device.

[0077] First, the pre-trained model is obtained by training with the simulated video set, and then the real video set, that is, the data set of real scenes, is used on the basis of the pre-trained model for correction training to obtain the final detection model. The simulated video set is relatively simple to obtain and has rich scene diversity. Training enables the detection model to have strong generalization ability. The real video set training enables the detection model to have high adaptability to specific real scenes. Through the training of the simulated video set and the real video set, the detection model achieves the optimal effect.

[0078] Step D specifically includes:

[0079] Step D1, select adjacent video frame image data from the acquired training video, record them as the first image and the second image, uniformly scale the first image and the second image to 640*640 pixel size, combine the 3-channel first image and the second image into a 6-channel third image, input the first image into the first convolution block, input the second image into the second convolution block, input the third image into the third convolution block, and send the outputs of the first convolution block, the second convolution block and the third convolution block into the feature fusion block for feature fusion to obtain the first fusion feature.

[0080] Step D2, extracting high-level features of camera pose parameters in the training video through a text feature extractor to obtain a second fusion feature.

[0081] Step D3, through the heterogeneous modality data fuser, the first fusion feature and the second fusion feature are fused to obtain an output vector.

[0082] Step D4, calculate the loss function, determine whether the loss function converges to the preset value, if less than the preset value, stop training and execute step E, if not less than the preset value, perform a second judgment. Whether the preset number of training times is reached, if greater than the preset number of training times, the preset number of training times is 100,000 times, stop training and execute step E, if not greater than the preset number of training times, repeat iteratively execute steps B to C.

[0083] The loss function is:

[0084] L=λ 1 L 1 +λ 2 L 2

[0085]

[0086]

[0087] Among them, h i represents the true value of the position element in the homography matrix, H i represents the predicted value of the position element in the homography matrix, i represents the position element, x represents the true value of the horizontal coordinate of the center point of the moving target, X represents the predicted value of the horizontal coordinate of the center point of the moving target, y represents the true value of the vertical coordinate of the center point of the moving target, Y represents the predicted value of the horizontal coordinate of the center point of the moving target, w represents the true value of the width of the moving target, W represents the predicted value of the width of the moving target, h represents the true value of the height of the moving target, and H represents the predicted value of the height of the moving target. 1 represents the weight of the pose loss, λ 2 Represents the weight of the position loss, satisfying λ 1 =0.6,λ 2 =5.

[0088] The established detection model is used to fuse the video image information and camera parameter information, and then the simulation video set and the real video set are obtained to train the detection model to obtain a trained detection model. The model takes the video and the corresponding camera parameters as input, and directly obtains the position of the moving target in the video through the forward reasoning analysis of the detection model, thereby realizing end-to-end moving target detection.

[0089] The advantage of training with a simple simulation data set first is that a massive data set can be automatically generated, allowing the detection model to learn some basic features. Since the real scene is more complex, fine-tuning is performed using the real scene data set to make the detection model more adaptable to the real scene.

[0090] Embodiment 2:

[0091] The steps of the second embodiment are basically the same as those of the first embodiment, except that in step A, when establishing the simulation video set, the number of different scenes is not less than 100, and the randomly generated geometric figures are not less than five.

[0092] Embodiment three:

[0093] The steps of embodiment 3 are basically the same as those of embodiment 1, except that, in step B, the number of hidden layer neurons of the multilayer perceptron of the text feature extractor is not less than 1024.

[0094] Embodiment 4:

[0095] The steps of the fourth embodiment are basically the same as those of the first embodiment, except that in step D, the weight λ of the posture loss in the loss function is 1 and the weight λ of the position loss 2 , satisfying 0.5<λ 1 <1.0, 3.0<λ 2 <7.0.

[0096] Embodiment five:

[0097] Reference Figure 3 A moving target detection method is provided, which uses a trained detection model to detect a moving target in a video stream to be detected, comprising the steps of:

[0098] Step S1, obtaining a video to be detected, selecting adjacent frame images of the video, recording the first frame of the adjacent frames as the first image, recording the second frame of the adjacent frames as the second image, and obtaining corresponding camera pose parameters;

[0099] Step S2, combining the first image and the second image to obtain a third image;

[0100] Step S3, inputting the first image, the second image, the third image and the corresponding camera pose parameters into the detection model to obtain the position information of the moving target.

[0101] The mobile target detection method proposed by the present invention and the prior art method are compared and tested according to the following experiment. Experimental object: 10 test videos under static background and dynamic background, each test video is 1 minute long, and the video resolution is 1920*1080 pixels. The mobile target detection algorithm of the present invention and the traditional mobile target detection method are used to obtain the mobile target detection results, and the corresponding false alarm rate and detection rate are statistically analyzed. The following table shows the test results of the mobile detection algorithm:

[0102] Table 1 Detection results of different detection methods under different backgrounds

[0103]

[0104] Experimental results analysis:

[0105] 1. Compared with the prior art methods, the method of the present invention has a higher detection rate and a lower false alarm rate, which fully demonstrates the effectiveness of the method of the present invention.

[0106] 2. Under static background, the false alarm rate index and detection rate index of the mobile target detection method of the present invention are higher than those of the traditional method. Under dynamic background, although the false alarm rate and detection rate of the method of the present invention are slightly reduced, compared with the detection result of the traditional method with a false alarm rate of up to 80.1%, the method of the present invention has strong adaptability in application scenarios with dynamic background.

[0107] The mobile target detection method of the present invention is of great significance. The technology has practical application value in the fields of video technology, virtual reality, navigation and guidance, traffic control, target tracking, human-computer interaction, etc.

[0108] Embodiment six:

[0109] A computer-readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to execute the moving target detection method proposed in the first embodiment.

[0110] Embodiment seven:

[0111] An electronic device includes a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the moving target detection method proposed in the first embodiment by executing the computer instructions.

[0112] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope of the present invention.

Claims

1. A detection model training method based on heterogeneous modal data fusion, It is characterized in that Includes steps: A. Obtain a training video set and store it as a label file. Specifically, storing the label file includes: processing the acquired training video, extracting two consecutive frames of images in the training video as an image pair, marking a plurality of key feature points in each frame, and calculating a homography transformation matrix between the image pairs according to the feature points to obtain a label file, wherein the label file includes the plurality of key feature points and the homography transformation matrix; B, extracting image features in the training video through an image feature extractor to obtain a first fusion feature, and extracting camera pose parameter features in the training video through a text feature extractor to obtain a second fusion feature; C, fusing the first fused feature and the second fused feature through a heterogeneous modality data fuser to obtain an output vector, where the output vector includes a homography matrix of adjacent frames of the training video; D. According to the acquired training video set, label file and output vector, the image feature extractor, text feature extractor and heterogeneous modal data fusion device are trained, the loss function is calculated, and steps B to C are iterated repeatedly until the loss function converges to a preset value, so as to obtain a detection model including the trained image feature extractor, text feature extractor and heterogeneous modal data fusion device.

2. According to claim 1, a detection model training method based on heterogeneous modal data fusion, It is characterized in that In step B, the image feature extractor includes a first convolution block, a second convolution block, a third convolution block and a feature fusion block, and the feature fusion block is connected to the first convolution block, the second convolution block and the third convolution block respectively; The first convolution block: used to perform convolution processing on the first frame of two consecutive frames in the training video; The second convolution block: used to perform convolution processing on the second frame of two consecutive frames in the training video; The third convolution block is used to perform convolution processing on the combined image of two consecutive frames in the training video.

3. According to the detection model training method based on heterogeneous modal data fusion according to claim 1, It is characterized in that In step B, the extraction of the camera pose parameter features is achieved through the encoder module and the first multi-layer perceptron connected in sequence.

4. According to claim 3, a detection model training method based on heterogeneous modal data fusion, It is characterized in that In step B, before extracting the high-level features of the camera pose parameters in the training video by the text feature extractor, the high-level features of the camera pose parameters in the training video are preprocessed, and the preprocessing is specifically as follows: The positions are encoded according to the position order of the parameters to obtain the corresponding position encoding feature vector. The encoding formula is: Among them, PE represents position encoding, pos represents the parameter position number, d_index represents the position of the encoded feature vector, and d represents the dimension of the encoded feature vector.

5. According to the detection model training method based on heterogeneous modal data fusion according to claim 1, It is characterized in that In step C, the heterogeneous modality data fuser includes a preprocessing block, a multiplication module, a weighted addition module and a second multilayer perceptron connected in sequence, and the preprocessing block is also connected to the weighted addition module.

6. A detection model training method based on heterogeneous modal data fusion according to claim 5, It is characterized in that Step C specifically includes: C1, performing normalization preprocessing on the first fusion feature and the second fusion feature to obtain the preprocessed first fusion feature and the second fusion feature; C2, multiplying the elements of the preprocessed first fusion feature and the second fusion feature to obtain a first heterogeneous modality data fusion feature vector; C3, performing weighted processing on the preprocessed first fusion feature and the second fusion feature, and performing addition processing on the first heterogeneous modality data fusion feature vector to obtain a second heterogeneous modality data fusion feature vector; C4, feature mapping is performed on the second heterogeneous modality data fusion feature vector through a second multi-layer perceptron to obtain an output vector.

7. The detection model training method based on heterogeneous modality data fusion according to claim 6, It is characterized in that In step C, the calculation formula of the second heterogeneous modality data fusion feature vector is: F 2 =F 1 +w 1 f 1 +w 2 f 2 Among them, F 1 represents the first heterogeneous modality data fusion feature vector, F 2 represents the second heterogeneous modality data fusion feature vector, f 1 represents the first fusion feature after preprocessing, w 1 represents the weight of the first fusion feature after preprocessing, f 2 represents the second fusion feature after preprocessing, w 2 Represents the weight of the second fusion feature after preprocessing.

8. The detection model training method based on heterogeneous modal data fusion according to claim 1, It is characterized in that In step D, the loss function is: L=λ 1 L 1 +λ 2 L 2 Among them, h i represents the true value of the position element in the homography matrix, H i represents the predicted value of the position element in the homography matrix, i represents the position element, x represents the true value of the horizontal coordinate of the center point of the moving target, X represents the predicted value of the horizontal coordinate of the center point of the moving target, y represents the true value of the vertical coordinate of the center point of the moving target, Y represents the predicted value of the horizontal coordinate of the center point of the moving target, w represents the true value of the width of the moving target, W represents the predicted value of the width of the moving target, h represents the true value of the height of the moving target, H represents the predicted value of the height of the moving target, λ 1 represents the weight of the pose loss, λ 2 Represents the weight of the position loss.

9. A method for detecting a moving target, It is characterized in that Using the detection model trained by the training method according to any one of claims 1 to 8 to detect a moving target in a video to be detected, the method comprises the following steps: S1, obtain a video to be detected, select adjacent frame images of the video, record the first frame of the adjacent frames as the first image, record the second frame of the adjacent frames as the second image, and obtain corresponding camera pose parameters; S2, combining the first image and the second image to obtain a third image; S3, inputting the first image, the second image, the third image and corresponding camera pose parameters into the detection model to obtain the position information of the moving target.

Citation Information

Patent Citations

  • Training method of incomplete information target recognition model and target recognition method

    CN115761444A

  • Target object position prediction and motion tracking

    US20200117952A1