A rapid positioning method and system based on monitoring records
By combining the target rapid feature extraction model, the reverse prediction positioning model and the monitoring prediction positioning model, the problem of insufficient data processing speed and analysis capability in the monitoring record positioning technology is solved, the recognition and tracking of the target is realized, and the recognition and tracking of the target is realized, providing higher target positioning accuracy and efficiency.
Patent Information
- Application Number
- CN202510051230.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Existing monitoring, recording and positioning technologies lack data processing speed and analysis capabilities in large-scale scenarios, making it difficult to conduct in-depth analysis based on event context information, affecting positioning accuracy and judgment capabilities. Especially in complex event processing, it is difficult to infer key data and potential risks before the event.
A rapid positioning method based on monitoring records is adopted. Through correlation analysis, reverse prediction and intelligent reasoning of multi-dimensional monitoring records, combined with the target rapid feature extraction model, the target reverse prediction positioning model and the monitoring prediction positioning model, the current position positioning and predictive positioning of the target are achieved.
It achieves fast and accurate target positioning and predictive positioning, can identify and track targets in complex environments, optimizes the target recognition capability of video surveillance, can identify and track targets in complex environments, provides higher target positioning recognition and tracking, and provides higher target positioning accuracy and efficiency.
Smart Images

Figure CN119991800B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of monitoring and positioning technology, and in particular to a rapid positioning method and system based on monitoring records. Background Art
[0002] Existing surveillance and recording location technologies primarily combine video surveillance and sensor data with data analysis algorithms to collect and process multi-dimensional data in real time, enabling rapid location of incidents or target objects. These technologies are widely used in security monitoring, emergency response, and event tracking. However, in large-scale surveillance scenarios, the data processing speed and analytical capabilities of existing technologies often fall short of the demands for rapid response. Furthermore, traditional analysis methods struggle to integrate the context of incidents for in-depth analysis, which impacts location accuracy and judgment. This is particularly true in complex incident processing, where existing technologies struggle to infer the specific circumstances or potential risks preceding an incident based on existing data, thereby inferring subsequent target location.
[0003] Therefore, in order to solve problems such as lack of real-time performance and accuracy, it is necessary to integrate and conduct in-depth analysis through analysis methods such as correlation analysis, reverse prediction, intelligent reasoning and event backtracking of multi-dimensional monitoring records. This will not only enable the rapid location of the incident area after the incident occurs, but also enable reverse inference of key data before the incident occurs, providing precise support for predicting potential dangers and optimizing emergency responses. Summary of the Invention
[0004] The present invention aims to provide a rapid positioning method and system based on monitoring records, which can realize the current position positioning and predictive positioning of the target.
[0005] A rapid positioning method based on monitoring records includes the following steps:
[0006] Obtain the current surveillance video of the target to be located; analyze the current surveillance video of the target to be located and the fast feature extraction model of the positioning target to obtain the current target monitoring features and the current target location; the current surveillance video of the target to be located contains n surveillance images of the target to be located; the fast feature extraction model of the positioning target includes an image preprocessing layer, a feature calculation layer, a feature output layer and a positioning recognition layer, wherein the feature calculation layer gradually extracts, fuses and refines the high-dimensional features of the enhanced target monitoring image through a multi-layer downsampling and decoder structure, and finally outputs the current target monitoring features, realizing an efficient feature calculation and fusion process;
[0007] Based on the current target monitoring features and the target reverse prediction positioning model, a prediction strategy is output to obtain the associated monitoring video of the target to be located. The target reverse prediction positioning model achieves accurate video positioning based on the target motion trajectory by combining motion feature recognition and reverse video retrieval strategy.
[0008] The current target location, current target monitoring features, and associated monitoring videos and monitoring prediction positioning models of the target to be located are used for predictive analysis to obtain the predicted target location. The monitoring prediction positioning model fuses historical target monitoring features and trajectory information, and combines 1D convolution and LSTM time series modeling to construct a predictive positioning framework that can accurately perform dynamic target positioning. Subsequent operations on the target are performed based on the current target location and predicted target location.
[0009] As a preferred technical solution of the present invention, the positioning target fast feature extraction model includes an image preprocessing layer, a feature calculation layer, a feature output layer and a positioning recognition layer;
[0010] The image preprocessing layer is used to stack and enhance the n monitoring images of the target to be located in the current monitoring video of the target to be located to obtain a fused enhanced target monitoring image;
[0011] The feature calculation layer is used to extract features from the fused enhanced target monitoring image to obtain the current target monitoring features;
[0012] The feature output layer is used to output the current target monitoring features;
[0013] The positioning and recognition layer is used to perform positioning and recognition based on the current target monitoring characteristics to obtain the current target positioning;
[0014] Specific steps for training the positioning recognition layer:
[0015] Collecting several groups of positioning and recognition training samples; each group of positioning and recognition training samples contains monitoring features and corresponding target positioning information; combining several groups of positioning and recognition training samples to obtain a positioning and recognition training set;
[0016] The positioning recognition training set is input into the CNN model for model training to obtain the initial positioning recognition layer; the initial positioning recognition layer is evaluated. If the initial positioning recognition layer passes the model evaluation, the initial positioning recognition layer is used as the positioning recognition layer in the positioning target rapid feature extraction model; otherwise, the positioning recognition training set is used to continue model training.
[0017] As a preferred technical solution of the present invention, the specific steps of performing feature extraction in the feature calculation layer include:
[0018] The feature calculation layer includes input layer, feature extraction layer, feature fusion layer and output layer;
[0019] The fused enhanced target monitoring image is received in the input layer, and is sent to the feature extraction layer through the downsampling path for feature calculation;
[0020] In the feature extraction layer, four high-dimensional feature extractions are performed on the fused enhanced target monitoring image to obtain the enhanced target monitoring image high-dimensional feature T1, the enhanced target monitoring image high-dimensional feature T2, the enhanced target monitoring image high-dimensional feature T3 and the enhanced target monitoring image high-dimensional feature T4;
[0021] Among them, the specific steps of performing four high-dimensional feature extractions are as follows: the feature extraction layer contains four layers of downsampling structure, and each time a layer of downsampling is performed, a high-dimensional feature extraction is performed on the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T1 of the enhanced target monitoring image is 1 / 4 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T2 of the enhanced target monitoring image is 1 / 8 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T3 of the enhanced target monitoring image is 1 / 16 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T4 of the enhanced target monitoring image is 1 / 32 of the fused enhanced target monitoring image;
[0022] In the feature fusion layer, there are four layers of decoders; each layer of decoders corresponds to a downsampling structure; in the decoder, the enhanced target monitoring image high-dimensional features T1 and the enhanced target monitoring image high-dimensional features T2 are fused to obtain the fused enhanced target monitoring image high-dimensional features T 12 ;
[0023] The enhanced target monitoring image high-dimensional feature T2 and the enhanced target monitoring image high-dimensional feature T3 are fused to obtain the fused enhanced target monitoring image high-dimensional feature T 23 ;
[0024] The enhanced target monitoring image high-dimensional feature T3 and the enhanced target monitoring image high-dimensional feature T4 are fused to obtain the fused enhanced target monitoring image high-dimensional feature T 34 ;
[0025] Finally, the high-dimensional features T of the fused enhanced target monitoring image are 12 , fusion enhancement target monitoring image high dimensional features T 23 and fusion enhancement of high-dimensional features of target monitoring images T 34 Perform feature fusion to obtain the current target monitoring features;
[0026] Output the current target monitoring features in the output layer.
[0027] As a preferred technical solution of the present invention, the target reverse prediction positioning model includes a motion feature recognition layer, a reverse prediction strategy layer and a video output layer;
[0028] The motion feature recognition layer is used to perform motion feature recognition on the current target monitoring features to obtain the motion trajectory features of the target to be located;
[0029] The reverse prediction strategy layer is used to perform reverse video retrieval based on the monitoring motion trajectory characteristics of the target to be located, and obtain the associated monitoring video of the target to be located;
[0030] The video output layer is used to output the associated surveillance video of the target to be located.
[0031] As a preferred technical solution of the present invention, the specific steps of performing reverse video retrieval in the reverse prediction strategy layer include:
[0032] Based on the monitoring motion trajectory features of the target to be located, M groups of monitoring video slices are extracted to obtain the monitoring video P of the associated target to be screened. m , m=1,2,…,M;
[0033] The surveillance video P of the associated target to be screened m Extract feature representation and obtain the feature representation P of the surveillance video of the associated target to be screened m ';
[0034] The feature representation of the surveillance video of the associated target to be screened P m 'Use 3D convolutional neural network to perform feature encoding and obtain the embedded vector L of the surveillance video of the associated target to be screened m ;
[0035] Based on the monitoring motion trajectory characteristics of the target to be located and the monitoring video embedding vector L of the associated target to be screened m Perform prediction matching to obtain the target surveillance video vector matching result J m ;
[0036] Based on the target surveillance video vector matching results J m For all the associated target surveillance videos P to be screened m Perform screening to obtain the surveillance video associated with the target to be located.
[0037] As a preferred technical solution of the present invention, the monitoring prediction and positioning model includes a correlation feature extraction layer, a feature connection layer and a positioning prediction layer;
[0038] The associated feature extraction layer is used to input the associated surveillance video of the target to be located into the positioning target fast feature extraction model for feature extraction, and obtain several groups of historical target monitoring features and historical target positioning;
[0039] The feature connection layer is used to connect several sets of historical target monitoring features with the current target monitoring features to obtain a prediction positioning feature network; and extract trajectory features from several sets of historical target positioning and current target positioning to obtain historical target trajectory features;
[0040] The positioning prediction layer is used to perform prediction analysis based on the prediction positioning feature network and historical target trajectory features to obtain the predicted target positioning;
[0041] The positioning prediction layer is constructed based on BP neural network training.
[0042] As a preferred technical solution of the present invention, the feature connection layer includes a prediction network construction layer and a trajectory feature extraction layer;
[0043] In the prediction network construction layer, a 1D convolutional layer is used to aggregate several groups of historical target monitoring features to obtain aggregated historical target monitoring features. The aggregated historical target monitoring features and the current target monitoring features are then concatenated to obtain a prediction positioning feature network.
[0044] In the trajectory feature extraction layer, a temporal network is established using the LSTM model and several sets of historical target positioning and current target positioning to obtain the historical target trajectory features;
[0045] Specific steps for training the positioning prediction layer:
[0046] Collect I groups of continuous positioning prediction training samples, where K is the total time steps included in the continuous positioning prediction training samples;
[0047] Using the formula As the loss function of the positioning prediction layer;
[0048] Among them, I is the total number of training samples, Y i is the actual position of the i-th training sample, Y i ′ is the predicted position of the i-th training sample, α is the corresponding weight coefficient; K is the total time step in the training sample, Z k is the predicted position at the kth time step, Z k+1 is the predicted position of the k+1th time step, and β is the corresponding weight coefficient;
[0049] Based on the loss function F total Perform model training to obtain the final positioning prediction layer.
[0050] A rapid positioning system based on monitoring records, comprising:
[0051] The monitoring feature extraction module includes a video extraction unit and a feature extraction unit; the video extraction unit is used to obtain the current monitoring video of the target to be located; the feature extraction unit is used to analyze the current monitoring video of the target to be located and the positioning target fast feature extraction model to obtain the current target monitoring features and the current target location; the current monitoring video of the target to be located contains n monitoring images of the target to be located; the positioning target fast feature extraction model includes an image preprocessing layer, a feature calculation layer, a feature output layer and a positioning recognition layer, wherein the feature calculation layer gradually extracts, fuses and refines the high-dimensional features of the enhanced target monitoring image through multi-layer downsampling and decoder structure, and finally outputs the current target monitoring features, realizing an efficient feature calculation and fusion process;
[0052] The monitoring and positioning prediction module includes an associated prediction unit and a positioning prediction unit; the associated prediction unit is used to output a prediction strategy based on the current target monitoring features and the target reverse prediction positioning model to obtain the associated monitoring video of the target to be located; the target reverse prediction positioning model realizes accurate video positioning based on the target motion trajectory by combining motion feature recognition and reverse video retrieval strategy; the positioning prediction unit is used to perform predictive analysis using the current target positioning, the current target monitoring features and the associated monitoring video of the target to be located and the monitoring prediction positioning model to obtain the predicted target positioning; the monitoring prediction positioning model constructs a prediction positioning framework by fusing historical target monitoring features and trajectory information, and combining 1D convolution and LSTM time series modeling, which can accurately perform dynamic target positioning; subsequent operations on the target are performed based on the current target positioning and predicted target positioning.
[0053] The present invention has the following advantages:
[0054] 1. The present invention can efficiently extract, fuse and refine the high-dimensional features of the target monitoring image through the multi-layer downsampling and decoder structure in the target fast feature extraction model, can quickly and accurately identify the target, and obtain detailed target features, thereby improving the accuracy and efficiency of target positioning; the target reverse prediction positioning model combines motion feature recognition and reverse video retrieval strategy, can accurately locate the target monitoring video according to the target's motion trajectory, and perform efficient video positioning, reducing positioning errors and optimizing the target tracking capability of video monitoring; the monitoring prediction positioning model can accurately perform dynamic target positioning by fusing historical target monitoring features and trajectory information, combining 1D convolution and LSTM time series modeling, which provides more reliable support for target tracking and positioning in dynamic scenes.
[0055] 2. The present invention gradually reduces the resolution of the feature map through four high-dimensional feature extractions in the feature extraction layer, and deeply extracts the multi-level features of the target monitoring image. This multiple downsampling and extraction method enables the model to capture complex information in the image from coarse to fine, which helps to extract more recognizable target features; the different resolutions of features in each layer help to extract multi-scale and multi-level image information, which is very effective for processing diverse target forms and different image details; through high-dimensional feature extraction in the feature extraction layer, the detailed features, spatial features and texture features of the target can be effectively extracted from the image. These high-dimensional features help to enhance the target recognition ability, especially in complex backgrounds and dynamically changing monitoring scenes, which can reduce noise interference and enhance the target's recognizability. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a schematic structural diagram of a rapid positioning system based on monitoring records adopted in an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0058] Example 1, a rapid positioning method based on monitoring records, comprising the following steps:
[0059] Obtain the current surveillance video of the target to be located; analyze the current surveillance video of the target to be located and the fast feature extraction model of the location target to obtain the current target monitoring features and the current target location; the current surveillance video of the target to be located contains n surveillance images of the target to be located; in actual use scenarios, the surveillance video is obtained through the video signal output by the camera;
[0060] The fast feature extraction model for positioning targets includes an image preprocessing layer, a feature calculation layer, a feature output layer, and a positioning recognition layer. The feature calculation layer gradually extracts, fuses, and refines the high-dimensional features of the target monitoring image through a multi-layer downsampling and decoder structure, and finally outputs the current target monitoring features, achieving an efficient feature calculation and fusion process.
[0061] The image preprocessing layer is used to stack and enhance the n monitoring images of the target to be located in the current monitoring video of the target to be located to obtain a fused enhanced target monitoring image;
[0062] The feature calculation layer is used to extract features from the fused enhanced target monitoring image to obtain the current target monitoring features;
[0063] The feature output layer is used to output the current target monitoring features;
[0064] The positioning and recognition layer is used to perform positioning and recognition based on the current target monitoring characteristics to obtain the current target positioning;
[0065] Specific steps for training the positioning recognition layer:
[0066] Collecting several groups of positioning and recognition training samples; each group of positioning and recognition training samples contains monitoring features and corresponding target positioning information; combining several groups of positioning and recognition training samples to obtain a positioning and recognition training set;
[0067] The positioning recognition training set is input into the CNN model for model training to obtain the initial positioning recognition layer; the initial positioning recognition layer is evaluated. If the initial positioning recognition layer passes the model evaluation, the initial positioning recognition layer is used as the positioning recognition layer in the positioning target fast feature extraction model; otherwise, the positioning recognition training set is used to continue model training;
[0068] The image preprocessing layer stacks and enhances the monitoring images of the target to be located, fusing the feature information of multiple images and laying a solid foundation for subsequent feature extraction. This enhancement can reduce noise, improve image quality, and provide more accurate input for target feature calculation. The feature calculation layer gradually extracts and fuses features of the image through multi-layer downsampling and decoder structure, so that the high-dimensional features of the image can be effectively extracted and refined. This process can better capture the detailed features of the target and improve the accuracy of target recognition. The positioning recognition layer performs positioning and recognition based on the target monitoring features obtained from the feature output layer, and can accurately identify the positioning information of the current target. By using the CNN model, this layer can effectively locate the target in space and improve the accuracy of target positioning.
[0069] The specific steps for feature extraction in the feature calculation layer include:
[0070] The feature calculation layer includes input layer, feature extraction layer, feature fusion layer and output layer;
[0071] The fused enhanced target monitoring image is received in the input layer, and is sent to the feature extraction layer through the downsampling path for feature calculation;
[0072] In the feature extraction layer, four high-dimensional feature extractions are performed on the fused enhanced target monitoring image to obtain the enhanced target monitoring image high-dimensional feature T1, the enhanced target monitoring image high-dimensional feature T2, the enhanced target monitoring image high-dimensional feature T3 and the enhanced target monitoring image high-dimensional feature T4;
[0073] Among them, the specific steps of performing four high-dimensional feature extractions are as follows: the feature extraction layer contains four layers of downsampling structure, and each time a layer of downsampling is performed, a high-dimensional feature extraction is performed on the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T1 of the enhanced target monitoring image is 1 / 4 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T2 of the enhanced target monitoring image is 1 / 8 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T3 of the enhanced target monitoring image is 1 / 16 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T4 of the enhanced target monitoring image is 1 / 32 of the fused enhanced target monitoring image;
[0074] In the feature fusion layer, there are four layers of decoders; each layer of decoders corresponds to a downsampling structure; in the decoder, the enhanced target monitoring image high-dimensional features T1 and the enhanced target monitoring image high-dimensional features T2 are fused to obtain the fused enhanced target monitoring image high-dimensional features T 12 ;
[0075] The enhanced target monitoring image high-dimensional feature T2 and the enhanced target monitoring image high-dimensional feature T3 are fused to obtain the fused enhanced target monitoring image high-dimensional feature T 23 ;
[0076] The enhanced target monitoring image high-dimensional feature T3 and the enhanced target monitoring image high-dimensional feature T4 are fused to obtain the fused enhanced target monitoring image high-dimensional feature T 34 ;
[0077] Finally, the high-dimensional features T of the fused enhanced target monitoring image are 12 , fusion enhancement target monitoring image high dimensional features T 23 and fusion enhancement of high-dimensional features of target monitoring images T 34 Perform feature fusion to obtain the current target monitoring features;
[0078] Output the current target monitoring features in the output layer;
[0079] In the feature extraction layer, through four high-dimensional feature extractions, the resolution of the feature map is gradually reduced, and the multi-level features of the target monitoring image are deeply extracted. This multiple downsampling and extraction method enables the model to capture complex information in the image from coarse to fine, which helps to extract more recognizable target features. The different resolutions of the features in each layer help to extract multi-scale and multi-level image information, which is very effective for processing diverse target morphologies and different image details. Through the high-dimensional feature extraction of the feature extraction layer, the detailed features, spatial features and texture features of the target can be effectively extracted from the image. These high-dimensional features help to enhance the recognition ability of the target, especially in complex backgrounds and dynamically changing monitoring scenes, which can reduce noise interference and enhance the recognizability of the target. The feature fusion layer fuses features at different levels through a multi-layer decoder structure, which can retain low-level detailed features while combining high-level abstract features to form a richer target description. The feature fusion strategy enables features of different resolutions to interact and enhance at multiple scales, improving the target's expression ability and enabling the model to comprehensively utilize low-level features and high-level features for more accurate positioning and recognition.
[0080] The feature calculation layer, through multiple feature extraction, fusion, and refinement enhancements, not only improves the expressiveness of target features and positioning accuracy, but also ensures the robustness and adaptability of the model through hierarchical information processing. This provides strong support for subsequent target positioning and recognition tasks, helping it to perform well in various complex monitoring scenarios.
[0081] Based on the current target monitoring features and the target reverse prediction positioning model, a prediction strategy is output to obtain the associated monitoring video of the target to be located. The target reverse prediction positioning model achieves accurate video positioning based on the target motion trajectory by combining motion feature recognition and reverse video retrieval strategy.
[0082] The target reverse prediction positioning model includes a motion feature recognition layer, a reverse prediction strategy layer and a video output layer;
[0083] The motion feature recognition layer is used to perform motion feature recognition on the current target monitoring features to obtain the motion trajectory features of the target to be located;
[0084] The reverse prediction strategy layer is used to perform reverse video retrieval based on the monitoring motion trajectory characteristics of the target to be located, and obtain the associated monitoring video of the target to be located;
[0085] The video output layer is used to output the associated surveillance video of the target to be located;
[0086] By combining motion feature recognition with reverse video retrieval strategy, the target reverse prediction positioning model can accurately locate the target in the video based on its motion trajectory. This method can infer the future or historical position of the target according to its dynamic changes, greatly improving the accuracy of video positioning. As key identification information, the motion trajectory can effectively associate the target with its specific position in the video through this model, reducing the ambiguity and error in traditional methods. The motion feature recognition layer extracts and identifies the motion features of the target monitoring features to obtain the target motion trajectory features, which not only reveals the motion state of the target in space, but also captures the target's changing trend in the time dimension, enhancing the model's understanding of the target's dynamic behavior. The reverse prediction strategy layer performs reverse video retrieval based on the target's motion trajectory features. This process can quickly locate video clips related to the current target from a large number of surveillance videos, improving the efficiency and relevance of video retrieval.
[0087] In complex environments, traditional video surveillance methods may cause positioning errors due to factors such as background interference or target occlusion. However, by incorporating motion trajectory features, the target reverse prediction positioning model can effectively reduce these interference factors and provide more stable and accurate tracking results. This model not only supports the positioning and tracking of a single target, but also supports dynamic video positioning of multiple targets by expanding the motion trajectory analysis to multiple targets. By combining the motion feature recognition layer and the reverse prediction strategy layer, it can handle complex multi-target scenarios and achieve more efficient video monitoring and analysis.
[0088] The specific steps of reverse video retrieval in the reverse prediction strategy layer include:
[0089] Based on the monitoring motion trajectory features of the target to be located, M groups of monitoring video slices are extracted to obtain the monitoring video P of the associated target to be screened. m , m=1,2,…,M;
[0090] The surveillance video P of the associated target to be screened m Extract feature representation and obtain the feature representation P of the surveillance video of the associated target to be screened m ';
[0091] The feature representation of the surveillance video of the associated target to be screened P m 'Use 3D convolutional neural network to perform feature encoding and obtain the embedded vector L of the surveillance video of the associated target to be screened m ;
[0092] Based on the monitoring motion trajectory characteristics of the target to be located and the monitoring video embedding vector L of the associated target to be screened m Perform prediction matching to obtain the target surveillance video vector matching result J m ;
[0093] Based on the target surveillance video vector matching results J m For all the associated target surveillance videos P to be screened m Perform screening to obtain surveillance videos associated with the target to be located;
[0094] By extracting feature representations from the surveillance videos of the associated targets to be screened and using 3D convolutional neural networks for feature encoding, spatial and temporal features can be efficiently extracted from video data. This feature encoding method can simultaneously capture the temporal relationship and spatial structure between video frames, providing a richer video representation. By matching the motion trajectory features of the monitored target to be located with the video embedding vector, the correlation between the video slice and the target can be accurately evaluated. This matching process not only relies on spatial information, but also fully considers dynamic changes in time, making video matching more accurate and intelligent. The application of feature extraction based on video slices and 3D convolutional neural networks can support video retrieval and target positioning in various scenarios, such as security monitoring, intelligent transportation, autonomous driving and other fields. Dynamic target behaviors in different scenarios can be effectively analyzed and located using this method.
[0095] The current target location, current target monitoring features, associated monitoring videos of the target to be located, and the monitoring prediction positioning model are used for predictive analysis to obtain the predicted target location. The monitoring prediction positioning model integrates historical target monitoring features and trajectory information, and combines 1D convolution and LSTM time series modeling to build a prediction positioning framework that can accurately locate dynamic targets.
[0096] The monitoring prediction and positioning model includes a correlation feature extraction layer, a feature connection layer, and a positioning prediction layer;
[0097] The associated feature extraction layer is used to input the associated surveillance video of the target to be located into the positioning target fast feature extraction model for feature extraction, and obtain several groups of historical target monitoring features and historical target positioning;
[0098] The feature connection layer is used to connect several sets of historical target monitoring features with the current target monitoring features to obtain a prediction positioning feature network; and extract trajectory features from several sets of historical target positioning and current target positioning to obtain historical target trajectory features;
[0099] The positioning prediction layer is used to perform prediction analysis based on the prediction positioning feature network and historical target trajectory features to obtain the predicted target positioning;
[0100] The positioning prediction layer is constructed based on BP neural network training;
[0101] The feature connection layer includes the prediction network construction layer and the trajectory feature extraction layer;
[0102] In the prediction network construction layer, a 1D convolutional layer is used to aggregate several groups of historical target monitoring features to obtain aggregated historical target monitoring features. The aggregated historical target monitoring features and the current target monitoring features are then concatenated to obtain a prediction positioning feature network.
[0103] In the trajectory feature extraction layer, a temporal network is established using the LSTM model and several sets of historical target positioning and current target positioning to obtain the historical target trajectory features;
[0104] By combining current target positioning, historical target monitoring features, and target trajectory information, the monitoring prediction positioning model can provide accurate dynamic target positioning. The model effectively captures the target's motion pattern through time series data modeling and can predict the target's future movement position. This dynamic positioning capability makes the model more accurate and forward-looking when processing target motion trajectory prediction. Through the association feature extraction layer and feature connection layer, several groups of historical target monitoring features are integrated with the current target monitoring features, providing richer background information for target prediction positioning. The introduction of historical data means that predictive positioning not only relies on the current state, but also can dynamically infer the target's historical behavior, thereby improving the robustness of the positioning results. In the prediction network construction layer, the historical target monitoring features are aggregated using a 1D convolutional layer, allowing multiple historical features to be efficiently integrated. This feature aggregation process can enhance the depth of feature representation, improve the model's ability to understand target features, and thus improve the accuracy of positioning prediction. The aggregated historical target monitoring features are spliced with the current target monitoring features to further enhance the model's comprehensive understanding of the target state. The combination of comprehensive information of the current state and historical behavior helps to provide more accurate predictions.
[0105] Specific steps for training the positioning prediction layer:
[0106] Collect I groups of continuous positioning prediction training samples, where K is the total time steps included in the continuous positioning prediction training samples;
[0107] Using the formula As the loss function of the positioning prediction layer;
[0108] Among them, I is the total number of training samples, Y i is the actual position of the i-th training sample, Y i ′ is the predicted position of the i-th training sample, α is the corresponding weight coefficient; K is the total time step in the training sample, Z k is the predicted position at the kth time step, Z k+1 is the predicted position of the k+1th time step, and β is the corresponding weight coefficient;
[0109] Based on the loss function F totalPerform model training to obtain the final positioning prediction layer;
[0110] Through the definition of the loss function, this method simultaneously optimizes the error between the predicted position and the true position and the smoothness of the predicted positions in the time series. This multi-dimensional loss design can avoid drastic fluctuations in the time series prediction results while ensuring prediction accuracy, thereby improving the stability and rationality of the prediction. The loss function comprehensively considers the global positioning error and the local changes in the time series, organically combining the two. This method that takes into account both global and local optimization can improve positioning accuracy while ensuring the logic and coherence of the predicted trajectory.
[0111] Perform subsequent operations on the target based on the current target positioning and predicted target positioning; for example, in a security system, it is necessary to track the current and future positions of the monitoring target in real time so that timely measures can be taken. If the predicted target positioning shows that the target is about to enter a restricted area, an alarm will be automatically triggered or security personnel will be notified; the camera will be controlled to dynamically adjust the angle and focal length to continuously track the movement trajectory of the target; or in an intelligent transportation system, the current and future positions of pedestrians and vehicles will be predicted to optimize traffic flow. According to the predicted position of the vehicle, the traffic light signals can be dynamically adjusted to reduce traffic congestion; the future trajectory of the vehicle can be predicted to detect potential collision risks and issue timely warnings to the driver.
[0112] Example 2, a rapid positioning system based on monitoring records, see Figure 1 Shown, including:
[0113] The monitoring feature extraction module includes a video extraction unit and a feature extraction unit; the video extraction unit is used to obtain the current monitoring video of the target to be located; the feature extraction unit is used to analyze the current monitoring video of the target to be located and the positioning target fast feature extraction model to obtain the current target monitoring features and the current target location; the current monitoring video of the target to be located contains n monitoring images of the target to be located; the positioning target fast feature extraction model includes an image preprocessing layer, a feature calculation layer, a feature output layer and a positioning recognition layer, wherein the feature calculation layer gradually extracts, fuses and refines the high-dimensional features of the enhanced target monitoring image through multi-layer downsampling and decoder structure, and finally outputs the current target monitoring features, realizing an efficient feature calculation and fusion process;
[0114] The monitoring and positioning prediction module includes an associated prediction unit and a positioning prediction unit; the associated prediction unit is used to output a prediction strategy based on the current target monitoring features and the target reverse prediction positioning model to obtain the associated monitoring video of the target to be located; the target reverse prediction positioning model realizes accurate video positioning based on the target motion trajectory by combining motion feature recognition and reverse video retrieval strategy; the positioning prediction unit is used to perform predictive analysis using the current target positioning, the current target monitoring features and the associated monitoring video of the target to be located and the monitoring prediction positioning model to obtain the predicted target positioning; the monitoring prediction positioning model constructs a prediction positioning framework by fusing historical target monitoring features and trajectory information, and combining 1D convolution and LSTM time series modeling, which can accurately perform dynamic target positioning; subsequent operations on the target are performed based on the current target positioning and predicted target positioning.
[0115] It should be understood that those skilled in the art may make improvements or modifications based on the above description, and all such improvements and modifications shall fall within the scope of protection of the appended claims. Any portion of this specification not described in detail is prior art known to those skilled in the art.
Claims
1. A rapid positioning method based on monitoring records, characterized in that: The following steps are involved: Get the current surveillance video of the target to be located; Based on the current monitoring video of the target to be located and the fast feature extraction model of the positioning target, the current target monitoring features and the current target positioning are obtained; The current surveillance video of the target to be located contains n surveillance images of the target to be located; The positioning target fast feature extraction model includes image preprocessing layer, feature calculation layer, feature output layer and positioning recognition layer; The image preprocessing layer is used to stack and enhance the n monitoring images of the target to be located in the current monitoring video of the target to be located to obtain a fused enhanced target monitoring image; The feature calculation layer is used to extract features from the fused enhanced target monitoring image to obtain the current target monitoring features; Based on the current target monitoring features and the target reverse prediction positioning model, the prediction strategy is output to obtain the associated monitoring video of the target to be located; The target reverse prediction positioning model includes a motion feature recognition layer, a reverse prediction strategy layer and a video output layer; The motion feature recognition layer is used to perform motion feature recognition on the current target monitoring features to obtain the motion trajectory features of the target to be located; The reverse prediction strategy layer is used to perform reverse video retrieval based on the monitoring motion trajectory characteristics of the target to be located, and obtain the associated monitoring video of the target to be located; The specific steps of reverse video retrieval in the reverse prediction strategy layer include: Based on the monitoring motion trajectory features of the target to be located, M groups of monitoring video slices are extracted to obtain the monitoring video P of the associated target to be screened. m , m=1, 2, …, M; The surveillance video P of the associated target to be screened m Extract feature representation and obtain the feature representation P of the surveillance video of the associated target to be screened m '; The feature representation of the surveillance video of the associated target to be screened P m 'Use 3D convolutional neural network to perform feature encoding and obtain the embedded vector L of the surveillance video of the associated target to be screened m ; Based on the monitoring motion trajectory characteristics of the target to be located and the monitoring video embedding vector L of the associated target to be screened m Perform prediction matching to obtain the target surveillance video vector matching result J m ; Based on the target surveillance video vector matching results J m For all the associated target surveillance videos P to be screened m Perform screening to obtain surveillance videos associated with the target to be located; Use the current target location, current target monitoring features, and the target to be located to associate the monitoring video and the monitoring prediction positioning model to perform predictive analysis to obtain the predicted target location; perform subsequent operations on the target based on the current target location and the predicted target location; The monitoring prediction and positioning model includes a correlation feature extraction layer, a feature connection layer, and a positioning prediction layer; The associated feature extraction layer is used to input the associated surveillance video of the target to be located into the positioning target fast feature extraction model for feature extraction, and obtain several groups of historical target monitoring features and historical target positioning; The feature connection layer is used to connect several sets of historical target monitoring features with the current target monitoring features to obtain a prediction positioning feature network; and extract trajectory features from several sets of historical target positioning and current target positioning to obtain historical target trajectory features; The positioning prediction layer is used to perform prediction analysis based on the predicted positioning feature network and historical target trajectory features to obtain the predicted target positioning.
2. A rapid positioning method based on monitoring records according to claim 1, characterized in that: The feature output layer is used to output the current target monitoring features; The positioning and recognition layer is used to perform positioning and recognition based on the current target monitoring characteristics to obtain the current target positioning; Specific steps for training the positioning recognition layer: Collecting several groups of positioning and recognition training samples; each group of positioning and recognition training samples contains monitoring features and corresponding target positioning information; combining several groups of positioning and recognition training samples to obtain a positioning and recognition training set; The positioning recognition training set is input into the CNN model for model training to obtain the initial positioning recognition layer; the initial positioning recognition layer is evaluated. If the initial positioning recognition layer passes the model evaluation, the initial positioning recognition layer is used as the positioning recognition layer in the positioning target rapid feature extraction model; otherwise, the positioning recognition training set is used to continue model training.
3. A rapid positioning method based on monitoring records according to claim 2, characterized in that: The specific steps for feature extraction in the feature calculation layer include: The feature calculation layer includes input layer, feature extraction layer, feature fusion layer and output layer; The fused enhanced target monitoring image is received in the input layer, and is sent to the feature extraction layer through the downsampling path for feature calculation; In the feature extraction layer, four high-dimensional feature extractions are performed on the fused enhanced target monitoring image to obtain the enhanced target monitoring image high-dimensional feature T1, the enhanced target monitoring image high-dimensional feature T2, the enhanced target monitoring image high-dimensional feature T3 and the enhanced target monitoring image high-dimensional feature T4; Among them, the specific steps of performing four high-dimensional feature extractions are as follows: the feature extraction layer contains four layers of downsampling structure, and each time a layer of downsampling is performed, a high-dimensional feature extraction is performed on the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T1 of the enhanced target monitoring image is 1 / 4 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T2 of the enhanced target monitoring image is 1 / 8 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T3 of the enhanced target monitoring image is 1 / 16 of the fused enhanced target monitoring image; the feature map resolution of the high-dimensional feature T4 of the enhanced target monitoring image is 1 / 32 of the fused enhanced target monitoring image; In the feature fusion layer, there are four layers of decoders; each layer of decoders corresponds to a downsampling structure; in the decoder, the enhanced target monitoring image high-dimensional features T1 and the enhanced target monitoring image high-dimensional features T2 are fused to obtain the fused enhanced target monitoring image high-dimensional features T 12 ; The enhanced target monitoring image high-dimensional feature T2 and the enhanced target monitoring image high-dimensional feature T3 are fused to obtain the fused enhanced target monitoring image high-dimensional feature T 23 ; The enhanced target monitoring image high-dimensional feature T3 and the enhanced target monitoring image high-dimensional feature T4 are fused to obtain the fused enhanced target monitoring image high-dimensional feature T 34 ; Finally, the high-dimensional features T of the fused enhanced target monitoring image are 12 , fusion enhancement target monitoring image high dimensional features T 23 and fusion enhancement of high-dimensional features of target monitoring images T 34 Perform feature fusion to obtain the current target monitoring features; Output the current target monitoring features in the output layer.
4. A rapid positioning method based on monitoring records according to claim 3, characterized in that: The video output layer is used to output the associated surveillance video of the target to be located.
5. A rapid positioning method based on monitoring records according to claim 4, characterized in that: The positioning prediction layer is constructed based on BP neural network training.
6. A rapid positioning method based on monitoring records according to claim 5, characterized in that: The feature connection layer includes the prediction network construction layer and the trajectory feature extraction layer; In the prediction network construction layer, a 1D convolutional layer is used to aggregate several groups of historical target monitoring features to obtain aggregated historical target monitoring features; The aggregated historical target monitoring features and the current target monitoring features are spliced together to obtain a prediction and positioning feature network; In the trajectory feature extraction layer, a temporal network is established using the LSTM model and several sets of historical target positioning and current target positioning to obtain the historical target trajectory features; Specific steps for training the positioning prediction layer: Collect I groups of continuous positioning prediction training samples, where K is the total time steps included in the continuous positioning prediction training samples; Using the formula As the loss function of the positioning prediction layer; in, is the total number of training samples, For the The actual location of the training samples, For the The predicted position of the training samples, is the corresponding weight coefficient; is the total time steps in the training sample, For the The predicted position of time steps, For the The predicted position of time steps, is the corresponding weight coefficient; Based on the loss function Perform model training to obtain the final positioning prediction layer.
7. A rapid positioning system based on monitoring records, characterized in that: The system applies a rapid positioning method based on monitoring records as described in any one of claims 1 to 6, including: The monitoring feature extraction module includes a video extraction unit and a feature extraction unit; the video extraction unit is used to obtain the current monitoring video of the target to be located; the feature extraction unit is used to analyze the current monitoring video of the target to be located and the positioning target fast feature extraction model to obtain the current target monitoring features and the current target location; the current monitoring video of the target to be located contains n monitoring images of the target to be located; the positioning target fast feature extraction model includes an image preprocessing layer, a feature calculation layer, a feature output layer and a positioning recognition layer, wherein the feature calculation layer gradually extracts, fuses and refines the high-dimensional features of the enhanced target monitoring image through multi-layer downsampling and decoder structure, and finally outputs the current target monitoring features, realizing an efficient feature calculation and fusion process; The monitoring and positioning prediction module includes an associated prediction unit and a positioning prediction unit; the associated prediction unit is used to output a prediction strategy based on the current target monitoring features and the target reverse prediction positioning model to obtain the associated monitoring video of the target to be located; the target reverse prediction positioning model realizes accurate video positioning based on the target motion trajectory by combining motion feature recognition and reverse video retrieval strategy; the positioning prediction unit is used to perform predictive analysis using the current target positioning, the current target monitoring features and the associated monitoring video of the target to be located and the monitoring prediction positioning model to obtain the predicted target positioning; the monitoring prediction positioning model constructs a prediction positioning framework by fusing historical target monitoring features and trajectory information, and combining 1D convolution and LSTM time series modeling, which can accurately perform dynamic target positioning; subsequent operations on the target are performed based on the current target positioning and predicted target positioning.
Citation Information
Patent Citations
Image processing method and system for intelligent security and protection monitoring
CN118887622A