Monitoring analysis method based on multi-view image region synthesis
By constructing a closed-loop self-calibration multi-view dynamic fusion framework, the problems of low target positioning accuracy and error accumulation caused by viewing angle differences and dynamic motion in the multi-view monitoring system are solved, and high-precision multi-view data fusion and self-calibration are achieved, which improves the robustness and real-timeness of the monitoring system in complex scenarios.
Patent Information
- Application Number
- CN202510620685.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing monitoring systems lack the elastic adjustment ability of space-time constraints in complex scenarios due to the differences in the coordinate system of multi-view sensors and dynamic movement of the target, resulting in the difficulty of fusion of multi-source data, target positioning drift and error accumulation, resulting in a decrease in tracking accuracy and insufficient robustness.
A multi-view dynamic fusion framework for closed-loop self-calibration is constructed. Through the adaptive generation of centralized dynamic anchor point sets and multi-view coordinate transformation parameters, combined with the spatiotemporal constraint mechanism of the elastic synthesis area, a Kalman filter is used to generate the target motion trajectory prediction coordinate sequence, and dynamically correct the coordinate transformation parameters and constraint conditions through the closed-loop feedback mechanism to achieve high-precision fusion and self-calibration of multi-view data.
Effectively unify the coordinate system reference of multi-source heterogeneous data, reduce cross-view angle positioning deviation, improve trajectory prediction stability and fault tolerance in complex motion scenarios, enhance the system's tracking sensitivity to target mutation motion, and broaden the application scope of multi-view angle monitoring technology.
Smart Images

Figure CN120495989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-view analysis, and in particular to a monitoring and analysis method based on multi-view image region synthesis. Background Art
[0002] Existing monitoring systems often use multi-view sensors to work together in complex scenarios. However, due to factors such as differences in coordinate systems between different viewpoints, dynamic target motion, and environmental interference, problems such as multi-source data fusion are difficult and target positioning drift occurs. Traditional methods usually rely on fixed coordinate transformation parameters or static anchor points for multi-view alignment, which makes it difficult to adapt to dynamic changes in the target motion trajectory and easily causes error accumulation.
[0003] In addition, existing technologies lack the ability to flexibly adjust to spatiotemporal constraints. In areas of overlapping perspectives or when the target moves rapidly, the predicted trajectory and the observed data are prone to deviations, resulting in reduced tracking accuracy and limiting the real-time and robustness of the monitoring system. Summary of the Invention
[0004] The present invention aims to provide a monitoring and analysis method based on multi-view image region synthesis to address the problems raised in the above-mentioned background technology. Specifically, it addresses the problems of low target positioning accuracy, error accumulation, and poor environmental adaptability caused by perspective differences and dynamic motion in multi-view monitoring systems.
[0005] To achieve the above objectives, the present invention provides the following technical solutions: a monitoring and analysis method based on multi-view image region synthesis. The core innovation of the present invention lies in the construction of a closed-loop self-calibrated multi-view dynamic fusion framework. First, a centralized dynamic anchor point set is generated from multi-view sensor data to eliminate local coordinate system deviations for each view. The rotation matrix and translation vector are calculated based on the singular value decomposition of the covariance matrix, achieving dynamic initialization of the multi-view coordinate transformation parameters.
[0006] Secondly, the concept of boundary parameters of elastic synthesis regions is introduced. Combined with the spatiotemporal constraints of the centralized dynamic anchor point set, the Kalman filter is used to generate the predicted coordinate sequence of the target motion trajectory. The prediction range is adaptively adjusted through the boundary parameters of the elastic synthesis region.
[0007] Subsequently, the multi-view observation coordinates and the predicted coordinates are aligned along the time axis through the target motion chain model, and the spatiotemporal deviation values, including the time deviation value and the space deviation value, are calculated and weightedly fused to form the correction instructions;
[0008] Finally, the coordinate transformation parameters and spatiotemporal constraints are dynamically modified through a closed-loop feedback mechanism to form self-calibration control and continuously improve positioning accuracy.
[0009] A further improvement of this technical solution is to calculate the center point of the dynamic anchor point of each perspective and subtract the corresponding center point from the dynamic anchor point of each perspective to eliminate the position offset between perspectives and enhance the robustness of the coordinate transformation parameters.
[0010] Further improvements to this technical solution integrate time constraints such as displacement, velocity, and acceleration, as well as cross-viewing distance and angle space constraints, to provide a dynamic adjustment basis for the elastic synthesis area.
[0011] A further improvement of this technical solution is to iteratively solve the optimal coordinate transformation parameters with the goal of minimizing the sum of squared errors to ensure the convergence and stability of the closed-loop correction process.
[0012] Compared with the prior art, the present invention has the following beneficial effects:
[0013] Through dynamic anchor point centralization and adaptive generation of multi-view coordinate transformation parameters, the coordinate system benchmark of multi-source heterogeneous data is effectively unified, reducing cross-view positioning deviations. Combined with the spatiotemporal constraint mechanism of the elastic synthesis area, the Kalman filter is used to dynamically predict the target trajectory and define the elastic boundary, improving the stability and fault tolerance of trajectory prediction in complex motion scenarios.
[0014] Furthermore, through a closed-loop feedback mechanism driven by spatiotemporal deviation values, coordinate transformation parameters and constraints are corrected in real time, forming a dynamic error suppression loop. This not only avoids the error accumulation problem caused by fixed parameters in traditional methods, but also enhances the system's tracking sensitivity to sudden changes in target motion.
[0015] Ultimately, multi-level technologies collaborated to achieve high-precision fusion and self-calibration optimization of multi-perspective data, enabling the monitoring system to remain highly robust and real-time in complex scenarios such as rapid target movement, perspective occlusion, or environmental interference, broadening the application scope of multi-perspective monitoring technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of the method steps of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0018] See also Figure 1 The present invention provides a technical solution: a monitoring and analysis method based on multi-view image region synthesis, comprising the following method steps:
[0019] S1. Based on the original target position data collected by the multi-view sensor, a centralized dynamic anchor point set is constructed to generate multi-view coordinate transformation parameters and multi-view observation position sequences;
[0020] S2. extracting the spatiotemporal constraints of the centralized dynamic anchor point set and generating a predicted coordinate sequence of the target motion trajectory within the elastic synthesis area, wherein the boundary parameters of the elastic synthesis area are defined by the spatiotemporal constraints of the centralized dynamic anchor point set;
[0021] S3. Align the multi-view observation position sequence with the predicted coordinate sequence along the time axis to construct a target motion chain model, where each chain node contains the mapping relationship between the predicted coordinates and the multi-view observation coordinates at the same moment, and calculate the spatiotemporal deviation value based on the spatiotemporal continuity of the chain nodes;
[0022] S4. Generate correction instructions based on the spatiotemporal deviation value, dynamically correct the multi-perspective coordinate transformation parameters, and the spatiotemporal constraints of the centralized dynamic anchor point set, so as to optimize the boundary parameters of the elastic synthesis area; wherein the boundary parameters of the elastic synthesis area are inherited from the spatiotemporal constraints of the centralized dynamic anchor point set; the spatiotemporal deviation value calculated by the target motion chain model drives the generation of correction instructions, and the correction instructions are fed back to S1 and S2 to form a closed-loop self-calibration control.
[0023] The specific process of each method step is as follows:
[0024] S1. The original position data collected by multi-view sensors, that is, the original position data of the target object is obtained from sensors with multiple viewpoints (such as cameras, lidar, etc.). The sensors are distributed in different positions and can capture the movement of the target from different angles. The original position data of the target collected by the sensor of each viewpoint i at time point t is , then the multi-view observation position sequence is expressed as , where i represents the index of the perspective and t represents the index of the time point;
[0025] Select key points from the original position data of the target in each perspective as dynamic anchor points and form an initial dynamic anchor point set , where k represents the index of the dynamic anchor point; the key point selection method includes but is not limited to feature extraction algorithm and clustering algorithm;
[0026] In order to simplify the calculation, the dynamic anchor points of each perspective are centralized. By calculating the center point of the dynamic anchor points under each perspective and subtracting the corresponding center point from the dynamic anchor points of each perspective, a centralized dynamic anchor point set is formed. , where the center point The calculation formula is as follows:
[0027] , where K is the total number of dynamic anchor points;
[0028] The calculation formula for subtracting the corresponding center point from the dynamic anchor point of each perspective is as follows:
[0029] ,in Indicates the result of centralization of the dynamic anchor point of each perspective.
[0030] For each pair of perspectives , construct the covariance matrix H, and perform singular value decomposition on the covariance matrix H. Based on the singular value decomposition results, calculate the rotation matrix ; According to the rotation matrix and the center point , calculate the translation vector , the calculation formula is as follows: .
[0031] The rotation matrix and translation vectors As the multi-view coordinate transformation parameters, the multi-view coordinate transformation parameters are used to centralize the dynamic anchor point set under different perspectives. Unified into a common reference coordinate system.
[0032] S2. Extract the spatiotemporal constraints of the centralized dynamic anchor point set, including:
[0033] Spatiotemporal constraints refer to the restrictions and relationships on the centralized dynamic anchor point set in time and space. Spatiotemporal constraints can understand the target's motion pattern and generate a more accurate predicted coordinate sequence of the target's motion trajectory within the elastic synthesis area.
[0034] Analyze the dynamic anchor points at different time points in the centralized dynamic anchor point set and find their temporal change patterns. For each perspective i, the centralized dynamic anchor point set , calculate the displacement, velocity and acceleration between adjacent time points as time constraints;
[0035] Analyze the dynamic anchor points under different perspectives in the centralized dynamic anchor point set and find their relative position relationship in space. For each pair of perspectives ,Calculate the relative position relationship between dynamic anchor points, including distance and angle, as spatial constraints;
[0036] Time constraints and space constraints are used as the spatiotemporal constraints of the centralized dynamic anchor point set.
[0037] The boundary parameters of the elastic synthesis area are defined based on the spatiotemporal constraints of the centralized dynamic anchor point set. The boundary parameters of the elastic synthesis area include the boundary parameters of the time range and the boundary parameters of the spatial range. The process of determining the boundary parameters of the time range is as follows:
[0038] Find the minimum and maximum values of all time points as the initial parameters of the time range boundary; dynamically adjust the time range boundary in the elastic synthesis area according to the speed and acceleration in the time constraint conditions. The specific adjustment formula is as follows:
[0039] ,in Indicates the lower bound of the time range in the adjusted elastic composite area; represents the lower bound of the time range in the initial elastic synthesis region; represents the modulus of the maximum velocity at all time points t and all view angles i; represents the modulus of the maximum acceleration at all time points t and all view angles i; Indicates the sampling period;
[0040] ,in Indicates the upper bound of the time range in the adjusted elastic synthesis area; represents the lower bound of the time range in the initial elastic synthesis region;
[0041] The lower bound of the time range in the adjusted elastic synthesis area and the upper bound of the time range in the elastic synthesis region , as the boundary parameter of the time range.
[0042] The process of determining the boundary parameters of the spatial range is as follows:
[0043] Find the maximum and minimum positions in the set of all centralized dynamic anchor points as the initial boundary parameters of the spatial range. According to the calculated distance and angle, adjust the initial boundary parameters of the spatial range to ensure that the relative position relationship under different perspectives is satisfied within the elastic synthesis area as the boundary parameters of the spatial range.
[0044] The dynamic model is used to generate a predicted coordinate sequence of the target motion trajectory within the elastic synthesis area. The dynamic model is a Kalman filter, which recursively estimates the state of the target through the state equation and observation equation, and combines historical data and current observations to generate a smooth and accurate predicted coordinate sequence, thereby depicting the future motion trajectory of the target within the elastic synthesis area as a predicted coordinate sequence. .
[0045] S3, multi-view observation position sequence With the predicted coordinate sequence Align by time axis to ensure that the observed data and predicted data corresponding to each time point t are synchronized;
[0046] Observe the position sequence according to the viewing angle With the predicted coordinate sequence , construct the target motion chain model, each chain node in the target motion chain model represents the mapping relationship between the predicted coordinates at the same time point and the multi-view observation coordinates, and the chain node Expressed as:
[0047] ;
[0048] The spatiotemporal deviation value is calculated based on the spatiotemporal continuity of the chain nodes. The spatiotemporal deviation value includes the time deviation value and the space deviation value. The time deviation value is calculated by The spatial deviation value is calculated by the changes in displacement, velocity and acceleration between the predicted coordinates of adjacent time points and the multi-view observation coordinates. The Euclidean distance between the predicted coordinates and the multi-view observation coordinates at the same time point is calculated; and the time deviation value and the space deviation value are weighted and summed to obtain the spatiotemporal deviation value.
[0049] S4. Generate correction instructions based on the spatiotemporal deviation value, dynamically correct the multi-view coordinate transformation parameters, and the spatiotemporal constraints of the centralized dynamic anchor point set to optimize the boundary parameters of the elastic synthesis area, specifically including:
[0050] In order to decide whether to generate a correction instruction, a threshold is set , if the spatiotemporal deviation value exceeds this threshold , then generate correction instructions, and according to the correction instructions, dynamically correct the multi-view coordinate transformation parameters, and optimize the multi-view coordinate transformation parameters according to the spatiotemporal deviation value through the least squares method. The specific steps are as follows:
[0051] Define the objective function as a multi-view observation position sequence The sum of square errors with the centralized dynamic anchor point set is then solved to minimize the sum of square errors, thereby finding the optimal rotation matrix and translation vector, which are used as the corrected multi-view coordinate transformation parameters and fed back to S1;
[0052] Based on the corrected multi-view coordinate transformation parameters, re-center the dynamic anchor point set The information is unified into a common reference coordinate system, fed back to S2, and the spatiotemporal constraints of the centralized dynamic anchor point set are modified to optimize the boundary parameters of the elastic synthesis area, forming a closed-loop self-calibration control.
[0053] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A monitoring and analysis method based on multi-view image region synthesis, characterized in that: The method steps are as follows: S1. Based on the original target position data collected by the multi-view sensor, a centralized dynamic anchor point set is constructed to generate multi-view coordinate transformation parameters and multi-view observation position sequences; S2. extracting the spatiotemporal constraints of the centralized dynamic anchor point set and generating a predicted coordinate sequence of the target motion trajectory within the elastic synthesis area, wherein the boundary parameters of the elastic synthesis area are defined by the spatiotemporal constraints of the centralized dynamic anchor point set; S3. Align the multi-view observation position sequence with the predicted coordinate sequence along the time axis to construct a target motion chain model, where each chain node contains the mapping relationship between the predicted coordinates and the multi-view observation coordinates at the same moment, and calculate the spatiotemporal deviation value based on the spatiotemporal continuity of the chain nodes; S4. Generate correction instructions based on the spatiotemporal deviation value, dynamically correct the multi-view coordinate transformation parameters, and the spatiotemporal constraints of the centralized dynamic anchor point set to optimize the boundary parameters of the elastic synthesis area; the spatiotemporal deviation value calculated by the target motion chain model drives the generation of correction instructions, and the correction instructions are fed back to S1 and S2 to form a closed-loop self-calibration control.
2. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The multi-view observation position sequence is acquired based on the target original position data collected by the multi-view sensor.
3. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The generation process of the centralized dynamic anchor point set is as follows: Select key points from the original position data of the target in each perspective as dynamic anchor points and form an initial dynamic anchor point set; The dynamic anchor points of each perspective are centralized by calculating the center point of the dynamic anchor points under each perspective and subtracting the corresponding center point from the dynamic anchor points of each perspective to form a centralized dynamic anchor point set.
4. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The generation process of the multi-view coordinate transformation parameters is as follows: For each pair of perspectives, a covariance matrix is constructed and singular value decomposition is performed on the covariance matrix. Based on the singular value decomposition results, the rotation matrix and translation vector are calculated; the rotation matrix and translation vector are used as multi-perspective coordinate transformation parameters.
5. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The process of extracting the spatiotemporal constraints of the centralized dynamic anchor point set includes: Analyze the dynamic anchor points at different time points in the centralized dynamic anchor point set. For each centralized dynamic anchor point set of each perspective, calculate the displacement, velocity, and acceleration between adjacent time points as time constraints. Analyze the dynamic anchor points in the centralized dynamic anchor point set at different viewing angles. For each pair of viewing angles, calculate the relative position relationship between the dynamic anchor points, including distance and angle, as a spatial constraint condition. Time constraints and space constraints are used as the spatiotemporal constraints of the centralized dynamic anchor point set.
6. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The boundary parameters of the elastic synthesis area include boundary parameters of a time range and boundary parameters of a space range.
7. The monitoring and analysis method based on multi-view image region synthesis according to claim 6, characterized in that: The process of determining the boundary parameters of the time range is as follows: Find the minimum and maximum values of all time points as the initial parameters of the time range boundary; dynamically adjust the boundary of the time range in the elastic synthesis area according to the speed and acceleration in the time constraint conditions, and use the lower boundary of the time range in the elastic synthesis area and the upper boundary of the time range in the elastic synthesis area after adjustment as the boundary parameters of the time range.
8. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The generation process of the predicted coordinate sequence is as follows: A dynamic model is used to generate a predicted coordinate sequence of the target motion trajectory within the elastic synthesis area. The dynamic model is a Kalman filter, which recursively estimates the target state through the state equation and observation equation, and generates a predicted coordinate sequence by combining historical data and current observations.
9. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The calculation process of the spatiotemporal deviation value specifically includes: According to the perspective observation position sequence and the predicted coordinate sequence, a target motion chain model is constructed. Each chain node in the target motion chain model represents the mapping relationship between the predicted coordinates and the multi-perspective observation coordinates at the same time point. The spatiotemporal deviation value is calculated based on the spatiotemporal continuity of the chain nodes. The spatiotemporal deviation value includes the time deviation value and the space deviation value. The time deviation value is calculated by calculating the changes in the displacement, velocity and acceleration between the predicted coordinates of adjacent time points in the chain node and the multi-view observation coordinates; the space deviation value is calculated by the Euclidean distance between the predicted coordinates and the multi-view observation coordinates at the same time point in the chain node; and the time deviation value and the space deviation value are weighted and summed to obtain the spatiotemporal deviation value.
10. The monitoring and analysis method based on multi-view image region synthesis according to claim 1, characterized in that: The dynamic correction process specifically includes: The objective function is defined as the sum of squared errors between the multi-view observation position sequence and the centralized dynamic anchor point set, and the rotation matrix and translation vector are solved to minimize the sum of squared errors, thereby finding the optimal rotation matrix and translation vector, which are used as the corrected multi-view coordinate transformation parameters and fed back to S1; Based on the corrected multi-view coordinate transformation parameters, the centralized dynamic anchor point set is unified into a common reference coordinate system, which is fed back to S2, and the spatiotemporal constraints of the centralized dynamic anchor point set are modified to optimize the boundary parameters of the elastic synthesis area, forming a closed-loop self-calibration control.
Citation Information
Cited By
Intelligent collaborative rail transit vehicle depot automatic maintenance assembly line system
CN121859211A
Intelligent collaborative rail transit depot automated maintenance assembly line system
CN121859211B