Image library construction and recognition algorithm under large-scale incomplete multi-view and multi-mode scene

By deploying sensors and cameras in traffic scenes, building a spatiotemporal alignment matrix and combining Kalman filtering and adversarial generation networks to generate multi-view candidate views, the problem of identifying traffic violations in multi-view and multi-modal scenarios is solved, and efficient target tracking and recognition is achieved.

CN120356099AActive Publication Date: 2025-07-22ZHIYE ELECTRONICS

Patent Information

Application Number
CN202510450344.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-22
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

In incomplete multi-view and multi-modal scenarios, the identification, evidence collection and triggering of traffic violations is complex, and traditional monitoring methods are difficult to achieve comprehensive coverage and effective integration of multi-view and multi-modal data.

Method used

By deploying multiple sensors and cameras in traffic scenes, building a spatiotemporal alignment matrix, unifying multimodal data to the same spatiotemporal coordinate system, combining Kalman filtering and adversarial generation network to generate multi-view candidate views, integrating target kinematics models and deep learning technology, achieving cross-camera target tracking and traffic violation identification.

Benefits of technology

It improves the accuracy of target recognition at multiple perspectives, reduces the recognition error caused by missing or insufficient perspectives, ensures the consistency of data in time and space dimensions, and provides rich image library content and accurate behavior recognition capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356099A_ABST
    Figure CN120356099A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of behavior recognition, in particular to an image library construction and recognition algorithm under a large-scale incomplete multi-view and multi-modal scene, which comprises the following steps of: deploying a plurality of sensors and cameras in a traffic scene, and unifying multi-modal data into the same space-time coordinate system; fusing the multi-modal data subjected to space-time alignment, predicting a target state and updating a measurement value; dynamically generating multi-view candidate views based on the fused data by using an adversarial generative network under space-time constraint, and complementing missing views in combination with a target kinematics model; a target kinematic model and a deep learning technology are fused, a cross-camera target is tracked, a traffic regulation ontology library is constructed, and traffic illegal behaviors are automatically identified and classified in combination with an image identification technology. According to the method, the consistency of different sensor data in time and space dimensions is ensured, and data chaos and errors caused by space-time differences are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of behavior recognition, and particularly to an image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario. Background Art

[0002] With the acceleration of the urbanization process and the rapid growth of traffic flow, traffic law enforcement faces unprecedented challenges. The modern urban traffic network is intricate and the traffic flow is huge. Traditional monitoring means are difficult to achieve full coverage of all road sections. Especially in some key areas, traffic violations occur frequently and more efficient monitoring means are needed. In the process of traffic law enforcement, data from different perspectives and different modalities need to be processed, and there are often correlations and complementarities between these data. How to effectively integrate and utilize these data has become a difficult problem.

[0003] Traditional traffic law enforcement means, such as manual patrols and on-site law enforcement, are difficult to meet the needs of modern traffic management. Especially in an incomplete multi-view and multi-modal scenario, the recognition, evidence collection, and triggering of traffic violations become particularly complex. Summary of the Invention

[0004] The present invention aims at the technical problems existing in the prior art and provides an image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario.

[0005] The technical solution of the present invention to solve the above technical problems is as follows: An image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario, specifically including the following steps: S101. Deploy a variety of sensors and cameras in the traffic scenario, collect their position, orientation, and time synchronization information, use this information to construct a spatio-temporal alignment matrix, and through matrix operations, unify the multi-modal data into the same spatio-temporal coordinate system; S102. Fuse the spatio-temporally aligned multi-modal data, predict the target state and update the measurement value, fuse different sensor data, and enhance the integrity of the image information; S103. Based on the fused data, use an adversarial generation network under spatio-temporal constraints to dynamically generate multi-view candidate views, combine with the target kinematic model, simulate the appearance and motion state of the target under different views, and complete the missing views; S104. Integrate the target kinematic model and deep learning technology, track the target across cameras, construct a traffic regulation ontology library, and combine with image recognition technology to automatically identify and classify traffic violations; In a preferred embodiment, in S101, the geographical location information of each sensor and camera is obtained using GPS, including longitude, latitude, and altitude. These are recorded and associated with the corresponding devices. A Network Time Protocol (NTP) server is equipped for all sensors and cameras to ensure time synchronization between devices, and the timestamp information of each device is recorded. For each sensor and camera, according to its position and orientation information, the transformation parameters from its own coordinate system to the unified coordinate system O-XYZ are calculated. The transformation parameters include a translation vector and a rotation matrix. The translation vector represents the position offset of the sensor and camera in the unified coordinate system. Subtracting the origin coordinates of the unified coordinate system from the position coordinates of the sensor and camera gives the translation vector. The calculation of the translation vector is based on the origin coordinates O=(O x ,O y ,O z ) of the unified coordinate system O-XYZ and the local coordinate system o-xyz of the sensor and camera. The coordinates of the origin of the sensor and camera in the unified coordinate system are P=(P x ,P y ,P z ). Then the calculation formula for the translation vector is as follows: ; where T is used to describe the position offset from the origin of the unified coordinate system to the origin of the local coordinate system. T x represents the position offset of the sensor and camera relative to the origin along the x-axis direction in the unified coordinate system. T y represents the position offset of the sensor and camera relative to the origin along the y-axis direction in the unified coordinate system. T z represents the position offset of the sensor and camera relative to the origin along the z-axis direction in the unified coordinate system. In practical applications, if the origin of the unified coordinate system is located at (0, 0, 0), then the translation vector is directly equal to the coordinates of the origin of the sensor and camera in the unified coordinate system, that is, T=(P x ,P y ,P z ); The rotation matrix represents the angular relationship between the orientation of the sensor and camera and the axes of the unified coordinate system. According to the horizontal and vertical angles of the sensor and camera, using the calculation formula of the rotation matrix, the rotation matrix is obtained. The rotation matrix is used to describe the rotation relationship of the local coordinate system relative to the unified coordinate system. The rotation angles are the yaw angle ψ, pitch angle θ, and roll angle Φ respectively. The translation vector and rotation matrix of each sensor and camera are combined into a 4×4 homogeneous transformation matrix, which is used as the spatio-temporal alignment matrix of the device. The general form of the homogeneous transformation matrix is: ; Among them, R 3×3 represents a rotation matrix, T 3×1 represents a translation vector, and 0 1×3 is a 1×3 zero vector.

[0006] Start all sensors and cameras to begin collecting multimodal data, including images, radar data, and speed data. At the same time, record the timestamp and corresponding device information for each time point. For each collected data point, use the corresponding spatio-temporal alignment matrix for coordinate transformation to convert it from the device coordinate system to a unified coordinate system. Multiply the coordinate vector of the data point by the spatio-temporal alignment matrix to obtain the coordinates in the unified coordinate system of the device. According to the timestamp information for each time point, perform time alignment on the multimodal data. Using the method of linear interpolation, unify the data collected by different devices at different times to the same time point. Suppose there are two adjacent time points t1 and t2 during data collection by the device, and the corresponding multimodal data values are v1 and v2 respectively. If we want to obtain the difference data value v at time point t, the corresponding specific calculation formula is as follows: ; Among them, v represents the difference data value at time point t, v1 represents the multimodal data value obtained by the sensor and other data collection devices at time point t1, and v2 represents the multimodal data value obtained by the sensor and other data collection devices at time point t2.

[0007] In a preferred embodiment, in S102, for the multimodal data collected by each sensor and camera, represent it in homogeneous coordinate form. Through matrix multiplication operations, convert the data from the local coordinate systems of the sensors and cameras to the unified spatio-temporal coordinate system. According to the motion characteristics of the targets in the traffic scene, define state variables and apply Kalman filtering for data fusion, which specifically includes the following steps: S1. Define state variables: The state variables of the vehicle target include its position (x, y, z), speed (v x , v y , v z ), and acceleration (a x , a y , a z ) in the unified spatio-temporal coordinate system, forming the state vector = [x, y, z, v x , v y , v z , a x , a y , a z . S1. Define state variables: According to the motion law of the vehicle target, establish the state transition equation: ; Among them, represents the state vector of the vehicle target at time k, and F k represents the state transition matrix, which describes the change relationship of the state of the vehicle target from time k - 1 to time k. represents the process noise; S2. Initialize the Kalman filter parameters: Estimate the initial state of the vehicle target according to the initial measurement value of a certain sensor in the multimodal data, describe the uncertainty of the initial state estimate through the initial covariance matrix P0, and set the process noise covariance matrix Q to quantify the statistical characteristics of the process noise ; S3. Prediction: According to the state transition equation and the state estimate of the vehicle target at the previous moment , calculate the state prediction value of the vehicle target at the current moment , calculate the covariance prediction value at the current moment , among which, represents the prediction of the state estimate uncertainty at the current moment before there is no new measurement data; S4. Update: Preprocess the measurement values of each sensor in the multimodal data after spatio-temporal alignment, including noise removal, outlier processing, and data format conversion operations, and then combine these measurement values into a measurement vector , establish a measurement equation according to the characteristics and measurement principle of the sensor , among which, H k represents the measurement matrix, which maps the state vector to the measurement space, represents the measurement noise. According to the covariance prediction value , measurement matrix H k and measurement noise covariance matrix R k , calculate the Kalman gain , use the measurement vector , Kalman gain K k and the vehicle target state prediction value , update the state estimate value at the current moment , according to the Kalman gain K k and covariance prediction value , update the covariance matrix at the current moment , where I represents the identity matrix; S5. Output: Use the updated state estimate value as the fused vehicle target state for subsequent view generation, target tracking, and behavior analysis processing. At the same time, use the covariance matrix P kSave it as the initial covariance matrix for the Kalman filter at the next moment, and repeat the prediction and update steps to achieve continuous fusion of multi-modal data.

[0008] In a preferred embodiment, the multi-modal data after Kalman filter fusion is used in S103, including the position, speed, and acceleration information of the vehicle target. Taking the fused data and the random noise vector as inputs, the generator generates candidate views from different perspectives according to the inputs. Multiple views are generated to cover different observation angles. A network structure is constructed using convolutional layers, deconvolutional layers, and fully connected layers. The convolutional layers are used to extract features, the deconvolutional layers are used to generate images, and the fully connected layers are used to integrate information. The discriminator receives the candidate views generated by the generator and the real multi-perspective images, extracts features and classifies the received images. Using the target kinematic model, based on the vehicle target state and historical motion information at the current moment, predict the future state of the target from different perspectives. If the current position of the target is (x1, y1), the speed is (vx1, vy1), the acceleration is (ax1, ay1), and the time interval is ∆t, the position (x2, y2) at the next moment can be calculated by the following formula: ; Among them, (x2, y2) represents the position coordinates of the target at the next moment. Input the predicted target state into the trained generator to generate candidate views from different perspectives. For the missing views, use the linear interpolation method for inference and completion. Suppose the target states (x1, y1) and (x2, y2) at the known adjacent perspectives δ1 and δ2 are known. The calculation formula for inferring the target state at the perspective δ is as follows: ; Among them, (B x , B y ) represents the position coordinates of the target at the perspective δ calculated by the linear interpolation method. According to the inferred target state, use the generator to generate the corresponding candidate views to complete the view library. Store the selected multi-perspective candidate views in the database and establish an index for each candidate view. The index information includes the view shooting time, perspective parameters, target category, and position information.

[0009] In a preferred embodiment, in S104, based on a target monitoring model of a convolutional neural network, a large amount of multi-view image data containing various targets is collected, and the targets are accurately labeled. The labeling information includes target categories and position boxes. The labeled data is used to train the convolutional neural network model. The model parameters are continuously adjusted through the backpropagation algorithm. The trained deep learning model is used to detect targets in the images captured by each camera. For the detected targets, by comparing the features of the targets and the position information in the images, a data association algorithm is used to associate the same target captured by different cameras, and a unique identifier is established for each target. The position predicted by the target kinematic model is fused with the actual position of the target detected by the deep learning model. When identifying a target, the target image T captured by the current camera A is matched with the image T B in the multi-view candidate view library. By calculating the similarity between the images, the candidate view most similar to the current target image is retrieved. The specific calculation formula is as follows: ; where sim represents the similarity, m represents the dimension of the feature vector, and respectively represent the i-th element of the feature vectors of T A and T B . All traffic regulations are comprehensively collected, and key concepts are extracted from the regulation clauses, including the constituent elements and judgment criteria of illegal acts. Using different traffic illegal acts, corresponding image features are extracted by the deep learning model. Video and image data containing traffic illegal acts are collected, and the collected data is labeled to mark different types of illegal acts and related features, including signal light status, vehicle position, and vehicle speed. The multi-view candidate views and real-time monitoring images are input, and multiple convolutional layers and pooling layers are used to extract the spatial features of the images. The output of the convolutional layer is flattened and connected to the fully connected layer for classifying the signal light status, and the signal light status and the relative position relationship of the vehicle are output. The trained model is integrated into the actual traffic recognition to identify traffic illegal acts in real time.

[0010] The beneficial effects of the present invention are as follows: By collecting the position, orientation, and time synchronization information of sensors and cameras, the present invention constructs a spatio-temporal alignment matrix to unify multi-modal data into the same spatio-temporal coordinate system, ensuring the consistency of different sensor data in the time and space dimensions, providing a solid foundation for subsequent accurate data processing and analysis, and avoiding data chaos and errors caused by spatio-temporal differences. Applying the Kalman filtering algorithm to fuse the multi-modal data after spatio-temporal alignment can effectively process the noise and uncertainty in the data, not only enhancing the accuracy and integrity of the image information, but also improving the quality of subsequent view generation, making the generated image closer to the real scene, which is conducive to more accurately identifying targets and behaviors. Dynamically generating multi-view candidate views using an adversarial generative network under spatio-temporal constraints, and combining with the target kinematic model to complete the missing views. At the same time, screening high-quality candidate views to construct a multi-view candidate view library, which greatly enriches the content of the image library, enabling the image library to cover a wider range of views and scenes, helping to improve the recognition accuracy of the target under different views and reducing the recognition errors caused by missing or insufficient views. Fusing the target kinematic model and deep learning technology, using the deep learning model to extract target features, and combining with the kinematic model to predict the target trajectory, realizing the accurate tracking of cross-camera targets. Based on the multi-view candidate view library, more abundant information is provided for target recognition, enabling the system to observe the target from multiple angles, thereby improving the recognition accuracy of the target under different views. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a flowchart of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0013] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0014] In the description of the present application, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without the use of these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but rather to be in line with the broadest scope consistent with the principles and features disclosed in the present application.

[0015] As Figure 1 , this embodiment provides: an image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario, specifically including the following steps: S101. Deploy a variety of sensors and cameras in the traffic scenario, collect their position, orientation, and time synchronization information, use this information to construct a spatio-temporal alignment matrix, and through matrix operations, unify the multi-modal data into the same spatio-temporal coordinate system; Furthermore, use GPS to obtain the geographical location information of each sensor and camera, including longitude, latitude, and altitude, record these, and associate them with the corresponding devices. Equip all sensors and cameras with a Network Time Protocol server to ensure time synchronization between devices, and record the timestamp information of each device. For each sensor and camera, calculate the conversion parameters from its own coordinate system to the unified coordinate system O-XYZ according to its position and orientation information. The conversion parameters include a translation vector and a rotation matrix; The translation vector represents the position offset of the sensor and the camera in the unified coordinate system. Subtract the origin coordinates of the unified coordinate system from the position coordinates of the sensor and the camera to obtain the translation vector. The calculation of the translation vector is based on the origin coordinates O=(O x , O y , O z ) of the unified coordinate system O-XYZ and the local coordinate system o-xyz of the sensor and the camera. The coordinates of the origin of the sensor and the camera in the unified coordinate system are P=(P x , P y , P z ). Then the calculation formula of the translation vector is as follows: ; where T is used to describe the position offset from the origin of the unified coordinate system to the origin of the local coordinate system, T x represents the position offset amount of the sensor and the camera relative to the origin along the x-axis direction in the unified coordinate system, Ty Represents the position offset of the sensor and camera along the y-axis relative to the origin in the unified coordinate system, T z Represents the position offset of the sensor and camera along the z-axis relative to the origin in the unified coordinate system. In practical applications, if the origin of the unified coordinate system is at (0, 0, 0), then the translation vector is directly equal to the coordinates of the origin of the sensor and camera in the unified coordinate system, i.e., T=(P x , P y , P z ); The rotation matrix represents the angular relationship between the orientation of the sensor and camera and the axes of the unified coordinate system. Based on the horizontal and vertical angles of the sensor and camera, the rotation matrix is obtained using the calculation formula of the rotation matrix. The rotation matrix is used to describe the rotation relationship of the local coordinate system relative to the unified coordinate system, and the rotation angles are the yaw angle ψ, pitch angle θ, and roll angle Φ respectively; Combine the translation vector and rotation matrix of each sensor and camera into a 4×4 homogeneous transformation matrix as the spatio-temporal alignment matrix of the device. The general form of the homogeneous transformation matrix is: ; Among them, R 3×3 represents the rotation matrix, T 3×1 represents the translation vector, and 0 1×3 is a 1×3 zero vector.

[0016] Start all sensors and cameras to begin collecting multimodal data, including images, radar data, and speed data. At the same time, record the timestamp and corresponding device information for each time point. For each collected data point, use the corresponding spatio-temporal alignment matrix for coordinate transformation to convert it from the device coordinate system to the unified coordinate system. Multiply the coordinate vector of the data point by the spatio-temporal alignment matrix to obtain the coordinates in the unified coordinate system of the device. According to the timestamp information of each time point, perform time alignment on the multimodal data. Using the method of linear interpolation, unify the data collected by different devices at different times to the same time point. Suppose there are two adjacent time points t1 and t2 when the device is collecting data, and the corresponding multimodal data values are v1 and v2 respectively. We want to obtain the difference data value v at time point t. Then the corresponding specific calculation formula is as follows: ; Among them, v represents the difference data value at time point t, v1 represents the multimodal data value obtained by the sensor and other data collection devices at time point t1, and v2 represents the multimodal data value obtained by the sensor and other data collection devices at time point t2.

[0017] It should be noted that there is a data point in the local coordinate system of the device for coordinate transformation, and its homogeneous coordinates are represented by a vector containing four elements. These four elements are the coordinate values of the data point in the x-axis direction, y-axis direction, and z-axis direction in the local coordinate system, and the four elements of the fixed position 1. When performing data transformation from the device's coordinate system, it is necessary to perform matrix multiplication on the vector representing the homogeneous coordinates of the data point and the spatio-temporal alignment matrix.

[0018] It should be noted that the rotation matrix R z (ψ) for rotation about the Z-axis is: ; The rotation matrix R y (θ) for rotation about the Y-axis is: ; The rotation matrix R x (Φ) for rotation about the X-axis is: ; S102. Fuse the multi-modal data that has undergone spatio-temporal alignment, predict the target state, and update the measurement values to fuse different sensor data and enhance the integrity of the image information. Furthermore, for the multi-modal data collected by each sensor and camera, represent it in the form of homogeneous coordinates. Through matrix multiplication, transform the data from the local coordinate systems of the sensors and cameras to the unified spatio-temporal coordinate system. According to the motion characteristics of the targets in the traffic scene, define state variables and apply Kalman filtering for data fusion, which specifically includes the following steps: S1. Define state variables: The state variables of the vehicle target include its position (x, y, z), velocity (v x , v y , v z ), and acceleration (a x , a y , a z ) in the unified spatio-temporal coordinate system, which form the state vector = [x, y, z, v x , v y , v z , a x , a y , a z . S1. Define state variables: According to the motion law of the vehicle target, establish the state transition equation: ; Among them, represents the state vector of the vehicle target at time k, and F kdenotes the state transition matrix, which describes the change relationship of the state of the vehicle target from time k-1 to time k. denotes the process noise; S2. Initialize the Kalman filter parameters: Estimate the initial state of the vehicle target based on the initial measurement value of a certain sensor in the multi-modal data, describe the uncertainty of the initial state estimate through the initial covariance matrix P0, and set the process noise covariance matrix Q to quantify the statistical characteristics of the process noise ; S3. Prediction: According to the state transition equation and the state estimate of the vehicle target at the previous moment , calculate the state prediction value of the vehicle target at the current moment , calculate the covariance prediction value at the current moment , where denotes the prediction of the uncertainty of the state estimate at the current moment before there is no new measurement data; S4. Update: Preprocess the measurement values of each sensor in the multi-modal data after spatio-temporal alignment, including noise removal, outlier processing, and data format conversion operations, and then combine these measurement values into a measurement vector , establish a measurement equation according to the characteristics and measurement principles of the sensor , where H k denotes the measurement matrix, which maps the state vector to the measurement space, denotes the measurement noise, according to the covariance prediction value , measurement matrix H k and measurement noise covariance matrix R k , calculate the Kalman gain , use the measurement vector , Kalman gain K k and the state prediction value of the vehicle target , update the state estimate value at the current moment , according to the Kalman gain K k and covariance prediction value , update the covariance matrix at the current moment , where I denotes the identity matrix; S5. Output: Use the updated state estimate value as the fused state of the vehicle target for subsequent view generation, target tracking, and behavior analysis processing. At the same time, save the covariance matrix P k as the initial covariance matrix for the next moment's Kalman filter, and repeat the prediction and update steps to achieve continuous fusion of multi-modal data.

[0019] S103. Based on the fused data, use the adversarial generative network under spatio-temporal constraints to dynamically generate multi-view candidate views, and combine with the target kinematic model to simulate the appearance and motion states of the target from different perspectives to complete the missing views. Further, use the multi-modal data fused by Kalman filtering, including the position, speed, and acceleration information of the vehicle target. Take the fused data and a random noise vector as inputs. The generator generates candidate views from different perspectives according to the inputs, generating multiple views to cover different observation angles. Use convolutional layers, deconvolutional layers, and fully connected layers to construct the network structure. The convolutional layers are used to extract features, the deconvolutional layers are used to generate images, and the fully connected layers are used to integrate information. The discriminator receives the candidate views generated by the generator and the real multi-view images, extracts features and classifies the received images. Use the target kinematic model to predict the future states of the target from different perspectives according to the current state of the vehicle target and historical motion information. If the current position of the target is (x1, y1), the speed is (vx1, vy1), the acceleration is (ax1, ay1), and the time interval is ∆t, the position (x2, y2) at the next moment can be calculated by the following formula: ; Among them, (x2, y2) represents the position coordinates of the target at the next moment. Input the predicted target state into the trained generator to generate candidate views from different perspectives. For the missing views, use the linear interpolation method for inference and completion. Suppose the target states (x1, y1) and (x2, y2) at the known adjacent perspectives δ1 and δ2 are known. The formula for inferring the target state at perspective δ is as follows: ; Among them, (B x , B y ) represents the position coordinates of the target at perspective δ calculated by the linear interpolation method. According to the inferred target state, use the generator to generate the corresponding candidate views to complete the view library. Store the selected multi-view candidate views in the database and establish an index for each candidate view. The index information includes the view shooting time, perspective parameters, target category, and position information.

[0020] Construct an image library by combining S101 to S103. Through the multi-sensor spatio-temporal alignment matrix, convert data in different coordinate systems into a unified world coordinate system, realize the time synchronization of multi-modal data, reduce the spatial coordinate conversion error, solve the spatio-temporal heterogeneity problem of multi-sensor data fusion, use the Kalman filtering algorithm to perform optimal estimation on multi-source data, reduce the error of target position measurement, generate virtual perspective images through spatio-temporal constraints, realize the monitoring ability without dead angles, and perform spatio-temporal joint completion based on the kinematic model and linear interpolation method to make up for the deficiency of missing view completion in traditional monitoring, providing a reliable data basis for subsequent target tracking and behavior recognition.

[0021] S104. Integrate the target kinematic model and deep learning technology to track targets across cameras, construct a traffic regulation ontology library, and combine image recognition technology to automatically identify and classify traffic violations; Furthermore, based on the target monitoring model of convolutional neural network, collect a large amount of multi-view image data containing various targets, accurately label the targets, and the labeling information includes target categories and position boxes. Use the labeled data to train the convolutional neural network model, continuously adjust the model parameters through the backpropagation algorithm, use the trained deep learning model to detect targets in the images captured by each camera. For the detected targets, by comparing the features of the targets and the position information in the images, use the data association algorithm to associate the same target captured by different cameras, establish a unique identifier for each target, and fuse the position predicted by the target kinematic model with the actual position of the target detected by the deep learning model. When identifying a target, the target image T captured by the current camera A is matched with the image T B in the multi-view candidate view library. By calculating the similarity between the images, retrieve the candidate view that is most similar to the current target image. The specific calculation formula is as follows: ; where sim represents the similarity, m represents the dimension of the feature vector, and respectively represent the T A and T BThe i-th element, comprehensively collect traffic regulations, extract key concepts from the regulations, including the constituent elements and judgment criteria of illegal acts. Utilize different traffic violations, use a deep learning model to extract corresponding image features, collect video and image data containing traffic violations, annotate the collected data, mark different types of illegal acts and related features, including signal light status, vehicle position, and vehicle speed. Input multi-view candidate views and real-time monitoring images, use multiple convolutional layers and pooling layers to extract the spatial features of the images, flatten the output of the convolutional layer, and connect it to a fully connected layer for classifying the signal light status, and output the signal light status and the relative position relationship of the vehicle. Integrate the trained model into actual traffic recognition to identify traffic violations in real time.

[0022] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0023] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0024] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0025] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0026] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks Figure 1 in the steps of the processes.

[0027] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0028] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An image library construction and recognition algorithm in large-scale incomplete multi-view and multi-modal scenarios, characterized in that, It includes the following steps: S101. Deploy multiple sensors and cameras in the traffic scenario, collect their location, orientation and time synchronization information, use this information to construct a spatio-temporal alignment matrix, and through matrix operations, unify the multi-modal data into the same spatio-temporal coordinate system; S102. Fusion the spatio-temporally aligned multi-modal data, predict the target state and update the measurement value, fuse different sensor data, and enhance the integrity of the image information; S103. Based on the fused data, use the adversarial generation network under spatio-temporal constraints to dynamically generate multi-view candidate views, combine with the target kinematic model, simulate the appearance and motion state of the target under different views, and complete the missing views; S104. Fusion the target kinematic model and deep learning technology, track the target across cameras, construct a traffic regulation ontology library, and combine with image recognition technology to automatically identify and classify traffic violations.

2. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 1, characterized in that, In the S101, use GPS to obtain the geographical location information of each sensor and camera, including longitude, latitude and altitude, record these, and associate them with the corresponding devices, equip all sensors and cameras with a network time protocol server to ensure time synchronization between devices, and record the timestamp information of each device. For each sensor and camera, according to its position and orientation information, calculate the conversion parameters from its own coordinate system to the unified coordinate system O-XYZ. The conversion parameters include the translation vector and the rotation matrix.

3. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 2, characterized in that, The translation vector represents the position offset of the sensor and the camera in the unified coordinate system. Subtract the origin coordinates of the unified coordinate system from the position coordinates of the sensor and the camera to obtain the translation vector. The calculation of the translation vector is based on the origin coordinates O=(O x ,O y ,O z ) of the unified coordinate system O-XYZ, and the local coordinate systems o-xyz of the sensor and the camera. The coordinates of the origins of the sensor and the camera in the unified coordinate system are P=(P x ,P y ,P z ). Then the calculation formula of the translation vector is as follows: ; Among them, T is used to describe the position offset from the origin of the unified coordinate system to the origin of the local coordinate system, and T x represents the position offset of the sensor and the camera along the x-axis direction relative to the origin in the unified coordinate system, and T y represents the position offset of the sensor and the camera along the y-axis direction relative to the origin in the unified coordinate system, and T z represents the position offset of the sensor and the camera along the z-axis direction relative to the origin in the unified coordinate system. In practical applications, if the origin of the unified coordinate system is located at (0, 0, 0), then the translation vector is directly equal to the coordinates of the origin of the sensor and the camera in the unified coordinate system, that is, T = (P x , P y , P z ).

4. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 2, characterized in that The rotation matrix represents the angular relationship between the orientation of the sensor and the camera and the coordinate axes of the unified coordinate system. According to the horizontal and vertical angles of the sensor and the camera, use the calculation formula of the rotation matrix to obtain the rotation matrix. The rotation matrix is used to describe the rotation relationship of the local coordinate system relative to the unified coordinate system, and the rotation angles are the yaw angle ψ, the pitch angle θ and the roll angle Φ respectively.

5. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 3, characterized in that, Combine the translation vector and the rotation matrix of each sensor and camera into a 4×4 homogeneous transformation matrix as the spatio-temporal alignment matrix of the device. The general form of the homogeneous transformation matrix is: ; where, R 3×3 represents a rotation matrix, T 3×1 represents a translation vector, and 0 1×3 is a 1×3 zero vector.

6. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 1, characterized in that Start all sensors and cameras to start collecting multi-modal data, including images, radar data, speed data, and at the same time record the timestamp and the corresponding device information at each time point. For each collected data point, use the corresponding spatio-temporal alignment matrix for coordinate transformation to convert it from the device coordinate system to the unified coordinate system. Multiply the coordinate vector of the data point by the spatio-temporal alignment matrix to obtain the coordinate in the unified coordinate system of the device. According to the timestamp information at each time point, perform time alignment on the multi-modal data. Use the method of linear interpolation to unify the data collected by different devices at different times to the same time point. Suppose there are two adjacent time points t1 and t2 when the device is collecting data, and the corresponding multi-modal data values are v1 and v2 respectively. We want to obtain the difference data value v at the time point t. Then the corresponding specific calculation formula is as follows: ; Among them, v represents the difference data value at time point t, v1 represents the multi-modal data value obtained by the sensor and other data acquisition devices at time point t1, and v2 represents the multi-modal data value obtained by the sensor and other data acquisition devices at time point t2.

7. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 1, characterized in that, In S102, for the multi-modal data collected by each sensor and camera, it is represented in homogeneous coordinate form. Through matrix multiplication operations, the data is transformed from the local coordinate systems of the sensor and camera to the unified spatio-temporal coordinate system. According to the motion characteristics of the targets in the traffic scene, state variables are defined, and Kalman filtering is applied for data fusion, which specifically includes the following steps: S1. Define state variables: The state variables of the vehicle target include its position (x, y, z), velocity (v x , v y , v z ), acceleration (a x , a y , a z ) in the unified space-time coordinate system, which form a state vector = [x, y, z, v x , v y , v z , a x , a y , a z . S1. Define state variables: According to the motion law of the vehicle target, establish a state transition equation: ; Among them, represents the state vector of the vehicle target at time k, and F k represents the state transition matrix, which describes the change relationship of the state of the vehicle target from time k-1 to time k, represents the process noise; S2. Initialize the Kalman filter parameters: Estimate the initial state of the vehicle target based on the initial measurement value of a certain sensor in the multi-modal data, describe the uncertainty of the initial state estimate through the initial covariance matrix P0, and set the process noise covariance matrix Q to quantify the statistical characteristics of the process noise ; ​ S3. Prediction: Estimate the vehicle target state according to the state transition equation and the vehicle target state at the previous moment , calculate the predicted value of the state of the vehicle target at the current moment , calculate the predicted value of the covariance at the current moment , where represents the prediction of the uncertainty of the state estimation at the current moment before new measurement data arrives; S4. Update: Preprocess each sensor measurement value in the spatio-temporally aligned multi-modal data, including noise removal, outlier handling, and data format conversion operations, and then combine these measurement values into a measurement vector. , and establish a measurement equation according to the characteristics and measurement principles of the sensors. , where H k represents the measurement matrix, which maps the state vector to the measurement space. represents the measurement noise. According to the covariance prediction value , measurement matrix H k and measurement noise covariance matrix R k , calculate the Kalman gain . Using the measurement vector , Kalman gain K k and the predicted value of the vehicle target state , update the state estimate value at the current moment . According to the Kalman gain K k and the covariance prediction value , update the covariance matrix at the current moment , where I represents the identity matrix. S5. Output the updated state estimate value As the fused vehicle target state, it is used for subsequent view generation, target tracking, and behavior analysis processing. At the same time, the covariance matrix P k is saved as the initial covariance matrix for the Kalman filter at the next moment. The prediction and update steps are repeated to achieve continuous fusion of multimodal data.

8. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 1, characterized in that In S103, the multi-modal data after Kalman filtering fusion is used, including the position, speed, and acceleration information of the vehicle target. The fused data and the random noise vector are used as inputs. The generator generates candidate views from different perspectives according to the inputs. Multiple views are generated to cover different observation angles. A network structure is constructed using convolutional layers, deconvolutional layers, and fully connected layers. The convolutional layers are used to extract features, the deconvolutional layers are used to generate images, and the fully connected layers are used to integrate information. The discriminator receives the candidate views generated by the generator and the real multi-perspective images, extracts features and classifies the received images. Using the target kinematic model, according to the vehicle target state and historical motion information at the current moment, the future state of the target from different perspectives is predicted. If the current position of the target is (x1, y1), the speed is (vx1, vy1), the acceleration is (ax1, ay1), and the time interval is ∆t, the position (x2, y2) at the next moment can be calculated by the following formula: ; Among them, (x2, y2) represents the position coordinates of the target at the next moment. The predicted target state is input into the trained generator to generate candidate views from different perspectives. For the missing views, linear interpolation methods are used for inference and completion. Suppose the target states (x1, y1) and (x2, y2) at the known adjacent perspectives δ1 and δ2 are known. The calculation formula for inferring the target state at perspective δ is as follows: ; Among them, (B x , B y ) represents the position coordinates of the target at the viewing angle δ calculated by the linear interpolation method. According to the inferred target state, the generator is used to generate corresponding candidate views to complete the view library. The selected multi-view candidate views are stored in the database, and an index is established for each candidate view. The index information includes the view capture time, viewing angle parameters, target category, and location information.

9. The image library construction and recognition algorithm in a large-scale incomplete multi-view and multi-modal scenario according to claim 1, characterized in that, In S104, based on the object monitoring model of the convolutional neural network, a large amount of multi-view image data containing various objects is collected, and the objects are accurately labeled. The labeling information includes the object category and the position box. The labeled data is used to train the convolutional neural network model, and the model parameters are continuously adjusted through the backpropagation algorithm. The trained deep learning model is used to detect objects in the images captured by each camera. For the detected objects, by comparing the features of the objects and the position information in the images, the data association algorithm is used to associate the same object captured by different cameras, and a unique identifier is established for each object. The position predicted by the object kinematic model is fused with the actual position of the object detected by the deep learning model. When identifying an object, the target image T captured by the current camera A is matched with the image T B in the multi-view candidate view library. By calculating the similarity between the images, the candidate view most similar to the current target image is retrieved. The specific calculation formula is as follows: ; where sim represents the similarity and m represents the dimension of the feature vector. and respectively represent the i-th elements of the feature vectors T A and T B Comprehensively collect traffic regulations, extract key concepts from the regulations, including the constituent elements and judgment criteria of illegal acts. Utilize different traffic illegal acts to extract corresponding image features using a deep learning model. Collect video and image data containing traffic illegal acts, annotate the collected data, mark different types of illegal acts and related features, including signal light status, vehicle position, and vehicle speed. Input multi-view candidate views and real-time monitoring images, use multiple convolutional layers and pooling layers to extract the spatial features of the images, flatten the output of the convolutional layer, and connect it to a fully connected layer for classifying the signal light status, outputting the signal light status and the relative position relationship of the vehicle. Integrate the trained model into actual traffic recognition to identify traffic illegal acts in real time.

Citation Information

Patent Citations

  • Vehicle re-identification method and device based on multi-modal information fusion

    CN111931627A

  • Multi-sensor space-time cooperative calibration method for fusion perception of camera and millimeter wave radar

    CN115018929A

  • Multi-target tracking method and device under aerial view angle

    CN115984586A

  • Incomplete multi-view data prediction method and system based on variational inference

    CN117152578A

  • Visual perception detection method and system

    CN118840633A

Cited By

  • Target intelligent monitoring system and method based on dual-optical data

    CN120580648A

  • Traffic violation snapshot control method and device, electronic equipment and storage medium

    CN121148146A

  • Moving object three-dimensional model reconstruction system and method based on cooperation of multiple unmanned aerial vehicles

    CN121414977A