An object monitoring method and device based on multi-source data and a storage medium
By synchronously collecting and associating multiple video streams, environmental parameters, and device location information, a unified spatial coordinate system mapping is established. Combined with dynamic scene analysis, the problems of stitching misalignment and inaccurate target recognition in panoramic video surveillance are solved, and high-quality panoramic video surveillance is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN STARCAM TECH
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-21
AI Technical Summary
Existing panoramic video surveillance methods fail to fully integrate multi-source data, resulting in inaccurate mapping between video feature points and real space, easy misalignment during stitching, lack of comprehensive data support for scene analysis, and lack of dynamic equipment control mechanisms, which affects system stability and battery life.
By simultaneously acquiring multiple video streams, environmental parameters, and device location information, performing time alignment and data association, establishing a mapping relationship between video feature points and a unified spatial coordinate system, and combining panoramic video streams for dynamic scene analysis, a device control strategy is generated to adjust the operating parameters of the image acquisition device.
It improves the stitching accuracy and target recognition accuracy of panoramic monitoring, enhances the real-time performance and stability of the system, and adapts to the dynamic adaptation needs of complex monitoring scenarios.
Smart Images

Figure CN121462725B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to an object monitoring method, device, and storage medium based on multi-source data. Background Technology
[0002] With the increasing demand for monitoring coverage and intelligence in fields such as agriculture and security, panoramic video surveillance systems have become a core means of achieving large-scale, blind-spot-free monitoring by deploying multiple image acquisition devices to collect multiple video streams and stitching them together. It is necessary to ensure the integrity of the panoramic image, the accuracy of target recognition, and the stability of system operation at the same time to meet the actual needs of real-time monitoring and risk warning.
[0003] Currently, existing panoramic video surveillance methods primarily acquire multiple video streams through multiple image acquisition devices, then stitch them together to form a panoramic video. However, existing technologies have several shortcomings. Firstly, they often focus solely on the acquisition and stitching of the video streams themselves, failing to fully integrate multi-source data closely related to the video streams. This results in inaccurate mapping between video feature points and real space, leading to misalignment and poor fusion issues in panoramic stitching. Furthermore, scene analysis lacks comprehensive data support, making it difficult to accurately identify target types and states. Secondly, existing systems lack dynamic device control mechanisms based on monitoring data and environmental conditions. They cannot flexibly adjust device operating parameters according to actual monitoring conditions and energy consumption, thus affecting the system's monitoring stability and battery life. Therefore, existing panoramic video surveillance methods suffer from insufficient stitching and analysis accuracy due to inadequate multi-source data collaborative processing, and instability and poor battery life due to the lack of dynamic device control, making it difficult to meet the high-quality monitoring needs of complex scenarios.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] This application provides an object monitoring method, device, and storage medium based on multi-source data, which can improve the stitching accuracy, real-time performance, and target recognition accuracy of panoramic monitoring, and realize dynamic adaptation and control of equipment.
[0006] In a first aspect, embodiments of this application provide an object monitoring method based on multi-source data, applied to a panoramic video surveillance system. The panoramic video surveillance system includes multiple multi-view image acquisition devices deployed within a monitoring area, an edge computing module communicatively connected to the multi-view image acquisition devices, a cloud server communicatively connected to the edge computing module, and multiple management devices communicatively connected to the cloud server. The multiple multi-view image acquisition devices are used to acquire multiple video streams from different perspectives within the monitoring area. The object monitoring method based on multi-source data includes:
[0007] Simultaneously acquire multiple video streams, environmental parameters associated with the video streams, and device location information, and perform time alignment and data association processing on the multiple video streams, the environmental parameters, and the location information to form associated data;
[0008] Feature matching is performed on multiple video streams in the associated data, and the location information and environmental parameters in the associated data are fused to establish a mapping relationship between video feature points and a unified spatial coordinate system;
[0009] The multiple video streams are stitched together in real time according to the mapping relationship to generate a panoramic video stream of the monitored area.
[0010] Dynamic scene analysis is performed based on the environmental parameters in the panoramic video stream and the associated data to identify the target types and states in the scene and obtain the analysis results.
[0011] Based on the analysis results, a device control strategy is generated, and control instructions for adjusting the working parameters of the multi-view image acquisition device are generated and issued based on the device control strategy.
[0012] Furthermore, in some embodiments of this application, the step of performing dynamic scene analysis based on environmental parameters in the panoramic video stream and the associated data to identify target types and states in the scene and obtain analysis results includes:
[0013] Based on the panoramic video stream, a dynamic background model of the monitored area is established;
[0014] The dynamic background model is used to extract foreground targets, and the morphological and motion features corresponding to the foreground targets are analyzed.
[0015] Based on the morphological features, the motion features, and the environmental parameters in the associated data, the foreground target is classified, the real-time state corresponding to the classification result is determined, and the analysis result is generated.
[0016] Furthermore, in some embodiments of this application, the step of classifying the foreground target based on the morphological features, the motion features, and the environmental parameters in the associated data, determining the real-time state corresponding to the classification result, and generating analysis results includes:
[0017] The morphological features and motion features of the foreground target are compared with the standard feature ranges of multiple predefined target types to generate comparison results;
[0018] Based on the comparison results, and combined with the auxiliary information for feature judgment from the environmental parameters in the associated data, the classification probability of the foreground target belonging to each category is calculated.
[0019] Determine whether the classification probability exceeds the classification threshold of the corresponding category, determine the final classification result of the foreground target, and use the key parameters in the motion features as the real-time state of the foreground target to form the analysis result.
[0020] Furthermore, in some embodiments of this application, the step of generating a device control strategy based on the analysis results, and generating and issuing control instructions for adjusting the operating parameters of the multi-view image acquisition device based on the device control strategy, includes:
[0021] If a specific type of target is identified in the analysis results, a video tracking strategy or alarm strategy corresponding to the specific type of target is generated;
[0022] Based on the environmental parameters in the associated data, and combined with the real-time energy consumption status of the panoramic video surveillance system, an equipment energy consumption control strategy is generated.
[0023] Based on the video tracking strategy, the alarm strategy, or the device energy consumption control strategy, generate corresponding control instructions;
[0024] The control command is sent to the multi-target image acquisition device to adjust the operating parameters of the multi-target image acquisition device.
[0025] Furthermore, in some embodiments of this application, the step of generating a device energy consumption control strategy based on environmental parameters in the associated data and combined with the real-time energy consumption status of the panoramic video surveillance system includes:
[0026] Obtain real-time light intensity and the remaining power of the panoramic video monitoring system;
[0027] If the real-time light intensity is detected to be lower than a preset light threshold and the remaining power is lower than a preset power threshold, a device energy consumption control strategy for entering emergency mode is generated.
[0028] Furthermore, in some embodiments of this application, the environmental parameters include object height distribution information; therefore, the real-time splicing of multiple video streams according to the mapping relationship further includes dynamically adjusting the weight coefficients of different video streams in the splicing process according to the object height distribution information.
[0029] Furthermore, in some embodiments of this application, the method further includes:
[0030] Real-time monitoring of image quality metrics of the panoramic video stream;
[0031] When the image quality index is detected to be lower than the preset standard, an adjustment command is generated to adjust the shooting angle of at least one camera in the multi-view image acquisition device, and the adjustment command is sent to the corresponding camera.
[0032] Furthermore, in some embodiments of this application, the edge computing module includes a positioning unit, an environmental sensing unit, and a preprocessing unit;
[0033] The positioning unit is used to acquire the device location information of the multi-view image acquisition device;
[0034] The environmental sensing unit is used to collect at least one environmental parameter, including light intensity, wind speed, temperature and crop height.
[0035] The preprocessing unit is used to perform frame synchronization and encoding on the multiple video streams, and to perform time alignment and data association processing on the synchronized video streams, the device location information and the environmental parameters to form associated data.
[0036] Furthermore, in some embodiments of this application, the cloud server is configured with a graphics processing unit for performing video feature matching and mapping calculations, and an intelligent analysis model for running the dynamic scene analysis.
[0037] Furthermore, in some embodiments of this application, the management device includes a mobile terminal and / or a fixed monitoring workstation for receiving and displaying the panoramic video stream, the analysis results, and / or alarm information triggered by the analysis results.
[0038] Secondly, this application also provides an object monitoring device based on multi-source data, comprising:
[0039] The acquisition module is used to simultaneously acquire multiple video streams, environmental parameters associated with the video streams, and device location information, and to perform time alignment and data association processing on the multiple video streams, the environmental parameters, and the location information to form associated data;
[0040] The mapping module is used to perform feature matching on multiple video streams in the associated data, and to fuse the location information and environmental parameters in the associated data to establish a mapping relationship between video feature points and a unified spatial coordinate system.
[0041] The stitching module is used to stitch the multiple video streams in real time according to the mapping relationship to generate a panoramic video stream of the monitored area;
[0042] The analysis module is used to perform dynamic scene analysis based on environmental parameters in the panoramic video stream and the associated data, identify the target type and state in the scene, and obtain analysis results.
[0043] The control module is used to generate a device control strategy based on the analysis results, and to generate and issue control instructions for adjusting the working parameters of the multi-view image acquisition device based on the device control strategy.
[0044] Thirdly, this application also provides a storage medium storing a computer program capable of being loaded by a processor and executing the object monitoring method based on multi-source data as described in the first aspect.
[0045] This application provides a method, device, and storage medium for object monitoring based on multi-source data. First, by synchronously acquiring multiple video streams, associated environmental parameters, and device location information, and performing time alignment and data association, the limitations of single video data are overcome, providing multi-dimensional and precisely correlated foundational data support for subsequent processing and avoiding analytical biases caused by isolated data. Second, video feature matching is performed based on the correlated data, and location information and environmental parameters are fused to establish a mapping relationship between video feature points and a unified spatial coordinate system. This enables video streams from different perspectives to achieve precise alignment based on a unified spatial reference, effectively solving the misalignment caused by perspective differences in traditional stitching. This addresses several issues, including ensuring the integrity and accuracy of the panoramic video stream after real-time stitching of multiple video streams; furthermore, by combining the panoramic video stream with environmental parameters to conduct dynamic scene analysis, fully utilizing the complementary value of multi-source data, and more comprehensively identifying the type and state of targets in the scene, thereby improving the accuracy of target recognition; finally, by generating equipment control strategies and issuing control commands based on the analysis results, the operating parameters of the multi-view image acquisition device can adapt to the real-time changes of the monitoring scene, thereby improving the overall real-time monitoring and target recognition accuracy of the panoramic video surveillance system, ensuring the adaptability and stability of system operation, and thus meeting the needs of dynamic adaptive monitoring in complex monitoring scenarios. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is an application environment diagram of the object monitoring method based on multi-source data provided in the embodiments of this application;
[0048] Figure 2 This is a flowchart illustrating the object monitoring method based on multi-source data provided in an embodiment of this application.
[0049] Figure 3 This is a schematic diagram of the process for establishing a mapping relationship provided in an embodiment of this application;
[0050] Figure 4 This is a schematic diagram of the process for generating analysis results provided in an embodiment of this application;
[0051] Figure 5 This is a schematic diagram of the process for generating control instructions provided in an embodiment of this application;
[0052] Figure 6 This is a schematic diagram of the structure of the object monitoring device based on multi-source data provided in the embodiments of this application;
[0053] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0054] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of systems and methods consistent with those detailed in the appended claims or with some aspects of this application.
[0055] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover descriptions such as non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0056] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0057] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0058] To address the aforementioned technical problems and overcome the shortcomings of existing technologies, this application provides an object monitoring method, device, and storage medium based on multi-source data, which can improve the stitching accuracy, real-time performance, and target recognition accuracy of panoramic monitoring, and achieve dynamic adaptation and control of equipment.
[0059] Figure 1 This is an application environment diagram of an object monitoring method based on multi-source data in one embodiment. (Refer to...) Figure 1 This multi-source data-based object monitoring method is applied to a panoramic video surveillance system. Taking farmland monitoring area as an example, multiple sets of cameras distributed on poles around the farmland are responsible for acquiring multiple video streams, capturing images of the farmland (crop area, pedestrian activity area) from different perspectives. The edge computing module integrates a positioning unit (to obtain its own deployment location), an environmental sensing unit (to collect environmental parameters such as corn height, farmland light / temperature, and terrain undulation), and a preprocessing unit (to perform frame synchronization and encoding on the video streams from multiple cameras, and associate the device location with environmental parameters). Mobile terminals and fixed monitoring workstations are responsible for receiving information and managing the data. The mobile terminals display the stitched panoramic video stream of the farmland in real time, while the fixed monitoring workstations display close-up monitoring images of the target (pedestrians) and analysis results (such as target location and movement status). The panoramic video and target analysis results received by the mobile terminals and fixed monitoring workstations are output by the cloud server after completing video feature matching / mapping calculations through the graphics processor and performing dynamic scene analysis (identifying pedestrian targets) through an intelligent analysis model.
[0060] This embodiment mainly uses the application of a multi-source data-based object monitoring method to a panoramic video surveillance system as an example. The panoramic video surveillance system includes multiple multi-view image acquisition devices deployed in the monitoring area, an edge computing module that communicates with the multi-view image acquisition devices, a cloud server that communicates with the edge computing module, and multiple management devices that communicate with the cloud server. The multiple multi-view image acquisition devices are used to acquire multiple video streams from different perspectives in the monitoring area.
[0061] Specifically, the panoramic video surveillance system provided in this embodiment includes multiple multi-view image acquisition devices, an edge computing module, a cloud server, and multiple management devices. The multiple multi-view image acquisition devices establish communication connections with the edge computing module to ensure that the acquired data can be transmitted to the edge computing module in real time. The edge computing module is connected to the cloud server via a communication link to enable further data transmission and interaction of processing instructions. The cloud server maintains communication connections with the multiple management devices for data display, instruction reception, and feedback.
[0062] For multiple multi-view image acquisition devices, which are deployed in the monitoring area, the core function is to acquire multiple video streams from different perspectives in the monitoring area. By shooting from multiple perspectives, a wider range of monitoring scenes can be covered, avoiding blind spots. For example, in the monitoring scene of farmland, video images of different areas and different heights of crops can be captured by multi-view image acquisition devices in different locations.
[0063] For edge computing modules, as the intermediate hub for data transmission and preliminary processing, they receive video streams transmitted by multi-view image acquisition devices, and simultaneously collect environmental parameters and device location information related to the video streams, completing preliminary collaborative data processing and providing a foundation for subsequent in-depth processing on cloud servers.
[0064] For cloud servers, they undertake core data processing and decision-making functions, and are responsible for feature matching, spatial mapping establishment, video stitching, scene analysis and control strategy generation of received data. They are the core carrier for realizing panoramic monitoring and intelligent control.
[0065] For multiple management devices, as human-computer interaction terminals, they are used to receive panoramic video streams, scene analysis results and other information transmitted from cloud servers. At the same time, they can display relevant monitoring data, making it convenient for users to keep track of the status of the monitored area in real time. For example, managers can remotely view the panoramic monitoring screen of farmland and the target recognition results through the management devices.
[0066] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of an object monitoring method based on multi-source data provided in this application. Specifically, the object monitoring method based on multi-source data provided in this application may include the following steps:
[0067] S1. Synchronously acquire multiple video streams, environmental parameters associated with the video streams, and device location information, and perform time alignment and data association processing on the multiple video streams, environmental parameters, and location information to form associated data;
[0068] Specifically, for step S1, three types of data are collected simultaneously: first, multiple video streams from different perspectives of the monitored area are collected using a multi-view image acquisition device; second, environmental parameters associated with these video streams are collected, including light intensity, crop height, and terrain undulation; and third, device location information, i.e., real-time deployment location data of the multi-view image acquisition device (e.g., latitude and longitude coordinates, altitude). The collected video streams, environmental parameters, and device location information are time-aligned to ensure that various data at the same point in time correspond to each other. Then, data association technology is used to bind the three together to form complete associated data, avoiding data isolation. For example, a video stream of a certain area of farmland collected at a certain moment, the light intensity of 1500 Lux in that area at that time, the crop height of 2.5 meters (corn), and the latitude and longitude (36°N, 118°E) of the multi-view image acquisition device are bound together to form a set of associated data.
[0069] S2. Perform feature matching on multiple video streams in the associated data, and fuse the location information and environmental parameters in the associated data to establish a mapping relationship between video feature points and a unified spatial coordinate system;
[0070] Specifically, in step S2, image feature points from multiple video streams are extracted from the formed associated data. These feature points include easily identifiable features such as field ridge edges and crop row boundaries in farmland videos. The extracted image feature points from different video streams are compared to find pairs of identical feature points from the same real-world target. For example, the starting point of a field ridge in video stream A and the starting point of the same field ridge in video stream B constitute a pair of identical feature points. Combining device location information (such as the latitude, longitude, and altitude of the acquisition device) and environmental parameters (such as terrain undulation and crop height) from the associated data, a unified spatial coordinate system is constructed. This coordinate system accurately corresponds to the real geographic space of the monitored area. Subsequently, the image coordinates of the found pairs of identical feature points are mapped one by one to this unified spatial coordinate system, establishing a precise mapping relationship between video feature points and real spatial locations, ensuring that video streams from different perspectives can be associated based on their real spatial locations.
[0071] S3. Based on the mapping relationship, multiple video streams are stitched together in real time to generate a panoramic video stream of the monitored area;
[0072] Specifically, for step S3, based on the established mapping relationship between video feature points and a unified spatial coordinate system, spatial alignment processing is performed on multiple video streams, unifying each frame of the video streams from different perspectives to the same spatial coordinate system for positioning; the spatially aligned frames are then integrated and stitched together to eliminate perspective differences and image breaks between different video streams, generating a continuous and complete panoramic video stream of the monitoring area. For example, in farmland monitoring, this step can stitch together video streams from four different directions—east, south, west, and north—into a panoramic view covering the entire farmland, allowing users to intuitively view the overall condition of the farmland.
[0073] S4. Perform dynamic scene analysis based on environmental parameters in panoramic video stream and associated data, identify target types and states in the scene, and obtain analysis results;
[0074] Specifically, for step S4, the generated panoramic video stream is used as the core analysis material, and dynamic scene analysis is carried out in conjunction with environmental parameters in the associated data. By dynamically monitoring the panoramic video stream, various targets appearing in the scene are identified and their types are distinguished. For example, in a farmland scene, human figures, crop swaying, scarecrows, etc. are identified, and the real-time status of the targets is determined, such as the direction of movement of human figures and the amplitude of crop swaying. Finally, a clear analysis result is formed, such as "A human figure has appeared in the northeast area of the farmland and is moving along the crop row" or "There is slight crop swaying in the southwest area of the farmland, with no abnormal targets."
[0075] S5. Generate equipment control strategies based on the analysis results, and generate and issue control instructions for adjusting the working parameters of the multi-view image acquisition device based on the equipment control strategies;
[0076] Specifically, for step S5, based on the obtained analysis results and combined with the actual needs of the monitoring scenario, a corresponding equipment control strategy is formulated; according to the equipment control strategy, specific control instructions are generated and sent to the multi-view image acquisition device to adjust its working parameters, such as adjusting the shooting angle, focal length, and shooting frame rate. For example, when the analysis results show that "there is an unidentified suspected target in a certain area of farmland", a control instruction is generated to "adjust the shooting angle of the multi-view image acquisition device corresponding to the area, focus on the suspected target area, and increase the shooting frame rate" to ensure accurate monitoring of the target.
[0077] This embodiment achieves panoramic coverage and precise image integration of the monitored area by synchronously acquiring and associating multiple video streams, environmental parameters, and device location information, combined with the establishment of a unified spatial coordinate system and real-time video stitching. Simultaneously, dynamic scene analysis based on panoramic video streams and environmental parameters improves the accuracy and timeliness of target recognition. Finally, through the generation and execution of dynamic device control strategies, the multi-view image acquisition device can adapt to real-time changes in the monitoring scene, comprehensively improving the comprehensiveness, accuracy, and flexibility of panoramic monitoring, and meeting the needs of complex monitoring scenarios for high-quality, intelligent monitoring.
[0078] Furthermore, in some embodiments, step S1, "synchronously acquiring multiple video streams, environmental parameters associated with the video streams, and device location information, and performing time alignment and data association processing on the multiple video streams, environmental parameters, and location information to form associated data," may specifically include:
[0079] S11. Obtain multiple video streams with synchronized timestamps from a multi-view image acquisition device;
[0080] Specifically, in step S11, multiple multi-view image acquisition devices deployed within the monitoring area are used to collect video data from different perspectives of the monitored scene. The video acquisition clocks of all acquisition devices are strictly synchronized to ensure that each frame of each video stream carries a unified standard timestamp. The multi-view image acquisition devices achieve time consistency of acquisition actions through preset synchronization protocols, such as hardware synchronization triggering or network time synchronization, avoiding video frame misalignment due to acquisition time differences between different devices. The timestamp is accurate to the millisecond level and is used to uniquely identify video frames from different perspectives acquired at the same time.
[0081] For example, in a farmland monitoring scenario, three multi-view image acquisition devices are deployed in the east, south, and west of the farmland to cover the entire area. After time synchronization configuration, the three devices simultaneously acquire video frames of the crop growth status in the east, the condition of the field ridges in the south, and the operation of the irrigation facilities in the west at the same time (e.g., 2025-12-18 09:30:00.123). Each video frame is labeled with the same timestamp, forming three synchronized video streams.
[0082] S12. Synchronously acquire environmental parameters and device location information collected by the edge computing module;
[0083] Specifically, in step S12, while acquiring multiple synchronized video streams in the same time dimension, two types of key data are extracted from the edge computing module connected to the multi-view image acquisition device: environmental parameters and device location information. This ensures that the acquisition time of these two types of data completely corresponds to the timestamp of the video stream. The edge computing module has built-in sensing and positioning units to collect environmental parameters related to the monitoring scene in real time, such as light intensity, crop height, and terrain flatness, as well as the real-time deployment location data of the multi-view image acquisition device itself (such as latitude, longitude, and altitude). The acquisition frequency is matched with the video stream frame rate to ensure data time synchronization.
[0084] For example, continuing with the above farmland monitoring scenario, while acquiring three video streams with a timestamp of 2025-12-18 09:30:00.123, the environmental parameters at that moment (illuminance 1300 Lux, corn crop height 2.4 meters, field flatness deviation ±5 cm) and the device location information of the three multi-view image acquisition devices (36.05°N 118.12°E, 36.05°N 118.13°E, and 36.06°N 118.12°E, respectively) are simultaneously acquired from the edge computing module.
[0085] S13. Bind multiple video streams, environmental parameters, and device location information under the same timestamp to form associated data;
[0086] Specifically, in step S13, using the timestamp as the unique association identifier, multiple video streams, environmental parameters, and device location information corresponding to the same timestamp are bound one by one, integrating them into a set of structurally complete and data-corresponding associated data to ensure the spatiotemporal consistency of various types of data. Through a data binding algorithm, all data marked with the same timestamp are identified and extracted, including multiple video frames, all environmental parameters at that moment, and the location information of all acquisition devices at that moment, and encapsulated into a unified data unit to avoid data mixing across different time dimensions, providing accurate multi-source data support for subsequent processing. For example, the three farmland video streams corresponding to the timestamp 2025-12-18 09:30:00.123, the environmental parameters of "light intensity 1300 Lux + corn height 2.4 meters + flatness ±5 cm", and the latitude and longitude location information of the three acquisition devices are bound to form a complete set of associated data; subsequent timestamps are bound according to this logic to form a continuous sequence of associated data.
[0087] This embodiment acquires multiple video streams, environmental parameters, and device location information through a strict time synchronization mechanism, and accurately binds them with timestamps as the core, constructing spatiotemporally consistent and complete associated data. This effectively avoids association deviations caused by time misalignment of multi-source data, providing reliable basic data support for subsequent video feature matching, spatial mapping establishment, panoramic stitching, and scene analysis, and improving the accuracy and efficiency of subsequent data processing.
[0088] Furthermore, such as Figure 3 As shown, in some embodiments, step S2, "performing feature matching on multiple video streams in the associated data, fusing location information and environmental parameters from the associated data, and establishing a mapping relationship between video feature points and a unified spatial coordinate system," may specifically include:
[0089] S21. Extract image feature points from each video stream in the associated data;
[0090] Specifically, in step S21, for each video stream included in the associated data, image feature points with stability, uniqueness, and recognizability are filtered and extracted from each of its image frames. These feature points can accurately represent the key information in the image, providing a foundation for subsequent matching. The extracted feature points need to have the characteristics of being resistant to changes in illumination and viewpoint shifts, and typically include key points in the image such as corner points, edge endpoints, and densely textured areas. Each frame of the image is scanned using a feature extraction algorithm to locate and record the pixel coordinates of these feature points and their corresponding feature description information, such as grayscale distribution, gradient direction, and neighborhood pixel features, ensuring that each feature point has a unique identification identifier.
[0091] For example, the associated data from farmland monitoring includes three video streams from different perspectives. Feature point extraction is performed on one of the video streams that captures the eastern area of the farmland. Feature points such as the intersections of field ridges, the inflection points of corn rows, and the endpoints of irrigation pipes are identified and extracted. At the same time, the pixel coordinates of each feature point and the grayscale gradient description information around that point are recorded.
[0092] S22. Match the image feature points extracted from different video streams to obtain pairs of feature points with the same name;
[0093] Specifically, in step S22, image feature points extracted from different video streams are compared across video streams. Through similarity analysis of feature description information, feature point combinations corresponding to the same real-world target are selected, i.e., pairs of identically named feature points. A feature matching algorithm is used to calculate the descriptor similarity (such as Euclidean distance, cosine similarity, etc.) between feature points in different video streams. A similarity threshold is set; when the similarity between two feature points exceeds the threshold, they are determined to be potential pairs of identically named feature points. Further, geometric consistency verification is performed to eliminate mismatched points caused by differences in viewpoint or interference factors, ultimately obtaining accurate pairs of identically named feature points.
[0094] For example, continuing with the above farmland scene, the "field ridge intersection point (189, 245)" extracted from the eastern video stream is compared with multiple feature points extracted from the southern video stream. It is found that the feature point with pixel coordinates (312, 198) in the southern video stream has a grayscale gradient description information that is similar to the field ridge intersection point to the threshold. Furthermore, geometric verification confirms that the two feature points correspond to the same field ridge intersection position in the farmland. Therefore, these two feature points are identified as a pair of feature points with the same name. Following this logic, the matching of feature points between all different video streams is completed to obtain multiple pairs of feature points with the same name.
[0095] S23. Based on the device location information and environmental parameters in the associated data, construct a three-dimensional spatial model of the monitoring area;
[0096] Specifically, for step S23, using the device location information in the associated data as the basic spatial reference, and integrating information from environmental parameters reflecting the real spatial characteristics of the monitored area, a three-dimensional spatial model capable of accurately reproducing the spatial structure of the monitored scene is constructed. First, the spatial origin and coordinate system reference of the three-dimensional model are determined using device location information (such as latitude, longitude, and altitude) to ensure the spatial positioning accuracy of the model. Then, environmental parameters, such as terrain undulation, crop height distribution, and ground flatness, are incorporated to supplement and calibrate the model's elevation dimension and vertical spatial distribution, enabling the model to truly reflect the terrain features and target distribution of the monitored area, forming a complete three-dimensional spatial structure.
[0097] For example, the device location information in the associated data is "the latitude and longitude of the multi-view image acquisition device are 36.05°N 118.12°E and 36.05°N 118.13°E, respectively, and the altitude is 52 meters." Based on this, an initial three-dimensional coordinate system is established. Combined with environmental parameters "field terrain undulation deviation ±6cm, average height of corn crops 2.3 meters, and field ridge height 0.3 meters", the elevation of the initial coordinate system is calibrated. The slight undulation of the ground, the three-dimensional distribution of corn, and the height difference of the field ridges are reproduced in the model, and finally a three-dimensional spatial model that fits the actual farmland scene is constructed.
[0098] S24. Map the image coordinates of the same feature point pairs to the three-dimensional space model to establish a mapping relationship;
[0099] Specifically, in step S24, a coordinate transformation algorithm is used to map the image pixel coordinates of each pair of corresponding feature points in different video streams to their real spatial coordinates in the 3D spatial model, forming a fixed mapping relationship between video feature points and a unified spatial coordinate system. Based on the coordinate system parameters of the 3D spatial model and the geometric relationship between the corresponding feature point pairs, a transformation equation between image pixel coordinates and 3D spatial coordinates is established. By solving the equation, the real spatial position corresponding to each image feature point is obtained. After the mapping of all corresponding feature point pairs is completed, a unified set of feature point and spatial coordinate mapping rules covering multiple video streams is formed, achieving accurate association between video images and real space.
[0100] For example, the previously determined "field ridge intersection point" corresponding feature point pair (eastern video stream coordinates (189,245) and southern video stream coordinates (312,198)) is transformed using a coordinate transformation algorithm to calculate their corresponding real spatial coordinates in the three-dimensional spatial model as (X:95m, Y:72m, Z:52.1m), thus completing the coordinate mapping of this feature point pair. Similarly, all corresponding feature point pairs are mapped from image coordinates to three-dimensional spatial coordinates, ultimately establishing a precise mapping relationship between video feature points and a unified spatial coordinate system.
[0101] It should be noted that the unified spatial coordinate system is the coordinate system used by the three-dimensional model that accurately reflects the real spatial structure of the monitored area after the initial three-dimensional geographic reference frame is established based on the device location information and after correction and structural constraints are incorporated by incorporating environmental parameters (terrain undulation, object height).
[0102] This embodiment extracts image feature points from multiple video streams and matches pairs of feature points with the same name. It then integrates device location information and environmental parameters to construct a realistic three-dimensional spatial model. Finally, it maps the image coordinates of the feature points to the three-dimensional space, achieving a precise association between video feature points and the real space of the monitored area. This provides core technical support for the subsequent spatial alignment and panoramic stitching of multiple video streams, effectively solving the matching deviation problem caused by inconsistent spatial references in video streams from different perspectives, and improving the spatial accuracy of subsequent video processing and scene analysis.
[0103] Furthermore, in some embodiments, step S23, "constructing a three-dimensional spatial model of the monitoring area based on the device location information and environmental parameters in the associated data," may specifically include:
[0104] S231. Establish a preliminary three-dimensional geographic reference framework using equipment location information as a spatial reference.
[0105] Specifically, for step S231, key spatial coordinate data is extracted from the device location information of the associated data. Based on this, the reference point of the three-dimensional coordinate system is determined, and a preliminary three-dimensional geographic reference framework that can reflect the approximate spatial range of the monitored area is constructed. First, the precise spatial parameters contained in the device location information are analyzed, focusing on extracting the latitude and longitude coordinates and altitude data of at least one multi-view image acquisition device. These data are the core basis for determining the spatial reference. Then, the physical point defined by the latitude and longitude coordinates and altitude is used as the spatial origin of the three-dimensional rectangular coordinate system. The three axes of the coordinate system are defined (for example, due east is the positive X-axis, due north is the positive Y-axis, and vertical upward is the positive Z-axis). The correspondence between the coordinate scale and the actual distance is clarified (e.g., 1 coordinate unit corresponds to 1 meter), thereby constructing a three-dimensional geographic reference framework that can preliminarily represent the spatial location of the monitored area.
[0106] For example, in a farmland monitoring scenario, the latitude and longitude of a multi-view image acquisition device can be determined from the device location information to be 36.08°N, 118.15°E, and an altitude of 55 meters. Using this point as the spatial origin (0,0,0), the X-axis points due east, the Y-axis points due north, and the Z-axis points vertically upward. One coordinate unit corresponds to one meter of actual distance. This establishes a preliminary three-dimensional geographic reference frame covering the farmland area, where the coordinates of any point correspond to its actual geographical location within the farmland.
[0107] S232. Based on the terrain undulation information in the environmental parameters, the elevation dimension of the preliminary three-dimensional geographic reference frame is corrected to obtain the corrected three-dimensional geographic reference frame.
[0108] Specifically, for step S232, topographic relief information reflecting the ground morphology is extracted from environmental parameters. Based on this information, a digital model that accurately represents the changes in ground elevation in the monitoring area is generated. This model is then used to replace the planar elevation in the initial framework, completing the elevation dimension correction. First, data related to topographic relief are selected from the environmental parameters, such as the relative height difference, slope, and flatness deviation of various points on the ground. This data directly reflects the true shape of the ground in the monitoring area. Then, using this data, a digital elevation model (DEM) is generated through interpolation calculations and terrain modeling algorithms. This model can mark the actual elevation of each geographical location within the monitoring area point by point. Finally, this digital elevation model is used to replace the default horizontal elevation plane in the initial three-dimensional geographic reference frame, so that the Z-axis (elevation) dimension of the frame can accurately match the changes in ground undulation, resulting in the corrected three-dimensional geographic reference frame.
[0109] For example, continuing with the farmland scenario described above, the topographic relief information of the farmland is obtained from environmental parameters: the ground in the eastern area is on average 2 meters higher than the origin, there is an irrigation ditch with a depth of 1.5 meters in the western area, and the overall ground flatness deviation is ±8cm. Based on this information, a digital elevation model is generated. In the model, the Z-axis coordinates within the eastern coordinate range (X: 0-50m, Y: 0-80m) of the farmland are marked as 55-57 meters, and the Z-axis coordinates within the western irrigation ditch location range (X: 60-70m, Y: 20-60m) are marked as 53.5-55 meters. This digital elevation model is then integrated into a preliminary three-dimensional geographic reference frame, replacing the original horizontal Z-axis plane, completing the elevation dimension correction so that the frame can accurately reflect the undulation of the farmland ground.
[0110] S233. Integrate the object height distribution information in the environmental parameters as a structured constraint in the vertical space into the corrected three-dimensional geographic reference frame to complete the construction of the three-dimensional spatial model;
[0111] Specifically, for step S233, the distribution information representing the height of the monitored objects is extracted from the environmental parameters. Within the 3D geographic reference frame with corrected elevation dimension, corresponding vertical height constraints are assigned to the monitored objects in different horizontal areas, thereby constructing a 3D spatial model that reflects the three-dimensional distribution of the monitored objects. First, the object height distribution information is obtained from the environmental parameters. This information needs to clearly define the actual height of different locations and types of monitored objects within the monitoring area (such as the average height of crops, the height of field facilities, etc.). Then, within the corrected 3D geographic reference frame, based on the object height distribution information, for each monitored object's horizontal area (defined by X-axis and Y-axis coordinates), a corresponding height range constraint is set in the Z-axis direction (i.e., the Z-axis coordinate interval from the bottom to the top of the object). Finally, by integrating these vertical spatial constraints, volumetric elements that can intuitively represent the three-dimensional shape and distribution location of the monitored objects are constructed within the frame, ultimately completing the construction of a complete 3D spatial model.
[0112] For example, the height distribution information of objects in the farmland is obtained from environmental parameters: the average height of corn planted in the eastern area is 2.6 meters, the average height of wheat planted on both sides of the western irrigation ditch is 1.3 meters, and the height of the irrigation tower in the middle of the farmland is 8 meters. Within the corrected 3D geographic reference frame, the Z-axis height constraint is set to "ground elevation +0 to +2.6 meters" for the eastern corn planting area (X:0-50m, Y:0-80m), "ground elevation +0 to +1.3 meters" for the western wheat planting area (X:80-120m, Y:20-60m), and "ground elevation +0 to +8 meters" for the irrigation tower location (X:60m, Y:40m). By incorporating these vertical constraints, a three-dimensional volume representation of corn, wheat, and irrigation tower is formed within the frame, ultimately constructing a 3D spatial model that fits the actual farmland scene.
[0113] This embodiment establishes a preliminary spatial benchmark based on device location information, corrects the elevation dimension by combining terrain undulation information, and incorporates object height distribution information as a vertical constraint. The constructed three-dimensional spatial model can accurately restore the terrain morphology and three-dimensional distribution characteristics of objects in the monitored area, realizing a precise correspondence between video feature points and real space. This provides a reliable spatial benchmark for subsequent video feature point mapping, multi-channel video stitching, and scene analysis, effectively improving the spatial accuracy and scene adaptability of subsequent data processing.
[0114] Furthermore, in some embodiments, step S231, "establishing a preliminary three-dimensional geographic reference frame using the device location information as a spatial reference," may specifically include:
[0115] S2311. Extract the latitude and longitude coordinates and altitude of at least one of the multi-view image acquisition devices from the device location information;
[0116] Specifically, in step S2311, the device location information recorded in the associated data is parsed to filter and extract precise spatial positioning data of at least one multi-view image acquisition device. The focus is on obtaining the two key parameters: latitude and longitude coordinates and altitude, providing a core basis for establishing a spatial benchmark. The device location information includes deployment location data related to the multi-view image acquisition device. The parsing process needs to remove redundant information, focusing on latitude and longitude (based on a geographic coordinate system) and altitude (based on an elevation benchmark) data that can characterize absolute spatial location. At least one acquisition device is selected as the benchmark because its spatial location can serve as the anchor point for the subsequent coordinate system, ensuring the spatial positioning accuracy of the initial three-dimensional framework. The parsed data needs to retain sufficient precision (e.g., latitude and longitude accurate to four decimal places, altitude accurate to 0.1 meters).
[0117] For example, in a farmland monitoring scenario, the device location information of a multi-view image acquisition device contains deployment data of three devices. From this data, the precise spatial data of one of the devices can be extracted: latitude and longitude of 36.0825°N, longitude of 118.1532°E, and altitude of 55.3 meters. This set of data will serve as the core benchmark data for establishing a preliminary three-dimensional geographic reference framework.
[0118] S2312. Using the points defined by the analyzed latitude and longitude coordinates and altitude as the spatial origin, establish a three-dimensional rectangular coordinate system as a preliminary three-dimensional geographic reference frame;
[0119] Specifically, in step S2312, the physical points corresponding to the resolved latitude, longitude, and altitude coordinates are set as the spatial origin of a three-dimensional rectangular coordinate system. The correspondence between the three axes of the coordinate system, the coordinate units, and the actual spatial distance is clarified, thus constructing a three-dimensional geographic reference framework that can initially represent the spatial range of the monitoring area. The spatial origin is the actual geographical location corresponding to the latitude, longitude, and altitude. The coordinate system axes are set according to conventional geographic spatial rules (e.g., the X-axis points due east, the Y-axis points due north, and the Z-axis is perpendicular to the ground and upwards, consistent with the altitude direction). At the same time, the mapping relationship between coordinate units and actual distances is defined (e.g., 1 coordinate unit corresponds to 1 meter of actual spatial distance), so that the coordinate system can be directly associated with the real geographic space. The resulting three-dimensional rectangular coordinate system is the preliminary three-dimensional geographic reference framework, which can cover the monitoring area and a certain range of surrounding space.
[0120] For example, continuing with the farmland scenario described above, using a specific point in the farmland corresponding to the analyzed coordinates "36.0825°N, 118.1532°E, altitude 55.3 meters" (such as the center point of the data acquisition device's mounting base) as the spatial origin (0,0,0), the X-axis is set to point due east, the Y-axis to point due north, and the Z-axis to point vertically upward, with 1 coordinate unit corresponding to 1 meter of actual distance. The resulting three-dimensional rectangular coordinate system can cover the farmland and a 100-meter radius around it. Any point within the farmland can be preliminarily characterized by its spatial location using (X,Y,Z) coordinates, forming a preliminary three-dimensional geographic reference framework.
[0121] This embodiment determines the spatial origin by analyzing the latitude and longitude coordinates and altitude of the multi-view image acquisition device, establishes a standardized three-dimensional rectangular coordinate system, and constructs a preliminary three-dimensional geographic reference frame with a clear spatial benchmark. This provides a stable and accurate basic spatial carrier for subsequent frame correction by combining information such as terrain and object height, ensuring the spatial positioning accuracy of subsequent three-dimensional spatial model construction, and laying the core benchmark for the association between video feature points and real space.
[0122] Furthermore, in some embodiments, step S232, "correcting the elevation dimension of the preliminary three-dimensional geographic reference frame based on the terrain undulation information in the environmental parameters to obtain the corrected three-dimensional geographic reference frame," may specifically include:
[0123] S2321. Obtain topographic relief information from environmental parameters used to characterize ground smoothness;
[0124] Specifically, for step S2321, key data reflecting differences in ground morphology within the monitored area are screened and extracted from the associated environmental parameters. This data directly characterizes the flatness of the ground and serves as the core basis for subsequent elevation correction. The terrain undulation information must include data such as the relative height changes, slope magnitude, and dimensions of local depressions or protrusions at different locations within the monitored area. This data is collected in real-time by sensing devices and recorded in the environmental parameters. The extraction process must ensure the completeness and accuracy of the data, covering the entire ground area of the monitored region, to avoid deviations in local terrain reconstruction due to missing data.
[0125] For example, in a farmland monitoring scenario, the topographic undulation information of the farmland can be obtained from environmental parameters: the eastern part of the farmland is 2.1 meters higher than a certain benchmark on average, there is a north-south irrigation ditch in the western part of the farmland, the bottom of the ditch is 1.6 meters deep relative to the benchmark, the ditch is 5 meters wide, the slope of the farmland in the middle is 3°, and the overall flatness deviation is ±7cm. These data fully reflect the undulation of the farmland.
[0126] S2322. Generate a digital elevation model of the monitored area based on terrain undulation information;
[0127] Specifically, in step S2322, the acquired terrain undulation information is used to process and model the data through terrain modeling algorithms, generating a digital elevation model (DEM) that can accurately label the actual elevation of each geographical location within the monitoring area. Algorithms such as interpolation and gridded modeling are employed to transform discrete terrain undulation data into a continuous elevation surface model. The model is based on two-dimensional plane coordinates (corresponding to the X and Y axes of the subsequent three-dimensional frame), assigning a unique elevation value (corresponding to the Z axis of the subsequent three-dimensional frame) to each plane coordinate point, forming a grid-like digital elevation model to ensure that the model can accurately reproduce the undulations of the ground.
[0128] For example, continuing with the farmland scenario described above, based on the extracted topographic relief information, the data is processed using the Kriging interpolation algorithm to generate a digital elevation model of the farmland. In the model, the elevation values of the eastern region (X: 0-60m, Y: 0-90m) are 55.2-57.3 meters, the western irrigation ditch region (X: 70-80m, Y: 10-70m) has an elevation value of 53.4-55.2 meters, and the central slope region (X: 30-90m, Y: 30-60m) gradually changes from 55.2 meters to 56.8 meters with the slope gradient, completely restoring the undulating shape of the farmland.
[0129] S2323. Replace the initial elevation plane in the preliminary three-dimensional geographic reference frame with a digital elevation model to complete the correction of the elevation dimension;
[0130] Specifically, in step S2323, the generated digital elevation model replaces the default horizontal elevation plane (i.e., the fixed plane along the Z-axis in the original frame) in the initial 3D georeferenced frame, ensuring that the frame's elevation dimension (Z-axis) accurately matches the actual terrain undulations of the monitored area, thus completing the correction. The initial 3D georeferenced frame's elevation plane is horizontal and cannot reflect the true terrain. During the replacement process, the elevation value corresponding to each (X, Y) coordinate in the digital elevation model is directly assigned to the Z-axis coordinate of the same (X, Y) coordinate point in the initial frame. This makes the frame's Z-axis no longer a fixed value but dynamically changes with the terrain undulations, ultimately forming a corrected 3D georeferenced frame with an elevation dimension consistent with the actual terrain.
[0131] For example, the initial 3D geographic reference frame has an elevation plane of Z=55.2 meters. After replacing it with the farmland digital elevation model generated above, the Z-axis coordinates of the eastern region (X: 0-60m, Y: 0-90m) in the frame become 55.2-57.3 meters, the Z-axis coordinates of the western irrigation ditch region become 53.4-55.2 meters, and the Z-axis coordinates of the central slope region gradually change with the slope. The elevation dimension of the frame perfectly matches the actual terrain of the farmland, and the elevation dimension correction is completed.
[0132] This embodiment generates a digital elevation model by extracting terrain undulation information and uses it to replace the horizontal elevation plane of the initial frame. This achieves accurate correction of the elevation dimension of the three-dimensional geographic reference frame, enabling the corrected frame to realistically reproduce the terrain undulation characteristics of the monitored area. This solves the problem that the initial frame cannot reflect the actual terrain and provides an elevation benchmark that fits the real scene for subsequent integration of object height information and construction of an accurate three-dimensional spatial model, thereby improving the accuracy of subsequent spatial mapping and video processing.
[0133] Furthermore, in some embodiments, step S233, "integrating the object height distribution information in the environmental parameters as a structured constraint in the vertical space into the corrected three-dimensional geographic reference frame to complete the construction of the three-dimensional spatial model," may specifically include:
[0134] S2331. Obtain the object height distribution information representing the height of the monitored object from the environmental parameters;
[0135] Specifically, for step S2331, data that clearly defines the height characteristics of various monitored objects within the monitoring area is filtered and extracted from the associated environmental parameters. This data must include the correlation information between object type, distribution location, and corresponding height, providing a precise basis for subsequent vertical spatial constraints. Monitored objects include naturally existing objects (such as crops and trees) and man-made facilities (such as irrigation equipment and fences) within the scene. The object height distribution information must clearly define the actual height (such as average height and height range) of objects at different locations and of different types. The data must cover the entire monitoring area to ensure that the height information of each key object is captured and corresponds one-to-one with its spatial location.
[0136] For example, in a farmland monitoring scenario, the height distribution information of objects can be obtained from environmental parameters: the average height of corn planted in the northern area of the farmland (east-west X: 0-80m, north-south Y: 0-50m) is 2.5 meters, the average height of wheat planted in the southern area (X: 0-80m, Y: 60-120m) is 1.2 meters, the height of the irrigation well house in the center of the farmland (X: 40m, Y: 60m) is 3.8 meters, and the height of the guardrail at the edge of the farmland (X: 80-100m, Y: 30-90m) is 1.5 meters.
[0137] S2332. Within the corrected three-dimensional geographic reference frame, based on the object height distribution information, assign corresponding height value constraints to the horizontal area where the monitored object is located along the vertical direction.
[0138] Specifically, for step S2332, based on the corrected 3D geographic reference frame and according to the acquired object height distribution information, a clear height range constraint is set in the vertical direction (Z-axis) for each monitored object's horizontal spatial area (defined by X-axis and Y-axis coordinates), i.e., the Z-axis coordinate interval from the bottom to the top of the object. The corrected frame has been calibrated for the elevation dimension (Z-axis) using terrain undulation information. The Z-axis coordinate at the bottom of the object is the actual ground elevation of that horizontal area; the Z-axis coordinate at the top of the object is "ground elevation + object height". The resulting Z-axis interval is the vertical height constraint of the object, ensuring that the constraint is consistent with the actual three-dimensional shape of the object and accurately corresponds to its spatial position.
[0139] For example, continuing with the above farmland scenario, in the corrected 3D geographic reference frame, the ground elevation Z of the northern corn planting area is 54.2-55.0 meters. Based on the corn height of 2.5 meters, a vertical height constraint "Z: ground elevation +0 to +2.5 meters" (i.e., actual Z: 54.2-57.5 meters) is assigned to this area. The ground elevation Z of the southern wheat planting area is 53.8-54.5 meters. Based on the wheat height of 1.2 meters, a constraint "Z: ground elevation +0 to +1.2 meters" (actual Z: 53.8-55.7 meters) is assigned to this area. The ground elevation Z of the location of the irrigation well house is 54.0 meters. Based on the height of 3.8 meters, a constraint "Z: 54.0 to 57.8 meters" is assigned to this area. The ground elevation Z of the area where the guardrail is located is 53.5-54.8 meters. Based on the height of 1.5 meters, a constraint "Z: ground elevation +0 to +1.5 meters" (actual Z: 53.5-56.3 meters) is assigned to this area.
[0140] S2333. Based on height constraints, construct volumetric elements representing the three-dimensional distribution of monitored objects in a three-dimensional spatial model;
[0141] Specifically, in step S2333, the set vertical height constraint is combined with the corresponding horizontal region (X and Y axis range). Within the corrected 3D geographic reference frame, volumetric elements that can intuitively represent the three-dimensional shape, spatial location, and distribution range of each monitored object are constructed, ultimately completing the construction of the 3D spatial model. The volumetric elements are presented through a combination of "horizontal region range + vertical height constraint." Each object corresponds to one or more continuous volumetric units. The boundaries of the units are defined by the horizontal range of the X and Y axes and the height constraint of the Z axis. After integrating the volumetric elements of all objects, a complete 3D spatial model reflecting the three-dimensional structure of the monitored scene is formed.
[0142] For example, based on the aforementioned height constraints, volume elements are constructed within the framework: the northern corn region forms continuous cuboid volume elements with dimensions "X: 0-80m, Y: 0-50m, Z: 54.2-57.5m", corresponding to the three-dimensional distribution of the entire cornfield; the southern wheat region forms volume elements with dimensions "X: 0-80m, Y: 60-120m, Z: 53.8-55.7m"; irrigation well houses form cuboid volume elements with dimensions "X: 38-42m, Y: 58-62m, Z: 54.0-57.8m"; and protective fences form elongated volume elements with dimensions "X: 80-100m, Y: 30-90m, Z: 53.5-56.3m". After integrating these volume elements, the three-dimensional distribution of crops and facilities in the farmland is fully reconstructed, completing the construction of the three-dimensional spatial model.
[0143] This embodiment acquires object height distribution information and assigns vertical height constraints, constructing the three-dimensional volume elements of the monitored object within a terrain-corrected three-dimensional framework. The resulting three-dimensional spatial model can accurately reproduce the three-dimensional distribution characteristics and spatial positional relationships of various objects within the monitored area, making up for the shortcomings of only considering terrain and lacking object three-dimensional information. This provides a more realistic spatial benchmark for subsequent accurate mapping of video feature points to real space, multi-channel video stitching, and target recognition, improving the scene adaptability and accuracy of subsequent data processing.
[0144] Furthermore, in some embodiments, step S3, "real-time stitching of multiple video streams according to the mapping relationship to generate a panoramic video stream of the monitored area," may specifically include:
[0145] S31. Based on the mapping relationship, transform each frame of the multi-channel video stream to a unified spatial coordinate system;
[0146] Specifically, for step S31, based on the established mapping relationship between video feature points and a unified spatial coordinate system, the geometric transformation rules for each video stream image frame are determined, and spatial coordinate transformation is performed on each frame image to ensure that all video frames are adapted to a unified spatial reference, laying the foundation for subsequent stitching. First, based on the mapping relationship, the geometric correspondence between the image coordinate system of each video stream and the unified spatial coordinate system is clarified, and the required geometric transformation parameters, such as the perspective transformation matrix or affine transformation matrix, are determined. Then, for each frame image of each video stream, pixel coordinate mapping is performed according to the corresponding geometric transformation parameters, transforming the image from the original viewpoint to the target viewpoint under the unified spatial coordinate system. Finally, through resampling technology, pixel holes generated during the transformation process are filled, pixel redundancy in overlapping areas is eliminated, and the transformed image is ensured to be complete and distortion-free under the unified coordinate system.
[0147] For example, in a farmland monitoring scenario, there are three video streams with different perspectives (east, south, and west). A mapping relationship between feature points of each video stream and a unified spatial coordinate system has been established. For a frame image of the eastern video stream, a perspective transformation matrix is determined based on the mapping relationship, and the pixel coordinates of feature points such as field ridges and corn rows in that frame are converted to coordinates in a unified coordinate system. For the frame image of the southern video stream at the same time, the coordinate transformation is completed through the corresponding affine transformation matrix, so that the spatial reference of the cornfield image from the southern perspective is consistent with that from the eastern perspective. The same process is applied to the frame image of the western video stream, and resampling is used to eliminate pixel holes in some areas after transformation. Finally, all three video frames are adapted to a unified spatial coordinate system.
[0148] S32. The overlapping areas of each frame image transformed to a unified spatial coordinate system are fused to generate a continuous panoramic video frame sequence, which constitutes a panoramic video stream;
[0149] Specifically, for step S32, firstly, the overlapping areas between different video frames in a unified spatial coordinate system are identified. Then, the pixel values of the overlapping areas are calculated using a reasonable weight allocation rule. The multi-transformed image frames are seamlessly fused to form a continuous and complete panoramic video frame sequence, which in turn forms a panoramic video stream. Image registration technology is used to detect overlapping areas of different video frames in a unified coordinate system (e.g., areas where adjacent viewpoint video frames overlap by 15%-30%). For each pixel position within the overlapping area, the pixel fusion weight from different source video frames is calculated based on the distance from that position to the center of each source video frame (the closer the distance, the greater the weight) or a predefined weight function (e.g., fade-in / fade-out weight rule). The corresponding pixel values of each source frame are weighted and averaged according to the fusion weight to generate naturally transitioning fused pixel values, avoiding stitching marks in the overlapping areas. All non-overlapping area pixels are integrated with the fused overlapping area pixels to form a single complete panoramic video frame. All panoramic frames are arranged in chronological order to form a continuous panoramic video stream.
[0150] For example, continuing with the farmland scene described above, the transformed eastern, southern, and western video frames overlap in a unified coordinate system: the eastern and southern frames overlap by 22% (corresponding to the southeastern intersection of the farmland), and the southern and western frames overlap by 18% (corresponding to the southwestern intersection of the farmland). For the overlapping area between the eastern and southern frames, the distance of each pixel to the center of the eastern frame and the center of the southern frame is calculated. Pixels closer to the center of the eastern frame are assigned a weight of 0.6-1.0, and pixels closer to the center of the southern frame are assigned a weight of 0.0-0.4. The weighted average is then used to fuse the pixel values. The overlapping area between the southern and western frames is processed similarly, eliminating edge breaks between different viewpoints after fusion. The fused complete frames are arranged in timestamp order to form a continuous panoramic video frame sequence covering the entire farmland, constituting a panoramic video stream.
[0151] This embodiment unifies multiple video frames into the same spatial coordinate system based on mapping relationships, and then performs weighted fusion on overlapping areas. This effectively solves problems such as inconsistent spatial references of video streams from different perspectives and obvious stitching marks in overlapping areas. The resulting panoramic video stream is continuous, complete, and has a natural transition. It achieves full coverage of the monitoring area while ensuring the continuity and integrity of the image, thus improving the visual effect and practical value of panoramic monitoring.
[0152] Furthermore, in some embodiments, step S31, "transforming each frame of the multi-video stream to a unified spatial coordinate system according to the mapping relationship," may specifically include:
[0153] S311. Based on the mapping relationship, determine the geometric transformation parameters required to map points in the image coordinate system of each video stream to a unified spatial coordinate system;
[0154] Specifically, for step S311, based on the established mapping relationship between video feature points and a unified spatial coordinate system, a sufficient number of corresponding feature point coordinate pairs are selected. The geometric transformation parameters required for coordinate mapping are obtained through mathematical solutions, providing a core basis for image transformation. First, feature point coordinate pairs between the unified spatial coordinate system and the image coordinate system of each video stream are extracted from the mapping relationship. Each video stream needs at least four sets of non-collinear corresponding feature point coordinates (ensuring the uniqueness of the transformation matrix solution). Then, based on the type of coordinate pair (perspective correspondence or affine correspondence), the least squares method is selected as the solution tool to construct an error equation. By minimizing the pixel coordinate mapping error, the eight coefficients of the perspective transformation matrix or the six coefficients of the affine transformation matrix are solved. These coefficients are the final geometric transformation parameters, directly determining the accuracy of the image transformation.
[0155] For example, in a farmland monitoring scenario, a video stream from an eastern perspective has been mapped to a unified spatial coordinate system. Four sets of feature point coordinate pairs are extracted: (150, 220) in the image coordinate system corresponds to (80, 60, 55) in the unified coordinate system; (280, 350) corresponds to (120, 90, 55); (100, 400) corresponds to (70, 110, 55); and (320, 200) corresponds to (130, 50, 55). The perspective transformation matrix is solved using the least squares method, yielding eight coefficients (e.g., a1=0.8, a2=0.1, ..., a8=25.3). This matrix represents the geometric transformation parameters for the video stream image transformation.
[0156] S312. Based on the geometric transformation parameters, perform perspective transformation or affine transformation on each frame of each video stream to obtain the transformed image;
[0157] Specifically, for step S312, based on the determined geometric transformation parameters, a mapping rule is established between the image pixels of each video stream and the pixels in the unified spatial coordinate system. Coordinate transformation is performed on all pixels of each frame, transforming the image from the original viewpoint to the target viewpoint in the unified spatial coordinate system, resulting in the transformed image. For each single frame of the video stream, a mapping equation is first established between the original image coordinate system (u,v) and the target image coordinate system (x,y) in the unified spatial coordinate system, based on the geometric transformation parameters (perspective transformation matrix or affine transformation matrix). Then, each pixel of the original image is traversed, and its original coordinates are substituted into the mapping equation to calculate the corresponding coordinates of the pixel in the target coordinate system. If the target coordinates are not integers, their floating-point coordinates are recorded to prepare for subsequent resampling, ensuring that all pixels complete the viewpoint transformation and form the initially transformed image.
[0158] For example, continuing with the farmland scene above, for a certain frame of the video stream from the eastern perspective, the pixel (200, 300) in the original image is substituted into the solved perspective transformation matrix, and its target coordinates in the unified spatial coordinate system are calculated as (95.6, 78.2) through the mapping equation. Following this logic, all 1920×1080 pixels of the frame are traversed to complete the coordinate transformation of each pixel, and the image after preliminary transformation is obtained. At this time, the image perspective has been transformed from the local perspective in the eastern region to the globally adapted perspective in the unified coordinate system.
[0159] S313. Resample the transformed image to eliminate pixel holes or overlaps caused by the transformation, and obtain image frames aligned in a unified spatial coordinate system;
[0160] Specifically, for step S313, for pixel holes (blank areas without pixel coverage) or pixel overlaps (areas where multiple original pixels are mapped to the same target coordinate) appearing in the transformed image, a resampling algorithm is used to fill in blank pixels and regularize overlapping pixels to ensure that the image is complete and distortion-free in a unified spatial coordinate system, ultimately obtaining an aligned image frame. For pixel holes, resampling methods such as bilinear interpolation and bicubic interpolation are used to calculate the pixel value of the blank position based on the existing effective pixel values around the hole, thus filling the hole; for pixel overlaps, the pixel values of multiple original pixels are integrated by averaging, weighted averaging, etc., to obtain a single, clear target pixel value; during the resampling process, the resolution and image quality of the image are maintained to ensure that the transformed image is spatially aligned with the transformed image frames of other video streams in a unified coordinate system, without misalignment or redundancy.
[0161] For example, in the image transformed from the eastern perspective video stream, there are pixel holes around the target coordinates (102.3, 85.7). Using bilinear interpolation, based on the gray values of the four adjacent effective pixels (102, 85), (102, 86), (103, 85), and (103, 86) (120, 125, 118, 123 respectively), the pixel gray value at the hole location is calculated to be 121.5, thus filling the hole. For the target coordinates (88.0, 66.0), where two original pixel mappings overlap, the average of the two original pixel gray values (130 and 132), 131, is taken as the final pixel value at that location. After resampling, the image is free of holes and overlaps, becoming a complete image frame aligned in a unified spatial coordinate system.
[0162] This embodiment solves the geometric transformation parameters by selecting feature point coordinates, performs precise perspective or affine transformation on each video frame, and then eliminates pixel holes and overlaps caused by the transformation through resampling. This achieves precise alignment of multi-channel video stream image frames to a unified spatial coordinate system, ensuring the integrity and distortion-free nature of the transformed image. This provides a high-quality image foundation for subsequent fusion of overlapping areas of multi-channel images and generation of panoramic video streams, improving the accuracy and efficiency of panoramic stitching.
[0163] Furthermore, in some embodiments, step S312, "performing perspective transformation or affine transformation on each frame of each video stream according to geometric transformation parameters to obtain the transformed image," may specifically include:
[0164] S3121. For the current image frame to be transformed, calculate the mapping relationship between the source image pixels and the target coordinate system pixels according to the corresponding geometric transformation parameters;
[0165] Specifically, in step S3121, taking the current single-frame image to be transformed as the processing object, and relying on pre-determined geometric transformation parameters (such as perspective transformation matrix and affine transformation matrix), a mathematical correspondence rule between the source image pixel coordinates and the target coordinate system (unified spatial coordinate system) pixel coordinates is established, clarifying the unique corresponding position of each source pixel in the target coordinate system. The geometric transformation parameters already include spatial transformation information such as scaling, rotation, translation, and projection. Based on these parameters, mapping equations are constructed, such as the homogeneous coordinate equation of perspective transformation and the linear equation of affine transformation. Through these equations, a one-way calculation from the source image pixel coordinates to the target coordinate system pixel coordinates can be realized, ensuring the uniqueness and accuracy of the mapping relationship, and providing a clear calculation basis for subsequent pixel-by-pixel transformation.
[0166] S3122. Based on the mapping relationship, traverse all pixels of the current image frame and transform each pixel from its original image coordinate position to a coordinate position defined by a unified spatial coordinate system;
[0167] Specifically, for step S3122, all pixels of the current image frame are scanned one by one according to a preset order (such as from left to right, from top to bottom). For each pixel, the established mapping relationship is applied, and its target coordinates in a unified spatial coordinate system are calculated, completing the coordinate transformation of a single pixel and ultimately achieving the perspective transformation of the entire image frame. The traversal process must cover all pixels of the image frame to ensure no omissions. For each pixel, its original coordinates are substituted into the mapping equation to accurately calculate the target coordinates. If the target coordinates are floating-point values, their high-precision values are retained to provide an accurate position reference for subsequent resampling. The entire traversal process keeps the grayscale / color information of the pixels unchanged, only changing their spatial coordinate positions.
[0168] This embodiment constructs a precise pixel mapping relationship for a single frame image and performs coordinate transformation by traversing all pixels one by one. This ensures that each pixel in the current image frame can accurately adapt to a unified spatial coordinate system, avoiding pixel misalignment or omission during the transformation process. It achieves a precise conversion from the original image perspective to the target spatial perspective, providing a pixel-level accurate foundation for subsequent resampling to eliminate pixel holes and overlaps, and for multi-channel image fusion and stitching, thereby improving the consistency and accuracy of image transformation.
[0169] Furthermore, in some embodiments, step S3122, "based on the mapping relationship, determining the geometric transformation parameters required to map points in each video stream image coordinate system to a unified spatial coordinate system," may specifically include:
[0170] Based on the mapping relationship, at least four sets of corresponding feature point coordinate pairs are obtained between the unified spatial coordinate system and the image coordinate system of each video stream.
[0171] Based on the feature point coordinate pairs, the coefficients of the perspective transformation matrix or the affine transformation matrix are solved by the least squares method, and the coefficients are used as geometric transformation parameters.
[0172] Specifically, for step S3122, from the established mapping relationship between video feature points and the unified spatial coordinate system, at least four sets of non-collinear feature point coordinate pairs are selected and extracted for each video stream. Each coordinate pair includes the pixel coordinates of the feature point in the video stream image coordinate system and its spatial coordinates in the unified spatial coordinate system, providing basic data for subsequent calculation of the transformation matrix. The mapping relationship clearly defines the correspondence between video feature points in the two coordinate systems. During extraction, the validity of the coordinate pairs must be ensured. The feature points must be stable and highly recognizable points in the image (such as corner points and edge endpoints), and the four sets of coordinate pairs cannot be collinear; otherwise, the geometric transformation matrix cannot be uniquely determined. Each set of coordinate pairs must accurately correspond to the same actual scene target to ensure the authenticity of the coordinate mapping. After extraction, the coordinates are organized according to a fixed format.
[0173] Using the obtained feature point coordinate pairs as known conditions, a mathematical model that minimizes error is constructed. The least squares method is used to solve for all coefficients of the perspective transformation matrix or affine transformation matrix; these coefficients constitute the geometric transformation parameters for image coordinate mapping. First, based on the viewing characteristics of the monitoring scene (such as whether the lens exhibits perspective distortion), the choice between perspective and affine transformation is determined. Perspective transformation requires solving for 8 coefficients, while affine transformation requires solving for 6. Then, an error equation for coordinate mapping is constructed, which minimizes the sum of squared deviations between the predicted coordinates calculated using the transformation matrix and the actual unified spatial coordinates. This equation is solved using the least squares method to obtain the matrix coefficients that minimize the deviation, ensuring that the transformation matrix accurately reflects the mapping relationship between the two coordinate systems. These coefficients together constitute the geometric transformation parameters.
[0174] This embodiment extracts at least four sets of non-collinear feature point coordinate pairs and uses the least squares method to solve for the coefficients of the perspective or affine transformation matrix. The resulting geometric transformation parameters can accurately characterize the mapping relationship between the video stream image coordinate system and the unified spatial coordinate system, effectively minimizing coordinate mapping errors. This provides a high-precision mathematical basis for the spatial transformation of subsequent image frames, ensuring accurate alignment of multiple video streams in the unified spatial coordinate system and laying a core foundation for the accuracy of panoramic video stitching.
[0175] Furthermore, in some embodiments, step S32, "fusion of overlapping regions of the images transformed to a unified spatial coordinate system to generate a continuous panoramic video frame sequence," may specifically include:
[0176] S321. Determine the overlapping area between adjacent image frames in a unified spatial coordinate system;
[0177] Specifically, in step S321, for the multiple image frames that have been transformed to a unified spatial coordinate system, the overlapping areas between adjacent viewpoint image frames are identified through feature point matching and spatial position comparison, clarifying the boundary range of the overlapping area and defining the target area for subsequent fusion processing. Using aligned feature points of the same name in adjacent image frames (such as field ridge bends, facility corners, etc.), the common distribution range of feature points in the two image frames is determined; edge detection technology is used to capture the continuous boundaries of image content, and combined with coordinate range comparison, the boundaries between overlapping and non-overlapping areas are accurately delineated; the X and Y axis ranges of the overlapping area are clarified using coordinate values in a unified spatial coordinate system, ensuring that subsequent processing is only performed on this specific area.
[0178] For example, in a farmland monitoring scenario, there are two adjacent image frames, one in the east and one in the south, under a unified spatial coordinate system. By comparing the same feature points such as "field ridge intersection" and "corn row endpoint" in the two frames, it was found that the two images share common feature point distributions within the coordinate range of X: 50-120m and Y: 40-90m. Combined with edge detection, it was confirmed that the image content within this range continuously overlaps. Finally, this coordinate interval was determined to be the overlapping area of the two image frames, while the remaining parts were non-overlapping areas.
[0179] S322. For each pixel position within the overlapping region, calculate the fusion weight of pixel values from different source image frames based on the distance from each pixel position to the center of each source image frame or a predefined weight function;
[0180] Specifically, in step S322, for each pixel within the overlapping region, a weight value from different source image frames is assigned using distance weighting or a predefined weighting rule. The sum of the weights is 1, ensuring that the fused pixel value retains the image information of each source frame while achieving a smooth transition. If a distance weighting rule is used, the closer the pixel is to the center of a source image frame, the greater the weight corresponding to that source frame (e.g., distance and weight are inversely proportional). If a predefined weighting function (e.g., a fade-in / fade-out function) is used, the weight is assigned according to the pixel's position in the overlapping region. Pixels closer to the edge of the first source frame are given a low weight in the first source frame, a high weight in the second source frame, and pixels closer to the middle transition line are given equal weights. The weight calculation must cover all pixels within the overlapping region to ensure that the weight assignment of each pixel accurately corresponds to its spatial position.
[0181] For example, continuing with the farmland scene above, the center coordinates of the eastern image frame are (80m, 60m), the center coordinates of the southern image frame are (90m, 70m), and the position of a pixel within the overlapping area is (75m, 55m). The distance from this pixel to the center of the eastern frame is calculated as √[(80-75)²+(60-55)²]≈7.07m, and the distance to the center of the southern frame is √[(90-75)²+(70-55)²]≈21.21m. The weights are inversely proportional to the distance: the weight of the eastern frame is 21.21 / (7.07+21.21)≈0.75, and the weight of the southern frame is 7.07 / (7.07+21.21)≈0.25. That is, the weight of this pixel originating from the eastern frame is 0.75, and the weight of it originating from the southern frame is 0.25.
[0182] S323. Based on the fusion weight, perform a weighted average calculation on the corresponding pixel values from different source image frames to generate fused pixel values, forming continuous panoramic video frames;
[0183] Specifically, in step S323, for each pixel position within the overlapping region, the pixel values (grayscale or RGB values) corresponding to different source image frames are multiplied by their respective weights and summed to obtain the fused pixel value. The fused overlapping region is then stitched together with the non-overlapping regions of the two image frames to form a single complete panoramic video frame. All panoramic video frames are arranged in timestamp order to form a continuous panoramic video frame sequence. The weighted average calculation must maintain the numerical range of pixel values (e.g., grayscale values 0-255) to avoid overflow distortion. The original pixel values of each source image frame are directly retained in the non-overlapping region without fusion processing. During stitching, it is ensured that there are no obvious discontinuities at the boundaries between the non-overlapping region and the fused overlapping region, and that the timestamps between frames are continuous, forming a smooth panoramic video frame sequence.
[0184] For example, continuing with the farmland scene described above, the grayscale value of the pixel (75m, 55m) in the overlapping area is 130 in the eastern frame and 125 in the southern frame. The weighted grayscale value after fusion is calculated as 130 × 0.75 + 125 × 0.25 = 97.5 + 31.25 = 128.75, rounded down to 129. After weighted fusion of all overlapping area pixels, this is stitched together with the non-overlapping areas of the eastern frame (X: 0-50m, Y: 0-120m) and the southern frame (X: 120-180m, Y: 0-120m) to form a complete panoramic frame covering X: 0-180m and Y: 0-120m. The panoramic frames fused from multiple image frames from the eastern, southern, and western regions at the same time are arranged chronologically to form a continuous panoramic video frame sequence.
[0185] This embodiment accurately identifies the overlapping areas of adjacent image frames, assigns fusion weights using distance or predefined weight rules, and then achieves seamless fusion of overlapping areas through weighted averaging. This effectively eliminates edge breaks and stitching marks when stitching multiple images. Combined with direct stitching of non-overlapping areas, the generated panoramic video frames are continuous, complete, and have natural transitions, improving the visual coherence and image integrity of the panoramic video stream and providing a high-quality panoramic image foundation for subsequent dynamic scene analysis.
[0186] Furthermore, such as Figure 4 As shown, in some embodiments, step S4, "performing dynamic scene analysis based on environmental parameters in the panoramic video stream and associated data, identifying target types and states in the scene, and obtaining analysis results," may specifically include:
[0187] S41. Based on the panoramic video stream, establish a dynamic background model of the monitored area;
[0188] Specifically, for step S41, based on continuous panoramic video streams, a dynamic background model is constructed using statistical modeling or adaptive learning to adapt to real-time changes in the background of the monitored area. This model is used to accurately distinguish between background areas and foreground targets that appear abnormally. Continuous video frames are collected from the panoramic video stream over a period of time (e.g., 30 seconds to 1 minute). Statistical analysis is performed on the pixel grayscale values and color information of each frame, recording the normal fluctuation range of each pixel (e.g., changes in illumination, slight swaying of crops causing pixel changes). Through an adaptive update mechanism, the model can dynamically adjust to adapt to the natural dynamic changes in the monitored scene (e.g., gradual changes in illumination caused by cloud movement, slight swaying of crops with the wind), ensuring that the model can always accurately represent the current background state, providing a reliable benchmark for subsequent foreground extraction.
[0189] For example, in farmland monitoring scenarios, 60 consecutive panoramic video frames (covering the entire farmland area) are collected, and the grayscale value variation range of each pixel is statistically analyzed. For cornfield areas, the grayscale fluctuation range caused by slight shaking of pixels is recorded (e.g., 120-135), and for field ridge areas, the stable grayscale value of static pixels is recorded (e.g., 150-155). The model parameters are updated in real time through an adaptive algorithm, so that the model can adapt to the overall grayscale increase caused by enhanced afternoon light, and finally establish a dynamic background model that can reflect the natural dynamic background of the farmland.
[0190] S42. Use a dynamic background model to extract foreground targets and analyze the morphological and motion characteristics of the foreground targets;
[0191] Specifically, in step S42, background subtraction or inter-frame subtraction is used to compare the current panoramic video frame with the dynamic background model, separating regions that differ significantly from the background (i.e., foreground targets). Feature analysis is then performed on the extracted foreground targets to extract key morphological and motion features. Pixel-level subtraction is performed between the current frame image and the dynamic background model to filter out pixel regions whose grayscale values or color information exceed the normal fluctuation range of the background. Noise interference is removed through morphological processing (such as erosion and dilation) to obtain a complete foreground target outline. Subsequently, morphological features are analyzed, including the target's outline shape, overall size, aspect ratio, and local details (such as the presence of obvious head or torso outlines). Simultaneously, the positional changes of the foreground target in consecutive frames are tracked, and motion features are analyzed, including movement speed, direction of movement, movement frequency, and trajectory shape, providing core feature basis for subsequent target classification.
[0192] For example, continuing with the aforementioned farmland scene, comparing the current panoramic video frame with the dynamic background model reveals two foreground target regions: Region A has an upright outline, with overall dimensions of approximately 0.5m wide × 1.8m high, an aspect ratio of approximately 3:1, and locally visible outlines resembling a head and torso; Region B has an irregular, dispersed outline, with overall dimensions of approximately 1.2m wide × 0.8m high, an aspect ratio close to 1:1. Tracking 10 consecutive frames reveals that Region A moves along a north-south direction at a speed of 0.3m / s, following a straight trajectory; Region B only fluctuates up and down in place, with a motion frequency of approximately 8Hz and no obvious displacement trajectory. Based on this, the morphological and motion features of the two targets can be extracted.
[0193] S43. Based on morphological features, motion features, and environmental parameters in the associated data, classify the foreground targets, determine the real-time status corresponding to the classification results, and generate analysis results;
[0194] Specifically, in step S43, the extracted morphological features, motion features, and environmental parameters (such as crop height, terrain type, and light intensity) from the associated data are fused and analyzed from multiple dimensions. Based on preset classification rules, the specific type of the foreground target is determined, and its real-time state is clarified, ultimately forming a structured analysis result. Environmental parameters provide a basis for scene adaptation for classification. For example, crop height can help distinguish between interfering targets matching the crop height (such as crop swaying) and abnormal targets exceeding that height (such as human figures). The multi-dimensional features are compared with preset target type standards (such as human figures, crop swaying, scarecrows, birds, etc.). If the feature matching degree exceeds a set threshold, it is determined to be the corresponding type. Simultaneously, motion features are combined to determine the target's real-time state (such as direction of movement, speed of movement, and whether it is stationary). Finally, the analysis result is generated in the form of "target type + real-time state". During classification, the estimated height of the foreground target (calculated from morphological features and mapping relationships) is compared with the height of local crops in the environmental parameters. If the target height is significantly higher than the crop height and has an upright form and walking movement characteristics, it is classified as 'human-shaped'. If the target height is similar to the crop height and the movement exhibits high-frequency, low-amplitude, random swaying characteristics, it is classified as 'crop swaying'.
[0195] For example, continuing with the above farmland scenario, the environmental parameters in the associated data are "average corn crop height 2.5m, light intensity 1200Lux, terrain is flat farmland". Analysis of region A: morphological features (upright outline, 3:1 aspect ratio, head-to-body division) match the standard features of "humanoid" with a 90% match rate; motion features (0.3m / s linear movement, no high-frequency fluctuations) conform to the movement patterns of humanoids; and the target height (1.8m) matches the corn height (2.5m), thus it is determined to be a "humanoid target," with a real-time status of "moving at a uniform speed along the north-south direction of the farmland". Analysis of region B: morphological features (irregularly dispersed outline, 1:1 aspect ratio) match the standard features of "corn swaying" with an 85% match rate; motion features (8Hz stationary fluctuations, no displacement) conform to the swaying patterns of crops with the wind; combined with the environmental parameters "corn height 2.5m", it is determined to be "crop swaying interference," with a real-time status of "no abnormal displacement, only natural swaying". Finally, analysis results containing two target types and their corresponding states are generated.
[0196] This embodiment establishes a background model that adapts to the dynamic changes of the monitoring scene, accurately extracts foreground targets and analyzes their core features, and then combines environmental parameters for multi-dimensional fusion classification. This effectively achieves accurate differentiation of different types of targets in the monitoring scene, avoids misjudgment caused by single feature judgment, and clarifies the real-time status of the targets. The generated analysis results are accurate and comprehensive, providing a reliable basis for the formulation of subsequent equipment control strategies and improving the intelligence level and accuracy of panoramic monitoring.
[0197] Furthermore, in some embodiments, step S43, "classifying the foreground target based on morphological features, motion features, and environmental parameters in the associated data, determining the real-time state corresponding to the classification result, and generating analysis results," may specifically include:
[0198] S431. Compare the morphological and motion features of the foreground target with the predefined standard feature ranges of various target types, and generate comparison results;
[0199] Specifically, for step S431, standard morphological feature ranges and standard motion feature ranges are pre-defined for various target types that may appear in the monitoring scene (such as human figures, swaying crops, scarecrows, birds, etc.). The extracted morphological and motion features of the foreground target are compared with these standard ranges one by one to quantify the degree of feature matching and generate corresponding comparison results. The standard feature ranges need to be preset based on scene characteristics. Morphological feature standards include the target's aspect ratio range, contour complexity threshold, and the proportion range of key parts (such as head-to-torso). Motion feature standards include the movement speed range, movement frequency range, and trajectory type (straight line / irregular). During comparison, the similarity between the foreground target features and standard features (such as Euclidean distance, cosine similarity) is calculated to quantify the degree of matching (such as a value between 0 and 1, with the closer to 1 indicating a higher degree of matching). Each target type corresponds to a set of matching degree data to form a comparison result.
[0200] For example, in a farmland monitoring scenario, three types of target standard feature ranges are predefined: ① Human figure: aspect ratio 2.5-3.5:1, moving speed 0.1-0.5m / s, movement frequency 0.5-2Hz; ② Crop swaying: aspect ratio 0.8-1.2:1, moving speed 0-0.05m / s, movement frequency 6-10Hz; ③ Scarecrow: aspect ratio 2-3:1, moving speed 0m / s, movement frequency 0Hz. A foreground target has the following characteristics: aspect ratio 3:1, moving speed 0.3m / s, and movement frequency 1Hz. Comparing it with these three standard ranges yields the following results: matching degree with human figure 0.92, matching degree with crop swaying 0.21, and matching degree with scarecrow 0.35.
[0201] S432. Based on the comparison results and combined with the auxiliary information of environmental parameters in the associated data for feature judgment, calculate the classification probability of the foreground target belonging to each category;
[0202] Specifically, in step S432, based on the feature matching degree, the auxiliary judgment role of environmental parameters in the associated data is incorporated. Through weighted calculation or a probability model, the classification probability of the foreground target belonging to each predefined category is obtained, with the sum of probabilities being 1. The auxiliary role of environmental parameters is reflected in scene adaptability correction. For example, crop height can limit the feature weights related to target height. When the target height is close to the crop height, the classification weight of crop swaying is increased. Light intensity can correct the matching reliability of morphological features. For example, under weak light, the morphological feature recognition is low, so its weight is appropriately reduced, and the weight of motion features is increased. Through the preset weight allocation rules, the feature matching degree is multiplied by the auxiliary coefficient of environmental parameters and then normalized to obtain the classification probability of each category.
[0203] For example, continuing with the above farmland scenario, the environmental parameters in the associated data are "average height of corn crops 2.5m, light intensity 1300Lux (strong light, high morphological feature recognition)". The height of the foreground target is 1.8m, which differs from the height of the corn. This reduces the auxiliary coefficient for the crop swaying category (0.3) and increases the auxiliary coefficient for the humanoid category (1.2). Due to sufficient light, the weight of morphological features is set to 0.6, and the weight of motion features is set to 0.4. Based on the matching degree from step 1, calculate the classification probabilities: Humanoid probability = (0.92 × 0.6 × 1.2 + 0.92 × 0.4 × 1.2) / normalization coefficient ≈ 0.88; Crop shaking probability = (0.21 × 0.6 × 0.3 + 0.21 × 0.4 × 0.3) / normalization coefficient ≈ 0.05; Scarecrow probability = (0.35 × 0.6 × 1.0 + 0.35 × 0.4 × 1.0) / normalization coefficient ≈ 0.07 (the normalization coefficient ensures that the sum of the three is 1).
[0204] S433. Determine whether the classification probability exceeds the classification threshold of the corresponding category, determine the final classification result of the foreground target, and use the key parameters in the motion features as the real-time state of the foreground target to form the analysis result.
[0205] Specifically, for step S433, a classification threshold is set for each predefined target type. The classification probability of each category is compared with the corresponding threshold, and the category with the highest probability exceeding the threshold is determined as the final classification result of the foreground target. Key parameters (such as movement speed, direction, and frequency) are extracted from motion features as the real-time state of the target. The classification results and real-time state are integrated to form a complete analysis result. The classification threshold needs to be preset based on the risk of scene misjudgment to ensure the recognition accuracy of core targets (such as human figures). If the probability of all categories does not exceed the threshold, it is judged as an "unknown target". The key parameters of the real-time state need to be able to intuitively reflect the target's dynamics, such as movement speed, movement direction, and whether it is stationary, and together with the classification result, they constitute a structured analysis result (such as "Target type: human figure; Real-time state: moving at a constant speed of 0.3m / s along the due north direction").
[0206] For example, continuing with the above farmland scenario, the preset classification threshold is 0.7. The classification probability of a human figure is 0.88 > 0.7, crop swaying is 0.05 < 0.7, and scarecrow is 0.07 < 0.7. Therefore, the final classification result is determined to be "human figure". Key parameters are extracted from the motion features: movement speed 0.3 m / s, movement direction due north, and movement frequency 1 Hz, as the real-time state. The final analysis result is: "Target type: human figure; Real-time state: moving at a constant speed of 0.3 m / s along the due north direction, movement frequency 1 Hz".
[0207] This embodiment compares the features of the foreground target with a predefined range of standard features, combines environmental parameters to assist in calculating the classification probability, and then determines the final classification result and extracts the real-time status through threshold judgment. This achieves accurate classification and dynamic representation of the foreground target, effectively reducing the risk of misjudgment caused by single feature judgment or lack of scene adaptation. The generated analysis results are both accurate and complete, providing accurate and reliable decision-making basis for the formulation of subsequent equipment control strategies and improving the intelligent recognition capability of the monitoring system.
[0208] Furthermore, such as Figure 5 As shown, in some embodiments, step S5, "generating a device control strategy based on the analysis results, and generating and issuing control instructions for adjusting the operating parameters of the multi-view image acquisition device based on the device control strategy," may specifically include:
[0209] S51. If a specific type of target is identified in the analysis results, a video tracking strategy or alarm strategy corresponding to the specific type of target will be generated;
[0210] Specifically, for step S51, the scene analysis results are first analyzed to determine whether there are any preset specific types of targets. If so, based on the target's type, real-time status (such as movement speed, location, and trajectory), and the core requirements of the monitoring scene, a targeted video tracking strategy or alarm strategy is generated to ensure that the strategy accurately adapts to the target's characteristics. Specific types of targets refer to targets that require focused attention in the monitoring scene, such as human figures or illegally intruding devices in farmland scenes, or unfamiliar vehicles or suspicious packages in park scenes. Their type must be combined with scene presets and their priority clearly defined. If the target is in motion and requires continuous monitoring of details, a video tracking strategy is generated; if the target poses a security risk (such as illegal intrusion or damage to facilities), an alarm strategy is generated simultaneously or separately. The strategy must clearly define the core control direction, such as a tracking strategy including "focusing on the target's location and adjusting shooting parameters for clear capture," or an alarm strategy including "triggering an alarm and pushing early warning information."
[0211] For example, in a farmland monitoring scenario, specific target types are preset as human figures (illegal intrusion) and large agricultural machinery (illegal operation). Analysis results show "Target type: Human figure; Real-time status: Located at the southwest boundary of the farmland (uniform spatial coordinates X: 150m, Y: 90m), moving at 0.5m / s along the farmland interior," classifying it as a specific target type requiring close monitoring. Based on this, a video tracking strategy is generated: ① Adjust the shooting angle of the nearest multi-view image acquisition device (No. 3) to the southwest direction, covering the target's current location and predicted movement trajectory range (X: 140-170m, Y: 80-110m); ② Adjust the focal length of this device from 15mm to 25mm to improve the clarity of the target image; ③ Increase the shooting frame rate from 25fps to 30fps to ensure continuous capture of the target's movements. Simultaneously, an alarm strategy is generated: ① Trigger an audible and visual alarm at the monitoring center; ② Push warning information to the management equipment of the administrators, including the target's location, real-time screenshot, and direction of movement.
[0212] S52. Based on the environmental parameters in the associated data and combined with the real-time energy consumption status of the panoramic video surveillance system, generate an equipment energy consumption control strategy;
[0213] Specifically, for step S52, environmental parameters (such as light intensity, temperature, wind speed, etc.) are extracted from the associated data. Simultaneously, the real-time energy consumption status of the panoramic video surveillance system is collected, including the power consumption of each multi-view image acquisition device, the load energy consumption of the edge computing module and cloud server. Through fusion analysis, a balance point between monitoring effectiveness and energy consumption optimization is found, generating a device energy consumption control strategy. Environmental parameters directly affect the energy consumption requirements of the device's operating parameters. For example, in strong light environments, there is no need to turn on supplementary lighting; exposure can be reduced to decrease energy consumption. In low light environments, supplementary lighting and energy consumption need to be balanced. The real-time energy consumption status must include the real-time power, cumulative energy consumption, and load rate of each device (such as the CPU utilization rate of the acquisition device and the bandwidth utilization rate of the transmission link). If the energy consumption exceeds a preset threshold or the device is in a low-load state, parameters need to be optimized without affecting the monitoring effect. The strategy must clearly define the specific direction of energy consumption optimization parameter adjustments to ensure that energy consumption is reasonable and controllable.
[0214] For example, continuing with the above farmland scenario, the environmental parameters in the associated data are "light intensity 1800 Lux (strong light), ambient temperature 28℃, wind force level 2"; the system's real-time energy consumption status is "current power of acquisition device 3 is 25W (load rate 70%), power of other acquisition devices is 18W (load rate 40%), and overall energy consumption exceeds the preset threshold by 5%". Based on this, the following device energy consumption control strategies are generated: ① Acquisition device 3 (tracking human-shaped target): Due to the strong light environment, adjust the exposure from the current 450 to 300, turn off the fill light function, and reduce power consumption; maintain a frame rate of 30fps to ensure tracking effect, without adjustment; ② Other low-load acquisition devices: reduce the shooting frame rate from 25fps to 20fps, keep the focal length at the current value, and the power can be reduced to about 15W; ③ Edge computing module: optimize the data transmission caching strategy to reduce the computing load during idle periods and reduce energy consumption.
[0215] S53. Generate corresponding control instructions based on video tracking strategies, alarm strategies, or equipment energy consumption control strategies;
[0216] Specifically, in step S53, the generated video tracking / alarm strategy and device power consumption control strategy are converted into standardized control commands that can be directly recognized and executed by the multi-view image acquisition device. The conversion process must strictly follow the device's communication protocol and command format to ensure that the command parameters are accurate and structurally complete. First, the core control requirements in each strategy are analyzed to clarify the types of device operating parameters to be adjusted (such as shooting angle, focal length, frame rate, exposure, fill light switch, etc.) and their specific values; then, according to the command protocol of the multi-view image acquisition device, the parameters are encoded (such as the numerical format and unit identifier of parameters such as angle and focal length) to construct a complete command frame containing "start bit, device identifier, parameter type, parameter value, check bit, and end bit"; if there are multiple strategies (such as tracking strategy and power consumption strategy), the commands need to be integrated to ensure that the parameter adjustments do not conflict (such as balancing the need to increase the frame rate for tracking and the need to reduce the frame rate for power consumption optimization, prioritizing the tracking needs).
[0217] For example, continuing with the above farmland scenario, the instruction protocol of acquisition device No. 3 requires parameters to be encoded in hexadecimal (1° corresponds to 0x01, 1mm corresponds to 0x01, 1fps corresponds to 0x01, 1 unit exposure corresponds to 0x01). Integrating video tracking strategy and energy consumption control strategy, the control instructions for acquisition device No. 3 are generated as follows: ① Angle adjustment instruction (target direction southwest, corresponding angle +20°): 0x01 (start position) + 0x03 (device identifier: No. 3) + 0x02 (parameter type: angle) + 0x14 (+20° encoding) + 0x4A (check bit) + 0xFF (end position); ② Focal length adjustment instruction (25mm): 0x01 + 0x03 + 0x04 (parameter type: angle) + 0x05 (parameter type: angle) + 0x06 (parameter type: angle) + 0x07 (parameter type: angle) + 0x08 (parameter type: angle) + 0x09 (parameter type: angle) + 0x0001 (parameter type: angle) + 0x ... ① Parameter type: focal length) + 0x19 (25mm encoding) + 0x50 + 0xFF; ② Exposure adjustment command (300): 0x01 + 0x03 + 0x05 (parameter type: exposure) + 0x12C (300 encoding) + 0x6D + 0xFF; ③ Frame rate adjustment command (30fps): 0x01 + 0x03 + 0x06 (parameter type: frame rate) + 0x1E (30fps encoding) + 0x5A + 0xFF. Simultaneously generate frame rate adjustment commands (20fps, encoding 0x14) for other low-load acquisition devices.
[0218] S54. Issue control commands to the multi-target image acquisition device to adjust the operating parameters of the multi-target image acquisition device;
[0219] Specifically, for step S54, the standardized control commands generated are accurately sent to the corresponding multi-view image acquisition devices via the pre-set communication links (such as wireless LAN, 4G / 5G, and industrial Ethernet) of the panoramic video monitoring system. During the sending process, command information (such as sending time, target device, and command content) is recorded to ensure the real-time and accuracy of command transmission, while also ensuring subsequent traceability. First, the control commands are matched with the corresponding devices to identify the multi-view image acquisition device for each command (distinguished by device identification). Then, a stable connection is established with the target device through the system's communication module, and the command is sent according to the pre-set transmission protocol. After the command is sent, feedback information from the device is received (such as "command received successfully" or "parameter adjustment completed"). If no feedback is received or the feedback is abnormal, a retransmission mechanism is triggered to ensure that the command is effectively executed. The entire process must avoid command transmission delays or mistransmissions to ensure timely adjustment of device operating parameters. Control commands are sent to the multi-target image acquisition device to adjust its operating parameters, including but not limited to the gimbal's pitch and azimuth angles, optical zoom magnification, digital zoom ratio, image sensor exposure time and gain, video encoding frame rate and bit rate, and the on / off state and intensity of the fill light.
[0220] For example, continuing with the above farmland scenario, using a wireless local area network (WiFi 6) as the communication link, the four control commands from acquisition device 3 and the frame rate adjustment commands from other low-load acquisition devices were sent out respectively: ① First, a command was sent to acquisition device 3, with the sending time recorded as 2025-12-18 14:25:00.345 and the device identifier as "Device_03"; ② One second later, device 3 responded with "Command received successfully, parameter adjustment started," and five seconds later, a confirmation message "Parameter adjustment completed, current status normal" was received; ③ Simultaneously, frame rate adjustment commands were sent to the other four low-load acquisition devices, and the sending information was recorded. Feedback indicating adjustment completion was received from each device within three seconds. Ultimately, the operating parameters of all target acquisition devices were adjusted according to the control commands.
[0221] This embodiment achieves dynamic and precise adaptation of the operating parameters of multi-view image acquisition devices by accurately identifying specific target types and generating targeted tracking or alarm strategies. It then optimizes energy consumption control strategies by combining environmental parameters and real-time system energy consumption status, and finally converts these into standardized commands for execution. On the one hand, this ensures effective tracking and risk warning of key targets, improving the security protection capabilities of the monitoring system. On the other hand, it optimizes equipment energy consumption and reduces system operating costs while ensuring monitoring effectiveness. Simultaneously, the precise issuance and execution of commands ensures the timeliness and reliability of control, enhancing the panoramic video surveillance system's adaptability to dynamic scenes and its overall operational efficiency.
[0222] Furthermore, in some embodiments, step S52, "based on environmental parameters in the associated data and combined with the real-time energy consumption status of the panoramic video surveillance system, generating a device energy consumption control strategy," may specifically include:
[0223] S521. Obtain real-time light intensity and the remaining power of the panoramic video surveillance system;
[0224] Specifically, for step S521, two key data points are accurately collected: the real-time light intensity of the monitored scene and the remaining power of the panoramic video monitoring system. These provide fundamental data support for subsequent threshold judgment and strategy generation. The real-time light intensity is directly extracted from the environmental parameters of the associated data. This data is collected in real-time by the system's accompanying light sensor and must retain sufficient accuracy (e.g., accurate to 1 Lux) to ensure it accurately reflects the current lighting conditions. The remaining power of the panoramic video monitoring system must cover the power status of the system's core components (multi-view image acquisition device, edge computing module, backup power supply, etc.). This data is aggregated and calculated by the system's energy management module and ultimately presented as a percentage of the total remaining power, ensuring the data comprehensively reflects the system's energy reserves.
[0225] For example, in a farmland monitoring scenario, the real-time light intensity is extracted from the associated environmental parameters as 400 Lux. The system energy management module summarizes the power of each core device: the remaining power of the three multi-view image acquisition devices are 22%, 20%, and 19%, respectively; the remaining power of the edge computing module is 21%; and the remaining power of the backup power supply is 17%. The total remaining power of the panoramic video monitoring system is calculated to be 18%.
[0226] S522. If the real-time light intensity is detected to be lower than the preset light threshold and the remaining power is lower than the preset power threshold, then generate a device energy consumption control strategy for entering emergency mode.
[0227] Specifically, for step S522, the acquired real-time light intensity and remaining system power are first compared with preset light intensity thresholds and power thresholds, respectively. Emergency mode is triggered only when both conditions are met simultaneously (both thresholds are below preset values), generating a targeted device energy consumption control strategy. The core objective is to minimize energy consumption while ensuring basic monitoring functions. The preset light intensity threshold needs to be set based on the basic imaging requirements of the monitoring scenario (e.g., in low-light environments, the device needs to activate supplementary lighting, which significantly increases energy consumption; therefore, the threshold is usually set to the minimum light intensity required to maintain basic imaging). The preset power threshold needs to be set based on the minimum energy requirements for emergency system operation, ensuring that the system can still maintain basic monitoring for a certain period after triggering emergency mode. The device energy consumption control strategy for emergency mode needs to specify concrete energy optimization measures, including disabling unnecessary functions and adjusting core device operating parameters to reduce power consumption.
[0228] For example, continuing with the above farmland scenario, the preset illumination threshold is 500 Lux (below this value, supplementary lighting is required to ensure image clarity), and the preset power threshold is 20% (below this value, the system's energy reserves are insufficient, requiring emergency power saving). Testing showed that the real-time illumination intensity was 400 Lux < 500 Lux, and the system's remaining power was 18% < 20%, both threshold conditions were met. Therefore, the following energy consumption control strategy was generated for the emergency mode: ① Turn off the supplementary lighting function of all multi-view image acquisition devices (to avoid high energy consumption from supplementary lighting); ② Reduce the shooting frame rate of all acquisition devices from 25fps to 15fps (to reduce the energy consumption of image acquisition and transmission); ③ Adjust acquisition devices 2 and 4 in non-core monitoring areas (such as areas without key targets at the edge of farmland) to "intermittent working mode" (acquiring one frame every 3 seconds, and sleeping the rest of the time); ④ Disable unnecessary data preprocessing functions in the edge computing module, retaining only basic image compression and key target recognition functions to reduce computing power consumption.
[0229] This embodiment achieves precise control of system energy consumption under the dual challenges of insufficient light and insufficient energy reserves by accurately acquiring real-time light intensity and remaining system power, triggering an emergency mode based on dual threshold judgment, and generating targeted energy consumption control strategies. This strategy, while shutting down unnecessary functions and optimizing the operating parameters of core equipment, ensures the normal operation of the system's core monitoring functions, effectively reducing system energy consumption, extending the system's continuous operation time in emergency situations, avoiding monitoring interruptions due to power depletion, and improving the reliability and adaptability of the panoramic video surveillance system under complex energy and lighting conditions.
[0230] Furthermore, in some embodiments, the environmental parameters include object height distribution information; then, real-time splicing of multiple video streams according to the mapping relationship also includes dynamically adjusting the weight coefficients of different video streams in the splicing process according to the object height distribution information.
[0231] Specifically, the key information representing the height and distribution of various objects within the monitored area is precisely filtered and extracted from the environmental parameters contained in the associated data. This provides a core basis for subsequent weight coefficient adjustments. The object height distribution information must clearly include two core aspects: first, the types of objects present in the monitored scene (such as crops, buildings, facilities, natural obstacles, etc.); and second, the actual height value of each type of object, as well as the specific distribution range of each object in a unified spatial coordinate system (i.e., the horizontal area defined by the X and Y axes). During the extraction process, the completeness and accuracy of the information must be ensured to avoid deviations in subsequent weight adjustments due to missing object height or distribution range information.
[0232] By comparing the shooting angle, imaging range, and extracted object height distribution information of each video stream, the imaging clarity and information completeness of each video stream for objects of different heights are determined, clarifying the adaptation advantages of each video stream for object regions of different heights. Video streams with different viewing angles produce different imaging effects for objects of different heights. For example, video streams deployed at higher positions capture the top details of tall objects (such as irrigation towers) more clearly, while video streams deployed at lower positions capture the bottom outline of short objects (such as wheat fields or guardrails) more completely. Through analysis, a correspondence between "video stream - adapted height object - imaging advantage" needs to be established to provide a judgment standard for adjusting the weighting coefficients.
[0233] Based on imaging adaptability analysis, differentiated weight coefficients are assigned to video streams covering different height objects within the monitored area. The core principle is that the video stream with a greater advantage in imaging objects of a certain height has a higher weight coefficient in the corresponding area. The weight coefficient directly determines the pixel contribution of the video stream in the overlapping area of the stitching; the higher the weight, the greater the proportion of the video stream's pixel information in the stitched image. The adjustment process needs to achieve regional dynamic adaptation, meaning that the weight coefficient of the same video stream can be different in areas corresponding to objects of different heights; for areas with objects of the same height, weight coefficients with a total sum of 1 are assigned based on the differences in imaging advantages of each video stream to ensure optimal fusion of pixel information in the stitched image.
[0234] For example, for objects at height (such as irrigation towers), video streams with clear upward-facing views are prioritized and given higher weight; for objects at low height (such as ground crops), video streams with clear downward-facing or eye-level views are prioritized. The weighting coefficients can be calculated based on models such as the object's height, the angle between the object's height and the optical axis of each camera, and the distance between them.
[0235] Based on dynamically adjusted weighting coefficients, the pixels in overlapping areas of each video stream are weighted and fused. Non-overlapping areas directly retain the pixel information of the corresponding video stream, ultimately generating a complete and clear panoramic video stream. During the stitching process, for each pixel location, its corresponding object height region is first determined, and then the weighting coefficients of each video stream for that region are called. For pixels in overlapping areas, the pixel values corresponding to different video streams are weighted and summed according to the weighting coefficients to obtain the fused pixel value. Non-overlapping areas directly use the original pixel values of the corresponding video streams, ensuring that the stitched image clearly presents details in object regions of various heights, without blurring or discontinuities.
[0236] This embodiment extracts object height distribution information and analyzes its imaging adaptability with various video streams, dynamically adjusting the stitching weight coefficients. This allows video streams with more obvious imaging advantages to receive higher weights in corresponding height object regions, effectively solving the problems of image blurring and detail loss that easily occur with objects of different heights when stitching with fixed weights. The final generated panoramic video stream can clearly present details in object regions of various heights, significantly improving the accuracy of video stitching and image quality, and providing a more reliable panoramic image foundation for subsequent dynamic scene analysis and target recognition.
[0237] Furthermore, in some embodiments, the method further includes:
[0238] S61. Real-time monitoring of image quality metrics for panoramic video streams;
[0239] Specifically, for step S61, real-time frame data of the panoramic video stream is continuously collected, and key indicators characterizing image quality are dynamically calculated and monitored to ensure timely capture of changes in image quality and provide data support for subsequent control decisions. The core image quality indicators to be monitored need to be clearly defined, including but not limited to image sharpness (reflecting the sharpness of image details), signal-to-noise ratio (reflecting the degree of noise interference in the image), contrast ratio (reflecting the difference in brightness levels in the image), and brightness uniformity (reflecting the consistency of brightness distribution in the image). During the monitoring process, sample frames need to be extracted from the panoramic video stream at fixed time intervals (e.g., every frame or every 3 frames), and standardized algorithms are used to calculate the values of each quality indicator. For example, sharpness is calculated using an edge gradient operator (the higher the gradient value, the better the sharpness), and the signal-to-noise ratio is calculated using the ratio of signal amplitude to noise amplitude (the higher the ratio, the less noise). At the same time, the continuity of monitoring must be ensured to avoid missing quality degradation due to intermittent monitoring.
[0240] S62. When the image quality index is detected to be lower than the preset standard, an adjustment command is generated to adjust the shooting angle of at least one camera in the multi-view image acquisition device, and the adjustment command is sent to the corresponding camera.
[0241] Specifically, for step S62, the real-time monitored image quality indicators are first compared with preset standards. If the indicators are lower than the standards, the image quality is determined to be substandard. Then, the camera corresponding to the substandard area is located, the optimal shooting angle adjustment parameters are calculated, standardized control commands are generated and accurately issued, and dynamic correction of the shooting angle is achieved. The preset standards need to be set in combination with the basic imaging requirements of the monitoring scene (such as calibrating the threshold of each quality indicator based on the minimum image quality requirements for target recognition and scene observation). The location of the substandard area needs to rely on the correspondence between the area of the panoramic video stream and the coverage of the multi-view image acquisition device, that is, to clarify which camera or several camera video streams are stitched together to form a certain blurry area in the panoramic image, and to lock the responsible camera. The shooting angle adjustment parameters need to be calculated according to the type of quality problem (such as the image edge blur caused by camera angle deviation, the camera orientation needs to be adjusted so that the problem area enters the clear imaging range of the center of the lens). The control commands need to comply with the camera's communication protocol and clearly include core information such as device identification, adjustment direction, and adjustment angle value to ensure that the camera can accurately identify and execute them.
[0242] This embodiment monitors the image quality indicators of the panoramic video stream in real time, enabling timely detection of image quality degradation caused by issues such as camera angle deviation. By accurately locating the responsible camera, calculating the optimal adjustment angle, and issuing control commands, the shooting angle can be quickly corrected, restoring clear imaging of the target area. This achieves dynamic closed-loop control of image quality, effectively avoiding the problem of substandard image quality affecting subsequent target recognition and scene analysis. It ensures the continuous stability and clarity of the panoramic video stream, providing crucial support for the reliable operation of the monitoring system.
[0243] Furthermore, in some embodiments, the edge computing module includes a positioning unit, an environmental sensing unit, and a preprocessing unit;
[0244] The positioning unit is used to acquire the device location information of the multi-view image acquisition device;
[0245] Specifically, the positioning unit has a built-in high-precision positioning module (such as GPS or BeiDou positioning module) that receives satellite signals or ground reference station auxiliary signals to collect the spatial position parameters of the multi-view image acquisition device in real time. The core data collected includes latitude and longitude coordinates and altitude, and the data accuracy must meet the requirements of spatial modeling (such as latitude and longitude accurate to 4 decimal places and altitude accurate to 0.1 meters). During the acquisition process, the positioning unit updates the position data at a fixed frequency (such as once every 1 second) and performs noise filtering and error correction on the collected raw data (such as removing outliers caused by signal interference) to ensure the stability and accuracy of the data. Finally, the processed device position information is output in a standardized format (such as JSON format) for subsequent units to use.
[0246] An environmental sensing unit is used to collect at least one environmental parameter, including light intensity, wind speed, temperature, and crop height.
[0247] Specifically, the environmental sensing unit accurately collects various environmental parameters within the monitored scene, including at least one of light intensity, wind speed, temperature, and crop height. This provides scene-adaptive data for subsequent 3D spatial model calibration, video stitching weight adjustment, and equipment energy consumption optimization. The environmental sensing unit consists of various dedicated sensing modules (such as light sensors, wind speed sensors, temperature sensors, and laser rangefinders), each deployed as needed within the monitored area or on a multi-view image acquisition device. During the acquisition process, each sensing module works synchronously at a preset frequency (e.g., light intensity and temperature are collected every 30 seconds, wind speed every minute, and crop height every hour) to ensure data timeliness. The collected raw data undergoes preprocessing such as signal amplification, analog-to-digital conversion, and outlier removal to transform it into directly usable digital data. Finally, various environmental parameters are integrated into a unified dataset, sorted by timestamp, and output to ensure the temporal correlation between the data and subsequent video streams and equipment location information.
[0248] The preprocessing unit is used to perform frame synchronization and encoding on multiple video streams, and to perform time alignment and data association processing on the synchronized video streams, device location information and environmental parameters to form associated data.
[0249] Specifically, the preprocessing unit undertakes the basic processing tasks of multiple video streams, completes frame synchronization and encoding, and simultaneously aligns the synchronized video streams with device location information and environmental parameters in time and associates them with data, ultimately forming structured associated data, which provides standardized data input for subsequent core processes such as 3D modeling and video stitching. The preprocessing unit's operation consists of three core stages: First, frame synchronization processing, which extracts the timestamps of each video stream and adjusts video streams from different sources and with different transmission delays to the same time base, ensuring accurate correspondence of scene images at the same moment (e.g., uniformly calibrating the frame timestamps of the three video streams to the millisecond level to ensure synchronization error ≤10ms); Second, video encoding processing, which uses efficient video encoding standards (such as H.264 and H.265) to compress and encode the synchronized video stream, reducing data transmission and storage pressure, while retaining key feature information of the video frames (such as edges and corners) without affecting subsequent feature extraction; Third, time alignment and data association, which uses the unified timestamp of the preprocessing unit as a reference to bind the synchronized encoded video stream, the device location information output by the positioning unit, and the environmental parameters output by the environmental sensing unit according to the time dimension, ensuring that video data, location data, and environmental data at the same time node correspond one-to-one, ultimately forming structured associated data containing "timestamp + video frame data + device location information + environmental parameters", which is stored in the local cache or transmitted to the subsequent processing module.
[0250] This embodiment achieves accurate acquisition, basic preprocessing, and structured association of location data from multi-view image acquisition devices, monitoring scene environment data, and multi-channel video stream data through the collaborative work of the positioning unit, environment sensing unit, and preprocessing unit of the edge computing module. The positioning unit ensures the accuracy of spatial reference data, the environment sensing unit provides environmental basis for scene adaptation, and the preprocessing unit solves the synchronization problem of multiple video streams and realizes time alignment and association of multi-source data.
[0251] Furthermore, in some embodiments, the cloud server is configured with a graphics processing unit for performing video feature matching and mapping calculations, and an intelligent analysis model for running dynamic scene analysis.
[0252] Specifically, the graphics processing unit (GPU) is dedicated to performing two core tasks: video feature matching and mapping calculation. Leveraging its efficient parallel computing capabilities, it enhances the processing efficiency and accuracy of these data-intensive tasks, providing core computing power support for the association between video streams and a unified spatial coordinate system. The GPU possesses numerous parallel processing cores, capable of processing massive amounts of video feature point data simultaneously, adapting to the high computing power requirements of video feature matching and mapping calculation. Its workflow consists of two steps: First, video feature matching: key feature points (such as corner points and edge endpoints) are extracted from the input video stream frames. Feature descriptors are generated using preset feature description algorithms (such as SIFT and SURF algorithms). Then, feature matching algorithms (such as brute-force matching and FLANN matching) are used to accurately match feature points from different video streams or video feature points with spatial reference feature points, filtering out homologous feature point pairs with high matching degrees. Second, mapping calculation: based on the matched feature point pairs, the mapping relationship between the video image coordinate system and the unified spatial coordinate system (such as perspective transformation matrix and affine transformation matrix) is calculated through matrix solving and other operations, providing a precise mathematical basis for subsequent spatial alignment and stitching of the video stream. Throughout the process, the graphics processing unit (GPU) significantly reduces the time required for feature matching and mapping calculations through a parallel computing architecture, ensuring real-time processing.
[0253] Intelligent analysis models are specifically designed for dynamic scene analysis. Based on pre-defined algorithm logic and model parameters, they analyze input panoramic video streams and related data to perform dynamic scene perception tasks such as foreground target extraction, feature analysis, and target classification, outputting accurate scene analysis results. These models are typically built on deep learning frameworks (such as convolutional neural networks and recurrent neural networks), possessing strong scene adaptability and target recognition capabilities. Their workflow revolves around dynamic scene analysis: first, panoramic video stream data is retrieved, and environmental parameters (such as light intensity and crop height) from related data are combined to optimize the analysis strategy; a background model of the monitored area is constructed using a dynamic background modeling algorithm, and background subtraction is used to separate foreground targets from the background; morphological features (outline, size, aspect ratio) and motion features (speed, direction, trajectory) of the extracted foreground targets are analyzed; finally, based on pre-defined classification rules and model training parameters, the foreground targets are accurately classified (such as human figures, crop swaying, and large agricultural machinery), and structured analysis results including target type, real-time status, and spatial location are output. The model has a certain degree of adaptability, allowing for fine-tuning of the analysis logic according to environmental parameters of different monitoring scenarios to improve analysis accuracy.
[0254] This embodiment leverages the parallel computing advantages of the graphics processing unit to significantly improve the efficiency and accuracy of video feature matching and mapping calculations, ensuring the real-time performance and reliability of the association between the video stream and the unified spatial coordinate system. The intelligent analysis model, through precise dynamic scene analysis, achieves effective extraction and classification of foreground targets, outputting high-quality scene analysis results.
[0255] Furthermore, in some embodiments, the management device includes a mobile terminal and / or a fixed monitoring workstation for receiving and displaying panoramic video streams, analysis results, and / or alarm information triggered by the analysis results.
[0256] Specifically, mobile terminals serve as portable information receiving and display carriers, enabling users to view panoramic video streams, analysis results, and alarm information anytime, anywhere, adapting to the monitoring needs of managers in scenarios such as mobile inspections and field work. Mobile terminals include portable devices with wireless communication and display capabilities, such as smartphones and tablets. They establish communication connections with the panoramic video monitoring system through pre-installed dedicated monitoring clients (APPs) or web interfaces (supporting wireless communication methods such as 4G, 5G, and WiFi). They possess core capabilities such as real-time information reception, instant display, and push notifications. Panoramic video streams can be presented in real-time playback, analysis results can be displayed in structured lists and combined text and graphics (such as target type, location, and status), and alarm information can be proactively pushed through pop-ups, ringtones, and vibrations to ensure that managers are informed of abnormal situations immediately. Simultaneously, simple interactive operations (such as clicking to view video details and marking alarms for processing) are supported, enhancing ease of use.
[0257] For example, in farmland monitoring scenarios, managers carry smartphones (mobile terminals) and connect to the system via the pre-installed "Farmland Panoramic Monitoring APP". The APP's homepage plays a real-time panoramic video stream covering the entire farmland. Clicking the "Analysis Results" module allows users to view text and image analysis results such as "2025-12-18 10:30:00, a human-shaped target appeared in the northeast of the farmland and moved in a southwest direction." When the system detects an alarm message about an illegal intrusion, the phone immediately pops up a red alarm window accompanied by vibration, and displays a real-time video screenshot of the alarm location, facilitating a quick response from managers during field inspections.
[0258] Fixed monitoring workstations serve as centralized monitoring and management platforms in fixed locations, enabling stable and comprehensive reception and visualization of panoramic video streams, analysis results, and alarm information. They are well-suited for the centralized control needs of fixed areas such as monitoring centers and duty rooms. A fixed monitoring workstation typically consists of a high-performance host, multiple high-definition monitors, an operator console, and dedicated monitoring software. It establishes a stable communication connection with the panoramic video monitoring system via a wired LAN or fiber optic cable, possessing high-bandwidth, low-latency data transmission capabilities. It supports multi-screen split-screen display (such as simultaneously displaying panoramic video streams, multiple raw video streams, and analysis result data panels). Panoramic video streams can be played in high definition in full-screen or multi-window format. Analysis results are summarized and displayed in the form of data tables, trend charts, and spatial distribution maps (such as statistically analyzing the frequency of target occurrences within a certain time period and marking the spatial location of targets within the monitored area). Alarm information is presented through full-screen pop-ups, audio-visual alerts, and alarm logs, facilitating centralized analysis and coordinated handling by management personnel. It also supports historical data backtracking (such as playing back panoramic video and querying historical analysis results), enhancing the comprehensiveness of monitoring management.
[0259] For example, a fixed monitoring workstation is deployed in the farmland monitoring center, consisting of one host and three high-definition monitors. The left monitor plays a full-screen panoramic video stream of the farmland, the middle monitor displays real-time analysis results in a data table format (including target type, occurrence time, location coordinates, status, etc.), and the right monitor displays a spatial distribution map of the monitored area (marking the real-time locations of all targets). When the system triggers an alarm, a full-screen red alarm window immediately pops up on the middle monitor, simultaneously playing a close-up video of the alarm location. The audible and visual alarm on the control panel sounds an alert, and the workstation automatically records the alarm log (including alarm time, type, and handling records), facilitating centralized viewing, analysis, and issuance of handling instructions by management personnel.
[0260] Mobile terminals or fixed monitoring workstations can be deployed independently or in combination for collaborative operation. When deployed independently, mobile terminals are suitable for distributed inspection scenarios, while fixed monitoring workstations are suitable for centralized management and control scenarios. When deployed in combination, both can simultaneously receive and display relevant information, achieving a collaborative mode of "centralized management and control + mobile response," thereby improving the flexibility and coverage of monitoring and management.
[0261] For example, a farmland monitoring system deploys both mobile terminals and fixed monitoring workstations. The fixed workstations at the monitoring center are responsible for 24-hour uninterrupted centralized monitoring, while management personnel carry mobile terminals for field inspections. When a fixed workstation triggers an alarm, the mobile terminal receives the alarm push notification simultaneously. Management personnel can view the alarm details in real time through the mobile terminal and quickly go to the scene to handle the situation without returning to the monitoring center, achieving a seamless connection between centralized control and mobile response.
[0262] This embodiment ensures that managers can obtain monitoring information in real time during mobile inspections and field work by using mobile terminals, thereby improving the timeliness and flexibility of monitoring response. On the other hand, fixed monitoring workstations meet the needs of centralized management and overall analysis of the monitoring center, ensuring the comprehensiveness and stability of monitoring management.
[0263] To better understand the object monitoring method based on multi-source data provided in this embodiment, this embodiment also provides a specific implementation of a panoramic video monitoring system. Taking a farmland video monitoring system based on multi-lens cameras as an example, the farmland video monitoring system based on multi-lens cameras includes a multi-lens camera module, a GPS positioning module, a terrain parameter acquisition module, a feature point matching module, a video stitching module, a power management module, and an intelligent analysis module.
[0264] The multi-lens module employs a combination of two PTZ cameras and one bullet camera to acquire multiple video streams. The PTZ cameras use wide-angle lenses, while the bullet camera uses medium-telephoto lenses, achieving shooting from different perspectives through a reasonable lens combination. The multi-lens module also includes a frame synchronization management unit, a video data rectification unit, and a video data splitting unit, used to control the initial frame acquisition time and video data channel classification of multiple video data streams.
[0265] The GPS positioning module is used to acquire the device's real-time location information and associate this location information with video data for storage. The terrain parameter acquisition module collects farmland terrain parameters in real time through sensors, including information such as flatness and crop height, and associates these parameters with video data for storage.
[0266] The feature point matching module uses the SIFT algorithm to extract and match farmland features from multiple video streams, including features such as field ridges and crop row boundaries. This module combines GPS location and terrain parameters to establish a three-dimensional coordinate system for the farmland, converting the pixel coordinates of the two video streams into actual geographic coordinates, and processing the stitching edges using a weighted fusion algorithm.
[0267] The video stitching module employs a feature-point-based stitching algorithm to stitch together multiple video streams in real time. This module dynamically adjusts the field-of-view weights based on crop height and calculates the movement frequency of crops and human figures using inter-frame difference, distinguishing different types of dynamic targets, thus improving stitching success rate and reducing false alarm rate.
[0268] The power management module monitors the real-time power generation of the solar panels, the remaining power of the lithium battery, and the power consumption status of the equipment. Combined with predictions based on sunlight intensity and historical data, it automatically adjusts power utilization according to sunlight levels. The module operates normally when there is sufficient sunlight and enters emergency mode when there is insufficient or no sunlight, shutting down non-essential functions to extend battery life.
[0269] The intelligent analysis module uses a Gaussian mixture model to create a dynamic background for farmland and extracts HOG and skeleton features from foreground targets to achieve human figure recognition and localization. This module also includes a scarecrow location calibration function, reducing false alarm rates through user-defined marking. Simultaneously, the module analyzes the pixel coverage of the stitched video in real time and automatically fine-tunes the secondary camera angle when the pixel grayscale variance in a certain area is less than a preset threshold, ensuring video continuity.
[0270] In a specific embodiment, the implementation method of the farmland video surveillance system based on multi-view lenses is as follows:
[0271] The multi-lens module consists of two PTZ cameras and one bullet camera. The PTZ cameras use wide-angle lenses with a focal length of 2.7-5mm and a field of view of 90°; the bullet camera uses a medium-telephoto lens with a focal length of 8-32mm and a field of view of 20°. The two PTZ cameras are installed on both sides of the blast furnace, and the bullet camera is installed in front of the blast furnace, achieving all-around visual coverage through a reasonable combination of lenses.
[0272] The GPS positioning module uses the NEO-M8N GPS module, with a positioning accuracy of 2.5 meters. The GPS module communicates with the video data acquisition board via serial port, transmitting location information to the data processing unit in real time.
[0273] The terrain parameter acquisition module includes a lidar sensor and an ultrasonic sensor. The lidar sensor is the Velodyne HDL-32E model, with a 360° scanning range and a maximum detection distance of 120 meters. The ultrasonic sensor is the Pepperl+Fuchs UC2000-30GM-IUR2-V15 model, with a detection range of 0.2-30 meters. The sensors communicate with the data acquisition board via a CAN bus to acquire terrain parameters in real time.
[0274] The feature point matching module uses the SIFT algorithm from the OpenCV library, with a matching threshold of 0.3. During the matching process, the video stream is first converted to grayscale, then SIFT feature points are detected and matched, and finally, a weighted fusion algorithm is used to process the stitching edges, with a weighting coefficient of 0.7.
[0275] The video stitching module employs GPU parallel acceleration for video stitching. It utilizes an NVIDIA Tesla T4 GPU card with 512MB of memory, and implements parallel processing of video frames through CUDA programming. During stitching, the field of view weight is dynamically adjusted based on crop height, with wheat crops receiving a weight of 0.8 and corn crops receiving a weight of 1.2.
[0276] The power management module uses a lithium battery pack with a rated capacity of 5000mAh, and monitors the battery status in real time through a BQ27Z561-R1 fuel gauge chip. The solar panel is controlled by a MAX15843 power management chip, with a maximum output current of 2A. Based on real-time light intensity and historical data, the module predicts the following: if the light intensity is greater than 1000Lux, it maintains full functionality; if the light intensity is between 500-1000Lux, it disables the secondary lens audio function, retaining only the main camera and human detection functions; if the light intensity is less than 500Lux, it enters emergency mode, disabling non-core functions.
[0277] The intelligent analysis module uses the HOG and skeleton algorithms from the OpenCV library to extract foreground features. The HOG algorithm uses 8x8 cells with a spacing of 5 and gradient templates in 9 directions. The skeleton algorithm uses the Canny edge detection algorithm to detect contours and extracts line segments through Hough transform to construct the skeleton. This module classifies scarecrows, swaying crops, and human figures using a Gaussian mixture model, with a classification threshold of 0.5. Simultaneously, it calculates the movement frequencies of crops and human figures using inter-frame differencing, with the crop movement frequency set at 6-10Hz and the human figure movement frequency at 0.5-2Hz.
[0278] It should be understood that, although Figure 2The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0279] To facilitate better implementation of the object monitoring method based on multi-source data in the embodiments of this application, the embodiments of this application also provide an object monitoring device based on multi-source data, which is based on the above-described object monitoring method based on multi-source data. The meanings of the terms used are the same as in the above-described object monitoring method based on multi-source data, and specific implementation details can be found in the descriptions in the method embodiments.
[0280] Please see Figure 6 , Figure 6 The structural diagram of the object monitoring device based on multi-source data provided in the embodiments of this application may specifically include:
[0281] The acquisition module 201 is used to simultaneously acquire multiple video streams, environmental parameters associated with the video streams, and device location information, and to perform time alignment and data association processing on the multiple video streams, environmental parameters, and location information to form associated data;
[0282] The mapping module 202 is used to perform feature matching on multiple video streams in the associated data, and to fuse the location information and environmental parameters in the associated data to establish a mapping relationship between video feature points and a unified spatial coordinate system.
[0283] The splicing module 203 is used to splice multiple video streams in real time according to the mapping relationship to generate a panoramic video stream of the monitored area;
[0284] Analysis module 204 is used to perform dynamic scene analysis based on environmental parameters in panoramic video stream and associated data, identify target types and states in the scene, and obtain analysis results;
[0285] The control module 205 is used to generate equipment control strategies based on the analysis results, and to generate and issue control instructions for adjusting the working parameters of the multi-view image acquisition device based on the equipment control strategies.
[0286] The object monitoring device based on multi-source data provided in this embodiment synchronously collects and associates multiple video streams, environmental parameters, and device location information to establish a mapping relationship between video feature points and a unified spatial coordinate system to achieve real-time panoramic video stitching. It combines multi-source data to conduct dynamic scene analysis to identify target types and states and generate device control strategies, thereby improving the stitching accuracy, real-time performance, and target identification accuracy of panoramic monitoring and realizing dynamic adaptation and control of the device.
[0287] For specific limitations regarding the object monitoring device based on multi-source data, please refer to the limitations of the object monitoring method based on multi-source data mentioned above, which will not be repeated here. Each module in the aforementioned object monitoring device based on multi-source data can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0288] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 7 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:
[0289] The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0290] The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.
[0291] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and object monitoring methods based on multi-source data by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.
[0292] The electronic device also includes a power supply 303 that supplies power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0293] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0294] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to realize various functions, as follows:
[0295] The system synchronously acquires multiple video streams, associated environmental parameters, and device location information. It then performs time alignment and data association processing on the video streams, environmental parameters, and location information to form associated data. Feature matching is performed on the multiple video streams within the associated data, and the location information and environmental parameters are fused to establish a mapping relationship between video feature points and a unified spatial coordinate system. Based on this mapping relationship, the multiple video streams are stitched together in real time to generate a panoramic video stream of the monitored area. Dynamic scene analysis is performed based on the panoramic video stream and environmental parameters in the associated data to identify target types and states within the scene, yielding analysis results. Based on the analysis results, a device control strategy is generated, and control commands are generated and issued to adjust the operating parameters of the multi-view image acquisition device.
[0296] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0297] This application embodiment achieves real-time panoramic video stitching by synchronously collecting and associating multiple video streams, environmental parameters, and device location information, establishing a mapping relationship between video feature points and a unified spatial coordinate system, combining multi-source data to conduct dynamic scene analysis to identify target types and states, and generating device control strategies, thereby improving the stitching accuracy, real-time performance, and target identification accuracy of panoramic monitoring, and realizing dynamic adaptation and control of the device.
[0298] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0299] Therefore, embodiments of this application provide a storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the object monitoring methods based on multi-source data provided in embodiments of this application. For example, the instructions can execute the following steps:
[0300] The system synchronously acquires multiple video streams, associated environmental parameters, and device location information. It then performs time alignment and data association processing on the video streams, environmental parameters, and location information to form associated data. Feature matching is performed on the multiple video streams within the associated data, and the location information and environmental parameters are fused to establish a mapping relationship between video feature points and a unified spatial coordinate system. Based on this mapping relationship, the multiple video streams are stitched together in real time to generate a panoramic video stream of the monitored area. Dynamic scene analysis is performed based on the panoramic video stream and environmental parameters in the associated data to identify target types and states within the scene, yielding analysis results. Based on the analysis results, a device control strategy is generated, and control commands are generated and issued to adjust the operating parameters of the multi-view image acquisition device.
[0301] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0302] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0303] Since the instructions stored in the storage medium can execute the steps of any of the object monitoring methods based on multi-source data provided in the embodiments of this application, the beneficial effects that any of the object monitoring methods based on multi-source data provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0304] The foregoing has provided a detailed description of an object monitoring method, apparatus, and storage medium based on multi-source data provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for object monitoring based on multi-source data, applied to a panoramic video surveillance system, characterized in that, The panoramic video surveillance system includes multiple multi-view image acquisition devices deployed within the monitoring area, an edge computing module communicatively connected to the multi-view image acquisition devices, a cloud server communicatively connected to the edge computing module, and multiple management devices communicatively connected to the cloud server. The multiple multi-view image acquisition devices are used to acquire multiple video streams from different perspectives in the monitored area; The object monitoring method based on multi-source data includes: Simultaneously acquire multiple video streams, environmental parameters associated with the video streams, and device location information, and perform time alignment and data association processing on the multiple video streams, the environmental parameters, and the location information to form associated data; Feature matching is performed on multiple video streams in the associated data, and the location information and environmental parameters in the associated data are fused to establish a mapping relationship between video feature points and a unified spatial coordinate system; The multiple video streams are stitched together in real time according to the mapping relationship to generate a panoramic video stream of the monitored area. Based on the panoramic video stream, a dynamic background model of the monitored area is established; the dynamic background model is used to extract foreground targets, and the morphological and motion characteristics of the foreground targets are analyzed. The morphological features and motion features of the foreground target are compared with the standard feature ranges of multiple predefined target types to generate comparison results. Based on the comparison results, combined with the auxiliary information of environmental parameters in the associated data for feature judgment, the classification probability of the foreground target belonging to each category is calculated. It is determined whether the classification probability exceeds the classification threshold of the corresponding category to determine the final classification result of the foreground target. The key parameters in the motion features are used as the real-time state of the foreground target to form the analysis result. Based on the analysis results, a device control strategy is generated, and control instructions for adjusting the working parameters of the multi-view image acquisition device are generated and issued based on the device control strategy.
2. The object monitoring method based on multi-source data according to claim 1, characterized in that, The process involves synchronously acquiring multiple video streams, environmental parameters associated with the video streams, and device location information. Time alignment and data association processing are then performed on the multiple video streams, the environmental parameters, and the location information to form associated data, including: The multi-channel video streams with synchronized timestamps are obtained from the multi-view image acquisition device; The environmental parameters and device location information collected by the edge computing module are acquired simultaneously. The multiple video streams, environmental parameters, and device location information under the same timestamp are bound together to form the associated data.
3. The object monitoring method based on multi-source data according to claim 1, characterized in that, The step of performing feature matching on multiple video streams in the associated data, and fusing the location information and environmental parameters in the associated data to establish a mapping relationship between video feature points and a unified spatial coordinate system includes: Extract image feature points from each video stream in the associated data; The image feature points extracted from different video streams are matched to obtain pairs of feature points with the same name; Based on the device location information and environmental parameters in the associated data, a three-dimensional spatial model of the monitoring area is constructed. The image coordinates of the corresponding feature point pairs are mapped to the three-dimensional spatial model to establish the mapping relationship.
4. The object monitoring method based on multi-source data according to claim 3, characterized in that, The step of constructing a three-dimensional spatial model of the monitoring area based on the device location information and environmental parameters in the associated data includes: Using the device location information as a spatial reference, a preliminary three-dimensional geographic reference framework is established; Based on the terrain undulation information in the environmental parameters, the elevation dimension of the preliminary three-dimensional geographic reference frame is corrected to obtain the corrected three-dimensional geographic reference frame. The object height distribution information in the environmental parameters is incorporated into the corrected three-dimensional geographic reference frame as a structured constraint in the vertical space to complete the construction of the three-dimensional spatial model.
5. The object monitoring method based on multi-source data according to claim 4, characterized in that, The establishment of a preliminary three-dimensional geographic reference framework based on the device location information includes: The latitude and longitude coordinates and altitude of at least one of the multi-view image acquisition devices are parsed from the device location information; Using the points defined by the analyzed latitude and longitude coordinates and altitude as the spatial origin, a three-dimensional rectangular coordinate system is established as the preliminary three-dimensional geographic reference frame.
6. The object monitoring method based on multi-source data according to claim 4, characterized in that, The step of correcting the elevation dimension of the preliminary three-dimensional geographic reference frame based on the terrain undulation information in the environmental parameters to obtain the corrected three-dimensional geographic reference frame includes: Obtain the terrain undulation information used to characterize ground flatness from the environmental parameters; Based on the terrain undulation information, a digital elevation model of the monitored area is generated; The initial elevation plane in the preliminary three-dimensional geographic reference frame is replaced by the digital elevation model to complete the correction of the elevation dimension.
7. The object monitoring method based on multi-source data according to claim 4, characterized in that, The step of incorporating the object height distribution information from the environmental parameters as a structured constraint in the vertical space into the corrected 3D geographic reference frame to complete the construction of the 3D spatial model includes: Obtain the object height distribution information representing the height of the monitored object from the environmental parameters; Within the corrected three-dimensional geographic reference frame, based on the object height distribution information, a corresponding height value constraint is assigned to the horizontal area where the monitored object is located along the vertical direction; Based on the height constraint, volumetric elements representing the three-dimensional distribution of the monitored objects are constructed in the three-dimensional spatial model.
8. The object monitoring method based on multi-source data according to claim 1, characterized in that, The step of stitching the multiple video streams in real time according to the mapping relationship to generate a panoramic video stream of the monitored area includes: Based on the mapping relationship, each frame of the multi-channel video stream is transformed to the unified spatial coordinate system; The overlapping regions of each frame image transformed to the unified spatial coordinate system are fused to generate a continuous panoramic video frame sequence, which constitutes the panoramic video stream.
9. The object monitoring method based on multi-source data according to claim 8, characterized in that, The step of transforming each frame image in the multi-channel video stream to the unified spatial coordinate system according to the mapping relationship includes: Based on the mapping relationship, determine the geometric transformation parameters required to map points in each video stream image coordinate system to the unified spatial coordinate system; Based on the geometric transformation parameters, perspective transformation or affine transformation is performed on each frame of each video stream to obtain the transformed image; The transformed image is resampled to eliminate pixel holes or overlaps caused by the transformation, resulting in an image frame aligned in the unified spatial coordinate system.
10. The object monitoring method based on multi-source data according to claim 9, characterized in that, The step of performing perspective transformation or affine transformation on each frame of each video stream according to the geometric transformation parameters to obtain the transformed image includes: For the current image frame to be transformed, the mapping relationship between the source image pixels and the target coordinate system pixels is calculated according to the corresponding geometric transformation parameters. Based on the mapping relationship, all pixels of the current image frame are traversed, and each pixel is transformed from its original image coordinate position to a coordinate position defined by the unified spatial coordinate system.
11. The object monitoring method based on multi-source data according to claim 9, characterized in that, The step of determining the geometric transformation parameters required to map points in each video stream image coordinate system to the unified spatial coordinate system based on the mapping relationship includes: Based on the mapping relationship, at least four sets of corresponding feature point coordinate pairs are obtained between the unified spatial coordinate system and the image coordinate systems of each video stream; Based on the feature point coordinate pairs, the coefficients of the perspective transformation matrix or the affine transformation matrix are solved by the least squares method, and the coefficients are used as the geometric transformation parameters.
12. The object monitoring method based on multi-source data according to claim 8, characterized in that, The step of fusing overlapping regions of images transformed to the unified spatial coordinate system to generate a continuous panoramic video frame sequence includes: Determine the overlapping region between adjacent image frames in the unified spatial coordinate system; For each pixel position within the overlapping region, the fusion weight of pixel values from different source image frames is calculated based on the distance from each pixel position to the center of each source image frame or a predefined weight function. Based on the fusion weights, the corresponding pixel values from different source image frames are weighted and averaged to generate fused pixel values, forming continuous panoramic video frames.
13. An object monitoring device based on multi-source data, characterized in that, include: The acquisition module is used to simultaneously acquire multiple video streams, environmental parameters associated with the video streams, and device location information, and to perform time alignment and data association processing on the multiple video streams, the environmental parameters, and the location information to form associated data; The mapping module is used to perform feature matching on multiple video streams in the associated data, and to fuse the location information and environmental parameters in the associated data to establish a mapping relationship between video feature points and a unified spatial coordinate system. The stitching module is used to stitch the multiple video streams in real time according to the mapping relationship to generate a panoramic video stream of the monitored area; The analysis module is used to establish a dynamic background model of the monitored area based on the panoramic video stream; extract foreground targets using the dynamic background model; and analyze the morphological and motion characteristics of the foreground targets. The morphological features and motion features of the foreground target are compared with the standard feature ranges of multiple predefined target types to generate comparison results. Based on the comparison results, combined with the auxiliary information of environmental parameters in the associated data for feature judgment, the classification probability of the foreground target belonging to each category is calculated. It is determined whether the classification probability exceeds the classification threshold of the corresponding category to determine the final classification result of the foreground target. The key parameters in the motion features are used as the real-time state of the foreground target to form the analysis result. The control module is used to generate a device control strategy based on the analysis results, and to generate and issue control instructions for adjusting the working parameters of the multi-view image acquisition device based on the device control strategy.
14. A storage medium, characterized in that, The computer program is stored and can be loaded by a processor and executed as described in any one of claims 1-12.
Citation Information
Patent Citations
Video stitching method and system for multi-camera monitoring
CN118678240A