Abnormity detection method and device based on video fusion, electronic equipment and medium

By performing anomaly detection and fusion on the video streams of the inspection site, generating anomaly reports and displaying them in the 3D virtual model, the problems of low efficiency, low accuracy and poor real-time performance in anomaly detection at the inspection site are solved, and efficient and accurate anomaly handling and location are achieved.

CN121767904APending Publication Date: 2026-03-31NUCTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies for detecting anomalies in inspection sites suffer from low efficiency, low accuracy, low real-time performance, and poor intuitiveness. They rely on manual inspections and suffer from poor communication, resulting in cumbersome and inefficient problem-solving processes.

Method used

By detecting abnormal events in the video stream of the inspection site, generating anomaly reports, and overlaying the panoramic video stream onto a 3D virtual model, the video stream is fused and displayed. Combined with wearable devices and remote assistance, the detection accuracy and processing efficiency are improved.

Benefits of technology

It achieves stable and objective real-time anomaly event identification, reduces false alarm and missed detection rates, improves detection accuracy and processing efficiency, provides an intuitive virtual-real fusion visual environment, and simplifies the anomaly handling process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767904A_ABST
    Figure CN121767904A_ABST
Patent Text Reader

Abstract

The invention provides an anomaly detection method and device based on video fusion, electronic equipment and a medium, and can be applied to the technical field of video processing and digital twinning. The method comprises the steps that abnormal event detection is conducted on at least one video stream, a detection result is obtained, the video stream is obtained through at least one image collection device arranged in an inspection site, and the image collection device corresponds to a virtual image collection device in a three-dimensional virtual model for the inspection site; under the condition that the target video stream with the target abnormal event exists, generating an abnormal report according to image acquisition information of a target image acquisition device for acquiring the target video stream and the target video stream; performing image fusion on the target video stream and an adjacent video stream having an overlapped acquisition area with the target video stream to obtain a panoramic video stream; and superposing the panoramic video stream to a collection area corresponding to the panoramic video stream in the three-dimensional virtual model, and displaying the three-dimensional virtual model after superposing the panoramic video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of video processing and digital twin technology, and more specifically, to an anomaly detection method, apparatus, electronic device, and medium based on video fusion. Background Technology

[0002] Currently, to ensure the daily safety of inspection sites, the commonly used method for anomaly detection is to deploy manual inspections. During each inspection, staff must strictly follow the established inspection plan, perform inspection tasks according to the designated route and standards, and immediately report and take preliminary emergency measures if any anomalies or potential hazards are discovered.

[0003] However, in practice, the success rate of manual inspections in detecting anomalies relies heavily on the personal experience of the inspectors and is easily affected by various external factors, leading to inconsistent anomaly detection quality. Furthermore, after an anomaly is discovered, staff must contact relevant maintenance or handling personnel. On the one hand, in some key areas, communication difficulties or untimely personnel deployment often make it difficult to contact the responsible party promptly, resulting in cumbersome and inefficient problem-solving processes. On the other hand, anomaly times are usually recorded only in written form, and maintenance personnel often struggle to accurately grasp the essence of the problem based solely on written descriptions, thus affecting the accuracy of their judgment and handling.

[0004] Therefore, the relevant technologies suffer from problems such as low efficiency, low accuracy, low real-time performance, and poor intuitiveness in detecting anomalies in inspection sites. Summary of the Invention

[0005] In view of this, the present disclosure provides an anomaly detection method, apparatus, electronic device and medium based on video fusion.

[0006] One aspect of this disclosure provides an anomaly detection method based on video fusion, comprising: performing anomaly event detection on at least one video stream to obtain a detection result, wherein the video stream is obtained using at least one image acquisition device installed at an inspection site, the image acquisition device corresponding to a virtual image acquisition device within a three-dimensional virtual model of the inspection site; if the detection result indicates that a target video stream in at least one video stream has a target anomaly event, generating an anomaly report for the target anomaly event based on image acquisition information from the target image acquisition device used to acquire the target video stream and the target video stream; performing image fusion on the target video stream and at least one adjacent video stream having an overlapping acquisition area with the target video stream to obtain a panoramic video stream; superimposing the panoramic video stream onto the acquisition area corresponding to the panoramic video stream in the three-dimensional virtual model, and displaying the three-dimensional virtual model after superimposing the panoramic video stream.

[0007] According to embodiments of this disclosure, generating an anomaly report for a target anomaly event based on image acquisition information from a target image acquisition device used to acquire a target video stream and the target video stream includes: determining the anomaly location of the target anomaly event in a three-dimensional virtual model based on the image acquisition information from the target image acquisition device used to acquire the target video stream; determining an anomaly time period based on at least one video frame in the target video stream where the target anomaly event exists, and extracting a video segment of the anomaly time period and a key video frame for identifying the target anomaly event from the target video stream; and generating an anomaly report for the target anomaly event based on the anomaly location, the anomaly time period, the video segment, the key video frame, and the anomaly type of the target anomaly event.

[0008] According to embodiments of this disclosure, determining the abnormal location of a target abnormal event in a three-dimensional virtual model based on image acquisition information of a target image acquisition device used to acquire a target video stream includes: determining the virtual position of the target image acquisition device used to acquire the target video stream in the three-dimensional virtual model based on the correspondence between the image acquisition device and the virtual image acquisition device; and determining the abnormal location of the target abnormal event in the three-dimensional virtual model based on the image acquisition information and the virtual position of the target image acquisition device, wherein the image acquisition information includes the acquisition orientation and acquisition field of view of the image acquisition device.

[0009] According to embodiments of this disclosure, image fusion is performed on a target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream, including: image distortion correction of the target video stream and at least one adjacent video stream to obtain a corrected target video stream and a corrected adjacent video stream; and stitching the corrected target video stream and the corrected adjacent video stream together to obtain a panoramic video stream.

[0010] According to embodiments of this disclosure, adjacent video streams are determined by the following method: determining at least one adjacent image acquisition device that has an overlapping acquisition area with the target image acquisition device based on the image acquisition information of at least one image acquisition device; and determining the video streams acquired by at least one adjacent image acquisition device as adjacent video streams.

[0011] According to embodiments of this disclosure, a panoramic video stream is obtained by stitching together a corrected target video stream and corrected adjacent video streams. This includes: determining multiple video frames to be stitched at the same time from the corrected target video stream and corrected adjacent video streams; extracting features from each of the multiple video frames to be stitched to obtain feature points for each of the multiple video frames to be stitched; matching feature points between the multiple video frames to be stitched to obtain multiple feature point sets, wherein the multiple feature points in each feature point set are feature points of the same object within an overlapping acquisition area in different video frames to be stitched; stitching the multiple video frames to be stitched together according to the multiple feature point sets to obtain panoramic video frames corresponding to a given time; and combining the panoramic video frames at multiple times according to the chronological order of the panoramic video frames at multiple times to obtain a panoramic video stream.

[0012] According to embodiments of this disclosure, displaying a three-dimensional virtual model overlaid with a panoramic video stream includes: displaying the three-dimensional virtual model overlaid with the panoramic video stream when the anomaly type of the target anomaly event is a first priority; displaying alarm information indicating the target anomaly event when the anomaly type of the target anomaly event is a second priority; and displaying the three-dimensional virtual model overlaid with the panoramic video stream in response to a triggering operation on the alarm information, wherein the first priority is greater than the second priority.

[0013] According to embodiments of this disclosure, the method further includes: acquiring and displaying multi-source data collected by the wearable device when the user wearing the wearable device is in a blind spot area within the inspection site; wherein the blind spot area is an area not captured by at least one image acquisition device, the multi-source data includes at least a video stream collected by the wearable device, and the wearable device is also used to display a three-dimensional virtual model and a location marker indicating the user's current location.

[0014] According to embodiments of this disclosure, the method further includes: in response to a remote assistance request triggered by a user through a wearable device, establishing a communication connection with a target terminal device corresponding to the remote assistance request when the user wearing the wearable device is in a blind spot area within the inspection site; sending multi-source data collected by the wearable device to the target terminal device through the communication connection, so as to display the multi-source data on the target terminal device; and sending inspection guidance data obtained from the target terminal device to the wearable device, wherein the inspection guidance data includes at least one of screen data and audio data used to guide the user to perform inspection operations.

[0015] According to embodiments of this disclosure, the method further includes: acquiring multi-source modeling data for the inspection site, wherein the multi-source modeling data is obtained using various types of modeling data acquisition devices; constructing a three-dimensional virtual model using the multi-source modeling data; and constructing a virtual image acquisition device corresponding to the image acquisition device in the three-dimensional virtual model based on the physical location of at least one image acquisition device set at the inspection site.

[0016] Another aspect of this disclosure provides an anomaly detection device based on video fusion, comprising: a detection module for detecting anomalies in at least one video stream and obtaining a detection result, wherein the video stream is obtained using at least one image acquisition device installed at an inspection site, and the image acquisition device corresponds to a virtual image acquisition device within a three-dimensional virtual model of the inspection site; a generation module for generating an anomaly report for the target anomaly event based on image acquisition information from the target image acquisition device used to acquire the target video stream and the target video stream, provided that the detection result indicates the presence of a target video stream in at least one video stream where a target anomaly event has occurred; a fusion module for performing image fusion on the target video stream and at least one adjacent video stream with an overlapping acquisition area with the target video stream to obtain a panoramic video stream; and a display module for overlaying the panoramic video stream onto the acquisition area corresponding to the panoramic video stream in the three-dimensional virtual model and displaying the three-dimensional virtual model after overlaying the panoramic video stream.

[0017] Another aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the methods described above.

[0018] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the methods described above.

[0019] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the methods described above.

[0020] In the embodiments of this disclosure, automated anomaly detection is performed using at least one video stream from the inspection site. This enables real-time analysis of the video stream with stable and objective standards, quickly and accurately identifying target anomalies, reducing false alarm and missed alarm rates, and significantly improving the detection accuracy and processing efficiency of target anomalies. After detecting a target anomaly, an anomaly report is automatically generated, avoiding reliability and communication issues caused by manual anomaly recording, facilitating subsequent accurate anomaly handling. Furthermore, the embodiments of this disclosure fuse the real target video stream with adjacent video streams to obtain a realistic panoramic video stream that facilitates observation of the target anomaly and its surrounding environment. This real panoramic video stream is then fused with a 3D virtual model and displayed, allowing users to intuitively and accurately understand the target anomaly itself and its surrounding environment, as well as the target anomaly's location within the entire inspection site, from the fused virtual and real 3D virtual model. This improves anomaly location time and efficiency, resulting in a better user experience. Attached Figure Description

[0021] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0022] Figure 1 An exemplary system architecture for applying a video fusion-based anomaly detection method according to embodiments of this disclosure is illustrated.

[0023] Figure 2 A flowchart illustrating an anomaly detection method based on video fusion according to an embodiment of the present disclosure is shown schematically.

[0024] Figure 3 The illustration schematically depicts a scene diagram of a three-dimensional virtual model of an overlaid panoramic video stream according to an embodiment of the present disclosure.

[0025] Figure 4 A flowchart illustrating the generation of an anomaly report according to an embodiment of this disclosure is shown schematically.

[0026] Figure 5 A flowchart illustrating a panoramic video stream obtained by stitching according to an embodiment of the present disclosure is shown.

[0027] Figure 6 A block diagram of a video fusion-based anomaly detection apparatus according to an embodiment of the present disclosure is shown schematically.

[0028] Figure 7 A block diagram of an electronic device suitable for implementing a video fusion-based anomaly detection method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0029] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0030] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0032] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0033] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0034] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.

[0035] In related technologies, manual inspection has several problems, mainly including the following: 1. Low efficiency and high cost: Manual inspection relies on personnel checking point by point, resulting in limited coverage and long time consumption, making it impossible to achieve high-frequency, large-scale, rapid inspection. Long-term investment of a large amount of manpower in repetitive labor leads to high labor and management costs. 2. Reliance on personal experience and unstable quality: Inspection results are easily affected by the fatigue and work status of the inspectors, leading to missed or incorrect inspections due to negligence or lack of expertise. Furthermore, the judgment standards of different inspectors are difficult to standardize, and their understanding and recording of "abnormalities" may be inconsistent, affecting the review and reliability of subsequent inspection data. 3. Inaccurate information recording and vague problem descriptions: Abnormal situations often rely on written descriptions from inspectors, making it difficult for maintenance personnel to accurately understand the actual situation and easily leading to misjudgments. 4. High safety risks: When involving high-risk areas (such as high temperature, high pressure, or toxic environments), manual inspection poses a potential threat to personnel safety. 5. Lack of data integration capabilities: Inspection data is isolated, making it difficult to conduct systematic statistical analysis and providing data support for predictive maintenance and management decisions. 6. Delayed response and lengthy processing: When abnormal events are discovered, they need to be reported at each level and coordinated with maintenance personnel for real-time handling, especially at night or in remote areas, resulting in long waiting times and high communication costs.

[0036] To at least partially address the problems of low efficiency, low accuracy, low real-time performance, and poor intuitiveness in anomaly detection in related technologies, embodiments of this disclosure provide an anomaly detection method based on video fusion, comprising: performing anomaly event detection on at least one video stream to obtain a detection result, wherein the video stream is obtained using at least one image acquisition device set up at the inspection site, and the image acquisition device corresponds to a virtual image acquisition device within a three-dimensional virtual model of the inspection site; if the detection result indicates that a target video stream in at least one video stream has a target anomaly event, generating an anomaly report for the target anomaly event based on the image acquisition information of the target image acquisition device used to acquire the target video stream and the target video stream; performing image fusion on the target video stream and at least one adjacent video stream with an overlapping acquisition area with the target video stream to obtain a panoramic video stream; superimposing the panoramic video stream onto the acquisition area corresponding to the panoramic video stream in the three-dimensional virtual model, and displaying the three-dimensional virtual model after superimposing the panoramic video stream.

[0037] Figure 1 This illustration schematically depicts an exemplary system architecture to which a video fusion-based anomaly detection method can be applied according to embodiments of this disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0038] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, an image acquisition device 104, a server 105, and a network 106. The network 106 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0039] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 106 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).

[0040] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0041] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0042] It should be noted that the anomaly detection method based on video fusion provided in this disclosure can generally be executed by server 105. Correspondingly, the anomaly detection device based on video fusion provided in this disclosure can generally be located in server 105. The anomaly detection method based on video fusion provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the anomaly detection device based on video fusion provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, the anomaly detection method based on video fusion provided in this disclosure can also be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103. Accordingly, the anomaly detection device based on video fusion provided in this embodiment can also be installed in the first terminal device 101, the second terminal device 102, and the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0043] For example, server 105 can acquire at least one video stream obtained using at least one image acquisition device 104 set up at the inspection site. Server 105 may pre-store a 3D virtual model of the inspection site, and the image acquisition device 104 set up at the inspection site corresponds to a virtual image acquisition device within the 3D virtual model of the inspection site. Server 105 can perform anomaly detection on at least one video stream in real time and obtain detection results. If the detection results indicate the presence of a target video stream in at least one video stream where a target anomaly event has occurred, an anomaly report for the target anomaly event is generated based on the image acquisition information of the target image acquisition device used to acquire the target video stream and the target video stream; image fusion is performed on the target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream; the panoramic video stream is then superimposed onto the acquisition area corresponding to the panoramic video stream in the 3D virtual model. Afterwards, server 105 can send the 3D virtual model after overlaying the panoramic video stream to any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103, so that inspection personnel can intuitively view the 3D virtual model after overlaying the panoramic video stream through the terminal device, and see in real time the panoramic video stream that shows the presence of target anomalies, has a wider field of view, and has an intuitive positional correlation with the 3D virtual model.

[0044] Alternatively, any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 can store the 3D virtual model and acquire at least one video stream in real time. By implementing the above-described anomaly detection method based on video fusion, a 3D virtual model superimposed with the panoramic video stream is displayed to the user.

[0045] It should be understood that Figure 1 The number of terminal devices, networks, image acquisition devices, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0046] Figure 2 A flowchart illustrating an anomaly detection method based on video fusion according to an embodiment of this disclosure is shown schematically. Figure 2 As shown, the method includes operations S210~S240.

[0047] In operation S210, abnormal event detection is performed on at least one video stream, and the detection results are obtained.

[0048] According to embodiments of this disclosure, the video stream is obtained using at least one image acquisition device located at the inspection site, the image acquisition device corresponding to a virtual image acquisition device within a three-dimensional virtual model of the inspection site.

[0049] Inspection sites can be locations that perform security checks, or locations inside and outside buildings. The image acquisition devices installed at the inspection sites can be various types of cameras, such as wide-angle cameras, fisheye cameras, and other cameras with a wide field of view, or telephoto cameras, standard cameras, etc.

[0050] A 3D virtual model can be constructed using multi-source modeling data for an inspection site. This 3D virtual model can be regarded as a model that realistically reproduces various objects in the inspection site, and the real textures and fine structures of various objects can be synchronously reproduced. For example, the 3D virtual model can reproduce the appearance, shape, and spatial topology of key facilities such as equipment, pipelines, buildings, and roads in the inspection site; it can also reproduce the overall environment such as the sky, space, and ground, forming an integrated "air-sky-ground" three-dimensional inspection site environment.

[0051] For example, a 3D virtual model may include a virtual image acquisition device. The spatial location of this virtual image acquisition device in the 3D virtual model is the same as the location of the physical image acquisition device within the inspection site. In other words, the physical image acquisition device corresponds to a virtual image acquisition device within the 3D virtual model of the inspection site. Through this correspondence, a one-to-one mapping exists between the physical image acquisition device and the virtual image acquisition device in the virtual space, facilitating the subsequent fusion and analysis of the video stream acquired by the physical image acquisition device with the 3D virtual model.

[0052] In one specific embodiment, anomaly detection can be performed on at least one video stream using a deep learning algorithm to obtain anomaly detection results. For example, anomaly detection can be performed on each video stream separately using a deep learning algorithm to obtain anomaly detection sub-results for each individual video stream, and then the anomaly detection sub-results can be combined to determine the final anomaly detection result. Alternatively, to meet the daily demand for wide-angle video viewing, at least one video stream can be first merged into a wide-angle video stream (including panoramic video stream), and then anomaly detection can be performed on the merged wide-angle video stream to determine the anomaly detection result.

[0053] For example, by utilizing various events that may occur at inspection sites, a model based on deep learning algorithms can be trained. This trained deep learning model can then accurately identify whether abnormal events exist within the input video stream. The deep learning model can be based on object detection algorithms (such as YOLO), convolutional neural networks (CNNs), or encoder-decoder models (Transformers), etc.

[0054] In operation S220, if the detection result indicates that there is a target video stream in at least one video stream where a target abnormal event has occurred, an abnormal report for the target abnormal event is generated based on the image acquisition information of the target image acquisition device used to acquire the target video stream and the target video stream.

[0055] The target anomaly is predetermined, and the target anomaly can differ across various inspection locations. For example, the target anomaly could be fireworks, not wearing a safety helmet, or exceeding area restrictions. Furthermore, the video stream can be acquired in real-time, and anomaly detection can be achieved by sampling multiple video frames from the video stream. This enables long-term, real-time, and efficient anomaly detection. The target video stream is the video stream containing the target anomaly, and the target image acquisition device is the image acquisition device used to acquire the target video stream.

[0056] Image acquisition information includes information used to correlate the position of the image acquisition device with the 3D virtual model. For example, image acquisition information may include the physical location of the image acquisition device, the acquisition orientation, and the acquisition field of view.

[0057] Since the detection of abnormal events in at least one video stream is performed in real time, in order to ensure that the detected target abnormal events are reported to maintenance personnel or handling personnel, an abnormal report for the target abnormal event can be generated based on the target video stream containing the target abnormal event and the image acquisition information of the target image acquisition device.

[0058] Anomaly reports can be automatically generated after a target anomaly event is detected. Anomaly reports can include structured text information, video, or image modalities, allowing maintenance or handling personnel to intuitively and quickly understand the target anomaly event upon receiving the report in real time.

[0059] In operation S230, image fusion is performed on the target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream.

[0060] Understandably, to ensure the daily safety of the inspection site, the video streams acquired by the image acquisition device can be displayed to relevant personnel in real time. In order to ensure the prominence of the target abnormal event among multiple video streams and the integrity of the surrounding environment of the target abnormal event, the target video stream including the target abnormal event can be fused with adjacent video streams to obtain a panoramic video stream, so as to display the panoramic video stream to relevant personnel.

[0061] In the embodiments of this disclosure, the acquisition areas of multiple image acquisition devices may overlap to ensure comprehensive detection of the entire inspection site. Therefore, overlapping acquisition areas may exist between the video streams acquired by the multiple image acquisition devices. To ensure the integrity of the environment surrounding the target anomaly, image fusion can be performed on the target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream.

[0062] When operating S240, the panoramic video stream is superimposed onto the acquisition area corresponding to the panoramic video stream in the 3D virtual model, and the 3D virtual model after superimposing the panoramic video stream is displayed.

[0063] While panoramic video streams can display the environment surrounding a target anomaly, they struggle to accurately represent the anomaly's overall location within the inspection area. Therefore, panoramic video streams can be overlaid onto the corresponding acquisition area within a 3D virtual model, creating a unified visual environment that blends the real and virtual worlds. After displaying the overlaid panoramic video stream to relevant personnel, they can intuitively and accurately understand the target anomaly and its surroundings through the realistic panoramic video stream. Furthermore, they can quickly and easily locate the target anomaly by observing its position within the 3D virtual model, eliminating the need for manual location checks of the image acquisition device.

[0064] In the embodiments of this disclosure, automated anomaly detection is performed using at least one video stream from the inspection site. This enables real-time analysis of the video stream with stable and objective standards, quickly and accurately identifying target anomalies, reducing false alarm and missed alarm rates, and improving the detection accuracy and processing efficiency of target anomalies. Furthermore, automated video stream detection can be automated 24 / 7, improving inspection efficiency. After detecting a target anomaly, anomaly reports are automatically generated, avoiding reliability and communication issues caused by manual anomaly recording, facilitating subsequent accurate anomaly handling. In addition, embodiments of this disclosure fuse the real target video stream with adjacent video streams to obtain a realistic panoramic video stream that facilitates observation of the target anomaly and its surrounding environment. This real panoramic video stream is then fused with a 3D virtual model and displayed, allowing users to intuitively and accurately understand the target anomaly itself and its surrounding environment, as well as the location of the target anomaly within the entire inspection site, achieving the effect of "knowing where the incident happened at a glance." This improves anomaly location time and efficiency, resulting in a better user experience.

[0065] Figure 3 The illustration schematically depicts a scene diagram of a three-dimensional virtual model of an overlaid panoramic video stream according to an embodiment of the present disclosure. Figure 3 The background is a 3D virtual model of the inspection site. Each room on each floor of the inspection site can be equipped with an image acquisition device, and correspondingly, a virtual image acquisition device is set in the 3D virtual model. The boxed areas can overlay the video stream from the image acquisition device in that room; or, overlay a panoramic video stream fused from multiple video streams in that room. Therefore, when displaying the 3D virtual model with the overlaid panoramic video stream to the user, the user can see the actual target anomaly and intuitively determine the location of the target anomaly within the entire inspection site.

[0066] According to embodiments of this disclosure, the method further includes: acquiring multi-source modeling data for the inspection site, wherein the multi-source modeling data is obtained using various types of modeling data acquisition devices; constructing a three-dimensional virtual model using the multi-source modeling data; and constructing a virtual image acquisition device corresponding to the image acquisition device in the three-dimensional virtual model based on the physical location of at least one image acquisition device set at the inspection site.

[0067] For example, multi-source modeling data may include: first modeling data of image or video modalities collected using drones, second modeling data of point cloud modalities determined using laser devices such as lidar, and third modeling data (such as building outlines, road centerlines, terrain data, etc.) collected using surveying devices such as levels. In one specific embodiment, the aforementioned first modeling data can be collected by aerial surveying methods such as oblique photogrammetry.

[0068] The aforementioned multi-source modeling data can be pre-stored in a database, and when constructing a 3D virtual model, it can be directly retrieved from the database and used for construction.

[0069] When constructing a 3D virtual model, the first, second, and third modeling data can be transformed using coordinates or corresponding coordinate transformation relationships can be established to ensure that the positional data within these three types of modeling data can be determined in a unified coordinate system. Next, the three types of modeling data can be cleaned, such as through noise reduction, image enhancement, and format conversion. Finally, the positional data within these three types of modeling data are registered and point cloud fused in a unified coordinate system to obtain a high-resolution 3D virtual model that integrates the three types of modeling data.

[0070] In one specific embodiment, the 3D virtual model can be imported into application software capable of associating geographical location with various types of information, forming a "digital twin" foundation. For example, the 3D virtual model can be imported into a Geographic Information System (GIS), enabling the 3D virtual model to establish associations with various sensors and image acquisition devices within the inspection site.

[0071] Since the location of the image acquisition device used to acquire video streams is relatively flexible in the inspection site (compared to buildings, fixed facilities, etc.), and its position can be flexibly increased, decreased, and changed according to actual needs, a three-dimensional virtual model can be constructed first based on the above multi-source modeling data, and then a virtual image acquisition device corresponding to the image acquisition device can be constructed on this basis.

[0072] For example, a virtual image acquisition device corresponding to the image acquisition device can be constructed in a three-dimensional virtual model based on the physical location of at least one image acquisition device set up at the inspection site.

[0073] It is understandable that for multiple sensors used to collect data in the inspection site, corresponding virtual sensors can also be constructed in the three-dimensional virtual model in the same way as image acquisition devices. The virtual sensors are synchronized with the state of the sensors. Thus, the spatial registration and attribute management of real data (such as video streams) can be realized through the three-dimensional virtual model.

[0074] In the embodiments of this disclosure, multi-source modeling data collected using various types of modeling data acquisition devices can be used to construct a high-precision real-scene 3D virtual model of the inspection site. This high-precision 3D virtual model is then imported into a GIS platform to form a digital twin foundation, creating a unified visual environment that blends the virtual and real worlds. This allows users to intuitively and quickly locate abnormal events by observing the position of the real panoramic video stream within the 3D virtual model.

[0075] Figure 4 A flowchart illustrating the generation of an anomaly report according to an embodiment of this disclosure is shown schematically. Figure 4 As shown, the method includes operations S410~S430.

[0076] In operation S410, the abnormal location of the target abnormal event in the three-dimensional virtual model is determined based on the image acquisition information of the target image acquisition device used to acquire the target video stream.

[0077] In operation S420, based on at least one video frame in the target video stream where a target abnormal event exists, an abnormal time period is determined, and a video segment of the abnormal time period and a key video frame used to identify the target abnormal event are extracted from the target video stream.

[0078] When operating S430, an anomaly report is generated based on the anomaly location, anomaly time period, video segment, key video frame, and anomaly type of the target anomaly event.

[0079] In the embodiments of this disclosure, if there is a target abnormal event in the target video stream acquired by the target image acquisition device, the acquisition area of ​​the target video stream can be aligned with the virtual area of ​​the three-dimensional virtual model based on the image acquisition information of the target image acquisition device; then, the abnormal position of the target abnormal event in the three-dimensional virtual model can be determined by the position of the target abnormal event in the target video stream.

[0080] For a target video stream containing a target anomaly, the detection results can characterize which video frames identified as having the target anomaly. Therefore, at least one video frame in the target video stream containing the target anomaly can be directly determined based on the detection results. Then, based on the timestamps of each of the at least one video frame in the target video stream, the abnormal time period in which the target anomaly occurred is determined.

[0081] For example, after identifying a target anomalous event from a target video stream, the timestamp of the first video frame in which the anomalous event was identified can be used as the start of the anomalous time period. If used for real-time anomalous event detection and notification, the time period including the preset duration before and after the timestamp of the first video frame is used as the basis for the anomalous time period. If it is necessary to detect the complete target anomalous event, the target time period consisting of multiple video frames in which the target anomalous event was identified can be determined, and the entire time period including the preset duration before and after the timestamp and the target time period is used as the basis for the anomalous time period.

[0082] To facilitate subsequent playback of the target anomaly event, video segments of the abnormal time period can be extracted from the target video stream, and key video frames that can identify the target anomaly event can be determined from these video segments. For example, the key video frame can be the first video frame of the video segment, or it can be a video frame within the video segment that clearly displays the target anomaly event. In the embodiments of this disclosure, techniques such as multimodal large modeling can be used to determine key video frames that clearly display the target anomaly event from the video segment based on the semantics of each video frame within the video segment.

[0083] For a target anomalous event, a multimodal anomaly report in a predefined format can be generated. This includes structured information such as the location of the anomaly, the time period of the anomaly, and the anomaly type of the target anomalous event identified by the anomaly event detection. In addition, the anomaly report also includes video clips in the video modality and key video frames in the image modality.

[0084] In one embodiment, the generated anomaly report can be stored in a central database for subsequent querying and statistical analysis based on structured location, time period, and anomaly type, facilitating the backtracking of the target anomaly event. Simultaneously, by querying the anomaly report of the target anomaly event, one can visually view the video segment in which the target anomaly event occurred, and intuitively identify the target anomaly event through key video frames between unplayed video segments.

[0085] In another embodiment, after detecting a target anomaly and generating an anomaly report, an alarm process can also be automatically triggered. For example, alarm information can be generated based on the anomaly report and transmitted to the terminal devices of relevant personnel, so that the relevant personnel can quickly and comprehensively understand the target anomaly directly based on the alarm information.

[0086] In another embodiment, while sending the alarm, a quick jump method such as a link can also be sent so that relevant personnel can quickly view the 3D virtual model after the panoramic video stream is overlaid, as well as the above-mentioned anomaly report, by operating the link.

[0087] In the embodiments of this disclosure, deep learning algorithms are used to perform real-time analysis of video streams, identify preset target abnormal events, and automatically generate an anomaly report containing images, videos, time, location and event type when a target abnormal event is identified, so that relevant personnel can intuitively, quickly and comprehensively trace back the target abnormal event based on the anomaly report.

[0088] According to embodiments of this disclosure, determining the abnormal location of a target abnormal event in a three-dimensional virtual model based on image acquisition information of a target image acquisition device used to acquire a target video stream includes: determining the virtual position of the target image acquisition device used to acquire the target video stream in the three-dimensional virtual model based on the correspondence between the image acquisition device and the virtual image acquisition device; and determining the abnormal location of the target abnormal event in the three-dimensional virtual model based on the image acquisition information and the virtual position of the target image acquisition device, wherein the image acquisition information includes the acquisition orientation and acquisition field of view of the image acquisition device.

[0089] For example, importing a 3D virtual model from a GIS can manage the correspondence between multiple image acquisition devices and multiple virtual image acquisition devices. For a target image acquisition device acquiring a target video stream, the virtual position of the target image acquisition device can be located in the 3D virtual model based on the aforementioned correspondence; that is, the virtual position of the target virtual image acquisition device corresponding to the target image acquisition device.

[0090] Image acquisition information can include the acquisition orientation and field of view of the image acquisition device. For example, the acquisition orientation can be determined based on the attitude of the image acquisition device, such as pitch angle, yaw angle, and roll angle. The field of view can be the field of view (FOV) of the image acquisition device. By using the acquisition orientation and field of view, the acquisition area of ​​the image acquisition device when it acquires the target abnormal event can be determined. Each position in this acquisition area can be mapped to a position in the three-dimensional virtual model.

[0091] For example, by determining the depth information of the target object in each video frame of the video stream, the physical location of the target object in the acquisition area can be determined; according to the mapping relationship between the acquisition area and the corresponding area in the 3D virtual model, the physical location can be mapped to the 3D virtual model, and the abnormal location of the target event in the 3D virtual model can be obtained.

[0092] In one specific embodiment, the target anomalous event can be directed at a moving object. By determining the depth information of the object targeted by the anomalous event in multiple video frames of the target video stream and arranging them according to the timestamps of the multiple video frames, a dynamic sequence of anomalous positions of the target anomalous event moving in the 3D virtual model can be determined. Continuous tracking of the target anomalous event can also be achieved by displaying the dynamic sequence of anomalous positions in the 3D virtual model.

[0093] In the embodiments of this disclosure, by directly determining the virtual location of the target abnormal event in the three-dimensional virtual model, the location of the target abnormal event can be intuitively and accurately determined when generating an anomaly report, achieving the effect of "knowing where the incident happened at a glance", facilitating the review and repair of the target abnormal event and shortening the search time.

[0094] According to embodiments of this disclosure, image fusion is performed on a target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream, including: image distortion correction of the target video stream and at least one adjacent video stream to obtain a corrected target video stream and a corrected adjacent video stream; and stitching the corrected target video stream and the corrected adjacent video stream together to obtain a panoramic video stream.

[0095] In the embodiments of this disclosure, image acquisition devices with a large viewing angle may exhibit distortion at the edges of the image. This distortion not only affects the accuracy of panoramic fusion but also the realism of the 3D virtual model after overlaying the panoramic video stream. Therefore, image distortion correction can be performed on the video streams involved during the panoramic video stream fusion stage.

[0096] For example, the distortion coefficients and intrinsic parameters of the image acquisition device for the target video stream and at least one adjacent video stream can be obtained. For each video frame of each video stream, the pixel coordinates of each pixel point within the video frame are calculated using the distortion correction formula to obtain the ideal pixel coordinates without distortion. The distortion correction formula is determined using a preset calibration relationship between the distortion coefficients, intrinsic parameters, and distortion.

[0097] In another embodiment, in addition to the target video stream and adjacent video streams, relevant personnel may have daily viewing needs. Therefore, each video stream accessed by the GIS system can undergo image distortion correction, ensuring the authenticity of the object outline and proportions in both the video stream corresponding to the 3D virtual model and the 3D virtual model (achieved through real video and high-resolution modeling, respectively). In this embodiment, the video stream targeted by the abnormal event detection can also be a distortion-corrected video stream.

[0098] According to embodiments of this disclosure, stitching together the corrected target video stream and the corrected adjacent video streams can be done by first stitching together video frames at the same time in the above video streams, and then combining the stitched panoramic video frames into a panoramic video stream in the time dimension.

[0099] In the embodiments of this disclosure, during the process of stitching together a panoramic video stream, the video stream is first subjected to distortion correction to eliminate distortion caused by the image acquisition device itself, resulting in a corrected video stream that matches the shape and proportion of the real object. Then, the corrected target video stream and the corrected adjacent video streams are stitched together to obtain a panoramic video stream, thereby reducing the problem of poor realism caused by distortion in the panoramic video stream and the problem of stitching marks caused by distortion.

[0100] According to embodiments of this disclosure, adjacent video streams are determined by the following method: determining at least one adjacent image acquisition device that has an overlapping acquisition area with the target image acquisition device based on the image acquisition information of at least one image acquisition device; and determining the video streams acquired by at least one adjacent image acquisition device as adjacent video streams.

[0101] Based on the acquisition orientation and field of view of each image acquisition device, the acquisition area of ​​each image acquisition device can be determined, and each position in the acquisition area can be mapped to a position in the three-dimensional virtual model.

[0102] For example, adjacent image acquisition devices whose acquisition areas overlap with the target image acquisition device can be determined directly based on the acquisition area of ​​each image acquisition device. Correspondingly, the video streams acquired by at least one adjacent image acquisition device are determined as adjacent video streams.

[0103] In the embodiments of this disclosure, the presence or absence of overlapping acquisition areas is used to determine adjacent image acquisition devices, and then the video streams acquired by the adjacent image acquisition devices are determined as adjacent video streams. This can quickly determine multiple adjacent video frames used to fuse with the target video stream into a wide-viewing angle, which can not only meet the needs of wide-viewing angle viewing, but also reduce the number of fused video streams (e.g., video streams that do not overlap with the acquisition area of ​​the target video stream).

[0104] Figure 5 A flowchart illustrating the process of stitching together a panoramic video stream according to an embodiment of this disclosure is shown. Figure 5 As shown, the method includes operations S510~S560.

[0105] In operation S510, multiple video frames to be spliced ​​at the same time are determined from the corrected target video stream and the corrected adjacent video streams.

[0106] During operation of S520, feature extraction is performed on multiple video frames to be stitched together, resulting in feature points for each of the multiple video frames to be stitched together.

[0107] During operation of S530, feature point matching is performed between multiple video frames to be stitched to obtain multiple feature point sets. In each feature point set, multiple feature points are feature points of the same object in the overlapping acquisition area in different video frames to be stitched.

[0108] In operation of S540, multiple video frames to be stitched are stitched together based on multiple feature point sets to obtain panoramic video frames corresponding to the time.

[0109] When operating the S550, panoramic video frames from multiple moments are combined in chronological order to obtain a panoramic video stream.

[0110] For example, multiple video frames to be stitched together at the same time can be stitched together first, according to the video frame dimension. When stitching along the video frame dimension, the corrected target video stream and the corrected adjacent video streams can be sampled separately to obtain multiple video frames to be stitched together at the same time. It should be noted that when acquiring at least one video stream, at least one video stream can be pre-aligned in time.

[0111] For multiple video frames to be stitched together at the same time, feature point extraction algorithms can be used to extract features from each frame. Each frame includes multiple feature points. These feature points are typically architectural corners, and the specific feature points depend on the feature point extraction algorithm. Algorithms can include Scale-Invariant Feature Transform (SIFT), SpeededUp Robust Features (SURF), and Oriented Fast and Rotated BRIEF (ORB), among others.

[0112] Understandably, for two video frames to be stitched together that have overlapping acquisition areas (belonging to the target video frame and the adjacent video frame, respectively), the relative transformation relationship between the two video frames can be determined by matching feature points within the overlapping acquisition areas. For multiple video frames to be stitched together, pairwise feature point matching can be performed, and feature points for the same object within the same overlapping acquisition area can be combined into a feature point set. Understandably, if only two video frames to be stitched together have an overlapping acquisition area, the feature point set includes two feature points from each of the two video frames; if three video frames to be stitched together have the same overlapping acquisition area, the feature point set includes two feature points from each of the three video frames, and so on.

[0113] For the feature point set after feature point matching, the feature point set can be used to register the video frames to be stitched. Then, the overlapping acquisition areas of the multiple registered video frames to be stitched are fused to obtain the panoramic video frame corresponding to that moment. In a specific embodiment, color correction, smoothing, and other operations can also be performed on the panoramic video frame to eliminate stitching artifacts.

[0114] After stitching together the video frames, a panoramic video stream can be obtained by combining the panoramic video frames from multiple moments in chronological order. In this embodiment, the multiple moments can be consecutive or discontinuous (e.g., obtained by sampling from different moments). For example, for discontinuous moments, the panoramic video stream obtained by combining the panoramic video frames from multiple moments in chronological order will have content discontinuity issues. Interpolation operations can be used to smooth the multiple panoramic video frames to obtain the final panoramic video stream.

[0115] In the embodiments of this disclosure, by first performing feature point extraction, matching, and stitching operations on the video frame dimensions, seamless wide-view panoramic video frames can be obtained, reducing monitoring blind spots. Furthermore, the panoramic video stream obtained by combining multiple panoramic video frames can display the dynamic changes of target anomalies in real time, facilitating the tracking of target anomalies.

[0116] According to embodiments of this disclosure, the image acquisition device corresponds to a virtual acquisition device in the 3D virtual model. A GIS system can be used to map the video stream acquired by the image acquisition device to the virtual acquisition device in the 3D virtual model. Therefore, even when no abnormal event is identified, the video stream and / or panoramic video stream can be overlaid with the 3D virtual model to form a unified visual environment that blends the virtual and real worlds. In this embodiment, users (operators, processors, and routine inspection personnel) can all perform routine inspections using the overlaid video stream and the 3D virtual model.

[0117] For example, user pose changes can be mapped to pose changes within a 3D virtual model. As the user interacts with the model, different positions and postures within the 3D virtual model are displayed, simulating a scenario where the user freely roams within the model. In one embodiment, if the user moves to a first position within the 3D virtual model, the corresponding video stream or panoramic video stream can be automatically overlaid onto the model to display a blended virtual and real image. In another embodiment, if the user moves to a second position within the 3D virtual model, the user can click on the icon of a virtual image acquisition device to display an image of the video stream corresponding to the virtual image acquisition device overlaid / fused panoramic video; alternatively, the video stream corresponding to the virtual image acquisition device can be processed to match the viewpoint of the second position before being displayed.

[0118] Therefore, by mapping the user's pose change operation to the pose change operation in the 3D virtual model and displaying the screen with the viewpoint corresponding to the pose change operation, seamless switching and precise positioning from macro layout (3D virtual model) to micro details (local video stream) can be achieved.

[0119] According to embodiments of this disclosure, displaying a three-dimensional virtual model overlaid with a panoramic video stream includes: displaying the three-dimensional virtual model overlaid with the panoramic video stream when the anomaly type of the target anomaly event is a first priority; displaying alarm information indicating the target anomaly event when the anomaly type of the target anomaly event is a second priority; and displaying the three-dimensional virtual model overlaid with the panoramic video stream in response to a triggering operation on the alarm information, wherein the first priority is greater than the second priority.

[0120] Alarm information can be content from anomaly reports or suggestive information extracted from them, such as "Attention! An anomaly of type XX has occurred in area XX. Click the link below to jump to the location of the target anomaly." Furthermore, alarm information can also be pushed to the terminal devices of personnel involved in the target anomaly via work orders.

[0121] For example, by using pre-defined push rules, the terminal devices of relevant personnel matching the target abnormal event can be identified, and alarm information can be sent through platform messages, application software (APP) push, SMS or email.

[0122] In one specific embodiment, the alarm information may further include the processing status, such as whether the alarm information has been received, whether the processing of the target abnormal event has been completed (processing or completed), and whether the processed target abnormal event has been accepted. Similar to storing anomaly reports, the alarm information and processing status of the target abnormal event are recorded synchronously to ensure closed-loop management of the entire process status of the target abnormal event. If the processing status is that the processed target abnormal event has been accepted, the alarm information can be archived as historical information in the central database.

[0123] As stated above, target anomalies can be of various types, such as fireworks, not wearing a safety helmet, or exceeding area restrictions. These different anomaly types can have different alarm priorities. Understandably, for first-priority (higher priority for the anomaly type) target anomalies, a 3D virtual model overlaid with the panoramic video stream is displayed simultaneously with the alarm message. This direct interface change alerts the user to the urgent anomaly and helps them quickly and intuitively identify the target anomaly. For second-priority (lower priority for the anomaly type) target anomalies, an alarm message is displayed to inform the inspection site of the anomaly, allowing alerts to be issued without disrupting the current tasks of relevant personnel. Relevant personnel can trigger the display of the 3D virtual model overlaid with the panoramic video stream by clicking on the alarm message or other trigger actions.

[0124] In the embodiments of this disclosure, after detecting a target abnormal event, alarm information is automatically pushed to users (operators, processors, and other relevant personnel), enabling the abnormal event detection to achieve a fully online closed loop of "automatic detection - reporting - allocation to relevant personnel - push to relevant personnel - processing - feedback." Furthermore, the entire process is linked through digital information, allowing for traceability of data at any intermediate stage. Therefore, the embodiments of this disclosure can solve the problems of long communication chains, slow response, and difficulty in tracing the processing process in traditional manual inspection methods, improving the overall efficiency of inspection and abnormal response.

[0125] According to embodiments of this disclosure, the method further includes: acquiring and displaying multi-source data collected by the wearable device when the user wearing the wearable device is in a blind spot area within the inspection site; wherein the blind spot area is an area not captured by at least one image acquisition device, the multi-source data includes at least a video stream collected by the wearable device, and the wearable device is also used to display a three-dimensional virtual model and a location marker indicating the user's current location.

[0126] Understandably, although at least one image acquisition device is installed within the inspection area, there will still be blind spots that are not captured. Inspection of these blind spots can be achieved by having inspection personnel wear wearable devices. These wearable devices can be glasses, watches, etc., allowing users (inspection personnel) to view a 3D virtual model and understand their current location within the entire model in real time. For example, based on the wearable device's position, a location marker corresponding to that location can be generated in the 3D virtual model and displayed through the wearable device, enabling real-time location awareness during human-machine collaborative inspections. Simultaneously, the wearable devices of the inspection personnel can be used to collect and display multi-source data on the blind spots, thus achieving effective inspection of these areas.

[0127] Multi-source data includes at least video streams collected by wearable devices; it may also include audio data, posture data, and other data. It is understood that wearable devices can integrate first-person perspective cameras, voice interaction, augmented reality displays, multiple sensors, and 4G / 5G / Wi-Fi wireless transmission functions. Multi-source data can be the audio and video communication involved in the aforementioned functions, enabling real-time acquisition and transmission of multi-source data.

[0128] In one embodiment, the wearable device can be smart glasses integrated with augmented reality (AR) technology, which blends a 3D virtual model with its surrounding environment. For example, after collecting multi-source data through AR glasses, the collected video stream can be directly overlaid onto the 3D virtual model, and the 3D virtual model with the overlaid video stream can be displayed through the AR glasses; alternatively, only the collected video stream can be displayed.

[0129] In the embodiments of this disclosure, for blind spots, multi-source data is collected and displayed through wearable devices, enabling human-machine collaborative mobile inspection. This inspection mode can serve as an effective extension and functional supplement to the automatic identification of fixed video streams, eliminating monitoring blind spots in time and space, improving inspection efficiency, and reducing inspection costs.

[0130] According to embodiments of this disclosure, the method further includes: in response to a remote assistance request triggered by a user through a wearable device, establishing a communication connection with a target terminal device corresponding to the remote assistance request when the user wearing the wearable device is in a blind spot area within the inspection site; sending multi-source data collected by the wearable device to the target terminal device through the communication connection, so as to display the multi-source data on the target terminal device; and sending inspection guidance data obtained from the target terminal device to the wearable device, wherein the inspection guidance data includes at least one of screen data and audio data used to guide the user to perform inspection operations.

[0131] For example, when inspection personnel wearing wearable devices encounter complex problems that they cannot solve independently, they can initiate a remote assistance request through AR glasses. When initiating a remote assistance request, the inspection personnel can specify the assisting user (backend expert or engineer), or the server can automatically assign an assisting user (backend expert or engineer) based on the remote assistance request. The assisting user's terminal device is also the target terminal device.

[0132] After receiving a remote assistance request initiated by the AR glasses, the server can establish a communication connection with the target terminal device. Through this communication connection, the server can send multi-source data collected by the wearable device to the target terminal device for display.

[0133] For example, the audio and video streams captured by AR glasses can be forwarded to the target terminal device for display. Through a server, an indirect communication connection can be established between the target terminal device and the AR glasses to help users understand the audio and video streams viewed through the AR glasses in real time (i.e., the first-person view of the scene).

[0134] After the audio and video stream is displayed on the target terminal device, the assistant user can generate inspection guidance data based on the audio and video stream and feed the inspection guidance data back to the AR glasses to remotely guide the inspection personnel. For example, the assistant user can conduct real-time audio and video conversations with the inspection personnel (generating audio data in the inspection guidance data), and the assistant user can also annotate the displayed audio and video stream (generating visual data in the inspection guidance data). For example, the assistant user can use the visual annotation tool to circle the fault location, mark the operation steps, etc.

[0135] Similar to handling target anomalies, initiating a remote assistance request can be considered a target anomaly (e.g., the type of remote assistance or the specific anomaly type handled by the remote assistance). Multi-source data collected via AR glasses, inspection guidance data, etc., can be used as information in the anomaly report. The physical location of the initiated remote assistance request is mapped to a virtual location in the 3D virtual model and recorded as the anomaly location. Similar to generating alarm information, this target anomaly also generates new alarm information, which is pushed to relevant personnel and the data storage center database, and integrated with the video stream processing of fixed image acquisition devices.

[0136] In the embodiments of this disclosure, real-time remote guidance through remote assistance requests can promptly and accurately identify and handle abnormal events, improving processing efficiency. In particular, for complex abnormal events, remote assistance can significantly improve processing efficiency in complex situations.

[0137] Figure 6 A block diagram of a video fusion-based anomaly detection apparatus according to an embodiment of the present disclosure is illustrated schematically. Figure 6As shown, the anomaly detection device 600 based on video fusion includes:

[0138] The detection module 610 is used to perform abnormal event detection on at least one video stream and obtain the detection result. The video stream is obtained using at least one image acquisition device set up at the inspection site. The image acquisition device corresponds to a virtual image acquisition device in the three-dimensional virtual model of the inspection site.

[0139] The generation module 620 is used to generate an anomaly report for the target anomaly event based on the image acquisition information of the target image acquisition device used to acquire the target video stream and the target video stream, when the detection result indicates that there is a target video stream in at least one video stream where a target anomaly event has occurred.

[0140] The fusion module 630 is used to perform image fusion on the target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream.

[0141] Display module 640 is used to overlay the panoramic video stream onto the acquisition area corresponding to the panoramic video stream in the three-dimensional virtual model, and to display the three-dimensional virtual model after overlaying the panoramic video stream.

[0142] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.

[0143] For example, any plurality of the detection module 610, generation module 620, fusion module 630, and display module 640 can be combined into one module / unit / subunit, or any one of these modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of the present disclosure, at least one of the detection module 610, generation module 620, fusion module 630, and display module 640 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the detection module 610, generation module 620, fusion module 630, and display module 640 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0144] It should be noted that the apparatus portion in the embodiments of this disclosure corresponds to the method portion in the embodiments of this disclosure. The description of the apparatus portion is specifically referred to in the method portion, and will not be repeated here.

[0145] Figure 7 A block diagram of an electronic device suitable for implementing a video fusion-based anomaly detection method according to an embodiment of the present disclosure is shown schematically. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0146] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0147] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0148] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.

[0149] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by processor 701, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0150] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0151] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0152] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.

[0153] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of this disclosure.

[0154] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0155] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0156] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0158] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. An anomaly detection method based on video fusion, characterized in that, The method includes: Anomaly detection is performed on at least one video stream to obtain a detection result, wherein the video stream is obtained using at least one image acquisition device set up at the inspection site, and the image acquisition device corresponds to a virtual image acquisition device in a three-dimensional virtual model of the inspection site; If the detection result indicates that at least one of the video streams contains a target video stream in which a target abnormal event has occurred, an abnormal report for the target abnormal event is generated based on the image acquisition information of the target image acquisition device used to acquire the target video stream and the target video stream. Image fusion is performed on the target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream; and The panoramic video stream is superimposed onto the acquisition area corresponding to the panoramic video stream in the three-dimensional virtual model, and the three-dimensional virtual model after superimposing the panoramic video stream is displayed.

2. The method according to claim 1, characterized in that, The step of generating an anomaly report for the target anomaly event based on the image acquisition information of the target image acquisition device used to acquire the target video stream and the target video stream includes: Based on the image acquisition information of the target image acquisition device used to acquire the target video stream, the abnormal location of the target abnormal event in the three-dimensional virtual model is determined; Based on at least one video frame in the target video stream containing the target anomalous event, an abnormal time period is determined, and a video segment of the abnormal time period and a key video frame for identifying the target anomalous event are extracted from the target video stream; and An anomaly report is generated based on the anomaly location, the anomaly time period, the video segment, the key video frame, and the anomaly type of the target anomaly event.

3. The method according to claim 2, characterized in that, The step of determining the abnormal location of the target abnormal event in the three-dimensional virtual model based on the image acquisition information of the target image acquisition device used to acquire the target video stream includes: Based on the correspondence between the image acquisition device and the virtual image acquisition device, the virtual position of the target image acquisition device used to acquire the target video stream in the three-dimensional virtual model is determined; and Based on the image acquisition information of the target image acquisition device and the virtual position, the abnormal position of the target abnormal event in the three-dimensional virtual model is determined, wherein the image acquisition information includes the acquisition orientation and acquisition field of view of the image acquisition device.

4. The method according to claim 1, characterized in that, The step of fusing images of the target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream includes: Image distortion correction is performed on the target video stream and at least one of the adjacent video streams to obtain the corrected target video stream and the corrected adjacent video streams; and The corrected target video stream and the corrected adjacent video streams are stitched together to obtain the panoramic video stream.

5. The method according to claim 4, characterized in that, The adjacent video streams are determined by the following method: Based on the image acquisition information of at least one of the image acquisition devices, at least one adjacent image acquisition device with an overlapping acquisition area with the target image acquisition device is determined. as well as The video stream acquired by at least one of the adjacent image acquisition devices is determined as the adjacent video stream.

6. The method according to claim 4, characterized in that, The step of stitching together the corrected target video stream and the corrected adjacent video streams to obtain the panoramic video stream includes: Multiple video frames to be spliced ​​at the same time are determined from the corrected target video stream and the corrected adjacent video streams; Feature extraction is performed on each of the multiple video frames to be stitched together to obtain the feature points of each of the multiple video frames to be stitched together; Feature point matching is performed between multiple video frames to be stitched to obtain multiple feature point sets, wherein the multiple feature points in each feature point set are feature points of the same object in the overlapping acquisition area in different video frames to be stitched. Based on multiple sets of feature points, multiple video frames to be stitched together are obtained to acquire a panoramic video frame corresponding to the stated time. The panoramic video stream is obtained by combining the panoramic video frames from multiple moments in chronological order.

7. The method according to any one of claims 1 to 6, characterized in that, The display of the 3D virtual model after overlaying the panoramic video stream includes: If the anomaly type of the target anomaly event is of the first priority, display the three-dimensional virtual model superimposed on the panoramic video stream; When the anomaly type of the target anomaly event is the second priority, an alarm message indicating the target anomaly event is displayed; and, in response to a triggering operation for the alarm message, a three-dimensional virtual model superimposed on the panoramic video stream is displayed, wherein the first priority is greater than the second priority.

8. The method according to claim 1, characterized in that, The method further includes: When a user wearing a wearable device is in a blind spot within the inspection area, the system acquires and displays multi-source data collected through the wearable device. The blind spot area is at least one area not captured by the image acquisition device. The multi-source data includes at least a video stream acquired by the wearable device. The wearable device is also used to display the three-dimensional virtual model and a location marker indicating the user's current location.

9. The method according to claim 8, characterized in that, The method further includes: In response to a remote assistance request triggered by the user through the wearable device, if the user wearing the wearable device is in a blind spot within the inspection site, a communication connection is established with the target terminal device corresponding to the remote assistance request. Through the communication connection, multi-source data collected by the wearable device is sent to the target terminal device so as to display the multi-source data on the target terminal device; The inspection guidance data obtained from the target terminal device is sent to the wearable device, wherein the inspection guidance data includes at least one of screen data and audio data for guiding the user to perform inspection operations.

10. The method according to claim 1, characterized in that, The method further includes: Acquire multi-source modeling data for the inspection site, wherein the multi-source modeling data is obtained using various types of modeling data acquisition devices; The three-dimensional virtual model is constructed using the multi-source modeling data; and Based on the physical location of at least one image acquisition device set up at the inspection site, a virtual image acquisition device corresponding to the image acquisition device is constructed in the three-dimensional virtual model.

11. A video fusion-based inspection device, characterized in that, The device includes: A detection module is used to detect abnormal events in at least one video stream and obtain detection results. The video stream is obtained using at least one image acquisition device set up at the inspection site, and the image acquisition device corresponds to a virtual image acquisition device in a three-dimensional virtual model of the inspection site. A generation module is configured to, when the detection result indicates that at least one of the video streams contains a target video stream in which a target abnormal event has occurred, generate an abnormal report for the target abnormal event based on the image acquisition information of the target image acquisition device used to acquire the target video stream and the target video stream. The fusion module is used to perform image fusion on the target video stream and at least one adjacent video stream with an overlapping acquisition area to obtain a panoramic video stream; The display module is used to overlay the panoramic video stream onto the acquisition area corresponding to the panoramic video stream in the three-dimensional virtual model, and to display the three-dimensional virtual model after overlaying the panoramic video stream.

12. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.

14. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.