Audio and video interaction method and system in live working scene based on WebRTC technology
By using WebRTC technology and lidar point cloud modeling in live operation scenarios, a three-dimensional spatial model is generated and labeled information is superimposed in the video stream, safe distance is calculated in real time and early warning is triggered, the problem that the existing technology cannot meet the security needs of complex and highly dynamic operation scenarios is solved, and efficient audio and video interaction and security monitoring are achieved.
Patent Information
- Application Number
- CN202510021086.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-27
AI Technical Summary
The existing live operation technology has technical shortcomings in real-time audio and video interaction and dynamic environment modeling, and cannot meet the security needs of complex and highly dynamic operation scenarios.
The audio and video interaction method in live operation scenarios based on WebRTC technology, by collecting real-time data of the operators and the environment, using the lidar module to generate a three-dimensional spatial model, and superimposing key area labeling information in the video stream, calculate the safe distance between the operators and the live body and the grounding body in real time. When the safety distance is lower than the preset threshold, an active warning is triggered.
It significantly improves the real-time and security of audio and video interaction in live operation scenarios. Through dynamic three-dimensional modeling and multi-modal early warning, the accuracy and safety of operations are improved, and the needs of complex and highly dynamic operation scenarios are met.
Smart Images

Figure CN120050266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power operation scenario modeling, and specifically to an audio-video interaction method and system in a live working scenario based on WebRTC technology. Background Art
[0002] With the rapid development of the power industry, live working has gradually become an important technical means to ensure the safe operation of the power grid. Live working can significantly improve the operation efficiency and economic benefits of the power grid by repairing and maintaining power equipment without stopping power supply. However, due to the complex and extremely risky working environment, how to ensure the safety of workers has always been the core issue in this field of research. In recent years, with the help of real-time audio-video interaction technology, live working has gradually developed towards remote collaboration and intelligence. Among them, real-time audio-video communication based on WebRTC (Web Real-Time Communication) technology has attracted much attention due to its low-latency and high-quality transmission performance. WebRTC technology supports high-frame-rate video streams and two-way voice interaction through peer-to-peer connections, and is widely used in fields such as video conferencing and remote collaboration, providing technical support for real-time information sharing in live working.
[0003] At the same time, significant progress has also been made in three-dimensional space modeling and point cloud processing technologies in the industrial field. As a high-precision three-dimensional data acquisition device, lidar can construct a point cloud model of the environment in real time, which is used to identify obstacles, locate key areas, and plan operation paths. The integration of these technologies provides a new solution for the live working scenario of the power system, and can generate a dynamic three-dimensional model of the working scenario in a complex environment. However, most of the existing technologies are limited to the modeling and analysis of static environments, and cannot meet the high-dynamic requirements of live working scenarios. Especially in terms of real-time interaction and safety warning, there is still a large room for improvement. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is that the existing live working technologies have technical shortcomings in real-time audio-video interaction and dynamic environment modeling, and cannot meet the safety requirements of complex and high-dynamic working scenarios.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: An audio-video interaction method in a live working scenario based on WebRTC technology, including: collecting real-time data of the operator and the working environment, collecting environmental point cloud data using a lidar module, and generating a three-dimensional space model of the working scenario; superimposing key area annotation information of the three-dimensional space model on the WebRTC video stream to achieve visual enhancement of the working environment; calculating the safety distance between the operator and the live conductor and the grounding body in real time, and when the safety distance is lower than the preset threshold, triggering an active warning through the audio-video stream.
[0007] As a preferred solution of the audio-video interaction method in a live working scenario based on WebRTC technology according to the present invention, wherein: the lidar module is fixed on the wearable device of the operator, and high-precision spatial coordinates of the key points of the operator are obtained through RTK positioning technology.
[0008] As a preferred solution of the audio-video interaction method in a live working scenario based on WebRTC technology according to the present invention, wherein: the generation of the three-dimensional space model of the working scenario includes fusing the point cloud data with the environmental image to generate a preliminary three-dimensional space model;
[0009] By analyzing the shape features and spatial positions of objects in the preliminary three-dimensional space model in real time, identifying the positions of the live conductor and the grounding body, and marking the dangerous areas;
[0010] Introducing a path optimization algorithm, generating the shortest path based on the point cloud model to avoid dangerous areas;
[0011] Based on real-time data, reconstructing the point cloud model in real time and updating the parameters of the objects, and embedding the real-time updated three-dimensional model information in the WebRTC video stream.
[0012] As a preferred solution of the audio-video interaction method in a live working scenario based on WebRTC technology according to the present invention, wherein: the dangerous area marking includes obtaining a three-dimensional coordinate point set of the working scenario according to the point cloud data, using voxel filtering to downsample the point cloud data, and retaining the key structural features;
[0013] Using the RANSAC plane segmentation algorithm, extracting a large-scale plane as the reference plane of the ground according to the normal vector distribution of the point cloud. When P i Satisfies:
[0014] |a·x i +b·y i +c·z i +d|<∈
[0015] P iBelonging to a plane point, where a, b, and c represent the components of the plane normal vector, calculated by a fitting algorithm; d represents the offset of the plane, calculated from the overall distribution of the point cloud data; x i , y i , z i The three-dimensional coordinates of point P in the point cloud; i
[0016] The remaining point cloud is grouped by a clustering algorithm (DBSCAN) to identify independent objects in the working environment, calculate geometric features for each clustering group, and match the object features with the known template data of live conductors and grounding conductors;
[0017] Using the electric field intensity sensor installed on the operator's device, the electric field intensity data in the working environment is collected in real time, and the electric field intensity data is matched with the point cloud coordinates to generate the electric field intensity value of each point;
[0018] Set the electric field intensity danger threshold. If any point in the point cloud satisfies E(x, y, z) > E threshold , then mark this point as a dangerous point;
[0019] Use the α-shape algorithm for all dangerous point sets to generate a three-dimensional boundary surrounding the dangerous points:
[0020]
[0021] Among them, E(x, y, z) represents the electric field intensity of point (x, y, z); E threshold Represents the dangerous electric field intensity threshold; Represents the three-dimensional boundary generated by the dangerous points; S represents a simple geometric unit in the point set; α represents the shape control parameter; label the output three-dimensional boundary as a visual marker for the dangerous area.
[0022] As a preferred solution of the audio and video interaction method in the live working scenario based on the WebRTC technology described in the present invention, wherein: the path optimization algorithm includes, according to the three-dimensional point cloud data of the working environment and the point set of the dangerous area, meshing the point cloud data and dividing it into a three-dimensional grid map, and the grid information includes occupancy marks and danger marks;
[0023] Obtain the current position and target point of the operator, and perform path optimization based on the A algorithm:
[0024] f(n) = g(n) + h(n)
[0025] The heuristic function h(n) uses the Euclidean distance:
[0026]
[0027] Among them, f(n) represents the evaluation function; g(n) represents the actual cost; (x n , y n , z n ) represents the coordinates of the target point; (x g , y g , z g ) represents the coordinates of the current position;
[0028] When it is detected that the path is blocked by obstacles or dangerous areas, the A algorithm is called again, starting from the position of the staff, and an updated path is generated;
[0029] According to the path point set, a path is generated in the form of a polyline and smoothed:
[0030]
[0031] Among them, P raw (i) represents the original path point; P smooth (i) represents the smoothed path point; λ represents the smoothing weight; the smoothed path is marked in the point cloud model, and the path is projected onto the real-time video stream using the WebRTC API.
[0032] As a preferred solution of the audio and video interaction method in the live working scenario based on the WebRTC technology of the present invention, wherein: calculating the safe distance between the operator and the live conductor and the grounding conductor in real time includes calculating the length of the shortest path from the operator to the dangerous target based on the three-dimensional space model;
[0033] Dynamically monitor the path length. When the path length is less than the safety threshold, trigger an alarm and display it in the form of a visual marker through the WebRTC video stream.
[0034] As a preferred solution of the audio and video interaction method in the live working scenario based on the WebRTC technology of the present invention, wherein: triggering the alarm and displaying it in the form of a visual marker through the WebRTC video stream includes:
[0035] Sending a voice prompt through the WebRTC audio channel;
[0036] Overlaying a warning marker in the video stream, marking the dangerous area and dynamically updating the warning information;
[0037] Combining voice and video multi-modal prompts to remind the operator to adjust the operation.
[0038] An audio and video interaction system in the live working scenario based on WebRTC technology adopting any of the methods of the present invention, wherein: a data acquisition module acquires real-time data of the operating personnel and the working environment, transmits it through WebRTC technology, and acquires environmental point cloud data by using a lidar module;
[0039] A three-dimensional modeling module generates a three-dimensional space model of the working scenario according to the acquired real-time data and point cloud data, and superimposes key area annotation information of the three-dimensional space model on the WebRTC video stream;
[0040] An interaction module calculates the safety distance between the operating personnel and the live body and the grounding body in real time. When the safety distance is lower than a preset threshold, it triggers an active warning through the audio and video stream.
[0041] A computer device includes: a memory and a processor; the memory stores a computer program, including: when the processor executes the computer program, the steps of any of the methods of the present invention are implemented.
[0042] A computer-readable storage medium stores a computer program thereon, including: when the computer program is executed by a processor, the steps of any of the methods of the present invention are implemented.
[0043] Advantages of the present invention: By combining WebRTC technology and lidar point cloud modeling, the present invention significantly improves the real-time performance and safety of audio and video interaction in the live working scenario. By generating a three-dimensional space model of the working scenario and dynamically superimposing key area annotation information on the video stream, the visualization of the working environment is more intuitive. At the same time, the safety distance between the operating personnel and the live body and the grounding body is calculated in real time, and the active warning function is triggered, comprehensively improving the accuracy and safety of the operation. Compared with the prior art, the present invention integrates efficient dynamic modeling, low-latency video interaction and real-time safety monitoring, providing a reliable guarantee for the intelligence and safety of live working. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0045] Figure 1 It is the overall flowchart of the audio and video interaction method in the live working scenario based on WebRTC technology provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0047] Example 1. Referring to Figure 1 , which is an embodiment of the present invention, provides an audio-visual interaction method in the live working scenario based on WebRTC technology, including:
[0048] S1: Collect the real-time data of the operator and the working environment, use the lidar module to collect the environmental point cloud data, and generate a three-dimensional space model of the working scenario.
[0049] Furthermore, the arrangement of the lidar module includes three schemes: single-point installation, multi-point distribution, and dynamic mobile collection. The specific arrangement method is selected according to the complexity of the working scenario and the spatial characteristics of the working environment. Among them, the single-point installation method is suitable for relatively open working environments. By fixing it on the wearable device of the operator, it can collect the point cloud data within a certain range around it; the multi-point distribution method arranges multiple lidar modules at key positions in the working environment, and through multi-source point cloud fusion technology, generates a more comprehensive three-dimensional space model; the dynamic mobile collection method is suitable for large or complex working scenarios. The lidar module is installed on a mobile device (such as a drone or a trajectory robot), and combined with path planning and motion synchronization technology, it can cover a wider working area.
[0050] In this embodiment, the lidar module adopts the single-point installation method and is installed on the helmet or shoulder wearable device of the operator. Through modular design, it realizes lightweight and high-sensitivity point cloud collection. This arrangement method is easy to operate and can meet the real-time three-dimensional modeling requirements in the conventional live working scenario. In an alternative solution, for complex working environments with occlusion or multiple obstacles, the multi-point distribution or dynamic mobile collection method can be adopted to optimize the coverage range through the spatial layout of the lidar module and improve the integrity and accuracy of the point cloud data.
[0051] Furthermore, the point cloud acquisition parameters of the lidar module (such as sampling frequency, resolution, and scanning angle) can be dynamically adjusted according to the operation requirements. For example, when high-precision modeling is required in the operation environment, the sampling frequency and resolution can be increased, and the scanning angle can be reduced simultaneously to concentrate on collecting data in key areas; while in large-scale scanning, the scanning angle and sampling range can be expanded to cover a larger spatial area. Similarly, the lidar module can combine the feedback information of the environmental sensor to optimize the acquisition parameters in real time to adapt to external factors such as environmental light intensity and electromagnetic interference, ensuring the stability and reliability of data acquisition.
[0052] Moreover, to improve the real-time performance and accuracy of point cloud data, the lidar module can integrate a high-performance data processing unit to complete point cloud denoising and filtering processing in real time through embedded algorithms. For example, voxel filtering (VoxelGrid Filter) is used to downsample the original point cloud data, reducing redundant points and retaining key structural features; statistical outlier removal is used to remove noise points and enhance the clarity of point cloud data. On the basis of real-time data processing, the lidar module can also combine attitude sensor data (such as accelerometers and gyroscopes) to correct the attitude of point cloud data, ensuring the spatial alignment accuracy between the point cloud and the operation scene.
[0053] Furthermore, the collected point cloud data not only includes the position information of the operators but also contains the environmental structure features in the operation scene, such as the geometric contours of live equipment and the three-dimensional distribution of surrounding obstacles. These data are integrated in real time through an embedded system to form a unified three-dimensional space model and transmitted to a remote collaboration center for experts to analyze. In practical applications, the integrity of point cloud data acquisition and the real-time update of the model directly affect the accuracy and safety of the operation. Therefore, the acquisition range, parameter settings, and real-time data processing capabilities of the lidar module are the keys to ensuring the three-dimensional modeling effect.
[0054] During the point cloud acquisition process, to meet the requirements of different operation environments, the lidar module can also combine various auxiliary technologies, such as environmental light intensity sensors and high-dynamic range imaging (HDR) technology, to improve the point cloud acquisition effect in low-light or high-reflectivity environments; use edge computing nodes to process point cloud data in real time and synchronize the processing results to a remote collaboration platform through a low-latency communication module, thereby realizing efficient digital modeling of the operation scene.
[0055] Based on the collected real-time data and point cloud data, the point cloud data is fused with the environmental image to generate a preliminary three-dimensional space model; by analyzing the shape features and spatial positions of objects in the preliminary three-dimensional space model in real time, the positions of live conductors and grounding conductors are identified, and dangerous areas are marked; a path optimization algorithm is introduced to generate the shortest path based on the point cloud model to avoid dangerous areas; based on the real-time data, the point cloud model is reconstructed in real time and the parameters of objects are updated, and the real-time updated three-dimensional model information is embedded in the WebRTC video stream.
[0056] S2: Superimpose the key area annotation information of the three-dimensional space model in the WebRTC video stream to achieve visual enhancement of the working environment.
[0057] Furthermore, according to the point cloud data in the working environment, a three-dimensional coordinate point set of the working scene is obtained, and the point cloud data is downsampled by a voxel grid filter. The voxel grid filter retains the key structural features while reducing the data volume and improving the subsequent processing efficiency.
[0058] The RANSAC plane segmentation algorithm is used to extract a large-scale plane as the reference plane of the ground according to the normal vector distribution of the point cloud. When P i satisfies:
[0059] |a·x i +b·y i +c·z i +d| < ∈
[0060] P i belongs to the plane points, where a, b, c represent the components of the plane normal vector, which are calculated by the fitting algorithm; d represents the offset of the plane, which is calculated by the overall distribution of the point cloud data; x i , y i , z i are the three-dimensional coordinates of point P in the point cloud. i
[0061] The remaining point cloud outside the plane is grouped by the DBSCAN clustering algorithm to identify independent objects (such as live conductors, grounding conductors, etc.) in the working environment. Geometric features (such as volume, bounding box size, and center point position) are calculated for each clustering group, and the object features are matched with known templates. For example, if the object shape and size match those of a live conductor (such as an insulator, line), it is marked as a live conductor; if the object is close to the ground plane and conforms to the size characteristics of the grounding device, it is marked as a grounding conductor.
[0062] An electric field intensity sensor installed on the operator's device is used to collect the electric field intensity data in the working environment in real time, and the electric field intensity data is matched with the point cloud coordinates to generate the electric field intensity value of each point.
[0063] Set the dangerous threshold of the electric field intensity. If any point in the point cloud satisfies E(x, y, z) > E threshold , then mark this point as a dangerous point. Use the α-shape algorithm to generate a three-dimensional boundary surrounding the dangerous points for all sets of dangerous points:
[0064]
[0065] where E(x, y, z) represents the electric field intensity at point (x, y, z); E threshold represents the dangerous electric field intensity threshold; S α represents the three-dimensional boundary generated by the dangerous points; S represents a simple geometric unit in the point set; α represents the shape control parameter; label the output three-dimensional boundary as a visual marker for the dangerous area.
[0066] Use WebGL to render the boundary of the dangerous area and synchronously overlay it with the video frame. Display the dangerous area in the video with a dynamic red boundary, and dynamically highlight the positions of the charged body and the grounding body in the three-dimensional model. Continuously collect point cloud and electric field data, and regularly update the boundary of the dangerous area (the refresh frequency can be set, for example, 10 times per second). The updated content includes: newly added dangerous points, boundary adjustment, and real-time display of electric field changes.
[0067] Furthermore, generate a safe path for the operator in the charged environment through a path optimization algorithm, avoid the dangerous area in real time, and make dynamic adjustments when the working environment changes. Finally, overlay the path visualization in the video stream.
[0068] According to the three-dimensional point cloud data of the working environment and the point set of the dangerous area, grid the point cloud data and divide it into a three-dimensional grid map. The grid information includes occupancy markers and danger markers. The occupancy marker indicates whether the grid is occupied (determined by the point cloud density), and the danger marker indicates whether the grid belongs to the dangerous area (determined by the point set of the dangerous area).
[0069] Obtain the current position and target point of the operator, and perform path optimization based on the A algorithm:
[0070] f(n) = g(n) + h(n)
[0071] The heuristic function h(n) uses the Euclidean distance:
[0072]
[0073] where f(n) represents the evaluation function; g(n) represents the actual cost; (x n , y n , z n ) represents the current position coordinates; (x g , y g, z g ) represents the coordinates of the target point.
[0074] The path search steps are expressed as:
[0075] Initialization:
[0076] Add the starting point to the open list, and the closed list is empty.
[0077] Iterative search:
[0078] Select the node with the smallest f(n) from the open list as the current node n;
[0079] Check whether the current node is the target point. If so, stop the search;
[0080] Otherwise, expand the neighborhood grid of the current node: ignore the occupied grid and the dangerous grid, calculate the f(n) value of the neighborhood grid and update the open list.
[0081] Path generation:
[0082] After the search is completed, generate a complete path by backtracking the parent node.
[0083] When it is detected that the path is blocked by an obstacle or a dangerous area, re - call the A algorithm, use the location of the staff as the starting point to generate an updated path; generate a path in the form of a polyline according to the path point set and perform smoothing processing:
[0084]
[0085] Among them, P raw (i) represents the original path point; P smooth (i) represents the smoothed path point; λ represents the smoothing weight; annotate the smoothed path in the point cloud model, project the path onto the real - time video stream using the WebRTC API, and convert the path points into pixel coordinates in the video through the video rendering interface to dynamically display the path update, and the operator can observe it in real - time.
[0086] It should be noted that through the annotation of key areas in the three - dimensional space model, accurate recognition and real - time visualization of the operation environment are achieved. Dynamically superimposing annotation information in the WebRTC video stream effectively improves the safety and operation intuitiveness in the live working scenario, and at the same time provides real - time updated basic data support for path planning, making the overall method more efficient and intelligent.
[0087] S3: Real - time calculate the safety distance between the operator and the live body and the grounding body. When the safety distance is lower than the preset threshold, trigger an active warning through the audio - video stream.
[0088] Furthermore, based on the three-dimensional space model and the real-time position of the operator, the Euclidean distance formula is used to calculate the shortest path length between the operator and the live body and the grounding body.
[0089]
[0090] Among them, d represents the distance between the operator and the target point; (x p , y p , z p ) represents the current position coordinates of the operator; (x t , y t , z t ) represents the coordinate of the target point (such as the boundary of the live body or the grounding body).
[0091] Extract the minimum value from the distance calculation results of all target points. Based on time series analysis, evaluate the movement trend between the operator and the dangerous target, and set the safety threshold d T , when d < d T , it is determined as a dangerous distance.
[0092] When a dangerous distance is detected, generate multimodal warning signals, including audio warning and video warning. Specifically, the audio warning sends a voice prompt through the WebRTC audio channel, such as "Danger! Please stay away from the live body"; the video warning dynamically superimposes warning marks in the WebRTC video stream, marks the dangerous area and updates the deviation direction of the operator's position.
[0093] Extract key warning feature parameters, mark the coordinates of dangerous points in the three-dimensional model, and superimpose the safe evacuation path of the operator according to the distance change trend. The update frequency of the warning signal is synchronized with the position change of the operator (such as 10 times per second).
[0094] Adjust the volume of the voice prompt in combination with the noise intensity of the working environment; strengthen the directive language through the voice content, such as "Move two meters to the left and stay away from the live body". By combining the dynamic update and real-time calculation of the three-dimensional space model, the method can accurately evaluate the distance between the operator and the dangerous target, and significantly improve the working safety through multimodal warning signals.
[0095] Embodiment 2: In an exemplary embodiment, an audio-video interaction system in a live working scenario based on WebRTC technology is further provided, including a data acquisition module, a three-dimensional modeling module, and an interaction module.
[0096] The data acquisition module is responsible for collecting real-time data of operators and the working environment, obtaining environmental point clouds using a lidar module and generating a three-dimensional space model. The lidar can be installed at a single point or distributed at multiple points according to requirements, and the attitude of the point cloud is corrected by combining attitude sensors (such as accelerometers and gyroscopes) to ensure the spatial alignment accuracy with the actual working scenario. At the same time, this module can automatically adjust the sampling parameters of the lidar according to external factors such as environmental light and electromagnetic interference to ensure the stability and integrity of data acquisition.
[0097] The data acquisition and processing module is docked with the real-time data fusion module through an embedded system, transmits the generated three-dimensional point cloud model and its key feature data to the next module, and at the same time returns the real-time optimization parameters to the sensor through a communication interface to form a data closed-loop.
[0098] The three-dimensional modeling module receives the point cloud data and electric field intensity information from the data acquisition and processing module, fuses the environmental image data, and generates a complete three-dimensional space model. Based on the model, key area analysis and annotation are carried out, including identifying the positions of live bodies and grounding bodies, delimiting dangerous areas, and updating object parameters in real time. On this basis, a safe path is generated through a path optimization algorithm to avoid dangerous areas, and the path is visualized in the model. At the same time, the module will dynamically analyze the distance between the real-time position of the operator and the dangerous target, and trigger a multi-modal warning signal when the safety distance is found to be lower than the threshold.
[0099] The real-time data fusion and analysis module is connected to the data acquisition module through a data bus to obtain real-time input, and at the same time interacts with the operator's head-mounted display device or mobile terminal through a network interface to provide real-time updated three-dimensional models and path guidance information.
[0100] The interaction module, based on the fused three-dimensional space model and real-time position data, is responsible for generating multi-modal warning signals and visualizing the environment. The audio warning reminds the operator of potential risks through voice prompts, and the video warning superimposes the boundaries of dangerous areas, highlights dangerous points, and the dynamic path of the operator in the WebRTC video stream. The module also intelligently adjusts the volume and instruction content of the voice prompt to ensure the effectiveness of the warning signal in a noisy environment. In addition, it displays the optimized safe path and changes in the working scenario in real time to assist the operator in quickly judging and taking countermeasures.
[0101] The warning and visualization interaction module works closely with the real-time data fusion and analysis module, receives data on dangerous area and path updates through an efficient communication protocol, and at the same time feeds back the visualization rendering results to the operator's display device. Synchronize warning signals through the WebRTC protocol to achieve low-latency audio and video transmission and dynamic picture overlay.
[0102] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0103] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.
[0104] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as necessary, and then storing it in a computer memory.
[0105] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0106] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. The audio and video interaction method in the live working scenario based on WebRTC technology is characterized by: include: Collect real-time data of workers and working environment, use LiDAR module to collect environmental point cloud data, and generate a three-dimensional spatial model of the working scene; Overlay key area annotation information of the 3D space model in the WebRTC video stream to enhance the visualization of the working environment; The safe distance between workers and live and grounded objects is calculated in real time. When the safe distance is lower than the preset threshold, an active warning is triggered through audio and video streams.
2. The audio and video interaction method in live working scenarios based on WebRTC technology according to claim 1, characterized in that: The laser radar module is fixed on the wearable device of the operator and obtains the high-precision spatial coordinates of the operator's key points through RTK positioning technology.
3. The audio and video interaction method in live working scenarios based on WebRTC technology as claimed in claim 2, characterized in that: Generating the three-dimensional space model of the operation scene includes fusing the point cloud data with the environment image to generate a preliminary three-dimensional space model; By real-time analysis of the shape characteristics and spatial positions of objects in the preliminary three-dimensional space model, the positions of charged and grounded objects can be identified and hazardous areas can be marked; Introducing a path optimization algorithm to generate the shortest path based on the point cloud model to avoid dangerous areas; Based on real-time data, the point cloud model is reconstructed in real time and the parameters of the object are updated, and the real-time updated 3D model information is embedded in the WebRTC video stream.
4. The audio and video interaction method in live working scenarios based on WebRTC technology as claimed in claim 3, characterized in that: The dangerous area marking includes obtaining a three-dimensional coordinate point set of the operation scene according to the point cloud data, downsampling the point cloud data using voxel filtering, and retaining key structural features; The RANSAC plane segmentation algorithm is used to extract a large-scale plane according to the normal vector distribution of the point cloud as the reference plane of the ground. i satisfy: |a·x i +b·y i +c·z i +d|<∈ P i belongs to a plane point, where a, b, c represent the components of the plane normal vector, which are calculated by the fitting algorithm; d represents the offset of the plane, which is calculated from the overall distribution of the point cloud data; x i ,y i , z i Point cloud P i The three-dimensional coordinates of the point; The remaining point clouds are grouped using a clustering algorithm (DBSCAN) to identify independent objects in the working environment, calculate geometric features for each cluster grouping, and match the object features with known template data of charged and grounded objects; Using the electric field strength sensor installed on the operator's equipment, the electric field strength data in the working environment is collected in real time, and the electric field strength data is matched with the point cloud coordinates to generate the electric field strength value of each point; Set the electric field strength danger threshold. If any point in the point cloud satisfies E ( x,y,z ) >E threshold , then mark the point as a dangerous point; Use the α-shape algorithm to generate a 3D boundary around all hazard points: Among them, E ( x,y,z ) Indicate point ( x,y,z ) The electric field strength; E threshold Indicates the dangerous electric field strength threshold; represents the three-dimensional boundary generated by the dangerous point; S represents a simple geometric unit in the point set; α represents the shape control parameter; the output three-dimensional boundary is labeled as a visual mark of the dangerous area.
5. The audio and video interaction method in live working scenarios based on WebRTC technology as claimed in claim 4, characterized in that: The path optimization algorithm includes, based on the three-dimensional point cloud data of the working environment and the point set of the dangerous area, gridding the point cloud data into a three-dimensional grid map, the grid information including occupancy marks and danger marks; Get the operator's current position and target point, and optimize the path based on algorithm A: f ( n 0 =g ( n 0 +h ( n ) Heuristic function h ( n ) Using Euclidean distance: Among them, f ( n ) represents the evaluation function; g ( n ) Indicates actual cost; ( x n ,y n , z n) Indicates the coordinates of the target point; ( x g ,y g , z g ) Indicates the current position coordinates; When it is detected that the path is blocked by obstacles or dangerous areas, the A algorithm is called again, and the updated path is generated with the location of the worker as the starting point; Generate a path in the form of a polyline based on a set of path points and perform smoothing: Among them, P raw( i ) represents the original path point; P smooth( i ) Represents the smoothed path point; λ represents the smoothing weight; the smoothed path is annotated in the point cloud model, and the path is projected into the real-time video stream using the WebRTC API.
6. The audio and video interaction method in live working scenarios based on WebRTC technology as claimed in claim 5, characterized in that: The real-time calculation of the safe distance between the operator and the charged object and the grounded object includes calculating the shortest path length from the operator to the dangerous target based on the three-dimensional space model; The path length is dynamically monitored. When the path length is less than the safety threshold, an early warning is triggered and displayed as a visual marker through the WebRTC video stream.
7. The audio and video interaction method in live working scenarios based on WebRTC technology as claimed in claim 6, characterized in that: The trigger warning is displayed in the form of a visual mark through the WebRTC video stream, including: Send voice prompts via WebRTC audio channel; Overlay warning signs on the video stream to mark dangerous areas and dynamically update warning information; Combine voice and video multi-modal prompts to remind operators to adjust operations.
8. An audio and video interaction system for live working scenarios based on WebRTC technology according to any one of the methods of claims 1 to 7, characterized in that: include, The data collection module collects real-time data of workers and working environment, transmits it through WebRTC technology, and uses the lidar module to collect environmental point cloud data; The 3D modeling module generates a 3D spatial model of the operation scene based on the collected real-time data and point cloud data, and overlays the key area annotation information of the 3D spatial model in the WebRTC video stream; The interactive module calculates the safe distance between the operator and the live and grounded objects in real time. When the safe distance is lower than the preset threshold, an active warning is triggered through the audio and video stream.
9. A computer device comprising: Memory and processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, the steps of the audio and video interaction method in the live working scenario based on WebRTC technology as described in any one of claims 1-7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the audio and video interaction method in the live working scenario based on WebRTC technology are implemented as described in any one of claims 1-7.
Citation Information
Cited By
Scene modeling method and electronic equipment
CN121074253A
Electric field operation safety control method and system
CN121147905A
Multimode test body lossless reinjection data synchronization method and device
CN121441925A