Robot video data stream processing method, electronic equipment, medium and device

Virtual video data is generated by using the spatial relationships of the robot's camera and positioning sensors. Only the differential information is transmitted, which solves the robot's video data transmission and storage needs, realizes data compression and load balancing, improves processing efficiency and reduces latency.

CN121691802APending Publication Date: 2026-03-17SHENZHEN ZHIDONG FUTURE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

With the increase in robot camera resolution and data transmission rate, as well as the increase in the number of robots, the demand for data transmission and storage has increased, especially in collaborative operation scenarios where the response latency requirements for data synchronization and processing have increased.

Method used

Virtual video information is generated by using the relative spatial relationship and positioning sensors based on the robot's camera, and differential video data is acquired. This data is then fused and processed in conjunction with a 3D model. Only the differential information is transmitted to compress the data size, and load balancing is achieved between the robot and the host computer.

Benefits of technology

It effectively compresses the size of video data, reduces response latency, improves processing efficiency, and meets the data transmission and storage requirements of collaborative work scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121691802A_ABST
    Figure CN121691802A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a robot video data stream processing method, electronic equipment, a medium and a device. A three-dimensional model of a given scene is combined with a positioning sensor of a first robot, virtual video information corresponding to a plurality of first cameras is simulated and generated, and first fusion virtual video data is acquired, so that only first difference video data representing difference information needs to be transmitted, and the first fusion virtual video data is acquired. In addition, first fused real video data can be obtained in combination with a first relative spatial relationship among a plurality of first cameras, and the scale of the video data is greatly compressed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, electronic device, medium and apparatus for processing robot video data streams. Background Technology

[0002] With the development of robotics, artificial intelligence, and virtual reality technologies, a large number of robots are deployed in application scenarios such as factory assembly lines and school laboratories to perform standardized tasks. Operators can remotely monitor and assist robot operations through virtual reality devices. A robot may be equipped with multiple cameras to collect real-time video data from different angles, enabling scene reconstruction and combining virtual reality and augmented reality technologies to improve the accuracy and efficiency of robot monitoring and assisted operation. However, the increasing resolution and data transmission rate of robot cameras, the growing number of cameras, and the increasing number of robots operating simultaneously have led to a significant increase in data transmission volume, placing higher demands on data transmission and storage. Furthermore, sometimes multiple robots need to work in coordination, such as multiple robots at different stations on an automotive assembly line or in a biochemical material processing flow. This requires synchronous transmission and processing of data from multiple robots while minimizing response latency, further increasing the demands on data transmission and storage.

[0003] To address this technical challenge, this application provides a method, electronic device, medium, and apparatus for processing robot video data streams. Summary of the Invention

[0004] In a first aspect, this application provides a method for processing robot video data streams. The method includes: processing real video information collected by each of the multiple first cameras based on a first relative spatial relationship between multiple first cameras associated with a first robot to obtain first fused real video data, wherein the multiple first cameras are respectively deployed on multiple first parts of the first robot, and the first relative spatial relationship is determined by referring to the physical model and pose detection of the first robot; determining a first position and a first orientation of the first robot in a given scene using positioning sensors deployed on the first robot; then determining the position and orientation of each of the multiple first cameras in the given scene by referring to the physical model and pose detection of the first robot; then generating virtual video information corresponding to each of the multiple first cameras using a 3D model of the given scene; processing the virtual video information corresponding to each of the multiple first cameras based on the first relative spatial relationship to obtain first fused virtual video data; and then obtaining first differential video data between the first fused real video data and the first fused virtual video data, wherein the first differential video data is used for video data transmission and video data storage associated with the first robot.

[0005] Through the first aspect of this application, by utilizing a 3D model of a given scene and combining it with the positioning sensors of a first robot, virtual video information corresponding to multiple first cameras is simulated and generated, and first fused virtual video data is acquired. Therefore, only the first differential video data representing the difference information needs to be transmitted. Furthermore, the first fused real video data can be acquired by combining the first relative spatial relationship between multiple first cameras, significantly compressing the video data size. This helps to achieve effects such as eliminating ghosting, deduplication, enhancing visual effects, or bypassing occlusions. Moreover, it achieves load balancing between the robot and the host computer, utilizing the richer and more economical computing and storage resources of the host computer, effectively improving overall processing efficiency and reducing response latency. This helps to meet the data transmission and data storage requirements of collaborative multi-robot operation scenarios such as factory assembly lines and school laboratories.

[0006] In one possible implementation of the first aspect of this application, the first relative spatial relationship includes a fixed portion and a non-fixed portion.

[0007] In one possible implementation of the first aspect of this application, the fixed portion of the first relative spatial relationship includes the relative spatial relationship between two cameras deployed in front of the head of the first robot and the relative spatial relationship between multiple cameras deployed in front of and behind the head of the first robot, and the non-fixed portion of the first relative spatial relationship includes the relative spatial relationship between multiple cameras deployed in the head, hands and chest of the first robot.

[0008] In one possible implementation of the first aspect of this application, the positioning sensor of the first robot is a lidar.

[0009] In one possible implementation of the first aspect of this application, the processing method is based on a first processor deployed on the first robot and a host processor communicatively connected to the first robot. The first processor is configured to determine the position and orientation of each of the plurality of first cameras in the given scene, send the position and orientation of each of the plurality of first cameras in the given scene to the host processor, and receive the first fused virtual video data from the host processor. The host processor is configured to receive the position and orientation of each of the plurality of first cameras in the given scene from the first processor, acquire the first fused virtual video data, and send the first fused virtual video data to the first processor.

[0010] In one possible implementation of the first aspect of this application, the three-dimensional model of the given scene is pre-collected and established using drones or manual detection methods. The three-dimensional model of the given scene includes the internal dimensions of buildings and the respective dimensions and fixed positional relationships of multiple objects fixed in the given scene. Furthermore, the first differential video data is used to indicate another object that exists in the given scene and is different from the multiple objects.

[0011] In one possible implementation of the first aspect of this application, the given scenario is a factory assembly line or a school laboratory.

[0012] In one possible implementation of the first aspect of this application, generating virtual video information corresponding to each of the plurality of first cameras using a 3D model of the given scene includes: using the position and orientation of each of the plurality of first cameras in the given scene to simulate and generate virtual video information simulated by each of the plurality of first cameras in the 3D model of the given scene.

[0013] In one possible implementation of the first aspect of this application, obtaining the first difference video data between the first fused real video data and the first fused virtual video data includes: using a pixel-level algorithm to remove duplicate pixel information or pixel information with similarity exceeding a preset threshold between the first fused real video data and the first fused virtual video data, thereby obtaining the first difference video data.

[0014] In one possible implementation of the first aspect of this application, obtaining the first differential video data between the first fused real video data and the first fused virtual video data includes: using an image feature extraction algorithm to extract image features of the first fused real video data and the first fused virtual video data in their respective regions of interest, and then identifying the differential image features in the regions of interest through feature comparison, thereby obtaining the first differential video data.

[0015] In one possible implementation of the first aspect of this application, the processing method further includes: reconstructing the first fused real video data using the three-dimensional model of the given scene and the first differential video data.

[0016] In one possible implementation of the first aspect of this application, the processing method further includes: when the motion state of the first robot is detected to be stable, reusing historical video data to replace the first fused real video data to generate the first differential video data, and transmitting the first differential video data to the host side; and when the motion state of the first robot is detected to be unstable, directly transmitting the first fused real video data to the host side.

[0017] In one possible implementation of the first aspect of this application, the processing method further includes: processing the real video information collected by each of the plurality of second cameras based on a second relative spatial relationship between the plurality of second cameras associated with the second robot to obtain second fused real video data, wherein the plurality of second cameras are respectively deployed on a plurality of second parts of the second robot, and the second relative spatial relationship is a relative spatial relationship between the plurality of second parts determined with reference to the physical model and pose detection of the second robot; using positioning sensors deployed on the second robot to determine a second position and a second orientation of the second robot in the given scene, then determining the position and orientation of each of the plurality of second cameras in the given scene with reference to the physical model and pose detection of the second robot, and then using a three-dimensional model of the given scene to generate virtual video information corresponding to each of the plurality of second cameras; processing the virtual video information corresponding to each of the plurality of second cameras based on the second relative spatial relationship to obtain second fused virtual video data, and then obtaining second differential video data between the second fused real video data and the second fused virtual video data, wherein the second differential video data is used for video data transmission and video data storage associated with the second robot.

[0018] In one possible implementation of the first aspect of this application, the processing method further includes: determining the relative position and orientation relationship between the first robot and the second robot based on the first position and the first orientation of the first robot in the given scene and the second position and the second orientation of the second robot in the given scene; and using the second differential video data to deduplicate and correct the first differential video data based on the relative position and orientation relationship between the first robot and the second robot.

[0019] In one possible implementation of the first aspect of this application, the relative position and orientation relationship between the first robot and the second robot indicates that the first robot and the second robot are in a front-back relationship or a left-right relationship, and the first robot and the second robot respectively correspond to different workstations in the given scenario.

[0020] Secondly, embodiments of this application also provide an electronic device, the electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method according to any of the above-mentioned implementations.

[0021] Thirdly, embodiments of this application also provide a computer-readable storage medium storing computer instructions that, when executed on a computer device, cause the computer device to perform a method according to any of the above-described implementations.

[0022] Fourthly, embodiments of this application also provide a computer program product, the computer program product including instructions stored on a computer-readable storage medium, which, when executed on a computer device, cause the computer device to perform a method according to any of the above-described aspects.

[0023] Fifthly, this application also provides a processing apparatus for robot video data streams. The processing apparatus includes: a communication module for communicatively connecting to a host side; and an end processor connected to the communication module. The end processor is used to: process the real video information collected by each of the plurality of first cameras based on a first relative spatial relationship between multiple first cameras associated with a first robot, thereby obtaining first fused real video data, wherein the plurality of first cameras are respectively deployed on multiple first parts of the first robot, and the first relative spatial relationship is a relative spatial relationship between the plurality of first parts determined with reference to the physical model and pose detection of the first robot; determine a first position and a first orientation of the first robot in a given scene using a positioning sensor deployed on the first robot, and then determine the position and orientation of each of the plurality of first cameras in the given scene with reference to the physical model and pose detection of the first robot. Then, the communication module is used to send the position and orientation of each of the plurality of first cameras in the given scene to the host side, and to receive first fused virtual video data from the host side through the communication module. The host side is used to: generate virtual video information corresponding to each of the plurality of first cameras using a 3D model of the given scene, and process the virtual video information corresponding to each of the plurality of first cameras based on the first relative spatial relationship to obtain the first fused virtual video data; and obtain first difference video data between the first fused real video data and the first fused virtual video data, the first difference video data being used for video data transmission and video data storage associated with the first robot.

[0024] Through the fifth aspect of this application, by utilizing a 3D model of a given scene and combining it with the positioning sensors of the first robot, virtual video information corresponding to multiple first cameras is simulated and generated, and first fused virtual video data is acquired. Therefore, only the first differential video data representing the difference information needs to be transmitted. Furthermore, the first fused real video data can be acquired by combining the first relative spatial relationship between the multiple first cameras, significantly compressing the video data size. This helps to achieve effects such as eliminating ghosting, deduplication, enhancing visual effects, or bypassing occlusions. Moreover, it achieves load balancing between the robot and the host computer, utilizing the richer and more economical computing and storage resources of the host computer, effectively improving overall processing efficiency and reducing response latency. This helps to meet the data transmission and data storage requirements of collaborative multi-robot operation scenarios such as factory assembly lines and school laboratories. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A flowchart illustrating a method for processing robot video data streams provided in an embodiment of this application; Figure 2 A schematic diagram of a robot video data stream processing device provided in an embodiment of this application; Figure 3 A schematic diagram illustrating a video data stream processing flow for multiple robots, provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation

[0027] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.

[0028] It should be understood that in the description of this application, "at least one" means one or more, and "multiple" means two or more. In addition, the words "first," "second," etc., unless otherwise stated, are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance or order.

[0029] Figure 1 This is a flowchart illustrating a method for processing robot video data streams provided in an embodiment of this application. Figure 1 As shown, the processing method includes the following steps.

[0030] Step S101: Based on the first relative spatial relationship between the multiple first cameras associated with the first robot, process the real video information collected by each of the multiple first cameras to obtain first fused real video data. The multiple first cameras are respectively deployed on multiple first parts of the first robot. The first relative spatial relationship is the relative spatial relationship between the multiple first parts determined with reference to the physical model and pose detection of the first robot.

[0031] Step S103: Using the positioning sensors deployed on the first robot, determine the first position and first orientation of the first robot in a given scene. Then, refer to the physical model and pose detection of the first robot to determine the position and orientation of each of the plurality of first cameras in the given scene. Then, use the three-dimensional model of the given scene to generate virtual video information corresponding to each of the plurality of first cameras.

[0032] Step S105: Based on the first relative spatial relationship, process the virtual video information corresponding to each of the plurality of first cameras to obtain first fused virtual video data. Then, obtain first difference video data between the first fused real video data and the first fused virtual video data. The first difference video data is used for video data transmission and video data storage associated with the first robot.

[0033] See Figure 1In step S101, based on the first relative spatial relationship between the multiple first cameras associated with the first robot, the real video information collected by each of the multiple first cameras is processed to obtain first fused real video data. The multiple first cameras are respectively deployed on multiple first parts of the first robot, and the first relative spatial relationship is determined with reference to the physical model and pose detection of the first robot. Here, the multiple first parts of the first robot can be the head, chest, hands, etc., and can be configured according to the specific structure and needs of the first robot, such as a full-body humanoid or half-body humanoid. In some embodiments, two cameras can be deployed in front of the head of the first robot, similar to human eyes. This allows for 3D scanning using a binocular algorithm, i.e., collecting depth information from directly in front of the first robot, thus providing a basis for robot operation. In some embodiments, a camera can also be deployed behind the head of the first robot, similar to the back of a human's head, to collect video information from behind the first robot. Additionally, cameras can be deployed on other parts of the first robot, such as the hands and chest, to provide video information from multiple cameras on a single robot. Furthermore, by utilizing the first relative spatial relationship between the multiple first cameras, the real video information captured by each of the multiple first cameras can be processed to obtain first fused real video data. This means that, for example, ghosting can be eliminated, deduplication can be performed, or visual effects can be enhanced. In some embodiments, the first relative spatial relationship includes a fixed relative spatial relationship between the multiple first cameras, such as a fixed relative spatial relationship between two cameras in front of the head of the first robot, or a fixed relative spatial relationship between cameras in front of and behind the head. In some embodiments, the first relative spatial relationship also includes a non-fixed relative spatial relationship between the multiple first cameras, such as a non-fixed relative spatial relationship between cameras on the head and hands of the first robot. As the first robot moves, the relative spatial relationship between its head and hands may change, which in turn causes a change in the relative spatial relationship between the cameras deployed on the head and hands. The first relative spatial relationship is the relative spatial relationship between the multiple first parts determined with reference to the physical model and pose detection of the first robot. In this way, by using the physical model of the first robot, such as structural parameters and part dimensions, as well as pose detection, such as detecting the first robot's bending and kneeling postures, the real video information collected by multiple cameras of a single robot can be processed. For example, by using a binocular algorithm to integrate the video data captured by the two cameras in front of the head into three-dimensional data with depth information, and by performing deduplication and error correction, etc.For example, by using cameras deployed on the head and hands, the influence of obstructions can be effectively bypassed. The head camera might capture real video information containing obstructions, thus failing to effectively identify the target object. However, the hand camera can bypass these obstructions and effectively identify the target object. As another example, using two head-mounted cameras, similar to human eyes, can identify most repetitive image features, allowing for deduplication algorithms to compress the data size to be transmitted. In some embodiments, using the real video information collected by multiple first cameras, image frames with the same timestamp can be fused using image fusion algorithms and timestamps to obtain fused image frames, thereby generating first fused real video data. Thus, based on the first relative spatial relationship between the multiple first cameras associated with the first robot, and referring to the robot's physical model and pose detection, as well as using cameras in specific locations such as the hands, the real video information collected by the multiple first cameras can be processed to obtain first fused real video data. This facilitates the compression of video data size and achieves effects such as ghosting elimination, deduplication, enhanced visual effects, or bypassing obstructions.

[0034] Continue reading Figure 1In step S103, the positioning sensors deployed on the first robot determine the first position and first orientation of the first robot in a given scene. Then, referring to the physical model and pose detection of the first robot, the positions and orientations of the multiple first cameras in the given scene are determined. Next, the 3D model of the given scene is used to generate virtual video information corresponding to each of the multiple first cameras. Here, the positioning sensors can be, for example, LiDAR, used to sense the surrounding environment, thereby accurately locating the position and orientation of the first robot in the given scene. Combined with the physical model and pose detection of the first robot itself, the positions and orientations of the cameras on the first robot in the given scene can be determined. Thus, based on the pre-established 3D model of the given scene, combined with the position and orientation of a specific camera, the image that the specific camera should capture in the 3D model of the given scene according to that position and orientation can be simulated and generated. That is, the 3D model of the given scene is used to generate virtual video information corresponding to each of the multiple first cameras. Here, the virtual video information corresponding to each of the multiple first cameras differs from the real video information collected by each of the multiple first cameras mentioned above. The real video information collected by each of the multiple first cameras is video data actually acquired by a visual sensor, which can include a video data stream composed of image frames acquired at different time points. Therefore, it reflects the surrounding environment as truly perceived by the first robot in the given scene, such as obstacles near the robot. Conversely, the virtual video information corresponding to each of the multiple first cameras is a video data stream composed of image frames virtually generated by an algorithm. It is simulated using a 3D model of the given scene, referencing the actual positions and orientations of the multiple first cameras in the given scene. It should be understood that the entire 3D model of the given scene can be stored on a host or cloud server, and the processing burden of generating the virtual video information corresponding to each of the multiple first cameras can also be distributed to the host or cloud server. In this way, the end processor of the first robot, that is, the computing and storage resources deployed on the first robot itself, only needs to determine the first position and first orientation of the first robot in the given scene through positioning sensors as the first robot moves and operates, then determine the position and orientation of specific cameras by referring to the physical model and pose detection of the first robot, and then send the position and orientation of specific cameras to the host or cloud server. Thus, after receiving the positions and orientations of the plurality of first cameras from the first robot in the given scene, the host or cloud server uses the three-dimensional model of the given scene to generate virtual video information corresponding to each of the plurality of first cameras.This means that the first robot itself only needs to determine the position and orientation of each of the multiple first cameras in the given scene, for example, by generating a series of parameter sequences composed of position and orientation, which greatly reduces the amount of data transmitted from the robot to the host (which may be a local host, computing node, or cloud platform, etc.).

[0035] Continue reading Figure 1In step S105, based on the first relative spatial relationship, the virtual video information corresponding to each of the plurality of first cameras is processed to obtain first fused virtual video data. Then, first differential video data between the first fused real video data and the first fused virtual video data is obtained. The first differential video data is used for video data transmission and video data storage associated with the first robot. Here, the processing burden of obtaining the first fused virtual video data is also distributed to the host or cloud server. Thus, the first robot itself only needs to determine the position and orientation of each of the plurality of first cameras in the given scene, for example, by generating a series of parameter sequences composed of position and orientation, and then sending the position and orientation of each of the plurality of first cameras in the given scene to the host or cloud server. Then, the host or cloud server is used to generate virtual video information corresponding to each of the plurality of first cameras using the 3D model of the given scene, and to process the virtual video information corresponding to each of the plurality of first cameras based on the first relative spatial relationship to obtain the first fused virtual video data. Therefore, from the robot end to the host end, only the first relative spatial relationship between multiple first cameras and the position and orientation of each first camera in the given scene need to be transmitted. This significantly reduces the data scale transmitted from the robot end to the host end (which can be a local host, computing node, or cloud platform, etc.). Furthermore, as mentioned above, the virtual video information corresponding to each of the multiple first cameras is virtually generated image and video information, representing the visual sensors on the robot. In the modeling of the virtually generated given scene, the reference image and video information that should be captured is determined by referring to the actual position and orientation of the real visual sensor on the first robot. Thus, by utilizing the first fused real video data representing the real image acquisition result and the first fused virtual video data representing the virtual image acquisition result, appropriate image processing algorithms can be used to obtain the difference information between the real image acquisition result and the virtual image acquisition result, thereby obtaining the first difference video data. In some embodiments, pixel-level subtraction can be used to treat identical or sufficiently similar pixels as repeats, thereby highlighting pixel-level differences. In some embodiments, the image information contained in the first fused real video data can be used as the real image, and the image information contained in the first fused virtual video data can be used as the reference image. A region of interest (ROI) is first delineated on the real image and the reference image, and then only the ROI is processed. For example, the image features of the real image and the reference image can be extracted by an image feature extraction model, and then the image features with sufficient differences can be identified by feature comparison.Thus, the first differential video data represents differential information, highlighting the changes brought about by the robot's operation—that is, the difference between the situation with and without robot operation. This first differential video data is used for video data transmission and storage associated with the first robot. Therefore, compared to transmitting the complete first fused real video data, only the first differential video data representing the differential information needs to be transmitted, effectively compressing the size of the video data to be transmitted. Thus, from the robot end to the host end, only the first relative spatial relationship between the multiple first cameras and the position and orientation of each of the multiple first cameras in the given scene need to be transmitted. Furthermore, from the host end to the robot end, only the first fused virtual video data needs to be transmitted. Because the first fused virtual video data is generated using an algorithm, compared to the first fused real video data based on the real video information collected by each of the multiple first cameras, it can achieve an excellent compression ratio while maintaining no information loss, and the data size can be further compressed through algorithm optimization and data format optimization. Therefore, by utilizing the aforementioned robot video data stream processing method, the transmission of difficult-to-compress real video data from the robot to the host computer is avoided. Instead, the robot sends a series of parameters (the first relative spatial relationship between multiple first cameras and the position and orientation of each first camera in the given scene) to the host computer. The host computer then uses its internal model to generate virtual video data suitable for compression and transmission. The robot then uses this virtual video data to extract discrepancies from the real video data. Finally, only the robot needs to transmit the first discrepancy video data to the host computer. At the host computer, using a pre-modeled 3D model of the given scene, the first fused real video data can be reconstructed based on the first discrepancy video data. This avoids transmitting the first fused real video data while simultaneously reproducing it on the host computer. Thus, using the first fused real video data, monitoring and assistance of robot operations can be effectively achieved, and it can also be combined with virtual reality and augmented reality technologies to enable remote monitoring by operators. As the resolution and data transmission rate of cameras on robots increase, the number of cameras increases, and the number of robots operating simultaneously increases, the scale of real video data increases. However, by using the robot video data stream processing method described above, the overall data scale to be transmitted is effectively compressed, and the processing load balancing between the robot end and the host end is achieved. The richer and more economical computing and storage resources on the host end can be utilized, effectively improving the overall processing efficiency and reducing response latency. This helps to meet the data transmission and data storage requirements of application scenarios such as factory assembly lines and school laboratories where multiple robots work together.

[0036] Continue reading Figure 1By referencing the robot's motion state during operation, a specific transmission mode can be set accordingly. For example, the robot might be in a relatively stable state, such as operating equipment at a specific workstation on an assembly line or operating a control panel in front of specific equipment in a laboratory. In this case, when the robot's positioning is relatively stable, image subtraction processing is used to keep the overall image data stream's burden on the system relatively light. However, when the robot's positioning is unstable, such as when the robot is continuously moving or its center of gravity is repeatedly shifting, a direct transmission mode can be switched to, where the camera directly transmits the actual image without subtraction processing. This is because when the robot is in a relatively stable state, reference images can be reused, as the camera's position and orientation may remain unchanged or change very little over a period of time.

[0037] In short, Figure 1 The robot video data stream processing method shown utilizes a 3D model of a given scene, combined with the positioning sensors of the first robot, to simulate and generate virtual video information corresponding to multiple first cameras and acquire first fused virtual video data. Therefore, only the first differential video data representing the differences needs to be transmitted. Furthermore, the first fused real video data can be obtained by combining the first relative spatial relationship between multiple first cameras, significantly compressing the video data size. This helps to achieve effects such as ghosting elimination, deduplication, enhanced visual effects, and bypassing occlusions. Moreover, it achieves load balancing between the robot and the host computer, utilizing the richer and more economical computing and storage resources of the host computer, effectively improving overall processing efficiency and reducing response latency. This helps meet the data transmission and data storage requirements of collaborative multi-robot operations, such as factory assembly lines and school laboratories.

[0038] See Figure 1 In one possible implementation, the first relative spatial relationship includes both fixed and non-fixed portions. This facilitates adaptation to the specific design and construction of the first robot, which may include any number of movable joints.

[0039] In some embodiments, the fixed portion of the first relative spatial relationship includes the relative spatial relationship between two cameras deployed in front of the head of the first robot and the relative spatial relationship between multiple cameras deployed in front of and behind the head of the first robot. The non-fixed portion of the first relative spatial relationship includes the relative spatial relationship between multiple cameras deployed on the head, hands, and chest of the first robot. Thus, in some embodiments, two cameras can be deployed in front of the head of the first robot, similar to human eyes. This allows for 3D scanning using a binocular algorithm, i.e., acquiring depth information directly in front of the first robot, thereby providing a basis for robot operation. In some embodiments, a camera can also be deployed behind the head of the first robot, similar to the back of a human's head, to acquire video information from behind the first robot. Additionally, cameras can be deployed on other parts of the first robot, such as the hands and chest, to provide video information from multiple cameras on a single robot. Furthermore, by utilizing the first relative spatial relationship between the multiple first cameras, the real video information acquired by each of the multiple first cameras can be processed to obtain first fused real video data. This means that, for example, ghosting can be eliminated, deduplication can be performed, or visual effects can be enhanced. In some embodiments, the first relative spatial relationship includes a fixed relative spatial relationship between multiple first cameras, such as a fixed relative spatial relationship between two cameras in front of the head of the first robot, or a fixed relative spatial relationship between cameras in front of and behind the head. In some embodiments, the first relative spatial relationship also includes a non-fixed relative spatial relationship between multiple first cameras, such as a non-fixed relative spatial relationship between cameras on the head and hands of the first robot. As the first robot moves, the relative spatial relationship between its head and hands may change, which in turn causes a change in the relative spatial relationship between the cameras deployed on the head and hands. The first relative spatial relationship is the relative spatial relationship between the multiple first parts determined with reference to the physical model of the first robot and pose detection. Thus, by utilizing the physical model of the first robot, such as structural parameters and part dimensions, and pose detection, such as detecting the first robot's bending posture, knee bending, etc., the real video information collected by multiple cameras of a single robot can be processed. For example, a binocular algorithm can be used to integrate the video data captured by the two cameras in front of the head into three-dimensional data with depth information, and deduplication and error correction can be performed. For example, by using cameras deployed on the head and hands, the influence of obstructions can be effectively bypassed. The real video information captured by the head camera may contain obstructions, thus failing to effectively identify the target object. However, by using the hand camera, the influence of obstructions can be bypassed, thus effectively identifying the target object.For example, by using two cameras deployed on the head, similar to human eyes, most of the repetitive image features can be identified, and deduplication algorithms can be used to compress the size of the data to be transmitted.

[0040] In one possible implementation, the positioning sensor of the first robot is a LiDAR (Light Detection and Ranging) sensor. Thus, LiDAR can be used to acquire positioning information. The end-processor of the first robot, i.e., the computing and storage resources deployed on the first robot itself, only needs to determine the first robot's initial position and orientation in a given scene through the positioning sensor as the first robot moves and operates. Then, it needs to determine the position and orientation of a specific camera by referring to the first robot's physical model and pose detection, and then send the position and orientation of the specific camera to the host or cloud server. This helps to compress the size of the data to be transmitted.

[0041] In one possible implementation, the processing method is based on a first processor deployed on the first robot and a host processor communicatively connected to the first robot. The first processor determines the position and orientation of each of the plurality of first cameras in the given scene, sends the position and orientation of each of the plurality of first cameras in the given scene to the host processor, and receives the first fused virtual video data from the host processor. The host processor, in turn, receives the position and orientation of each of the plurality of first cameras in the given scene from the first processor, acquires the first fused virtual video data, and sends the first fused virtual video data to the first processor. Thus, compared to transmitting complete first fused real video data, only first differential video data representing differential information needs to be transmitted, effectively compressing the size of the video data to be transmitted. Therefore, from the robot end to the host end, only the first relative spatial relationship between the plurality of first cameras and the position and orientation of each of the plurality of first cameras in the given scene need to be transmitted; and from the host end to the robot end, only the first fused virtual video data needs to be transmitted. Because the first fused virtual video data is generated using an algorithm, compared to the first fused real video data based on real video information collected by multiple first cameras individually, it has an excellent compression ratio while maintaining no information loss. Furthermore, the data size can be further compressed through algorithm optimization and data format optimization. Therefore, using the above-described robot video data stream processing method avoids transmitting difficult-to-compress real video data from the robot to the host. The robot sends a series of parameters (the first relative spatial relationship between multiple first cameras and the position and orientation of each first camera in the given scene) to the host. The host then uses its internal model to generate virtual video data suitable for compression and transmission. The robot then uses the virtual video data to extract the difference information from the real video data. Finally, only the robot needs to transmit the first difference video data to the host. Finally, on the host, using a pre-modeled 3D model of the given scene, the first fused real video data can be reconstructed based on the first difference video data. This avoids transmitting the first fused real video data and also enables its reproduction on the host. In this way, by utilizing the first fusion of real video data, it is possible to effectively monitor and assist robot operations, and also to combine virtual reality and augmented reality technologies to enable remote monitoring by operators.As the resolution and data transmission rate of cameras on robots increase, the number of cameras increases, and the number of robots operating simultaneously increases, the scale of real video data increases. However, by using the robot video data stream processing method described above, the overall data scale to be transmitted is effectively compressed, and the processing load balancing between the robot end and the host end is achieved. The richer and more economical computing and storage resources on the host end can be utilized, effectively improving the overall processing efficiency and reducing response latency. This helps to meet the data transmission and data storage requirements of application scenarios such as factory assembly lines and school laboratories where multiple robots work together.

[0042] In one possible implementation, the 3D model of the given scene is pre-collected and established using drones or manual inspection methods. The 3D model includes the internal dimensions of buildings and the dimensions and fixed positional relationships of multiple objects fixed within the given scene. Furthermore, the first differential video data is used to indicate another object present in the given scene that differs from the multiple objects. Thus, the multiple objects can correspond to devices, structures, obstacles, etc., fixed within the given scene. By highlighting the differential information, another object present in the given scene that differs from the multiple objects can be identified, such as movable consumables used in robot operations, or other robots, which helps to reduce the data transmission scale.

[0043] In one possible implementation, the given scenario is a factory assembly line or a school laboratory. This helps to meet the data transmission and data storage requirements of application scenarios involving the coordinated operation of multiple robots, such as factory assembly lines and school laboratories.

[0044] In one possible implementation, generating virtual video information corresponding to each of the plurality of first cameras using a 3D model of the given scene includes: simulating virtual video information captured by each of the plurality of first cameras within the 3D model of the given scene using the positions and orientations of the respective cameras within the given scene. Thus, based on a pre-established 3D model of the given scene, combined with the position and orientation of a specific camera, it is possible to simulate and generate the image that the specific camera should capture within the 3D model of the given scene according to its position and orientation; that is, to generate virtual video information corresponding to each of the plurality of first cameras using the 3D model of the given scene. Here, the virtual video information corresponding to each of the plurality of first cameras differs from the actual video information captured by each of the plurality of first cameras mentioned above. The actual video information captured by each of the plurality of first cameras is video data actually acquired by a visual sensor, which may include a video data stream composed of image frames acquired at different time points, thus reflecting the surrounding environment actually perceived by the first robot in the given scene, such as obstacles near the first robot. In contrast, the virtual video information corresponding to each of the multiple first cameras is a video data stream composed of image frames virtually generated by an algorithm. This data is simulated using a 3D model of the given scene, referencing the actual positions and orientations of each of the multiple first cameras within that scene. Because the first fused virtual video data is generated using an algorithm, compared to the first fused real video data based on the actual video information collected by the multiple first cameras, it achieves an excellent compression ratio while maintaining information integrity. Furthermore, the data size can be further compressed through algorithm optimization and data format optimization.

[0045] In one possible implementation, obtaining the first difference video data between the first fused real video data and the first fused virtual video data includes: using a pixel-level algorithm to remove duplicate pixel information or pixel information with a similarity exceeding a preset threshold between the first fused real video data and the first fused virtual video data, thereby obtaining the first difference video data. Thus, by utilizing the first fused real video data representing the real image acquisition result and the first fused virtual video data representing the virtual image acquisition result, a suitable image processing algorithm can be used to obtain the difference information between the real image acquisition result and the virtual image acquisition result, thereby obtaining the first difference video data.

[0046] In one possible implementation, acquiring the first difference video data between the first fused real video data and the first fused virtual video data includes: using an image feature extraction algorithm to extract image features of each of the first fused real video data and the first fused virtual video data in a region of interest; then, identifying the difference image features in the region of interest through feature comparison, thereby obtaining the first difference video data. Thus, by utilizing the first fused real video data representing the real image acquisition result and the first fused virtual video data representing the virtual image acquisition result, a suitable image processing algorithm can be used to obtain the difference information between the real image acquisition result and the virtual image acquisition result, thereby acquiring the first difference video data.

[0047] In one possible implementation, the processing method further includes: reconstructing the first fused real video data using a 3D model of the given scene and the first differential video data. Thus, on the host side, using a pre-modeled 3D model of the given scene, the first fused real video data can be reconstructed based on the first differential video data. This avoids transmitting the first fused real video data while simultaneously enabling its reproduction on the host side. Therefore, using the first fused real video data, monitoring and assistance of robot operations can be effectively achieved, and it can also be combined with virtual reality and augmented reality technologies to enable remote monitoring by operators.

[0048] In one possible implementation, the processing method further includes: when the motion state of the first robot is detected to be stable, reusing historical video data to replace the first fused real video data to generate the first differential video data, and transmitting the first differential video data to the host side; and when the motion state of the first robot is detected to be unstable, directly transmitting the first fused real video data to the host side. Thus, by referring to the robot's motion state during operation, a specific transmission mode can be set accordingly. For example, the first robot may be in a relatively stable state, such as operating equipment at a specific workstation on an assembly line, or operating a control panel in front of specific equipment in a laboratory. Thus, when the robot's positioning is relatively stable, image subtraction processing is used, which keeps the overall image data stream's burden on the system relatively light. When the robot's positioning is unstable, such as when the robot is continuously moving, or when the robot's center of gravity is repeatedly changing, a direct transmission mode can be switched, with the camera directly transmitting the real image without subtraction processing. This is because when the robot is in a relatively stable state, reference images can be reused, since the camera's position and orientation may remain unchanged or change very little over a period of time.

[0049] In one possible implementation, the processing method further includes: processing the real video information collected by each of the plurality of second cameras based on a second relative spatial relationship between the plurality of second cameras associated with the second robot to obtain second fused real video data, wherein the plurality of second cameras are respectively deployed on a plurality of second parts of the second robot, and the second relative spatial relationship is the relative spatial relationship between the plurality of second parts determined with reference to the physical model and pose detection of the second robot; using positioning sensors deployed on the second robot to determine a second position and a second orientation of the second robot in the given scene, then determining the position and orientation of each of the plurality of second cameras in the given scene with reference to the physical model and pose detection of the second robot, and then using a three-dimensional model of the given scene to generate virtual video information corresponding to each of the plurality of second cameras; processing the virtual video information corresponding to each of the plurality of second cameras based on the second relative spatial relationship to obtain second fused virtual video data, and then obtaining second differential video data between the second fused real video data and the second fused virtual video data, wherein the second differential video data is used for video data transmission and video data storage associated with the second robot. Thus, referring to the above-described robot video data stream processing method, second differential video data can be obtained for the video data transmission and storage associated with the second robot. Furthermore, the processing method for the video data streams associated with the first and second robots can be easily extended to applications where any number of robots operate in the same given scenario, significantly reducing the data transmission scale while maintaining necessary information. This effectively improves overall processing efficiency and reduces response latency, helping to meet the data transmission and storage requirements of collaborative multi-robot operation scenarios such as factory assembly lines and school laboratories.

[0050] In some embodiments, the processing method further includes: determining the relative position and orientation relationship between the first robot and the second robot based on the first position and first orientation of the first robot in the given scene and the second position and second orientation of the second robot in the given scene; and using the second differential video data to deduplicate and correct the first differential video data based on the relative position and orientation relationship between the first robot and the second robot. Thus, based on the above-described optimized design for video data transmission from multiple cameras for a single robot, further optimization is made here for the case of multiple robots operating synchronously. For example, in a factory assembly line with multiple workstations, each with a robot, the overlap and correlation between the video data captured by adjacent robots can be utilized. The relative positional relationships between adjacent robots, such as front-back or left-right relationships, can be referenced. This allows for further operations such as deduplication, error correction, and improved accuracy using the video data output by each adjacent robot. For example, adjacent robots can be considered as two objects with a relatively fixed spatial relationship, which can further compress the overall video data size. In addition, the relative orientation between the camera behind the head of the previous robot and the camera in front of the head of the next robot can be used to perform processes such as error correction and mutual calibration in combination with the 3D model of the given scene.

[0051] In some embodiments, the relative position and orientation between the first robot and the second robot indicate that they are in a front-to-back or left-to-right relationship, and that the first robot and the second robot correspond to different workstations in the given scenario. This further optimizes the simultaneous operation of multiple robots by utilizing the relative positional relationships between adjacent robots, which helps in operations such as deduplication, error correction, and improved accuracy.

[0052] Figure 2 This is a schematic diagram of a robot video data stream processing device provided in an embodiment of this application. Figure 2As shown, the processing device 200 includes: a communication module 201 for communicatively connecting to the host side; and an end processor 203 connected to the communication module 201. The end processor 203 is used to: process the real video information collected by each of the multiple first cameras based on a first relative spatial relationship between multiple first cameras associated with the first robot, thereby obtaining first fused real video data, wherein the multiple first cameras are respectively deployed on multiple first parts of the first robot, and the first relative spatial relationship is determined by referring to the physical model and pose detection of the first robot; using positioning sensors deployed on the first robot, determine the first position and first orientation of the first robot in a given scene, and then, referring to the physical model and pose detection of the first robot, determine the position and orientation of each of the multiple first cameras in the given scene, and then... Then, the communication module 201 transmits the position and orientation of each of the plurality of first cameras in the given scene to the host side, and receives first fused virtual video data from the host side through the communication module 201, wherein the host side is configured to: generate virtual video information corresponding to each of the plurality of first cameras using a 3D model of the given scene, and process the virtual video information corresponding to each of the plurality of first cameras based on the first relative spatial relationship, thereby obtaining the first fused virtual video data; and obtain first difference video data between the first fused real video data and the first fused virtual video data, the first difference video data being used for video data transmission and video data storage associated with the first robot.

[0053] Figure 2 The processing device 200 shown utilizes a 3D model of a given scene, combined with the positioning sensors of the first robot, to simulate and generate virtual video information corresponding to multiple first cameras and acquire first fused virtual video data. Therefore, it only needs to transmit the first differential video data representing the differences. Furthermore, it can combine the first relative spatial relationship between the multiple first cameras to acquire the first fused real video data, significantly compressing the video data size. This helps to achieve effects such as eliminating ghosting, deduplication, enhancing visual effects, or bypassing occlusions. Moreover, it achieves load balancing between the robot and the host computer, utilizing the richer and more economical computing and storage resources of the host computer, effectively improving overall processing efficiency and reducing response latency. This helps meet the data transmission and data storage requirements of collaborative multi-robot operation scenarios such as factory assembly lines and school laboratories.

[0054] Figure 3This is a schematic diagram illustrating a video data stream processing flow for multiple robots, provided in an embodiment of this application. (See attached diagram.) Figure 3 As shown, three robots operate in the same given scene. Each robot has multiple parts, on which cameras are deployed. Robot A310 includes front head A312 and rear head A314, robot B320 includes front head B322 and rear head B324, and robot C330 includes front head C332 and rear head C334. Further optimizations are made for the scenario of multiple robots operating synchronously. For example, in a factory assembly line with multiple workstations, each with a robot, the overlap and correlation between the video data captured by adjacent robots can be utilized. The relative positional relationships between adjacent robots, such as front-to-back or left-to-right relationships, can be referenced. This allows for further operations such as deduplication, error correction, and improved accuracy using the video data output by each adjacent robot. For instance, adjacent robots can be considered as two objects with a relatively fixed spatial relationship, further compressing the overall video data size. Additionally, the relative orientation between the camera behind the head of the preceding robot and the camera in front of the head of the following robot can be used to perform processes such as error correction and mutual calibration, combined with 3D modeling of a given scene. Figure 3 For example, arrows indicate how to utilize the relative positional relationships between adjacent robots to optimize the video data output by each robot. For instance, video data captured by the camera deployed at B322 in front of robot B320's head can be used together with video data captured by the camera deployed at A312 in front of robot A310's head for optimization, due to the overlap and correlation between the video data. Similarly, the positions of A314 behind robot A310's head and B324 behind robot B320's head, B324 behind robot B320's head and C334 behind robot C330's head, and C332 in front of robot C330's head and B322 in front of robot B320's head can all be used for optimization in multi-robot operation applications. Thus, further optimization for simultaneous multi-robot operations, utilizing the relative positional relationships between adjacent robots, helps with operations such as deduplication, error correction, and improved accuracy.

[0055] Figure 4This is a schematic diagram of a computing device 400 provided in an embodiment of this application. The computing device 400 includes one or more processors 410, a communication interface 420, and a memory 430. The processors 410, communication interface 420, and memory 430 are interconnected via a bus 440. Optionally, the computing device 400 may further include an input / output interface 450, which is connected to input / output devices for receiving user-set parameters, etc. The computing device 400 can be used to implement some or all of the functions of the device embodiment or system embodiment described above in this application; the processor 410 can also be used to implement some or all of the operation steps of the method embodiment described above in this application. For example, the specific implementation of various operations performed by the computing device 400 can be referred to the specific details in the above embodiments, such as the processor 410 being used to execute some or all of the steps or operations in the above method embodiments. For example, in the embodiments of this application, the computing device 400 can be used to implement some or all of the functions of one or more components in the above-described device embodiments. In addition, the communication interface 420 can be used specifically for communication functions necessary to implement the functions of these devices and components, and the processor 410 can be used specifically for processing functions necessary to implement the functions of these devices and components.

[0056] It should be understood that, Figure 4 The computing device 400 may include one or more processors 410, and the multiple processors 410 may collaboratively provide processing power in a parallel connection mode, a serial connection mode, a serial-parallel connection mode, or an arbitrary connection mode; or the multiple processors 410 may form a processor sequence or a processor array; or the multiple processors 410 may be divided into a main processor and an auxiliary processor; or the multiple processors 410 may have different architectures, such as adopting a heterogeneous computing architecture. Furthermore, Figure 4 The structural and functional descriptions of the computing device 400 shown are exemplary and non-limiting. In some exemplary embodiments, the computing device 400 may include... Figure 4 The diagram shows more or fewer components, or combinations of some components, or splitting of some components, or different arrangements of components.

[0057] The processor 410 can have various specific implementations. For example, it can include one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), or a data processing unit (DPU). This application embodiment does not impose specific limitations. The processor 410 can also be a single-core or multi-core processor. The processor 410 can be a combination of a CPU and hardware chips. The aforementioned hardware chips can be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The aforementioned PLDs can be complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), or any combination thereof. The processor 410 can also be implemented using logic devices with built-in processing logic, such as FPGAs or digital signal processors (DSPs). The communication interface 420 can be a wired interface or a wireless interface, used to communicate with other modules or devices. The wired interface can be an Ethernet interface, a local interconnect network (LIN), etc., and the wireless interface can be a cellular network interface or a wireless LAN interface, etc.

[0058] Memory 430 may be non-volatile memory, such as read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Memory 430 may also be volatile memory, which may be random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM). The memory 430 can also be used to store program code and data, so that the processor 410 can call the program code stored in the memory 430 to execute some or all of the operation steps in the above method embodiments, or to execute the corresponding functions in the above device embodiments. Furthermore, the computing device 400 may include, compared to... Figure 4 The number of components displayed may be more or less, or there may be different component configurations.

[0059] Bus 440 can be a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc. Bus 440 can be divided into address bus, data bus, control bus, etc. In addition to the data bus, bus 440 can also include a power bus, control bus, and status signal bus. However, for clarity, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0060] The methods and devices provided in this application are based on the same inventive concept. Since the principles by which the methods and devices solve problems are similar, the embodiments, implementation methods, examples, or methods of implementation of the methods and devices can be referred to each other, and repeated details will not be repeated. This application also provides a system comprising multiple computing devices, the structure of each computing device of which can refer to the structure of the computing devices described above. The functions or operations achievable by this system can refer to the specific implementation steps in the above method embodiments and / or the specific functions described in the above device embodiments, and will not be repeated here.

[0061] This application also provides a computer-readable storage medium storing computer instructions. When these computer instructions are executed on a computer device (such as one or more processors), they can implement the method steps described in the above method embodiments. The specific implementation of the above method steps by the processor of the computer-readable storage medium can refer to the specific operations described in the above method embodiments and / or the specific functions described in the above device embodiments, and will not be repeated here.

[0062] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. This application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Embodiments of this application can be implemented wholly or partially by software, hardware, firmware, or any other combination. When implemented in software, the above embodiments can be implemented wholly or partially as a computer program product. This application can take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless network communication, microwave, etc.) means. Computer-readable storage media can be any available medium that a computer can access, or a data storage device such as a server or data center that contains one or more sets of available media. Available media can be magnetic media (such as floppy disks, hard disks, and magnetic tapes), optical media, or semiconductor media. Semiconductor media can be solid-state drives, random access memory, flash memory, read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, or any other suitable form of storage medium.

[0063] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. Each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0064] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. The steps in the methods of the embodiments of this application can be adjusted in order, combined, or deleted according to actual needs; the modules in the systems of the embodiments of this application can be divided, combined, or deleted according to actual needs. If these modifications and variations of the embodiments of this application fall within the scope of the claims of this application and their equivalents, then this application also intends to include these modifications and variations.

Claims

1. A method of processing a robotic video data stream, the method comprising: The processing method comprises: processing real video information collected by each of a plurality of first cameras associated with a first robot based on a first relative spatial relationship between the plurality of first cameras, wherein the plurality of first cameras are respectively disposed at a plurality of first parts of the first robot, and the first relative spatial relationship is a relative spatial relationship between the plurality of first parts determined with reference to a physical model and pose detection of the first robot; determining a first position and a first orientation of the first robot in a given scene by using a positioning sensor disposed on the first robot, then determining a position and an orientation of each of the plurality of first cameras in the given scene with reference to the physical model and the pose detection of the first robot, and then generating virtual video information corresponding to each of the plurality of first cameras by using a three-dimensional model of the given scene; processing the virtual video information corresponding to each of the plurality of first cameras based on the first relative spatial relationship, thereby obtaining first fused virtual video data, then obtaining first difference video data between the first fused real video data and the first fused virtual video data, and the first difference video data is used for video data transmission and video data storage associated with the first robot.

2. The treatment method according to claim 1, characterized in that, The first relative spatial relationship comprises a fixed part and a non-fixed part.

3. The treatment method according to claim 2, characterized in that, The fixed part of the first relative spatial relationship comprises a relative spatial relationship between two cameras disposed in front of a head of the first robot and a relative spatial relationship between a plurality of cameras disposed in front of and behind the head of the first robot, and the non-fixed part of the first relative spatial relationship comprises a relative spatial relationship between a plurality of cameras disposed on the head, the hand and the chest of the first robot.

4. The treatment method of claim 1, wherein The positioning sensor of the first robot is a laser radar.

5. The treatment method of claim 1, wherein The processing method is implemented based on a first processor disposed on the first robot and a host processor communicatively connected with the first robot, wherein the first processor is configured to determine the position and the orientation of each of the plurality of first cameras in the given scene, send the position and the orientation of each of the plurality of first cameras in the given scene to the host processor, and receive the first fused virtual video data from the host processor, and the host processor is configured to receive the position and the orientation of each of the plurality of first cameras in the given scene from the first processor, obtain the first fused virtual video data, and send the first fused virtual video data to the first processor.

6. The treatment method of claim 1, wherein The three-dimensional model of the given scene is pre-acquired and established by means of unmanned aerial vehicles or artificial detection means, and includes the internal dimensions of the building and the dimensions and fixed position relationships of the plurality of objects in the given scene, and the first difference video data is used to indicate the presence of another object in the given scene that is different from the plurality of objects.

7. The treatment method of claim 1, wherein The given scene is a factory assembly line or a school laboratory.

8. The treatment method of claim 1, wherein, Generating the virtual video information corresponding to each of the plurality of first cameras by means of the three-dimensional model of the given scene includes: Simulating the virtual video information obtained by simulating the shooting of each of the plurality of first cameras in the three-dimensional model of the given scene by means of the position and orientation of each of the plurality of first cameras in the given scene.

9. The treatment method of claim 1, wherein, Obtaining the first difference video data between the first fused real video data and the first fused virtual video data includes: Removing the repeated pixel information or pixel information with a similarity exceeding a preset threshold between the first fused real video data and the first fused virtual video data by means of a pixel-level algorithm, thereby obtaining the first difference video data.

10. The treatment method of claim 1, wherein Obtaining the first difference video data between the first fused real video data and the first fused virtual video data includes: Extracting the image features in the region of interest of each of the first fused real video data and the first fused virtual video data by means of an image feature extraction algorithm, and then identifying the difference image features in the region of interest by feature comparison, thereby obtaining the first difference video data.

11. The treatment method of claim 1, wherein, The processing method further includes: Reconstructing the first fused real video data by means of the three-dimensional model of the given scene and the first difference video data.

12. The treatment method of claim 1, wherein, The processing method further includes: When it is detected that the motion state of the first robot is a stable state, multiplexing historical video data to replace the first fused real video data for generating the first difference video data, and transmitting the first difference video data to the host side, and when it is detected that the motion state of the first robot is a non-stable state, directly transmitting the first fused real video data to the host side.

13. The treatment method of claim 1, wherein, The processing method further includes: Processing the real video information collected by each of the plurality of second cameras associated with the second robot based on a second relative spatial relationship between the plurality of second cameras, thereby obtaining second fused real video data, wherein the plurality of second cameras are respectively disposed at a plurality of second parts of the second robot, and the second relative spatial relationship is a relative spatial relationship between the plurality of second parts determined with reference to the physical model and the pose detection of the second robot. determining a second position and a second orientation of the second robot in the given scene by using a positioning sensor deployed on the second robot, and then determining the position and the orientation of each of the plurality of second cameras in the given scene with reference to the physical model and the pose detection of the second robot, and then generating the corresponding virtual video information of each of the plurality of second cameras by using the three-dimensional model of the given scene; processing the corresponding virtual video information of each of the plurality of second cameras based on the second relative spatial relationship, thereby obtaining second fused virtual video data, and then obtaining second difference video data between the second fused real video data and the second fused virtual video data, the second difference video data being used for video data transmission and video data storage associated with the second robot.

14. The processing method according to claim 13, characterized by, The processing method further comprises: determining a relative position and orientation relationship between the first robot and the second robot based on the first position and the first orientation of the first robot in the given scene and the second position and the second orientation of the second robot in the given scene; de-duplicating and error correcting the first difference video data by using the second difference video data based on the relative position and orientation relationship between the first robot and the second robot.

15. The processing method according to claim 14, characterized in that, The relative position and orientation relationship between the first robot and the second robot indicates that the first robot and the second robot conform to a front-rear relationship or a left-right relationship, and the first robot and the second robot correspond to different stations of the given scene, respectively.

16. An electronic device, comprising: The electronic device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the method according to any one of claims 1 to 15 when executing the computer program.

17. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and when the computer instructions are executed on a computer device, the computer device executes the method according to any one of claims 1 to 15.

18. An apparatus for processing a robotic video data stream, the apparatus comprising: The processing apparatus comprises: a communication module configured to be communicatively connected with a host side; an end processor connected with the communication module, the end processor being configured to: process real video information collected by each of a plurality of first cameras associated with a first robot based on a first relative spatial relationship between the plurality of first cameras, thereby obtaining first fused real video data, wherein the plurality of first cameras are respectively deployed at a plurality of first positions of the first robot, and the first relative spatial relationship is a relative spatial relationship between the plurality of first positions determined with reference to a physical model and a pose detection of the first robot. determining a first position and a first orientation of the first robot in a given scene by using a positioning sensor deployed on the first robot, then determining positions and orientations of the plurality of first cameras in the given scene with reference to a physical model of the first robot and pose detection, then sending the positions and orientations of the plurality of first cameras in the given scene to the host side through the communication module, and receiving first fused virtual video data from the host side through the communication module, wherein the host side is configured to generate virtual video information corresponding to each of the plurality of first cameras by using a three-dimensional model of the given scene, and process the virtual video information corresponding to each of the plurality of first cameras based on the first relative spatial relationship, thereby obtaining the first fused virtual video data; obtaining first difference video data between the first fused real video data and the first fused virtual video data, the first difference video data being used for video data transmission and video data storage associated with the first robot.