Video remote transmission method and system
By realizing remote sharing of the field of view images at the operating end, multi-person voice calls and real-time three-dimensional model sharing in the video remote transmission system, the problem of low efficiency in video remote communication guidance and maintenance in the prior art is solved, and the efficiency and reliability of remote guidance are improved.
Patent Information
- Application Number
- CN202510117330.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the simulation of maintenance and disassembly equipment is concentrated on single-person operation simulation. The visualization method of the maintenance and disassembly process is single, which makes it difficult for video connection experts to observe the scope of maintenance and disassembly, resulting in low efficiency of video remote communication guiding maintenance and disassembly.
By obtaining the video data corresponding to the field of view of the operation end and sending it to the expert end, multi-person voice calls are built based on the expert end and the operation end. The expert end builds a three-dimensional model in real time and shares it to the operation end, realizing remote field of view sharing, multi-person voice calls and 3D system operations.
It improves the efficiency of video remote communication guidance, maintenance and disassembly, realizes remote vision sharing, multi-person voice calls and 3D system operations, and ensures the reliability and security of data transmission.
Smart Images

Figure CN120034614A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data transmission, and in particular to a method and system for remote video transmission. Background Art
[0002] With the high-tech development of equipment, its structure is becoming increasingly complex and its functions are becoming more powerful. Various information technologies and intelligent technologies are used throughout it, making its fault diagnosis and maintenance more complicated and specialized. The structure of modern equipment is becoming increasingly complex and the degree of automation is becoming higher and higher. Many equipment integrate many advanced technologies such as machinery, electronics, automatic control, and computers. Various components in the equipment are interrelated and interdependent, which greatly increases the difficulty of equipment fault diagnosis.
[0003] As shown in Chinese patent CN101411148A, a method for promoting the effective use of wireless capacity in a wireless communication system is provided, and the wireless communication system includes a communication station operable to receive through multiple sub-channels of the wireless communication system. At this point, the data in packet format has an overall sequence, and it is divided into a plurality of parts that can be received by the communication station through corresponding sub-channels, wherein each part has a corresponding link sequence. The method includes detecting missing packet format data at the communication station based on the overall sequence or at least one link sequence. Then, when the missing packet format data is detected based on the overall sequence, a timing time period is set, and when the communication station fails to receive the missing packet format data before the timing step times out, a request is made to retransmit the missing packet format data.
[0004] However, in the prior art, the simulation of maintenance and disassembly of equipment is focused on the operation simulation of a single person, and the visualization method of the maintenance and disassembly process is single, which makes it difficult for experts connected by video to fully observe the scope of maintenance and disassembly, resulting in low efficiency of video remote communication guidance of maintenance and disassembly. Summary of the invention
[0005] 1. Problem to be solved
[0006] Based on this, it is necessary to provide a video remote transmission method and system that can improve the efficiency of video remote communication guidance maintenance and disassembly in response to the above technical problems.
[0007] 2. Technical solution
[0008] In a first aspect, the present application provides a video remote transmission method. The method comprises:
[0009] Obtain the video data corresponding to the visual field image of the operator end and send it to the expert end;
[0010] Build multi-person voice calls based on the expert side and the operation side;
[0011] The expert side builds the 3D model in real time and shares it to the operator side.
[0012] In one embodiment, obtaining video data corresponding to the field of view image of the operator end and sending it to the expert end includes:
[0013] Get video data, split it into groups and add it to the preset encoding pool;
[0014] When the destination end state information is undecodable, retransmitting the video data packet, wherein the retransmission is a video packet combination;
[0015] When the confirmation information is received, the grouped video data is moved out of the preset encoding pool.
[0016] In one embodiment, retransmitting the video data packet includes:
[0017] When a packet received from the source arrives at the coding layer via TCP, the coding layer generates a random linear combination of all packets in the coding window and sends it to the destination;
[0018] Calculate the number of decoded packets based on the difference between the maximum sequence number of packets received by the receiver and the number of packets seen;
[0019] If the difference between the current event and the historical retransmission event is greater than the preset timer timeout value, the sender retransmits DIFF number of coded packets, where the retransmission packets are composed of a linear combination of the first DIFF original packets in the coding window;
[0020] If the difference between the current event and the historical retransmission event is not greater than the preset timer timeout value, then compare whether the currently received DIFF value is greater than the historical DIFF value;
[0021] If the currently received DIFF value is greater than the historical DIFF value, the video data packet is retransmitted.
[0022] In one embodiment, sending to the expert terminal includes:
[0023] The corresponding encryption and decryption keys are negotiated and exchanged between the sender and the destination;
[0024] The sender sends the data packet and performs XOR operation with the key to perform encryption operation and send the value to the destination;
[0025] The destination end performs a preset row-column transformation on the data grouping based on the need;
[0026] Perform XOR operation with the key to perform decryption operation.
[0027] In one embodiment, establishing a multi-person voice call based on an expert end and an operation end includes:
[0028] Capture images of operators and compress image data;
[0029] When receiving a connection request, connect with the expert end until obtaining a receiving identifier;
[0030] When a transmission signal is received, the compressed image data is transmitted directly until a stop signal is received.
[0031] In one of the embodiments, the connection between the nodes is established based on a P2P mechanism;
[0032] Eliminate voice delay jitter based on buffering strategy;
[0033] Reduce the number of voice packets transmitted on the network based on the high compression rate Speex algorithm;
[0034] Reduce the transmitted voice packets based on silence detection and denoising algorithms;
[0035] The speech is encoded based on a variable bit rate approach.
[0036] In a second aspect, the present application also provides a video remote transmission system. The system includes:
[0037] A remote sharing module is used to obtain the video data corresponding to the visual field image of the operator end and send it to the expert end;
[0038] Multi-person voice module, used to build multi-person voice calls based on the expert side and the operator side;
[0039] The 3D sharing module is used by the expert side to build 3D models in real time and share them to the operator side.
[0040] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0041] Obtain the video data corresponding to the visual field image of the operator end and send it to the expert end;
[0042] Build multi-person voice calls based on the expert side and the operation side;
[0043] The expert side builds the 3D model in real time and shares it to the operator side.
[0044] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0045] Obtain the video data corresponding to the visual field image of the operator end and send it to the expert end;
[0046] Build multi-person voice calls based on the expert side and the operation side;
[0047] The expert side builds the 3D model in real time and shares it to the operator side.
[0048] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0049] Obtain the video data corresponding to the visual field image of the operator end and send it to the expert end;
[0050] Build multi-person voice calls based on the expert side and the operation side;
[0051] The expert side builds the 3D model in real time and shares it to the operator side.
[0052] 3. Beneficial effects
[0053] This application adopts the above method. Remote operation guidance is an important function in the induced maintenance system. During the remote guidance process, the system realizes remote guidance operation in three ways: remote vision sharing, multi-person voice call, and 3D system operation. During the data transmission process, the data is encrypted and retransmitted to ensure the reliability and security of the transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A diagram showing an application environment of a video remote transmission method in an embodiment;
[0055] Figure 2 A schematic diagram of all possible events that may occur within one time slot in one embodiment;
[0056] Figure 3 is a schematic diagram of a redundancy factor retransmission state under normal circumstances in one embodiment;
[0057] Figure 4 A flowchart of a TCP / IP protocol data unit encapsulation process in one embodiment;
[0058] Figure 5 An encryption block diagram of an AES algorithm in one embodiment;
[0059] Figure 6 A flowchart of a message transmission process in one embodiment;
[0060] Figure 7 is an operation flow chart of a video acquisition terminal in one embodiment;
[0061] Figure 8 A structural block diagram of a video remote transmission system in one embodiment;
[0062] Fig. 9 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0063] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0064] The video remote transmission method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the video remote transmission method includes a on-site maintenance guidance terminal and a remote expert terminal. The remote expert terminal includes three devices: Machine 1, Machine 2, and Machine 3. Among them, Machine 1 is responsible for remote connection and data transmission, Machine 2 is responsible for the display of remote videos, and Machine 3 is responsible for making 3D collaborative operation processes. Among them, Machine 1 is also connected to Machine 2 and Machine 3 through the TCP protocol.
[0065] During the video transmission process, the device at the on-site end is responsible for the image acquisition, encoding and transmission of the operator's head-mounted camera; Machine 1 at the remote expert terminal is responsible for the reception and forwarding of data; Machine 2 is responsible for the display of images; during remote 3D collaborative operations, Machine 3 at the remote expert terminal is responsible for the production of maintenance processes and the generation of process instruction data; Machine 1 is responsible for the forwarding of process instruction data; the device at the on-site end is responsible for the reception of data, the reception and parsing of maintenance process instruction data, and the display of three-dimensional process animations; the 3D model operation instruction sharing process of the expert terminal for sharing operation instructions based on the 3D model.
[0066] The method includes the following steps:
[0067] Step 202, obtain video data corresponding to the field of view image of the operation end and send it to the expert end.
[0068] Among them, the expert can view the image in the operator's field of view at the maintenance site in real time through the network;
[0069] Step 204, construct a multi-person voice call based on the expert end and the operation end.
[0070] Step 206, the expert end constructs a three-dimensional model in real time and shares it to the operation end.
[0071] Among them, the expert can produce three-dimensional operation process flows remotely and share them to the maintenance site in real time to guide the maintenance personnel to perform maintenance operations.
[0072] In the above-mentioned video remote transmission method, remote operation guidance is an important function in the induced maintenance system. During the remote guidance process, combined with AR remote assisted maintenance, the system realizes remote guidance operation in three ways: remote vision sharing, multi-person voice call, and 3D system operation. During the data transmission process, the data is encrypted and retransmitted to ensure the reliability and security of the transmission.
[0073] In one embodiment, referring to Figure 2 , because the transmitted data includes video streams, voice streams and other data with large transmission volume, therefore, during the transmission process, ensuring the integrity of the data is an important condition, including:
[0074] In a military network environment with uncertain factors, the transmission of large amounts of video and audio data may not achieve satisfactory performance using traditional methods. Even if an ideal congestion control protocol can prevent any congestion packet loss during transmission, packet loss caused by non-congestion (including random packet loss with a fixed bit error rate and burst packet loss due to bad weather or signal shielding) will still seriously affect transmission performance. Moreover, the higher the link utilization of a congestion control algorithm, the greater the impact of non-congestion packet loss on the algorithm. Therefore, how to deal with erroneous packet loss is an important issue in wireless network transmission.
[0075] The core idea of the network coding method is to transmit a combination of multiple data packets instead of a single packet. Therefore, the transmission is the process of adding an original data packet to the coding pool and removing the data from the pool after receiving the confirmation. It has a good ability to cope with transmission errors. The encoded data transmission is independent of each other and is not affected by any other packet loss. Therefore, combining network coding with TCP can greatly improve the robustness and effectiveness of TCP in data transmission.
[0076] In network coding, packets are considered as vectors over a finite field Fq of size q, and pk is the kth original packet sent by the source. When the sending window moves so that the source can send a new packet, it sends a random linear combination of all unacknowledged packets instead of sending an original data packet.
[0077] The source generates a new packet in each time slot in sequence based on a slotted time system (the smaller numbered packet comes first). A newly arrived packet triggers a transmission, which is a linear combination of all packets in the coding window (including the newly arrived packet). Since the probability of randomly generating two linearly independent vectors over a finite field Fq of size q is greater than 1-1 / q, it can be assumed that the range of the field is large enough that each combination received can cause the new packet to be seen. Assume that there is perfect feedback from the receiver to the sender (no errors, no losses, no delays). The sender removes pk from the cache queue after receiving the confirmation that pk has been seen. Retransmissions are not triggered by the arrival of new packets in the coding window. They occur after the transmission and the feedback are received, in the same time slot, with probability r. Therefore, retransmissions are a combination of all old packets in the coding window. The following figure shows the order of events in a time slot. Assume that a new transmission containing a new packet arrives before the receiver sends feedback, and the feedback for the packet occurs in the same time slot. The feedback arrives before the retransmission occurs, and the retransmission occurs at the end of the same time slot, not occupying new time. Non-congested packet losses are independent of each other and occur with probability e.
[0078] In one embodiment, Figure 3 As shown, the specific status of the receiver and the destination include:
[0079] (State of the destination) The destination has two states: D (decodable) state and U (undecodable) state. If it can decipher all the current packets, the destination is in the D state. If packet loss occurs during transmission, the destination will not be able to decipher the packets that arrive later, and the destination is said to be in the U state. When the destination is in the D state, it will still be in the decipherable state after receiving a new combination without packet loss, because only one new packet is added each time it is encoded. If the source sends a retransmission packet when the destination is in the D state, this is a redundant transmission. If packet loss occurs, the destination will change to the U state. Only when it receives enough retransmission packets to make the number of combined packets greater than or equal to the vector dimension can it become decipherable and decipher all the packets. The initial state of the destination is the D state. Therefore, during the entire transmission process, the state of the receiver is composed of a series of D chains and U chains. When the first packet loss occurs in the D state, a U chain begins, and the U chain ends when the number of packet losses is less than or equal to the number of retransmissions.
[0080] Let x represent the number of retransmitted packets that the current receiver needs to receive to become decodable. It is a Markov chain with an absorbing wall, and its state space is I: {0, 1, 2, ...}. When it reaches state 0, the U chain ends. The subsequent packet loss is considered to be the beginning of the next U chain. Let the length of chain Ui be L. Therefore, these L packets need to be decoded at the end of the U chain, and their contribution to the total delay is:
[0081] L+(L-1)+(L-2)…+1=(1+L)L / 2;
[0082] Among them, the length of chain Ui is L;
[0083] Therefore, the sum of the decoding delays for all packets is:
[0084] T=f(L 1 )+f(L 2 )+…+f(L k )+C;
[0085] Among them, T is the sum of the decoding delays of all packets, k is the total number of U chains in this transmission, N is the total number of packets transmitted at one time, and C is the total length of D chains; all U chains follow the same distribution. Therefore, T is related to the length of the U chain L. The shorter L is, the shorter T is.
[0086] In one embodiment, Figure 5 As shown in the figure, in order to shorten the decoding delay and the number of redundancies, a feedback-based network coding retransmission mechanism is proposed. Taking the TCP system as the background, the number of packets that need to be retransmitted by the receiver to become decodable is obtained by confirming the difference between the sequence numbers of the packets sent and the packets seen. Specifically, it includes:
[0087] First, after a packet from the source reaches the coding layer from TCP, the coding layer generates a random linear combination of all packets in the coding window and sends it to the destination. Using the seen mechanism, the maximum sequence number of the packets received by the receiver minus the number of packets seen is the number of packets required for decoding. Therefore, the receiver needs to maintain two variables: the number of coded packets seen, SEEN_CNT; the maximum sequence number MAX_SEQ contained in the received coded packets. In each ACK, the destination not only embeds the sequence number of the packet that has not been seen for the longest time in the packet header, but also adds the difference between MAX_SEQ and SEEN_CNT, called DIFF. Assume that the source sends the following three linear combinations, but the second transmission of y is lost during the transmission process, and the destination only receives combinations x and z. Therefore, after the receiver receives z, p1 and p2 can be seen, the SEEN_CNT value is 2, and the maximum sequence number MAX_SEQ is 3, so DIFF = 3-2 = 1. This means that the destination needs 1 retransmitted packet to switch to the D state. At the sending end, if the DIFF of the received ACK is greater than 0, the sender makes the following judgment to decide whether to retransmit.
[0088] First, check whether the difference between the current time and the last retransmission time is greater than the timer timeout value (set to a value slightly larger than RTT). If it is greater than the timer timeout value, the sender retransmits DIFF number of coded packets, which are composed of linear combinations of the first DIFF original packets in the coding window. This can avoid multiple retransmissions within one RTT; if not, the sender compares whether the currently received DIFF value is greater than the last DIFF value (LAST_DIFF). If it is indeed greater, it means that a new packet loss has occurred, so retransmit (DIFF-LAST_DIFF), and the retransmitted packet is composed of a linear combination of the first (DIFF-LAST_DIFF) original packets in the current coding window.
[0089] If the retransmitted packet is composed of all packets in the current window, as designed in the TCP protocol, although the decoding delay is reduced, it is still much higher than one RTT. This is mainly due to the delay in transmission. During the transmission of a packet, the encoded packets sent later may still be lost, and the confirmation sent back does not include the number of packets lost later, but the encoding window contains these newly generated packets. Therefore, when the source receives the confirmation and prepares to retransmit, the actual DIFF value is already greater than the value displayed in the confirmation packet, and the number of retransmitted packets is still not enough to make the receiver end the U chain and change to the D state. When the source receives the confirmation, it has actually lost 2 packets, but the sender still only retransmits one packet. If the retransmitted encoded packet consists of p2, p3 and p4 in the current encoding window, then the receiver will still be in the unsolvable U state, the U chain length is 4, and all packets are unsolvable. The improved method only retransmits the linear combination of the first DIFF (or DIFF-LAST_DIFF) packets in the encoding window, which can make some packets decrypted. In this way, although the U chain has not ended, the smallest undecipherable packet is moved backward, which reduces the average decoding delay. In addition, encoding and decoding fewer packets can also reduce the complexity and time of the operation. Therefore, the DIFF value received in the first confirmation is 1, so the retransmitted coded packet consists of only p2, thus successfully decoding p1 and p2.
[0090] In this embodiment, the algorithm not only avoids redundant transmission when the receiver is in a decomposable state, but also when it is in an undecomposable state, it can effectively retransmit a suitable number of packets so that the receiver can change from the U state to the D state as quickly as possible, and decode some packets in advance. This greatly reduces decoding delay and redundancy, and most importantly, it does not improve performance at the expense of throughput. In theory, regardless of the correlation of the coefficient vector, each received packet is valid and updateable, so the throughput of the algorithm is optimal. In fact, the algorithm is not affected by the coefficient generation algorithm and the size of the finite field. Even if the encoding algorithm generates some linearly related coefficient vectors, the method can still work effectively. And unlike SNC, this method does not change the nature of the seen packet confirmation, so the expectation of the queue length is still Ω(1-ε)^(-1), and the queue length will not increase due to the increase in decoding speed.
[0091] Compared with the traditional retransmission mechanism, the algorithm focuses on the total amount of packet loss, which greatly reduces the overhead of the packet header. This method better hides the packet loss from the congestion algorithm. It can not only handle random packet loss at a constant rate, but also handle sudden continuous packet loss that may occur due to other reasons. At the same time, the mechanism retains the end-to-end semantics of TCP, and the encoding operation is only performed at the end node, achieving the purpose of hiding the packet loss from the congestion control algorithm. The traditional method roughly compensates for the error packet loss by taking the retransmission ratio as the inverse of the transmission success rate, while the congestion packet loss is still retransmitted by the congestion algorithm. However, since the retransmission algorithm at the coding layer cannot distinguish the cause of packet loss, the above retransmission algorithm is not suitable for algorithms that may generate a lot of congestion. Compared with the traditional algorithm that evenly retransmits the packet loss, the method proposed in this project can retransmit the required number of packets at the time when retransmission is required without hiding the congestion packet loss. However, for congestion methods such as VCP and MLCP, which basically have zero congestion packet loss, there is no need to adopt the above strategy. This can better handle sudden and unknown packet losses in environments without a fixed bit error rate, and has lower decoding delay.
[0092] It can be seen that the feedback-based retransmission method has a lower decoding delay than the quantitative retransmission algorithm. However, if the retransmission method used by these methods is based on combining all the packets in the current coding window, its performance will be deeply affected by the bit error rate. Therefore, a feedback-based network coding retransmission mechanism is proposed. It is based on the seen confirmation method used in network coding, and accurately knows the number of retransmitted packets required for decoding based on the available implicit information. In addition, the coding rules are changed to decode some packets as early as possible. This method can not only deal with random packet loss, but also effectively cope with sudden continuous packet loss.
[0093] In one embodiment, since the TCP / IP protocol was not designed with security issues in mind, its security vulnerabilities are increasingly exposed after it is widely used and are exploited by many network attacks. The specific steps for encrypting the protocol include:
[0094] Reference Figure 4, the TCP / IP protocol family corresponds to a four-layer protocol model, also known as the DARPA model. The four layers in this model are, from top to bottom, the application layer, the transport layer, the internet layer, and the network interface layer. Each layer corresponds to one or more layers in the seven-layer open system interconnection (OSI) reference model. When data is transmitted, it is first encapsulated from top to bottom by adding a header or a trailer to the upper layer data. The application layer is the topmost layer of the TCP / IP protocol model. The data unit is initially generated by the application program of the application layer, such as the application, FTP, or email. After the application layer data reaches the transport layer, a transport layer header is added before the data (the transport layer header has a TCP header and a UDP header), which is called a message. If it is a TCP protocol header, it is called a segment. After that, the transport layer data continues to be transmitted to the IP layer. At the IP layer, an IP header is added before the message / segment, which is called a packet / packet, and continues to be transmitted downward. The network interface layer includes the data link layer and the physical layer in the OSI seven-layer model. When the IP packet reaches this layer, a frame header and a frame tail are added to the front and back of the data, respectively, which are called frames. Then, the data is transmitted to the physical medium and the bit stream (0, 1) is transmitted through current pulses. When the data reaches the destination, the opposite operation is performed, that is, decapsulation, and the header and tail messages are removed in turn to obtain the final application data.
[0095] This method comprehensively utilizes symmetric encryption (AES algorithm) to establish a message sending model, which can effectively solve the security problems faced by the system, such as information tampering, theft, and forgery.
[0096] Symmetric encryption algorithms are mainly used to ensure the confidentiality of data. The communicating parties will share their single key, which is the only key for encryption or decryption during information transmission. Symmetric encryption algorithms are still widely used, and the cryptography community has also conducted in-depth research on symmetric encryption algorithms. The Data Encryption Standard (DES) algorithm is the most commonly used symmetric encryption algorithm. It is a block cipher algorithm using a 56-bit key developed by IBM under the instruction of the National Security Agency (NSA) of the United States. According to the different encryption methods of plaintext messages, symmetric encryption algorithms can be roughly divided into two categories, namely block ciphers and stream ciphers. Block ciphers, that is, message packets, all have a fixed length, and the length of the ciphertext packet obtained after encryption is the same as the length of the plaintext packet. The AES algorithm belongs to this type of block cipher algorithm. Whether it is the packet when the information is input, the ciphertext packet obtained after encryption, or the packet used in the encryption and decryption process, it is 128 bits. For the AES algorithm, the key length K can be 128 bits, 192 bits, and 256 bits. Encryption algorithms such as Figure 5 shown.
[0097] In the model, users A and B are the clients of the system. They have each other's digital certificates or public keys. The whole process of secure information transmission is as follows: Figure 6 As shown:
[0098] In this embodiment, through the above information encryption concept, through the designed encryption authentication model, an implementation model of a secure instant messaging system is established; the model includes two levels of authentication: client-to-client two-way authentication, before sending data, the clients at both ends must negotiate the keys required for encryption and decryption, and exchange the keys securely;
[0099] The two-way authentication between the server and the client uses public key technology as the authentication cryptographic technology. The public key technology in the model also has the function of encrypting shared keys and digital signatures, which not only solves the key transmission problem for data communication between clients, but also solves the identity authentication between the server and the client, and between the clients.
[0100] In one embodiment, remote video transmission encodes the image obtained by the on-site user end through the network, transmits it to the expert end, and then decodes it to obtain the video image. Video encoding reduces the network transmission volume while ensuring the visual effect. This solution uses DIVX as a video encoder to compress and transmit the video.
[0101] In this embodiment, as a system developed for low bandwidth, system consumption is also a key issue that DIVX needs to consider. The minimum requirements for DIVX hardware configuration are 300MHz CPU, 64M memory, and 8M graphics card;
[0102] The MPEG-1 standard compresses video data into a standard data stream of 1-2Mbps, but the compression results in a lower resolution and unsatisfactory image quality. In order to obtain higher resolution and provide broadcast-level video and CD-quality audio, the MPEG-2 standard has a transmission rate of 3-10Mbps. MPEG-4, which is designed for low bandwidth, reduces the transmission rate to 4.8-6.4kbits / s while retaining image quality close to that of MPEG-2.
[0103] MPEG-4 uses object-oriented compression. The transmission rate of MPEG-4 can be set in network transmission, and the image quality can also be changed accordingly within a certain range. Users can make different settings according to different requirements for recording time, number of transmission channels and clarity, which greatly improves the adaptability and flexibility of the system. At the same time, when there is bit error or packet loss in transmission, MPEG-4 is less affected and can recover quickly.
[0104] In one embodiment, Figure 7 As shown in the figure, the video transmission process is realized. The video is captured from the camera, encoded and transmitted to the remote expert end through the encoder, and then the decoded image sequence is subjected to virtual-real fusion and other operations. The camera video capture adopts Microsoft's DirectShow technology, and its image is saved as Bitmap type in the buffer. Before encoding at the sender, the following pre-operations must be performed: First, obtain the camera-related parameters through the graphics device interface. Then, set the encoder's bit rate, key frame and other control information. Finally, obtain the buffer data segment head address through the pointer. At this point, the encoder starts running. Similar to the start of data transmission, the receiving end also needs to perform corresponding pre-operations. After completion, the decoder starts working and data begins to be received.
[0105] In one embodiment, Figure 8 As shown in the figure, a multi-person voice call specifically includes:
[0106] Multi-person voice calls are a necessary interactive method for remote guidance. The system uses a P2P transmission method, and each node is both a client and a server. Assuming that node 2 is a server and node 1 is a client, node 2 can connect to node 1 for voice communication. Among them, the listening port of node 2 is used to establish a signaling connection, and the listening port is used to establish a media connection. Signaling is mainly used to feedback whether node 1 and node 2 can establish voice communication;
[0107] The media is mainly used to transmit voice. Node 1 and Node 2 represent a machine respectively. The P2P mechanism is adopted. It is both a client and a server. The connection and establishment of signaling are realized by the TCP protocol, and the voice communication is also realized by the TCP protocol. Due to the characteristics of TCP retransmission, some problems will be caused to the transmission of voice, which will greatly affect the quality of voice. Therefore, some methods must be adopted to reduce the negative impact caused by the TCP retransmission mechanism, so as to improve the quality of voice.
[0108] It is worth mentioning that the system uses the following methods to solve the TCP retransmission problem, thereby improving the voice quality:
[0109] Use buffering strategies to eliminate voice delay jitter;
[0110] The high compression rate Speex algorithm is used to reduce the number of voice packets transmitted on the network and improve the quality of voice communication. The data transmission rate is about 2.2kbps-40kbps.
[0111] Silence detection and denoising algorithms filter out silence and noise in voice, reduce the transmitted voice packets, and thus improve voice quality;
[0112] The variable bit rate encoding method makes it possible to obtain higher voice quality with fewer bits.
[0113] In one embodiment, the operation instruction sharing based on the 3D model specifically includes:
[0114] Both the on-site and maintenance ends of the system have 3D digital maintenance manuals as data support. Whether the front-line maintenance personnel use the manual to check how to repair, or the rear experts provide guidance to the front-line, they must use this 3D digital maintenance manual. This manual contains all 3D models of product parts and maintenance tools. These 3D models can not only be used to preset specific maintenance actions, but also for interactive operations.
[0115] Therefore, when the front and rear collaborate, they can interact based on the 3D model in the manual. The 3D scenes on both sides establish real-time communication, and only simple command information such as viewpoint posture and model operation information needs to be transmitted to control the other scene to maintain synchronization. When the expert is guiding the operation in the 3D scene, his operation instructions and viewpoint information are transmitted to the maintainer in real time. A section of maintenance process instruction code, based on these instructions, synchronizes the operation model and viewpoint in the 3D scene, so that the maintainer can see the same action from the same perspective. The amount of data for this instruction is very small, generally 1-10K bytes, so it adds almost no additional pressure on communication.
[0116] Remote operation guidance is an important function in the induced maintenance system. In the remote guidance process, combined with AR remote assisted maintenance, the system realizes remote guidance operation in three ways: remote vision sharing, multi-person voice call, and 3D system operation. During the data transmission process, the data is encrypted and retransmitted to ensure the reliability and security of the transmission. The main research points of this solution include:
[0117] In this embodiment, the TCP protocol is used for data transmission. Based on the protocol, data verification and retransmission are implemented to ensure the reliability and real-time performance of data transmission when the network conditions are not optimal. The AES method is used to encrypt and decrypt the data transmitted over the network to ensure the security of data transmission over the network. The image data collected by the user camera at the on-site maintenance end is encoded and transmitted to the expert end, thereby realizing remote field of view sharing. Multi-person voice calls are realized between the expert end and the on-site end, and real-time voice communication is realized through the integration of the server and the client. The maintenance process data at the expert end is encoded and transmitted to the on-site end, and real-time three-dimensional augmented reality operation guidance is performed for on-site maintenance personnel.
[0118] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0119] Based on the same inventive concept, the embodiment of the present application also provides a video remote transmission system for implementing the video remote transmission method involved above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more video remote transmission system embodiments provided below can refer to the limitations of the video remote transmission method above, and will not be repeated here.
[0120] In one embodiment, Fig. 9 As shown, a video remote transmission system is provided, including: a remote sharing module, a multi-person voice module and a three-dimensional sharing module, wherein:
[0121] A remote sharing module is used to obtain the video data corresponding to the visual field image of the operator end and send it to the expert end;
[0122] Multi-person voice module, used to build multi-person voice calls based on the expert side and the operator side;
[0123] The 3D sharing module is used by the expert side to build 3D models in real time and share them to the operator side.
[0124] In one embodiment, the remote sharing module is also used to: obtain video data and split it into groups and add it to a preset encoding pool; when the destination status information is undecodable, retransmit the video data group, and the retransmission is a video group combination; when a confirmation message is received, the grouped video data is removed from the preset encoding pool.
[0125] In one embodiment, the remote sharing module is also used for: when a packet received from the source arrives at the coding layer via TCP, the coding layer generates a random linear combination of all packets in a coding window and sends it to the destination; the number of decoded packets is calculated based on the difference between the maximum sequence number of the packet received by the receiver and the number of packets seen; if the difference between the current event and the historical retransmission event is greater than a preset timer timeout value, the sender retransmits a DIFF number of coding packets, and the retransmission packets are composed of a linear combination of the first DIFF original packets in the coding window; if the difference between the current event and the historical retransmission event is not greater than a preset timer timeout value, the currently received DIFF value is compared to see if it is greater than the historical DIFF value; if the currently received DIFF value is greater than the historical DIFF value, the video data packet is retransmitted.
[0126] In one embodiment, the remote sharing module is also used for: encrypting and decrypting corresponding keys based on negotiation between the sending end and the destination end and exchanging them; the sending end sends data packets and performs an XOR operation with the key to perform an encryption operation and sends the value to the destination end; the destination end performs a preset row-column transformation on the data packet based on the default; and performs an XOR operation with the key to perform a decryption operation.
[0127] In one embodiment, the multi-person voice module is also used to: collect images of operators and compress image data; connect with the expert end when a connection request is received until a receiving identifier is obtained; and send compressed image data when a sending signal is received and directly receive a stop signal.
[0128] In one embodiment, the multi-person voice module is also used to: establish connections between nodes based on a P2P mechanism; eliminate voice delay jitter based on a buffering strategy; reduce the number of voice packets transmitted on the network based on a high compression rate speex algorithm; reduce the transmitted voice packets based on a silence detection and denoising algorithm; and encode voice based on a variable bit rate method.
[0129] Each module in the above video remote transmission system can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0130] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a video remote transmission method is implemented.
[0131] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input system connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a video remote transmission method is implemented.
[0132] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0133] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0135] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0137] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0138] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0139] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A video remote transmission method, characterized in that: The method comprises: Obtain the video data corresponding to the visual field image of the operator end and send it to the expert end; Build multi-person voice calls based on the expert side and the operation side; The expert side builds the 3D model in real time and shares it to the operator side.
2. The video remote transmission method according to claim 1, characterized in that: The step of acquiring video data corresponding to the visual field image of the operator end and sending the video data to the expert end includes: Get video data, split it into groups and add it to the preset encoding pool; When the destination end state information is undecodable, retransmitting the video data packet, wherein the retransmission is a video packet combination; When the confirmation information is received, the grouped video data is moved out of the preset encoding pool.
3. The video remote transmission method according to claim 2, characterized in that: The retransmission of the video data grouping comprises: When a packet received from the source arrives at the coding layer via TCP, the coding layer generates a random linear combination of all packets in the coding window and sends it to the destination; Calculate the number of decoded packets based on the difference between the maximum sequence number of packets received by the receiver and the number of packets seen; If the difference between the current event and the historical retransmission event is greater than the preset timer timeout value, the sender retransmits DIFF number of coded packets, where the retransmission packets are composed of linear combinations of the first DIFF original packets in the coding window; If the difference between the current event and the historical retransmission event is not greater than the preset timer timeout value, then compare whether the currently received DIFF value is greater than the historical DIFF value; If the currently received DIFF value is greater than the historical DIFF value, the video data packet is retransmitted.
4. The video remote transmission method according to claim 1, characterized in that: The sending to the expert terminal includes: The corresponding encryption and decryption keys are negotiated and exchanged between the sender and the destination; The sender sends the data packet and performs XOR operation with the key to perform encryption operation and send the value to the destination; The destination end performs a preset row-column transformation on the data grouping based on the need; Perform XOR operation with the key to perform decryption operation.
5. The video remote transmission method according to claim 1, characterized in that: The multi-person voice call based on the expert end and the operator end includes: Capture images of operators and compress image data; When receiving a connection request, connect with the expert end until obtaining a receiving identifier; When a transmission signal is received, the compressed image data is transmitted directly until a stop signal is received.
6. The video remote transmission method according to claim 5, characterized in that: The method further comprises: Establish connections between nodes based on P2P mechanism; Eliminate voice delay jitter based on buffering strategy; Reduce the number of voice packets transmitted on the network based on the high compression rate Speex algorithm; Reduce the transmitted voice packets based on silence detection and denoising algorithms; The speech is encoded based on a variable bit rate approach.
7. A video remote transmission system, characterized in that: The system comprises: A remote sharing module is used to obtain the video data corresponding to the visual field image of the operator end and send it to the expert end; Multi-person voice module, used to build multi-person voice calls based on the expert side and the operator side; The 3D sharing module is used by the expert side to build 3D models in real time and share them to the operator side.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Delaying retransmission requests in multi-carrier systems
CN101411148A