Real-time remote guidance system and method based on online video key information fusion

By integrating video and hand operation information in the remote online guidance system, the problem of intuition in the text description in the prior art and deviation in voice communication and understanding of speech in the existing technology is solved, and fast and intuitive remote operation guidance is achieved, reducing communication costs.

CN120050515APending Publication Date: 2025-05-27SHANGHAI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510128932.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When providing instrument and equipment operation guidance, existing remote online guidance methods have unintuitive text descriptions, comprehension biases and cross-lingual barriers in voice communication, resulting in operation difficulties and high communication costs.

Method used

By establishing a key information fusion system based on online video between the guidance end and the guidance end, two camera devices are used to capture the instrument and the guide personnel's hand operation movements in real time, extract the hand contour and fuse it into the instrument and equipment video, and transmit it to the guidance end in real time, real-time interactive video online remote auxiliary guidance.

Benefits of technology

Through intuitive body language guidance, the system reduces the time for operators to review manuals and understand text descriptions, reduces communication costs and errors, and achieves fast and efficient remote operation guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050515A_ABST
    Figure CN120050515A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time remote guidance system and method based on online video key information fusion. The real-time remote guidance system and method are used for helping other personnel to operate instruments and equipment on line by experts or equipment skilled operators through remote immersive videos. According to the invention, two controllers and camera devices at different geographic positions are utilized, one camera device is used for shooting instruments and equipment needing to be operated by a guided person, the other camera device is used for shooting hand actions of the guided person, and the main controller carries out real-time fusion on contents shot by the two camera devices, so that the real-time fusion of the contents is realized. A fusion result is displayed on a video display device of an auxiliary controller used by a guided person, and clear and visual remote online guidance is provided for an unskilled operator. By adopting the technical scheme of the invention, the learning cost of technicians in operating instruments and equipment is saved to a certain extent, and the communication cost between the instructor and the instructed person is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is a real-time remote guidance system and method based on the fusion of key information of online videos, belonging to the technical field of instrument and equipment operation, and is used for experts or skilled operators of equipment to assist other personnel remotely through immersive online videos. Background Art

[0002] When an operator encounters difficulties in operating an instrument or equipment, they can usually refer to the instruction manual of the instrument or equipment or seek help from others. Although the instruction manual can provide the operator with detailed usage steps, the manual is often too technical and the content is too extensive. For non-professionals, referring to the manual is not only difficult to understand but also very time-consuming. In an emergency, it is difficult to obtain detailed operation steps from the manual in a short time, which may lead to accidents. Seeking help from others is a relatively intuitive method. However, when the instructor is not in the same location, remote online guidance is required, and the means mainly include text communication, voice communication, and video demonstration, etc. Text communication can achieve remote online guidance, enabling unskilled operators to obtain concise and clear operation steps in a timely manner. However, the content described in text is far less intuitive than voice or video communication, resulting in a longer understanding time and prone to errors in the guidance content. Voice communication can conduct real-time voice communication and naturally convey important information. However, voice communication has poor intuitiveness, is prone to understanding deviations, and there are communication difficulties caused by accents, cross-languages, etc. Video demonstration can conduct remote online guidance. The image information in the video can clearly and intuitively reflect the difficulties and problems. However, there are also problems such as the description method, professional terms, and time delay in the operation of the equipment, which bring difficulties in understanding to a certain extent and increase the communication cost between the two parties.

[0003] In order to achieve remote real-time guidance, the present invention provides a system and method for real-time remote guidance based on the fusion of key information of online videos. First, a video is taken from the terminal of the instrument or equipment to be operated, that is, the guided end, and transmitted online to the guiding end. The guiding personnel perform virtual operations on the instrument or equipment in front of the display device, and use a camera device to capture the hand movements of the guiding personnel. The hand information of the guiding personnel is extracted from the real-time video and fused onto the instrument or equipment to be operated. Through the cut-out gesture actions, it is transmitted to the guided end to guide the operator in real time on how to operate the instrument or equipment. In this way, when the operator is not familiar with the instrument or equipment or in an emergency and needs guidance from an absent expert or skilled operator, the expert or the skilled operator of the equipment can directly conduct fast remote video guidance using the system and method proposed by the present invention.

[0004] The present invention combines hand movements and images of instruments and equipment, and provides guidance to operators using intuitive and easily understandable body language, avoiding problems such as the time-consuming process of referring to huge manuals, the non-intuitiveness of text descriptions, the understanding deviation of voice and pure video descriptions, and cross-language issues, reducing the learning cost of operators and the communication cost between the instructor and the trainee, thereby achieving real-time and rapid solution to the problems of operators operating instruments and equipment. Summary of the Invention

[0005] The object of the present invention is to provide a system and method for real-time remote guidance based on the fusion of key information in online videos, aiming at the deficiencies of existing technologies. By simultaneously shooting the instrument and equipment to be operated and the hand operation actions of the instructor using two camera devices, fusing the two pieces of information and transmitting them to the video display device of the trainee, real-time remote guidance is achieved.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A system and method for real-time remote guidance based on the fusion of key information in online videos, the system comprising an instructor-end system and a trainee-end system; the instructor-end system includes: a main controller, a first camera device, and a first video display device; the trainee-end system includes: a sub-controller, a second camera device, and a second video display device. In the instructor-end system, the main controller is used to connect the sub-controller of the trainee-end system, and the first camera device is used to shoot the first video display device used by the instructor; the instructor conducts hand movement guidance in front of the first video display device, which is shot by the first camera device and transmitted to the main controller; the main controller extracts the hand contour and integrates it into the instrument and equipment displayed on the first video display device; the fused image is sent to the second video display device of the trainee-end system in real time. In the trainee-end system, the sub-controller is used to connect the main controller in the instructor-end system; the second camera device is used to shoot the instrument and equipment; the sub-controller transmits the video of the instrument and equipment to be operated to the instructor-end system; the second video display device of the trainee-end system is used to display the fused image in the instructor-end system.

[0008] On the basis of the completion of the construction of the hardware device, a method is designed and developed to extract the hand contour in a complex environment and integrate the hand operation actions of the instrument device and the instructor. First, the second camera device is used to photograph the instrument device that the person to be instructed needs to operate and transmit it to the video display device of the instructor. The first camera device is placed in front of the first video display device of the instructor and photographs the image of the first video display device. The instructor performs virtual operations on the instrument device in the video display device in front of the video display device, and these operation actions are photographed by the first camera device. Using the hand shape extraction method, the hand shape of the instructor in the real-time communication video is extracted from the complex environmental background, and a certain transparency is set for the hand shape. It is fused with the target video, that is, the video of the instrument device to be operated, in proportion, and weighted summation is performed on the transparency to obtain the final fused video. The fused video is transmitted to the video display device of the person to be instructed through the network. The person to be instructed can quickly and efficiently understand the operation of the instrument device according to the operation position and hand shape action on the instrument device in the transmitted video.

[0009] The operation interface of the instructor-end system mainly includes: ① the real-time display area of the device instrument, ② the selected connected person-to-be-instructed-end system, ③ the IP address of the person-to-be-instructed-end system to be connected, ④ the connection button and the button to agree to the connection request of the other party, ⑤ the connection status display, ⑥ the disconnection button;

[0010] The operation interface of the person-to-be-instructed-end system mainly includes: ① the real-time display area of the fusion result, ② the selected connected instructor-end system, ③ the IP address of the instructor-end system to be connected, ④ the connection button and the button to agree to the connection request of the other party, ⑤ the connection status display, ⑥ the disconnection button.

[0011] The detailed description of the interface operation is as follows:

[0012] (1-1) The device instrument display area in the instructor-end system interface: The online video of the instrument device to be operated is displayed in the display area;

[0013] (1-2) The selected connected person-to-be-instructed-end system in the instructor-end system interface: One of the connection methods, which can be connected by selecting the ID number of the person-to-be-instructed-end system to be connected and the person-to-be-instructed-end system;

[0014] (1-3) The IP address of the person-to-be-instructed-end system to be connected in the instructor-end system interface: One of the connection methods, which can be connected by the IP of the person-to-be-instructed-end system to be connected and the person-to-be-instructed-end system;

[0015] (1-4) Connection button and button to approve the connection request from the other party in the guiding end system interface: After selecting the connection method, use this button to connect to the guided end system. When the other party approves, end-to-end real-time video transmission can be carried out; when a connection request is sent from the guided end system, also use this button to approve the connection;

[0016] (1-5) Connection status display in the guiding end system interface: Used to observe the connection situation;

[0017] (1-6) Disconnection button in the guiding end system interface: Used to close the network connection between the guiding end system and the guided end system; when a connection request is sent from the other party, also use this button to reject the connection;

[0018] (1-7) Real-time display area for the fusion result in the guided end system interface: The video of the fusion result is displayed in real time on the display area;

[0019] (1-8) Guiding end system selected for connection in the guided end system interface: One of the connection methods, can connect to the guiding end system by selecting the ID number of the guiding end system to be connected;

[0020] (1-9) IP address of the guiding end system to be connected in the guided end system interface: One of the connection methods, can connect to the guiding end system by the IP of the guiding end system to be connected;

[0021] (1-10) Connection button in the guided end system interface: After selecting the connection method, use this button to connect to the guiding end system. When the other party approves, end-to-end real-time video transmission can be carried out; when a connection request is sent from the guiding end system, also use this button to approve the connection;

[0022] (1-11) Connection status display in the guided end system interface: Used to observe the connection situation;

[0023] (1-12) Disconnection button in the guided end system interface: Used to close the network connection between the guiding end system and the guided end system; when a connection request is sent from the other party, also use this button to reject the connection.

[0024] Based on the above system device, the technical solution of a real-time remote guidance system and method based on the fusion of key information in online video is as follows:

[0025] (2-1) The guided end system captures the real-time video of the instrument and equipment to be operated through the second imaging device, sends it to the main controller of the guiding end system through the network, and displays it on the first video display device controlled by the main controller;

[0026] (2-2) In the guiding end system, the main controller captures the first video display device through the first imaging device, extracts the edge of the first video display device using the Canny edge detection algorithm, takes the area of the first video display device as the working area, and cuts off all non-working areas in the video, making the working area fill the entire screen. This screen is the background screen B1;

[0027] (2-3) In the guiding end system, the main controller captures the gesture video of the operation demonstration in front of the first video display device through the first imaging device. Its image is the foreground image F1. Using the hand shape recognition method, the hand shape of the guiding personnel is recognized and extracted from the complex environment, and the complex background is removed, and the position of the palm center in the working area of the first video display device is recorded. The method is as follows:

[0028] Construct a deep learning network model for hand extraction. The main network is MobileNetV3. A Squeeze-and-Excitation module is added to the feature extraction encoder in MobileNetV3 to recalibrate the feature channels, enhance the sensitivity of the network to key features, and enhance the expression ability of the model. Its formula is:

[0029] SE: S = σ(W 2 δ(W1AvgPool(X))) (1)

[0030] Among them, σ and δ represent the sigmoid and ReLU functions respectively, W 1 and W 2 are the weights of the fully connected layer, and X represents the input;

[0031] In the deep learning network model for hand extraction, a simplified LR-ASPP module is added to efficiently extract multi-scale context information and improve the feature expression ability for segmenting key content in the video. Its formula is:

[0032] aspp 1 = ReLU(BN(W 1 *X)) (2)

[0033] aspp 2 = σ(W 2 *GAP(X)) (3)

[0034] Y = aspp 1 ⊙aspp 2 (4)

[0035] Among them, X represents the input, W1 and W2 represent weights, * represents the convolution operation, σ represents the sigmoid operation, ⊙ represents the Hadamard product operation, BN represents the batch normalization operation, GAP represents the global average pooling operation, and Y represents the output;

[0036] In the deep learning network model for hand extraction, a recurrent decoder is added, mainly including the convolutional long short-term memory network ConvLSTM and the convolutional gated recurrent unit ConvGRU, which are used to learn features in time and space, maintain the state information between video frames, and ensure the spatio-temporal dependence between video frames;

[0037] A lightweight feature enhancement module CBAM is added to the recurrent decoder of the deep learning network model for hand extraction, and spatial and channel attention are used to enhance the features, thereby improving the ability to refine local and global features. The formula is:

[0038] M c = σ(MLP(AvgPool(X)) + MLP(MaxPool(X))) (5)

[0039] F 1 = M c X (6)

[0040] M s = σ(Conv([AvgPool(F 1 );MaxPool(F 1 )])) (7)

[0041] F out = M s F 1 (8)

[0042] Among them, X is the input feature map, AvgPool represents the average pooling operation, MaxPool represents the maximum pooling operation, and σ represents the sigmoid operation;

[0043] A depth-guided filtering module is added to the deep learning network model for hand extraction, which is used to improve the spatial resolution of images and videos while maintaining edge and detail information. The formula is:

[0044] O i = a k I i + b k (9)

[0045] Among them, the output image O is generated based on the linear relationship within the local regions of the input feature map I and the feature map G of the guidance image. a k and b kis the linear coefficient calculated on window k where pixel i is located, and can be obtained by minimizing the following cost function:

[0046]

[0047] where ω k represents the window centered on pixel k, and ∈ is a regularization parameter;

[0048] Using the coordinates of each point on the outer contour of the hand shape, calculate and record the position of the palm center in the working area;

[0049] (2-4) The hand-shaped image pixels extracted from the foreground image F1, the non-hand-shaped image part has its pixel values deleted and is given white, and this image is the foreground image F2;

[0050] (2-5) According to the convex hull size of the instrument image displayed on the first video display device controlled by the main controller in the guiding end system, automatically scale the hand shape in the foreground image F2 so that the size of the hand shape is 1 / 4 of the convex hull of the instrument image, obtaining the foreground image F3. In F3, the pixels in the area where the hand shape is located are the pixels of the scaled-down hand shape, and the pixel values outside the area where the hand shape is located are 0. The position of the palm center in F3 is the same as that in F2;

[0051] (2-6) Extract the outer contour of the hand shape in the foreground image F3, set the image pixels outside the outer contour to 0, and set the image pixels on and inside the outer contour to 255, obtaining hand_mask;

[0052] (2-7) In the background image B1, use the inverse mask umask = -hand_mask, and obtain the background image BUH without the hand shape through bitwise operation. The formula is as follows:

[0053] BUH = B1 ∧ (umask) (11)

[0054] where ∧ is the bitwise AND operation;

[0055] (2-8) In the background image B1, use the mask mask = hand_mask, and obtain the hand-shaped image BH through bitwise operation. The formula is as follows:

[0056] BH = B1 ∧ (mask) (12)

[0057] where ∧ is the bitwise AND operation;

[0058] (2-9) Perform bitwise operation and transparency operation on BUH, BH, and F3 to obtain the fused image Combined_Image, as shown below:

[0059] Combined_Image = BUH ∨ (α * BH + β * F3 + λ) (13)

[0060] Wherein, ∨ is a bitwise OR operation, α and β are corresponding weighting factors used to control the transparency of the hand positions in BH and F3 in the output image, and can be set to 0.5 respectively. λ is a scalar value added to the final result and can be set to 0;

[0061] (2 - 10) In the guiding - end system, the main controller transmits the fused video to the guided - end system through the network and displays it on the second video display device of the guided - end system, thereby realizing interactive video online remote assistance guidance.

[0062] Compared with the prior art, the present invention has the following obvious outstanding substantive features and remarkable advantages:

[0063] (1) The hardware devices included in the assistance system

[0064] The assistance system includes a guiding - end system and a guided - end system; the guiding - end system includes: a main controller, a first camera device, and a first video display device; the guided - end system includes: a sub - controller, a second camera device, and a second video display device. In the guiding - end system, the main controller is used to connect to the sub - controller of the guided - end system, and the first camera device is used to capture the first video display device used by the guiding personnel; the guiding personnel perform hand - movement guidance in front of the first video display device, which is captured by the first camera device and transmitted to the main controller; the main controller extracts the hand contour and integrates it into the instrument equipment displayed on the first video display device; the fused image is sent to the second video display device of the guided - end system in real - time. In the guided - end system, the sub - controller is used to connect to the main controller in the guiding - end system; the second camera device is used to capture the instrument equipment; the sub - controller transmits the video of the instrument equipment that needs to be operated to the guiding - end system; the second video display device of the guided - end system is used to display the fused image in the guiding - end system.

[0065] (2) The operation interfaces of the assistance system

[0066] The operation interface of the guiding - end system mainly includes: ① a real - time display area for equipment and instruments, ② selection of the connected guided - end system, ③ the IP address of the guided - end system to be connected, ④ connection buttons and buttons for agreeing to the connection request from the other party, ⑤ connection - status display, ⑥ stop - connection buttons.

[0067] The operation interface of the guided - end system mainly includes: ① a real - time display area for the fusion result, ② selection of the connected guiding - end system, ③ the IP address of the guiding - end system to be connected, ④ connection buttons and buttons for agreeing to the connection request from the other party, ⑤ connection - status display, ⑥ stop - connection buttons.

[0068] The detailed description of the interface operations is as follows:

[0069] (1-1) Device and instrument display area in the instructor-end system interface: The online video of the instrument and equipment to be operated is displayed in the display area;

[0070] (1-2) Selected connected trainee-end system in the instructor-end system interface: One of the connection methods, which can be connected to the trainee-end system by selecting the ID number of the trainee-end system to be connected and the trainee-end system;

[0071] (1-3) IP address of the trainee-end system to be connected in the instructor-end system interface: One of the connection methods, which can be connected to the trainee-end system by the IP of the trainee-end system to be connected and the trainee-end system;

[0072] (1-4) Connection button and button to agree to the other party's connection request in the instructor-end system interface: After selecting the connection method, connect to the trainee-end system through this button. When obtaining the consent of the other party, end-to-end real-time video transmission can be carried out; when the trainee-end system sends a connection request, also agree to the connection through this button;

[0073] (1-5) Connection status display in the instructor-end system interface: Used to observe the connection situation;

[0074] (1-6) Disconnection button in the instructor-end system interface: Used to close the network connection between the instructor-end system and the trainee-end system; when the other party sends a connection request, also disagree to the connection through this button;

[0075] (1-7) Real-time fusion result display area in the trainee-end system interface: The video of the fusion result is displayed in real time in the display area;

[0076] (1-8) Selected connected instructor-end system in the trainee-end system interface: One of the connection methods, which can be connected to the instructor-end system by selecting the ID number of the instructor-end system to be connected and the instructor-end system;

[0077] (1-9) IP address of the instructor-end system to be connected in the trainee-end system interface: One of the connection methods, which can be connected to the instructor-end system by the IP of the instructor-end system to be connected and the instructor-end system;

[0078] (1-10) Connection button in the trainee-end system interface: After selecting the connection method, connect to the instructor-end system through this button. When obtaining the consent of the other party, end-to-end real-time video transmission can be carried out; when the instructor-end system sends a connection request, also agree to the connection through this button;

[0079] (1-11) Connection status display in the trainee-end system interface: Used to observe the connection situation;

[0080] (1 - 12) Stop connection button in the guided - end system interface: used to close the network connection between the guiding - end system and the guided - end system; when a connection request is sent by the other party, it is also through this button to reject the connection.

[0081] (3) Technical solution of the auxiliary system

[0082] The guided - end system captures the real - time video of the instrument to be operated through the second camera device, sends it to the main controller of the guiding - end system through the network, and displays it on the first video display device controlled by the main controller; in the guiding - end system, the main controller captures the first video display device through the first camera device, uses the canny edge detection algorithm to extract the edge of the first video display device, takes the area of the first video display device as the working area, and cuts off all non - working areas in the video, so that the working area fills the whole screen, and this screen is the background image B1; the main controller in the guiding - end system captures the gesture video of the operation demonstration in front of the first video display device through the first camera device, its image is the foreground image F1, uses the hand - shape recognition method to recognize and extract the hand - shape of the guiding personnel from the complex environment, removes the complex background, and records the position of the palm center in the working area of the first video display device; for the pixel of the hand - shape image extracted from the foreground image F1, the pixel value of the non - hand - shape image part is deleted and given white, and this image is the foreground image F2; according to the convex hull size of the instrument image displayed on the first video display device controlled by the main controller in the guiding - end system, automatically scale the hand - shape in the foreground image F2 so that the size of the hand - shape is 1 / 4 of the convex hull of the instrument image, and get the foreground image F3. In F3, the pixels in the area where the hand - shape is located are the pixels of the scaled - down hand - shape, and the pixel values outside the area where the hand - shape is located are 0, and the position of the palm center in F3 is the same as that in F2; extract the outer contour of the hand - shape in the foreground image F3, set the image pixels outside the outer contour to 0, and set the image pixels on and inside the outer contour to 255 to get hand_mask; in the background image B1, use the inverse mask umask = - hand_mask, and obtain the background image B without the hand - shape through bitwise operation UH ; in the background image B1, use the mask mask = hand_mask, and obtain the image B of the hand - shape through bitwise operation H ; for B UH ,B H ,F3 perform bitwise operation and transparency operation to obtain the fused image Combined_Image; the main controller in the guiding - end system transmits the fused video to the guided - end system through the network and displays it on the second video display device of the guided - end system, thus realizing the interactive video online remote assistance guidance. Description of the drawings

[0083] Figure 1 Schematic diagram of the auxiliary system for real-time remote guidance of key information fusion in online videos in the present invention;

[0084] Figure 2 Operation interfaces of the guiding end system and the guided end system in the auxiliary system;

[0085] The operation interface of the guiding end system mainly includes: ① Real-time display area of equipment and instruments, ② Selection of the connected guided end system, ③ IP address of the guided end system to be connected, ④ Connection button and button to agree to the connection request from the other party, ⑤ Display of connection status, ⑥ Disconnection button;

[0086] The operation interface of the guided end system mainly includes: ① Real-time display area of fusion results, ② Selection of the connected guiding end system, ③ IP address of the guiding end system to be connected, ④ Connection button and button to agree to the connection request from the other party, ⑤ Display of connection status, ⑥ Disconnection button;

[0087] Figure 3 Flowchart of the real-time remote guidance method of the present invention. Specific embodiments

[0088] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.

[0089] The above solution will be further described below in conjunction with specific implementation examples. The preferred embodiments of the present invention are described in detail as follows:

[0090] Embodiment 1:

[0091] See Figure 1The embodiment of the present invention provides a system of real-time remote guidance based on online video key information fusion, including a guidance end system and a guided end system; the guidance end system includes: a main controller, a first camera device, and a first video display device; the guided end system includes: a sub-controller, a second camera device, and a second video display device. In the guidance end system, the main controller is used to connect to the sub-controller of the guided end system, and the first camera device is used to shoot the first video display device used by the instructor; the instructor guides hand movements in front of the first video display device, which is shot by the first camera device and transmitted to the main controller; the main controller extracts the hand contour and integrates it into the instrument displayed by the first video display device; the fused picture is sent to the second video display device of the guided end system in real time. In the guided end system, the sub-controller is used to connect to the main controller in the guidance end system; the second camera device is used to shoot the instrument; the sub-controller transmits the video of the instrument to be operated to the guidance end system; the second video display device of the guided end system is used to display the fused picture in the guidance end system.

[0092] Embodiment 2:

[0093] See also Figure 1 and Figure 2 This embodiment is basically the same as the first embodiment, except that:

[0094] The operation interface of the guidance end system mainly includes: ① real-time display area of ​​equipment and instruments, ② selection of the guided end system to be connected, ③ the IP address of the guided end system to be connected, ④ connection button and button to agree to the other party's connection request, ⑤ connection status display, ⑥ stop connection button;

[0095] The operation interface of the guided end system mainly includes: ① real-time display area of ​​fusion results, ② selection of the guiding end system to be connected, ③ the IP address of the guiding end system to be connected, ④ connection button and button to agree to the other party's connection request, ⑤ connection status display, ⑥ stop connection button.

[0096] The interface operations are described in detail as follows:

[0097] (1-1) Equipment and instrument display area in the guidance end system interface: the online video of the equipment and instrument to be operated is displayed in the display area;

[0098] (1-2) Selecting the guided end system to be connected in the guiding end system interface: a connection method in which the connection can be made by selecting the guided end system ID number and the guided end system to be connected;

[0099] (1-3) The IP address of the guided end system to be connected in the guiding end system interface: a connection method in which the guided end system can be connected through the IP address of the guided end system to be connected;

[0100] (1-4) Connection button and button to agree to the connection request from the other party in the guiding end system interface: After selecting the connection method, use this button to connect to the guided end system. When the other party's consent is obtained, end-to-end real-time video transmission can be carried out; when a connection request is sent from the guided end system, also use this button to agree to the connection;

[0101] (1-5) Connection status display in the guiding end system interface: Used to observe the connection situation;

[0102] (1-6) Disconnection button in the guiding end system interface: Used to close the network connection between the guiding end system and the guided end system; when a connection request is sent from the other party, also use this button to decline the connection;

[0103] (1-7) Real-time display area for the fusion result in the guided end system interface: The video of the fusion result is displayed in real time on the display area;

[0104] (1-8) Guiding end system selected for connection in the guided end system interface: One of the connection methods, can connect to the guiding end system by selecting the ID number of the guiding end system to be connected;

[0105] (1-9) IP address of the guiding end system to be connected in the guided end system interface: One of the connection methods, can connect to the guiding end system by the IP of the guiding end system to be connected;

[0106] (1-10) Connection button in the guided end system interface: After selecting the connection method, use this button to connect to the guiding end system. When the other party's consent is obtained, end-to-end real-time video transmission can be carried out; when a connection request is sent from the guiding end system, also use this button to agree to the connection;

[0107] (1-11) Connection status display in the guided end system interface: Used to observe the connection situation;

[0108] (1-12) Disconnection button in the guided end system interface: Used to close the network connection between the guiding end system and the guided end system; when a connection request is sent from the other party, also use this button to decline the connection.

[0109] Example 3:

[0110] Refer to Figure 1 and Figure 3 , this embodiment of the present invention provides a real-time remote guidance method based on the fusion of key information of online videos:

[0111] (2-1) The guided-end system captures the real-time video of the instrument to be operated through the second imaging device, sends it to the main controller of the guiding-end system through the network, and displays it on the first video display device controlled by the main controller;

[0112] (2-2) The main controller in the guiding-end system captures the first video display device through the first imaging device, extracts the edge of the first video display device using the canny edge detection algorithm, takes the area of the first video display device as the working area, and cuts off all non-working areas in the video, making the working area fill the entire screen. This screen is the background screen B1;

[0113] (2-3) The main controller in the guiding-end system captures the gesture video of the operation demonstration in front of the first video display device through the first imaging device. Its image is the foreground image F1. Using the hand shape recognition method, the hand shape of the guiding personnel is recognized and extracted from the complex environment, and the complex background is removed. The position of the palm center in the working area of the first video display device is recorded. The method is as follows:

[0114] Construct a deep learning network model for hand extraction. The main network is MobileNetV3. A Squeeze-and-Excitation module is added to the feature extraction encoder in MobileNetV3 to recalibrate the feature channels, enhance the sensitivity of the network to key features, and enhance the expression ability of the model. Its formula is:

[0115] SE: S = σ(W 2 δ(W 1 AvgPool(X))) (1)

[0116] σ and δ represent the sigmoid and ReLU functions respectively. W 1 and W 2 are the weights of the fully connected layer, and X represents the input;

[0117] In the deep learning network model for hand extraction, a simplified LR-ASPP module is added to efficiently extract multi-scale context information and improve the feature expression ability for segmenting key content in the video. Its formula is:

[0118] aspp 1 = ReLU(BN(W 1 *X)) (2)

[0119] aspp 2 = σ(W 2 *GAP(X)) (3)

[0120] Y = aspp 1 ⊙aspp 2(4)

[0121] Among them, X represents the input, W1 and W2 represent weights, * represents the convolution operation, σ represents the sigmoid operation, ⊙ represents the Hadamard product operation, BN represents the batch normalization operation, GAP represents the global average pooling operation, and Y represents the output;

[0122] In the deep learning network model for hand extraction, a recurrent decoder is added, mainly including the convolutional long short-term memory network ConvLSTM and the convolutional gated recurrent unit ConvGRU, which are used to learn features in time and space, maintain the state information between video frames, and ensure the spatio-temporal dependence between video frames;

[0123] A lightweight feature enhancement module CBAM is added to the recurrent decoder of the deep learning network model for hand extraction, and spatial and channel attention are used to enhance the features, thereby improving the refinement ability for local and global features. The formula is:

[0124] M c =σ(MLP(AvgPool(X)) + MLP(MaxPool(X))) (5)

[0125] F 1 =M c X (6)

[0126] M s =σ(Conv([AvgPool(F 1 );MaxPool(F 1 )])) (7)

[0127] F out =M s F 1 (8)

[0128] Among them, X is the input feature map, AvgPool represents the average pooling operation, MaxPool represents the maximum pooling operation, and σ represents the sigmoid operation;

[0129] A depth-guided filtering module is added to the deep learning network model for hand extraction, which is used to improve the spatial resolution of images and videos while maintaining edge and detail information. The formula is:

[0130] O i =a k I i +b k (9)

[0131] Among them, the output image O is generated based on the linear relationship within the local regions of the input feature map I and the feature map G of the guidance image, and ak and b k are linear coefficients calculated on window k where pixel i is located, and can be obtained by minimizing the following cost function:

[0132]

[0133] where, ω k represents the window centered on pixel k, and ∈ is a regularization parameter;

[0134] Using the coordinates of each point on the outer contour of the hand shape, calculate and record the position of the palm center in the working area;

[0135] (2 - 4) For the hand - shaped image pixels extracted from the foreground image F1, the non - hand - shaped image part has its pixel values deleted and is given white, and this image is the foreground image F2;

[0136] (2 - 5) According to the convex hull size of the instrument image displayed on the first video display device controlled by the main controller in the guiding - end system, automatically scale the hand shape in the foreground image F2 so that the size of the hand shape is 1 / 4 of the convex hull of the instrument image, obtaining the foreground image F3. In F3, the pixels in the area where the hand shape is located are the pixels of the scaled - down hand shape, and the pixel values outside the area where the hand shape is located are 0. The position of the palm center in F3 is the same as that in F2;

[0137] (2 - 6) Extract the outer contour of the hand shape in the foreground image F3, set the image pixels outside the outer contour to 0, and set the image pixels on and inside the outer contour to 255, obtaining hand_mask;

[0138] (2 - 7) In the background image B1, use the inverse mask umask=-hand_mask, and obtain the background image BUH without the hand shape through bitwise operation. The formula is as follows:

[0139] BUH = B1 ∧ (umask) (11)

[0140] where, ∧ is the bitwise AND operation;

[0141] (2 - 8) In the background image B1, use the mask mask = hand_mask, and obtain the hand - shaped image BH through bitwise operation. The formula is as follows:

[0142] BH = B1 ∧ (mask) (12)

[0143] where, ∧ is the bitwise AND operation;

[0144] (2 - 9) Perform bitwise operation and transparency operation on BUH, BH, and F3 to obtain the fused image Combined_Image, as follows:

[0145] Combined_Image = BUH ∨ (α * BH + β * F3 + λ) (13)

[0146] Wherein, ∨ is a bitwise OR operation, α and β are corresponding weight factors used to control the transparency of the hand positions in BH and F3 in the output image, and can be set to 0.5 respectively. λ is a scalar value added to the final result and can be set to 0;

[0147] (2 - 10) In the guiding end system, the main controller transmits the fused video to the guided end system through the network and displays it on the second video display device of the guided end system, thereby realizing interactive online remote assisted guidance of the video.

[0148] The above has described the embodiments of the present invention in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments and can be variously changed according to the purpose of the invention of the present invention. Any changes, modifications, substitutions, combinations, or simplifications made based on the spirit and principle of the technical solution of the present invention shall be equivalent replacement methods. As long as they meet the invention purpose of the present invention and do not deviate from the technical principle and inventive concept of the present invention using the remote online action assistance system and method based on the fusion of key contents of online videos, they all fall within the protection scope of the present invention.

Claims

1. A real-time remote guidance system based on online video key information fusion, characterized in that: It includes the guiding end system and the guided end system; The guiding end system includes: a main controller, a first camera device, and a first video display device; the guided end system includes: a sub-controller, a second camera device, and a second video display device; In the guidance end system, the main controller is used to connect to the sub-controller of the guided end system, and the first camera device is used to shoot the first video display device used by the instructor; the instructor guides hand movements in front of the first video display device, which is shot by the first camera device and transmitted to the main controller; the main controller extracts the hand contour and integrates it into the instrument displayed by the first video display device; the fused picture is sent to the second video display device of the guided end system in real time; in the guided end system, the sub-controller is used to connect to the main controller in the guidance end system; the second camera device is used to shoot the instrument; the sub-controller transmits the video of the instrument to be operated to the guidance end system; the second video display device of the guided end system is used to display the fused picture in the guidance end system.

2. The real-time remote guidance system based on online video key information fusion according to claim 1 is characterized in that: The operation interface of the guidance end system includes: ① real-time display area of ​​equipment and instruments, ② selection of the guided end system to be connected, ③ the IP address of the guided end system to be connected, ④ connection button and button to agree to the other party's connection request, ⑤ connection status display, ⑥ stop connection button.

3. The real-time remote guidance system based on online video key information fusion according to claim 1 is characterized in that: The operation interface of the guided end system mainly includes: ① real-time display area of ​​fusion results, ② selection of the guiding end system to be connected, ③ the IP address of the guiding end system to be connected, ④ connection button and button to agree to the other party's connection request, ⑤ connection status display, ⑥ stop connection button.

4. The system for real-time remote guidance based on online video key information fusion according to claim 1, characterized in that: The technical scheme of the method is as follows: Step S1, the guided end system captures a real-time video of the instrument to be operated through a second camera device, sends the video to a main controller of the guiding end system through a network, and displays the video on a first video display device controlled by the main controller; Step S2, the main controller in the guidance end system captures the first video display device through the first camera device, extracts the edge of the first video display device, takes the first video display device area as the working area, and cuts off all non-working areas in the video so that the working area fills the entire screen, which is the background screen B1; Step S3, the main controller in the guidance end system captures the gesture video of the operation demonstration in front of the first video display device through the first camera device, and the image is the foreground image F1. The hand shape recognition method is used to recognize and extract the hand shape of the instructor from the complex environment, remove the complex background, and record the position of the palm of the hand in the working area of ​​the video display device: Step S4, extracting the hand image pixels from the foreground image F1, deleting the pixel values ​​of the non-hand image parts and assigning them white, and this image is the foreground image F2; Step S5, according to the size of the convex hull of the instrument image displayed on the first video display device controlled by the main controller in the guidance end system, the hand shape in the foreground image F2 is automatically scaled so that the size of the hand shape is 1 / 4 of the convex hull of the instrument image, and the foreground image F3 is obtained. In F3, the pixels in the area where the hand shape is located are the pixels of the reduced hand shape, and the pixel values ​​outside the area where the hand shape is located are 0, and the position of the palm of the hand in F3 is consistent with the position of the palm of the hand in F2; Step S6, extracting the outer contour of the hand in the foreground image F3, setting the image pixels outside the outer contour to 0, and setting the image pixels on and within the outer contour to 255, to obtain hand_mask; Step S7: Use the reverse mask umask=-hand_mask in the background image B1 to obtain the background image B1 without the hand shape through bit operation. UH ; Step S8: Use mask=hand_mask in the background image B1 to obtain the hand image B through bitwise operation. H ; Step S9, performing bitwise operation and transparency operation on BUH, BH, and F3 to obtain a fused image Combined_Image; Step S10: The main controller in the guiding end system transmits the fused video to the guided end system through the network, and displays it on the second video display device of the guided end system, thereby realizing interactive video online remote auxiliary guidance.

5. The method for real-time remote guidance based on online video key information fusion according to claim 4 is characterized in that: Using the hand shape recognition method, the instructor's hand shape is identified and extracted from a complex environment through a deep learning network model for hand extraction. The deep learning network model for hand extraction includes: the main network is MobileNetV3, a feature extraction encoder with a Squeeze-and-Excitation module, an LR-ASPP module, a loop decoder and a deep guided filtering module; among them, a lightweight feature enhancement module CBAM is added to the loop decoder, and the features are enhanced by using spatial and channel attention.

Citation Information

Patent Citations

  • Real-time 2D gesture estimation method based on loop architecture and coordinate system regression

    CN115953839A

  • Generation method and device of substation maintenance guidance video and processor

    CN118354033A

  • Security facility auditing system, method and device and computer readable storage medium

    CN119027067A

  • 3D entity digital magnifying glass system having 3D visual instruction function

    CN1973311A