Intelligent glasses and mobile phone double-video-stream synchronous acquisition converged communication method and system

By enabling wireless collaboration between smart glasses and mobile phones, simultaneous acquisition and fusion of dual-view video were achieved, solving the problem of collaborative work between two independent devices and enhancing the immersiveness and interactivity of video communication and live streaming.

CN122002007APending Publication Date: 2026-05-08SERXIN DIGITAL TECH (JIANGSU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SERXIN DIGITAL TECH (JIANGSU) CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies are insufficient to enable collaborative work between smart glasses and mobile phones, which are two independent devices. They cannot meet the requirements of simultaneous dual-view acquisition, low-latency fusion, and stable transmission, and cannot adapt to the diverse usage needs of users in various scenarios.

Method used

A dedicated communication link is established via Wi-Fi to enable device pairing and communication parameter adaptation between smart glasses and mobile phones. Third-person and first-person perspective videos are collected and synchronized using timestamp and frame-level alignment technology. Multiple layout modes and real-time marking tools are supported for video stream fusion and transmission.

Benefits of technology

It achieves simultaneous capture and seamless fusion of dual-view video, enhancing the immersiveness and interactivity of video communication and live streaming, and providing a rich visual experience and flexible content display capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002007A_ABST
    Figure CN122002007A_ABST
Patent Text Reader

Abstract

The invention discloses a smart glasses and mobile phone dual-video stream synchronous acquisition converged communication method and system, and relates to the technical field of video communication and live broadcast, and the method comprises the following steps: S1, the smart glasses and a mobile terminal establish a special communication link through Wi-Fi, and complete equipment pairing and communication parameter adaptation; and S2, the mobile terminal sends a synchronous acquisition instruction through the link, triggers a second image acquisition module of the mobile terminal and a first image acquisition module of the intelligent glasses to be started at the same time, and respectively acquires view angle videos of the third person and the first person. According to the intelligent glasses and mobile phone dual-video stream synchronous acquisition converged communication method and system, synchronous capture and seamless fusion of dual-view videos are realized, the immersion and interactivity of video communication and live broadcast are remarkably enhanced, and through wireless cooperative work of the intelligent glasses and the mobile phone, the real-time performance of the video communication and live broadcast is improved. The visual angle limitation of traditional single camera equipment is broken, and richer and more comprehensive visual experience is provided for audiences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video communication and live streaming technology, specifically to a method and system for simultaneous acquisition and fusion communication of dual video streams from smart glasses and mobile phones. Background Technology

[0002] With the continuous development of video communication and live streaming technologies, their application scenarios have widely covered multiple fields such as teaching demonstrations, product evaluations, and outdoor activities. In actual use, users have gradually developed a core need for simultaneous dual-perspective display. They want to present a first-person real-time view of what they see, intuitively conveying the scene and details, while also showing their third-person perspective within the environment, allowing viewers to have a more comprehensive understanding of the overall situation. This complementary dual-perspective display method can significantly enhance content expressiveness and interactive experience.

[0003] Currently, most mainstream video capture solutions rely on a single camera device, resulting in a relatively fixed field of view, which makes it difficult to meet the demand for simultaneous dual-view presentation. Even some solutions that support multi-camera switching are mostly limited to different cameras on the same device, significantly restricting the expansion of the field of view. Furthermore, existing technologies lack a solution that can easily enable collaborative work between two independent devices such as smart glasses and mobile terminals. Mature solutions for simultaneous capture of dual video streams, low-latency fusion, and stable transmission have not yet been developed, failing to fully adapt to users' diverse needs in various scenarios. To address this, we propose a method and system for simultaneous capture and fusion communication of dual video streams between smart glasses and mobile phones. Summary of the Invention

[0004] To address the aforementioned technical issues, this paper provides a method and system for simultaneous acquisition and fusion communication of dual video streams from smart glasses and mobile phones. This technical solution solves the problems of existing technologies that rely heavily on a single camera device or multiple cameras on the same device, resulting in limited field of view, difficulty in achieving simultaneous presentation of dual perspectives, lack of collaborative solutions for dual independent devices such as smart glasses and mobile terminals, and the absence of mature solutions for simultaneous acquisition of dual video streams, low-latency fusion, and stable transmission, thus failing to meet the diverse needs of users in various scenarios.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for simultaneous acquisition and fusion communication of dual video streams from smart glasses and mobile phones includes the following steps: S1. The smart glasses and the mobile terminal establish a dedicated communication link via Wi-Fi to complete device pairing and communication parameter adaptation. S2. The mobile terminal sends a synchronous acquisition command through this link, triggering its own second image acquisition module and the first image acquisition module of the smart glasses to start simultaneously, and acquire third-person and first-person perspective videos respectively; S3: The smart glasses encode the first video stream in real time and transmit it to the mobile terminal via Wi-Fi. S4. The mobile terminal receives the first video stream and the second video stream it has collected, and uses local timestamp and frame-level alignment technology to synchronize and calibrate the two video streams. S5 and mobile terminals offer multiple layout modes and support users in switching the display priority of the main and secondary screens; S6. Users can select the vector marking tool through touch operation in the mobile terminal preview or live broadcast interface, and the marking is superimposed on the video screen in real time. S7. The mobile terminal re-encodes the fused video stream after synchronization, layout, and marking processing, and sends it to the receiving end via the network.

[0006] Preferably, S1 includes: The smart glasses activate the Wi-Fi module, switch to AP mode or Wi-Fi Direct mode, and broadcast a connection signal containing the device model and unique identifier; the mobile terminal activates the Wi-Fi function, searches for signals of nearby dedicated communication devices, filters out the broadcast signal of the target smart glasses, and initiates a connection request. After receiving a connection request from a mobile terminal, the smart glasses send pairing verification information to the mobile terminal. The verification method is either comparing the device's unique identifier or having the user enter a preset password. Once the verification is successful, the device pairing and binding is completed. Once paired successfully, both parties will automatically negotiate the transmission protocol. Synchronize the system clock references of both devices; The system negotiates the video encoding format, detects the current network environment, sets the data transmission bandwidth threshold and latency compensation parameters, generates a communication parameter configuration file, and synchronizes it to both devices.

[0007] Preferably, S2 includes: Based on the established dedicated communication link, the mobile terminal generates a synchronous acquisition command that includes acquisition resolution, frame rate, and exposure parameters, and embeds a local reference timestamp. The mobile terminal transmits the synchronous acquisition command to the smart glasses via a dedicated link, and simultaneously starts the command transmission timing; After receiving the command, the smart glasses decode and extract the configuration information and the reference timestamp. The main processing chip triggers the first image acquisition module to start and initializes the acquisition state according to the configuration parameters. After the smart glasses start collecting data, they send a ready signal to the mobile terminal. Upon receiving the signal, the mobile terminal immediately activates its second image acquisition module. The mobile terminal determines the start time difference between the two parties by the timing difference. If the difference exceeds the set threshold, the synchronization command is resent. After both acquisition modules are working stably, the smart glasses acquire first-person perspective video, and the mobile terminal acquires third-person perspective video.

[0008] Preferably, S3 includes: The main processing chip of the smart glasses receives the raw video data transmitted by the first image acquisition module and starts the hardware encoding engine; The encoding engine performs frame-level compression on the raw video data according to the negotiated format, and assigns a unique frame number and acquisition timestamp to each video frame, which are then associated with the corresponding video frame. Start the transmission buffer queue and store the encoded video frames into the queue in order of frame number; Based on the real-time transmission status of the Wi-Fi link, the video frames in the queue are segmented and processed, and each segment is transmitted after adding a checksum. A data retransmission mechanism is set up so that if the mobile terminal reports that a certain segment is lost, the smart glasses will retrieve the corresponding segment from the cache queue and retransmit it.

[0009] Preferably, S4 includes: The mobile terminal receives the first video stream transmitted by the smart glasses through a dedicated link, and simultaneously collects its own second video stream, establishing frame data storage buffers for the two video streams respectively. Extract the acquisition timestamp from each frame of data from the two video streams to generate a first timestamp sequence and a second timestamp sequence; Calculate the time difference between corresponding frames in two timestamp sequences, set a synchronization threshold, and filter out frame data whose time difference exceeds the threshold. For frames exceeding the threshold, frame delay buffering or redundant frame discarding is used for adjustment, and the frame-level timestamps of the first video stream and the second video stream are matched and aligned frame by frame.

[0010] Perform image misalignment correction on the two aligned video streams.

[0011] Preferably, S5 includes: The mobile terminal control APP has multiple preset layout modes, and corresponding layout selection controls are set in the user interface; Users select the target layout mode through touch controls. After receiving the selection command, the APP reads the screen resolution parameters of the mobile terminal. The display areas for the two video streams are allocated based on the selected layout mode and screen resolution. The screen size of the two video streams is adaptively scaled to maintain the original aspect ratio; The adjusted two video streams will be rendered and displayed according to the selected layout mode.

[0012] Preferably, S5 further includes: The mobile terminal interactive interface is configured with a main and secondary screen switching control, which associates the display priority indicators of the first and second video streams. When a user triggers the switching control, the APP receives the switching command and reads the display status of the current main and secondary screens. Swap the display priorities of the two video streams, switch the original secondary screen to the main screen and expand it to the corresponding display area, and switch the original main screen to the secondary screen and adjust its size;

[0013] Update the interface display status.

[0014] Preferably, S6 includes: The mobile terminal preview or live broadcast interface features a marker toolbar containing various vector marker tools, which users can select by touch. Users touch and slide to draw and mark trajectories on the interface. The APP captures touch coordinate data in real time and generates a continuous sequence of trajectory points. Convert the trajectory point sequence into Bézier curve vector graphics data and remove redundant coordinate points; Provides options for customizing marker styles; The processed vector marker data is superimposed onto the corresponding coordinate positions of the video frame in real time; Provides mark undo and clear functions.

[0015] Preferably, S7 includes: The mobile terminal re-encodes the fused video stream that has undergone synchronization calibration, layout adjustment, and marker overlay. Real-time monitoring of current network bandwidth status; dynamically adjusting the encoding bitrate based on bandwidth data. Select the corresponding transmission protocol; The encoded fused video stream is encapsulated according to the protocol format and transmitted in segments to the receiving end. Monitor network transmission status in real time and adjust video resolution according to bandwidth conditions; If the receiving end reports data loss, a partial retransmission mechanism is triggered to retransmit the lost fragments.

[0016] A dual-video stream synchronous acquisition and fusion communication system for smart glasses and mobile phones includes: Communication connection module, video acquisition and synchronization module, video fusion and processing module, user interaction module, encoding and transmission module; The communication connection module is used to establish a dedicated link between smart glasses and mobile terminals via Wi-Fi, complete device pairing and communication parameter adaptation, and provide a foundation for dual video stream transmission; The video acquisition synchronization module is electrically connected to the communication connection module and is used to receive the synchronization command from the mobile terminal, triggering both image acquisition modules to start simultaneously, acquire dual-view video respectively and achieve synchronous transmission. The video fusion processing module is electrically connected to the video acquisition synchronization module. It is used to receive two video streams, calibrate them using timestamp and frame-level alignment technology, and complete video fusion in combination with the selected layout mode. The user interaction module is electrically connected to the video fusion processing module, and is used to provide layout selection and main / sub screen switching functions, and supports users to add vector markers and overlay them in real time through touch operation; The encoding and transmission module is electrically connected to the user interaction module and is used to re-encode the fused video stream and send it to the other end of the call or the receiving end of the live streaming platform via the network.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a method and system for synchronous acquisition and fusion communication of dual video streams using smart glasses and mobile phones. This achieves synchronous capture and seamless fusion of dual-view video, significantly enhancing the immersiveness and interactivity of video communication and live streaming. Through the wireless collaborative work of smart glasses and mobile phones, it breaks the perspective limitations of traditional single-camera devices, allowing users to simultaneously view the intuitive scene from a first-person perspective and the overall environment from a third-person perspective, providing viewers with a richer and more comprehensive visual experience. Its unique layout mode and dynamic main and secondary screen switching function further enhance the flexibility and adaptability of content display. The built-in finger marking tool allows users to add graphic marks to the video screen in real time, greatly enhancing the interactivity and accuracy of live explanations. Through optimized synchronization mechanisms and timestamp alignment technology, low-latency synchronization of the two video streams is ensured, providing users with a smooth and stable video communication experience. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the method steps of the present invention; Figure 2 This is a block diagram of the overall system architecture of the present invention. Detailed Implementation

[0019] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0020] Reference Figure 1 As shown, the method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones includes the following steps: S1. The smart glasses and the mobile terminal establish a dedicated communication link via Wi-Fi to complete device pairing and communication parameter adaptation. S1 includes: The smart glasses activate the Wi-Fi module, switch to AP mode or Wi-Fi Direct mode, and broadcast a connection signal containing the device model and unique identifier; The Wi-Fi module's working mode switching is achieved through the configuration of the mode control register inside the module. When switching to AP mode, the smart glasses act as an access point, allocate IP address ranges, and start DHCP service to assign unique IPs to subsequent mobile terminals. When switching to Wi-Fi Direct mode, the smart glasses, as the group owner, initiates the creation of a P2P group, setting the group ID and channel parameters. The broadcast signal is sent in the form of beacon frames with a transmission period of 100ms. The beacon frame carries information such as the device's unique identifier, device model, and supported communication protocol types. The mobile terminal enables Wi-Fi, searches for signals from nearby dedicated communication devices, and filters out the target smart glasses' broadcast signal by parsing the device model and unique identifier fields in the beacon frame, initiating a connection request. The connection request is in the form of a probe request frame, carrying the mobile terminal's own device information and supported encryption methods.

[0021] After receiving a connection request from a mobile terminal, the smart glasses send pairing verification information to the mobile terminal. The verification method is either device unique identifier comparison or user input of a preset password. If the verification is successful, the device pairing and binding is completed. The pairing verification information is transmitted using EAP authentication frames. If the device unique identifier comparison method is used, the smart glasses extract the unique identifier of the mobile terminal from the connection request and compare it with the locally stored list of trusted devices. If the comparison is successful, an authentication success response frame is generated. If the user inputs a preset password, the smart glasses generate a random challenge code and send it to the mobile terminal along with the verification request frame. Upon receiving the code, the mobile terminal prompts the user to enter the preset password, performs a hash operation on the password and challenge code, and returns the result. The smart glasses compare the received result with the result of the same local operation; if they match, the verification is successful. After successful verification, both parties negotiate a session key via a four-way handshake protocol. Subsequent communication data is encrypted using the AES-128 encryption algorithm, completing the device pairing and binding.

[0022] After successful pairing, both parties automatically negotiate the transmission protocol. The transmission protocol negotiation adopts a negotiation request-response mechanism. The smart glasses, as the initiator, send a protocol negotiation request frame, which carries the supported transmission protocol types and parameters. Supported transmission protocols include UDP, TCP, and RTP. UDP is suitable for low-latency scenarios, TCP is suitable for scenarios with high reliability requirements, and RTP is used for real-time video streaming.

[0023] During the negotiation process, both parties initially select the RTP+UDP combined protocol. If the network packet loss rate exceeds 5%, the system automatically switches to the RTP+TCP combined protocol. Protocol parameter configuration includes RTP timestamp increment, payload type, and UDP port number, while TCP parameters include connection timeout and retransmission count.

[0024] The system clock reference of both devices is synchronized. The clock reference is obtained using the Network Time Protocol. The standard timestamp is obtained first from the public network NTP server accessed by the mobile terminal. If the mobile terminal has no public network connection, the UTC time obtained by the GPS module inside the smart glasses is used as the reference clock. If there is no GPS module, the system clock of the mobile terminal is used as the reference.

[0025] The synchronization algorithm uses an improved NTP algorithm, which calculates network latency and clock offset through multiple bidirectional timestamp interactions. The specific steps are as follows: the smart glasses send a synchronization request with local timestamp T1 to the mobile terminal; the mobile terminal records the local timestamp T2 upon receiving the request and sends a response frame with T2 and T3 to the smart glasses; the smart glasses record the local timestamp T4 upon receiving the response frame.

[0026] Network latency (Where T1 is the local timestamp of the smart glasses sending the synchronization request, T2 is the local timestamp of the mobile terminal receiving the synchronization request, T3 is the local timestamp of the mobile terminal sending the response frame, and T4 is the local timestamp of the smart glasses receiving the response frame), clock offset (in the formula) (This refers to the clock offset between the mobile terminal and the smart glasses).

[0027] The system clock of the smart glasses is calibrated based on the calculated clock offset. After calibration, a clock synchronization maintenance mechanism is activated, performing a bidirectional timestamp check every 500ms. If a clock offset exceeding 10ms is detected, the calibration process is re-executed. Error compensation uses a linear interpolation algorithm to dynamically compensate for clock deviations within the two calibration intervals, ensuring clock synchronization accuracy.

[0028] The system negotiates the video encoding format, detects the current network environment, sets the data transmission bandwidth threshold and latency compensation parameters, generates a communication parameter configuration file, and synchronizes it to both devices.

[0029] The video encoding format negotiation supports three standards: H.264, H.265 (HEVC), and VP9. The negotiation priority is H.265 > H.264 > VP9, ​​prioritizing the format with higher encoding efficiency. If the smart glasses or mobile terminal do not support H.265, it will be downgraded to H.264. If neither is supported, VP9 will be selected.

[0030] The encoding parameter selection logic is determined based on the device hardware performance and acquisition resolution: when the acquisition resolution is 1080P, the target bitrate for H.265 encoding is set to 4-6Mbps, and for H.264 it is set to 6-8Mbps; when the resolution is 720P, the target bitrate for H.265 is 2-3Mbps, and for H.264 it is 3-4Mbps; the frame rate is 30fps by default, and can be adjusted to 25fps or 20fps according to hardware performance.

[0031] Network environment detection involves continuously sending 10 probe data packets and calculating the round-trip time, packet loss rate, and available bandwidth. Available bandwidth is calculated using a throughput-based estimation algorithm. ; In the formula For available bandwidth, For packet loss rate, This represents the actual measured throughput.

[0032] The bandwidth threshold is set to 70% of the available bandwidth, that is: ; In the formula A data transmission bandwidth threshold is set to ensure sufficient bandwidth is reserved to cope with network fluctuations. The delay compensation parameter is determined based on the network delay 'd', and the delay compensation time... In the formula The delay compensation time is d, where d is the network latency and 20ms is the reserved buffer time. If the latency exceeds 100ms, a dynamic latency compensation adjustment mechanism is activated. This reduces the amount of data by lowering the video encoding bitrate, thereby reducing transmission latency. The generated communication parameter configuration file is in JSON format and includes information such as encoding format, bitrate, frame rate, transmission protocol type, port number, bandwidth threshold, and latency compensation time. It is synchronized to both devices via an encrypted channel and stored in the local configuration directory.

[0033] S2. The mobile terminal sends a synchronous acquisition command through this link, triggering its own second image acquisition module and the first image acquisition module of the smart glasses to start simultaneously, and acquire third-person and first-person perspective videos respectively; S2 includes: Based on the established dedicated communication link, the mobile terminal generates a synchronous acquisition command that includes acquisition resolution, frame rate, and exposure parameters, and embeds a local reference timestamp. The timestamp generation is based on the synchronized system clock reference and adopts the Unix timestamp format. The generation principle is to use the number of seconds of the reference clock as a basis and add a millisecond-level offset to ensure the uniqueness and accuracy of the timestamp.

[0034] The synchronous acquisition command adopts a binary encoding format, and the command structure is: command header + command length + acquisition parameter field + timestamp field + check code.

[0035] In the acquisition parameter field, the acquisition resolution is represented by 2 bytes, the frame rate is represented by 1 byte, and the exposure parameters include the exposure time and ISO sensitivity.

[0036] The mobile terminal transmits the synchronization acquisition command to the smart glasses via a dedicated link, and simultaneously initiates a command transmission timing mechanism. The timing utilizes a high-precision timer internal to the mobile terminal, with a timer accuracy of 1ms. The start time is recorded as follows: , The local time at which the synchronous data collection command is sent to the mobile terminal.

[0037] After receiving the command, the smart glasses decode and extract the configuration information and the reference timestamp. The main processing chip triggers the first image acquisition module to start and initializes the acquisition state according to the configuration parameters. After receiving the command, the smart glasses first verify the verification code. If the verification is successful, the command field is parsed and the acquisition parameters and the reference timestamp are extracted.

[0038] The main processing chip sends configuration commands to the first image acquisition module via the I2C interface to set the resolution and frame rate parameters, and configures the exposure time and ISO parameters of the image signal processor via the SPI interface to complete initialization. After initialization, the main processing chip generates a hardware trigger signal to trigger the first image acquisition module to start acquisition.

[0039] After the smart glasses start collecting data, they send a ready signal to the mobile terminal. Upon receiving the signal, the mobile terminal immediately activates its second image acquisition module. The ready signal is in binary format and includes the smart glasses' local startup timestamp. It is transmitted to the mobile terminal via a dedicated link.

[0040] The mobile terminal records the moment it receives the ready signal and simultaneously triggers the hardware of its second image acquisition module to activate the acquisition module and begin acquisition, recording the start time.

[0041] The mobile terminal determines the startup time difference between the two parties by using the timing difference. If the difference exceeds a set threshold, the synchronization command is resent. The startup time difference is calculated using two methods: First, calculate the time difference between when the mobile terminal sends the command and when it receives the ready signal: ; Round-trip time for instruction transmission; Second, calculate the time difference between the actual start times of both parties: ; In the formula The startup time difference between the two-end acquisition modules Here is the estimated one-way delay for command transmission, where This is an estimated one-way delay for command transmission. The set startup time difference threshold is 33ms. If the time exceeds 33ms, the mobile terminal will regenerate and resend the synchronization acquisition command, up to a maximum of 3 times. If the requirement is still not met, the user will be prompted that the communication link is abnormal.

[0042] The correction algorithm employs a delayed start mechanism. If the smart glasses start too early, the mobile terminal starts its own data acquisition module after a delay. If the smart glasses start too late, the mobile terminal resends the synchronization command, which carries the corrected reference timestamp, thus delaying the smart glasses' start.

[0043] After both acquisition modules are working stably, the smart glasses acquire first-person perspective video, and the mobile terminal acquires third-person perspective video.

[0044] S3: The smart glasses encode the first video stream in real time and transmit it to the mobile terminal via Wi-Fi. S3 includes: The main processing chip of the smart glasses receives the raw video data transmitted by the first image acquisition module and starts the hardware encoding engine; The raw video data is in YUV420 format and is transmitted to the buffer of the main processing chip via the MIPI CSI-2 interface. The main processing chip transmits the data to the hardware encoding engine via the AXI bus and configures the working mode and parameters of the encoding engine.

[0045] The hardware encoding engine works as follows: First, the raw YUV data is preprocessed, and then intra-frame prediction and inter-frame prediction are performed. Intra-frame prediction uses prediction blocks of 4×4, 8×8, and 16×16 sizes, while inter-frame prediction uses motion estimation and motion compensation techniques, with a search range of ±32 pixels. Next, transform coding and quantization are performed. The quantization parameter (QP) is dynamically adjusted according to the target bitrate. When the bitrate is too high, the QP value is increased, and when the bitrate is too low, the QP value is decreased. Finally, entropy coding is performed to generate the encoded bitstream.

[0046] The encoding engine performs frame-level compression on the raw video data according to the negotiated format, and assigns a unique frame number and acquisition timestamp to each video frame, which are then associated with the corresponding video frame. The frame sequence number uses a 32-bit unsigned integer, starting from 0 and incrementing sequentially. Each frame is assigned a unique sequence number, and keyframes, prediction frames, and bidirectional prediction frames are numbered sequentially. The acquisition timestamp is generated based on the synchronized system clock, accurate to milliseconds, and corresponds one-to-one with the acquisition time of the video frame. It is embedded in the encoded bitstream via SEI messages and transmitted along with the encoded data.

[0047] The association mechanism stores the frame sequence number and timestamp in the header information of the encoded frame, forming a key-value pair, which facilitates subsequent synchronization calibration.

[0048] A transmission buffer queue is started, and encoded video frames are stored in the queue in order of frame number. The transmission buffer queue adopts a circular queue structure with a queue length of 30 frames. The enqueue and dequeue operations of the queue are thread-safe through a mutex lock. The queue management strategy is as follows: the first-in, first-out (FIFO) principle is adopted. When the queue is full, the earliest enqueued non-critical frame is discarded to ensure that the queue has enough space to store new encoded frames. The frame data in the queue is checked periodically. If a frame data has been in the queue for more than 500ms, it is forcibly dequeued and discarded to avoid queue blocking.

[0049] Based on the real-time transmission status of the Wi-Fi link, the video frames in the queue are processed into segments, and each segment is transmitted after adding a checksum. The segment size is dynamically adjusted according to the current network bandwidth. When the bandwidth is ≥4Mbps, the segment size is set to 1460 bytes. When the bandwidth is between 2-4 Mbps, the fragment size is set to 730 bytes; when the bandwidth is <2 Mbps, the fragment size is set to 365 bytes. During fragmentation processing, a fragment header is added to each fragment, containing information such as frame sequence number, fragment sequence number, total number of fragments, and fragment length. The checksum uses the CRC-32 algorithm, calculated on the fragment header and fragment data, and the checksum field is added to the fragment tail.

[0050] A data retransmission mechanism is set up so that if the mobile terminal reports that a certain segment is lost, the smart glasses retrieve the corresponding segment from the cache queue and retransmit it. After receiving the segment, the mobile terminal verifies each segment. If the verification passes, it stores the segment. If a segment of a frame is found to be missing, a packet loss feedback frame is generated. The frame carries the sequence number of the missing frame and the segment sequence number and is sent to the smart glasses through a dedicated link.

[0051] Packet loss detection employs a sliding window mechanism. The mobile terminal maintains a receiving window and checks the continuity of fragments within the window. The retransmission trigger condition is: receiving packet loss feedback from the mobile terminal, or failing to receive a corresponding acknowledgment frame within a timeout period after the smart glasses send a fragment. The timeout period is dynamically adjusted based on network latency; the greater the latency, the longer the timeout period.

[0052] After receiving a retransmission request, the smart glasses search for the corresponding fragment of the corresponding frame in the transmission buffer queue, prioritize retransmitting the fragment of the key frame, add a retransmission mark to the retransmitted fragment, and retransmit the fragment up to 3 times. If it is still unsuccessful after 3 retransmissions, the frame data corresponding to the fragment is discarded.

[0053] S4. The mobile terminal receives the first video stream and the second video stream it has collected, and uses local timestamp and frame-level alignment technology to synchronize and calibrate the two video streams. S4 includes: The mobile terminal receives the first video stream transmitted by the smart glasses through a dedicated link, and simultaneously collects its own second video stream, establishing frame data storage buffers for the two video streams respectively. Both the receive buffer for the first video stream and the capture buffer for the second video stream adopt a circular buffer structure, with a buffer size set to 60 frames. Read and write operations of the buffers are synchronized using semaphores to avoid read and write conflicts. Each buffer stores the corresponding frame data, frame sequence number, capture timestamp, and frame type information.

[0054] Extract the acquisition timestamp from each frame of data from the two video streams to generate a first timestamp sequence and a second timestamp sequence.

[0055] Calculate the time difference between corresponding frames in two timestamp sequences, set a synchronization threshold, and filter out frame data whose time difference exceeds the threshold. The time difference calculation adopts a dynamic matching algorithm, which traverses the two timestamp sequences, finds the frame pair that satisfies the minimum, takes the frame pair as the corresponding frame, and calculates the time difference.

[0056] The synchronization threshold is determined based on the video frame rate. The synchronization threshold is set to 16.7ms for 30fps video, 20ms for 25fps, and 25ms for 20fps video. When the threshold is reached, the frame pair is determined to be out of sync.

[0057] For frames exceeding the threshold, frame delay buffering or redundant frame discarding is used for adjustment, and the frame-level timestamps of the first video stream and the second video stream are matched and aligned frame by frame. The frame delay buffering method is suitable for situations where a frame from one video stream arrives early. The early-arriving frame is stored in a delay buffer and retrieved for subsequent processing after the corresponding frame from the other video stream arrives. The delay buffering time is no more than 500ms; if it exceeds this time, the frame is discarded.

[0058] The redundant frame dropping method is applicable when a video stream frame arrives late. When the delay time exceeds the synchronization threshold, the delayed frame in the video stream is dropped and the next arriving frame is directly matched.

[0059] The alignment algorithm employs linear interpolation frame completion technology. If there is a difference in the number of frames between two video streams, interpolation is performed on the stream with fewer frames. The completed frame data is generated based on the pixel information of the preceding and following frames, ensuring that the number of frames in the two aligned video streams is consistent. Timestamps are matched frame by frame. Synchronization threshold.

[0060] Perform image misalignment correction on the two aligned video streams.

[0061] Image misalignment correction includes two techniques: spatial matching and temporal interpolation. Spatial matching uses a feature point matching algorithm to extract SIFT feature points from the aligned frames of the two video streams, calculate the Euclidean distance between the feature points, select matching feature point pairs, and remove abnormal matching points through a random sampling consensus algorithm to obtain a spatial transformation matrix. This matrix is ​​then used to perform spatial transformation on the image of one of the video streams to achieve spatial alignment between the two images.

[0062] Temporal interpolation is suitable for situations with slight time misalignment. It uses a bilinear interpolation algorithm to interpolate the pixel values ​​of video frames in the time domain, compensating for the misalignment caused by small time differences and ensuring that the video movements of the two video streams are synchronized.

[0063] S5 and mobile terminals offer multiple layout modes and support users in switching the display priority of the main and secondary screens; S5 includes: The mobile terminal control APP has multiple preset layout modes, including picture-in-picture, split-screen, and full-screen switching modes, and corresponding layout selection controls are set in the user interface.

[0064] Users select the target layout mode via touch controls. After receiving the selection command, the APP reads the screen resolution parameters of the mobile terminal.

[0065] Based on the selected layout mode and screen resolution, the display areas for the two video streams are allocated; the display area allocation uses a coordinate mapping calculation method: If picture-in-picture mode is selected, the coordinates of the main screen display area are (0, 0, W, H), and the coordinates of the secondary screen display area are... ; If you select the split-screen mode, the coordinates of the main screen and the secondary screen display areas are as follows: and ; If you select the split-screen mode, the coordinates are as follows: and During the allocation process, ensure that the display area coordinates are integer pixels to avoid image blurring.

[0066] The screen size of the two video streams is adaptively scaled to maintain the original aspect ratio; Adaptive scaling uses a proportional scaling algorithm, assuming the original video resolution is... ( The original width of the video. (The original height of the video), the target display area resolution is... ( Define the width of the target display area. Calculate the scaling ratio for the target display area height. , ,Pick ( To determine the final scaling ratio, take... and (The smaller value in the range to avoid image stretching) results in a scaled video resolution of [value missing]. .

[0067] If the scaled video resolution is smaller than the target display area resolution, it will be centered within the display area with a black border around it; if the scaled video resolution is larger than the target display area resolution, the excess portion will be cropped proportionally to ensure the video image is complete and free from stretching or distortion.

[0068] The adjusted two video streams are rendered and displayed according to the selected layout mode. The rendering process uses layer compositing technology, with the main image as the bottom layer and the secondary image as the top layer. The images are composited using the GPU's layer mixer, and the transparency of the secondary image is set during compositing. The composited image data is then output to the screen through the display driver chip. Simultaneously, a display synchronization mechanism is activated to ensure that the rendering frame rate matches the video frame rate, preventing stuttering or tearing.

[0069] The S5 also includes: The mobile terminal's interactive interface features a main and secondary screen switching control, which associates the display priority indicators of the first and second video streams. The priority indicators are stored as Boolean variables in the app's memory and are updated in real time.

[0070] When a user triggers the switching control, the app receives the switching command and reads the current display state of the primary and secondary screens. The switching command is obtained through touch event listening; upon receiving the command, the app reads the priority flag in memory. Determine the video stream corresponding to the current main and secondary screens.

[0071] Swap the display priorities of the two video streams, switch the original secondary screen to the main screen and expand it to the corresponding display area, and switch the original main screen to the secondary screen and adjust its size; During the exchange process, the priority flag value is first updated, and then the display area coordinates are recalculated and allocated according to the new priority flag. The adaptive scaling algorithm is then re-executed on both video streams to adjust the screen size.

[0072] Update the interface display status. To ensure a smooth transition during the switching process, a fade-in / fade-out animation effect is used, with a transition time set to 300ms. During the transition, both the old and new layouts are rendered simultaneously, gradually decreasing the transparency of the old layout and increasing the transparency of the new layout. After the transition ends, only the new layout is rendered, completing the interface display status update.

[0073] S6. Users can select the vector marking tool through touch operation in the mobile terminal preview or live broadcast interface, and the marking is superimposed on the video screen in real time. S6 includes: The mobile terminal preview or live broadcast interface features a marker toolbar containing various vector marker tools. Users can select the target tool by touching it. The toolbar is a floating window located on the right side of the screen. Each tool has a unique icon, which is highlighted when the user touches it, indicating the currently selected tool type.

[0074] Users touch and slide to draw a marked trajectory on the interface. The APP captures the touch coordinate data in real time and generates a continuous trajectory point sequence. The touch coordinate data is obtained through the system's touch event callback function. Each touch event includes the x and y coordinates of the touch point and a timestamp. The APP collects coordinate data every 10ms to generate a trajectory point sequence.

[0075] The trajectory point sequence is converted into Bézier curve vector graphics data, and redundant coordinate points are removed. The Bézier curve is generated using a cubic Bézier curve mathematical model. For a trajectory point sequence P, three adjacent points are taken as a group, with the first point as the starting point, the third point as the ending point, and the second point as the control point, to generate a cubic Bézier curve. The curve equation is: ; In the formula Let be any point on the Bézier curve. For parameters and , As the starting point of the curve, As control points, (the endpoint of the curve), where .

[0076] Redundant coordinate point removal uses the Douglas-Peucker algorithm, with a set threshold. Pixels (As a distance threshold from a point to the fitted line), iterate through the sequence of trajectory points, calculate the distance from each point to the fitted line, and if the distance is less than a certain threshold... If the point is not cleared, then the key turning point is retained, reducing the amount of data while ensuring trajectory accuracy.

[0077] It provides options for customizing marker styles, including line color, line width, and line type. Users can select or input parameters via touch controls, and the app stores the style parameters as vector graphic attributes.

[0078] The processed vector marker data is superimposed onto the corresponding coordinate positions of the video frame in real time; the superimposition process uses coordinate mapping technology to convert the screen touch coordinates into normalized coordinates of the video frame. ; ; In the formula , For normalized coordinates, , The screen coordinates of the touch point. , The screen's width and height are given in pixels. Then, based on the coordinates of the video image's display area on the screen, the normalized coordinates are converted into the video image's pixel coordinates. ; ; In the formula , To mark the pixel coordinates on the video frame, , The width and height pixel values ​​of the video on the screen are set to ensure that the marked position is consistent with the user's drawing position. Then, the Bézier curve vector graphics are drawn onto the corresponding coordinate positions of the video frame using the GPU's vector renderer, achieving real-time overlay.

[0079] It provides mark undo and clear functions; the undo function uses a stack structure to store the mark data drawn each time. When the user triggers the undo command, the mark data drawn last time is popped from the stack and deleted; the clear function clears all mark data in the stack, refreshes the video screen, and removes all marks.

[0080] S7. The mobile terminal re-encodes the fused video stream after synchronization, layout, and marking processing, and sends it to the receiving end via the network.

[0081] S7 includes: The mobile terminal re-encodes the merged video stream after synchronization calibration, layout adjustment, and marker overlay. Re-encoding utilizes the mobile terminal's built-in hardware encoding engine. The encoding format is selected based on the receiver's support, prioritizing H.265 encoding; if the receiver does not support it, H.264 encoding is used. Initial encoding parameters are determined based on the merged video resolution: the initial target bitrate for 1080P merged video is 6Mbps, and for 720P it is 3Mbps, while the frame rate remains consistent with the original video.

[0082] Real-time monitoring of current network bandwidth status; dynamically adjusting the encoding bitrate based on bandwidth data. Dynamic bitrate adjustment uses a PID control algorithm, with the target bitrate set to 1. The actual available network bandwidth is Bitrate adjustment amount: ; In the formula This is the bitrate adjustment amount. This is the proportionality coefficient. The integral coefficient is... These are the differential coefficients. This represents the actual available network bandwidth. The current encoding bitrate, The integral term of the error. This is the differential term of the error; in (Proportion coefficient) (Integral coefficient) (Differential coefficients). When At that time, increase the bitrate ; when At that time, reduce the bit rate The bitrate adjustment range is 50%-150% of the initial value to avoid excessive bitrate fluctuations that could cause drastic changes in image quality. Network bandwidth detection uses a periodic sending of probe packets, once every 1 second. The probe packet size is 1KB, and the available bandwidth is obtained by calculating the transmission rate of the probe packets.

[0083] Select the corresponding transmission protocol; select the protocol according to the transmission scenario: for real-time live streaming scenarios, prioritize the RTMP protocol, with the default port 1935, which supports low-latency transmission. For video-on-demand or file transfer scenarios, select the HTTP-FLV protocol and transmit via HTTP port 80 or 443; if the receiving end is a local area network device, select the RTP+UDP protocol to improve transmission efficiency. The protocol selection logic is determined by reading the transmission mode set by the user or by automatically detecting the receiver type.

[0084] The encoded fused video stream is encapsulated according to the protocol format and transmitted in segments to the receiving end. The encapsulation format is determined by the selected protocol: The RTMP protocol encapsulates the video stream into an RTMP message. The message header contains information such as message type, message length, and timestamp, and the message body is the encoded video frame data; The HTTP-FLV protocol encapsulates the video stream into an FLV file format, which includes FLV header, tags, and other structures. The tags are divided into audio tags, video tags, and script tags, and the video stream corresponds to the video tag. The RTP+UDP protocol segments video frames and encapsulates them into RTP packets. The RTP packet header includes information such as version number, payload type, sequence number, timestamp, and synchronization source identifier. The segment size is consistent with the segmentation strategy in S3, dynamically adjusted according to network bandwidth, and a checksum is added to each segment to ensure data integrity.

[0085] The system monitors network transmission status in real time and adjusts video resolution based on bandwidth conditions. Network transmission status monitoring metrics include packet loss rate, latency, and jitter. If the packet loss rate exceeds 10% for 5 consecutive seconds, latency exceeds 500ms, or jitter exceeds 100ms, the video resolution is reduced, such as from 1080P to 720P, or from 720P to 480P. If the packet loss rate is below 2% and latency is below 200ms for 5 consecutive seconds, the original resolution is restored. During resolution adjustment, the initial encoding bitrate is adjusted simultaneously to ensure that the adjusted bitrate matches the resolution.

[0086] If the receiving end reports data loss, a partial retransmission mechanism is triggered to retransmit the lost fragments. The receiving end reports the lost fragment information via RTCP protocol or HTTP response message, including the frame sequence number and fragment sequence number corresponding to the lost fragment. After receiving the feedback, the mobile terminal searches for the corresponding fragment data in the encoding buffer, prioritizes retransmitting fragments of key frames, adds a retransmission flag to the retransmitted fragments, sets the retransmission timeout to 300ms, and allows a maximum of 3 retransmissions. If the retransmission is still unsuccessful after 3 retransmissions, the receiving end is notified to discard the frame corresponding to the fragment, and subsequent frame interpolation technology is used to compensate for the missing image.

[0087] Reference Figure 2 As shown, the smart glasses and mobile phone dual video stream synchronous acquisition and fusion communication system includes: Communication connection module, video acquisition and synchronization module, video fusion and processing module, user interaction module, encoding and transmission module; The communication connection module is used to establish a dedicated link between smart glasses and mobile terminals via Wi-Fi, complete device pairing and communication parameter adaptation, and provide a foundation for dual video stream transmission; The video acquisition synchronization module is electrically connected to the communication connection module and is used to receive the synchronization command from the mobile terminal, triggering both image acquisition modules to start simultaneously, acquire dual-view video respectively and achieve synchronous transmission. The video fusion processing module is electrically connected to the video acquisition synchronization module. It is used to receive two video streams, calibrate them using timestamp and frame-level alignment technology, and complete video fusion in combination with the selected layout mode. The user interaction module is electrically connected to the video fusion processing module, and is used to provide layout selection and main / sub screen switching functions, and supports users to add vector markers and overlay them in real time through touch operation; The encoding and transmission module is electrically connected to the user interaction module and is used to re-encode the fused video stream and send it to the other end of the call or the receiving end of the live streaming platform via the network.

[0088] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A method for smart glasses and mobile phone dual video stream synchronous acquisition fusion communication, characterized in that, Includes the following steps: S1. The smart glasses and the mobile terminal establish a dedicated communication link via Wi-Fi to complete device pairing and communication parameter adaptation. S2. The mobile terminal sends a synchronous acquisition command through this link, triggering its own second image acquisition module and the first image acquisition module of the smart glasses to start simultaneously, and acquire third-person and first-person perspective videos respectively; S3: The smart glasses encode the first video stream in real time and transmit it to the mobile terminal via Wi-Fi. S4. The mobile terminal receives the first video stream and the second video stream it has collected, and uses local timestamp and frame-level alignment technology to synchronize and calibrate the two video streams. S5 and mobile terminals offer multiple layout modes and support users in switching the display priority of the main and secondary screens; S6. Users can select the vector marking tool through touch operation in the mobile terminal preview or live broadcast interface, and the marking is superimposed on the video screen in real time. S7. The mobile terminal re-encodes the fused video stream after synchronization, layout, and marking processing, and sends it to the receiving end via the network.

2. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 1, characterized in that, S1 includes: The smart glasses activate the Wi-Fi module, switch to AP mode or Wi-Fi Direct mode, and broadcast a connection signal containing the device model and unique identifier; the mobile terminal activates the Wi-Fi function, searches for signals of nearby dedicated communication devices, filters out the broadcast signal of the target smart glasses, and initiates a connection request. After receiving a connection request from a mobile terminal, the smart glasses send pairing verification information to the mobile terminal. The verification method is either comparing the device's unique identifier or having the user enter a preset password. Once the verification is successful, the device pairing and binding is completed. Once paired successfully, both parties will automatically negotiate the transmission protocol. Synchronize the system clock references of both devices; The system negotiates the video encoding format, detects the current network environment, sets the data transmission bandwidth threshold and latency compensation parameters, generates a communication parameter configuration file, and synchronizes it to both devices.

3. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 2, characterized in that, S2 includes: Based on the established dedicated communication link, the mobile terminal generates a synchronous acquisition command that includes acquisition resolution, frame rate, and exposure parameters, and embeds a local reference timestamp. The mobile terminal transmits the synchronous acquisition command to the smart glasses via a dedicated link, and simultaneously starts the command transmission timing; After receiving the command, the smart glasses decode and extract the configuration information and the reference timestamp. The main processing chip triggers the first image acquisition module to start and initializes the acquisition state according to the configuration parameters. After the smart glasses start collecting data, they send a ready signal to the mobile terminal. Upon receiving the signal, the mobile terminal immediately activates its second image acquisition module. The mobile terminal determines the start time difference between the two parties by the timing difference. If the difference exceeds the set threshold, the synchronization command is resent. After both acquisition modules are working stably, the smart glasses acquire first-person perspective video, and the mobile terminal acquires third-person perspective video.

4. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 3, characterized in that, S3 includes: The main processing chip of the smart glasses receives the raw video data transmitted by the first image acquisition module and starts the hardware encoding engine; The encoding engine performs frame-level compression on the raw video data according to the negotiated format, and assigns a unique frame number and acquisition timestamp to each video frame, which are then associated with the corresponding video frame. Start the transmission buffer queue and store the encoded video frames into the queue in order of frame number; Based on the real-time transmission status of the Wi-Fi link, the video frames in the queue are segmented and processed, and each segment is transmitted after adding a checksum. A data retransmission mechanism is set up so that if the mobile terminal reports that a certain segment is lost, the smart glasses will retrieve the corresponding segment from the cache queue and retransmit it.

5. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 4, characterized in that, S4 includes: The mobile terminal receives the first video stream transmitted by the smart glasses through a dedicated link, and simultaneously collects its own second video stream, establishing frame data storage buffers for the two video streams respectively. Extract the acquisition timestamp from each frame of data from the two video streams to generate a first timestamp sequence and a second timestamp sequence; Calculate the time difference between corresponding frames in two timestamp sequences, set a synchronization threshold, and filter out frame data whose time difference exceeds the threshold. For frames exceeding the threshold, frame delay buffering or redundant frame discarding is used for adjustment, and the frame-level timestamps of the first video stream and the second video stream are matched and aligned frame by frame. Perform image misalignment correction on the two aligned video streams.

6. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 5, characterized in that, S5 includes: The mobile terminal control APP has multiple preset layout modes, and corresponding layout selection controls are set in the user interface; Users select the target layout mode through touch controls. After receiving the selection command, the APP reads the screen resolution parameters of the mobile terminal. The display areas for the two video streams are allocated based on the selected layout mode and screen resolution. The screen size of the two video streams is adaptively scaled to maintain the original aspect ratio; The adjusted two video streams will be rendered and displayed according to the selected layout mode.

7. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 6, characterized in that, The S5 also includes: The mobile terminal interactive interface is configured with a main and secondary screen switching control, which associates the display priority indicators of the first and second video streams. When a user triggers the switching control, the APP receives the switching command and reads the display status of the current main and secondary screens. Swap the display priorities of the two video streams, switch the original secondary screen to the main screen and expand it to the corresponding display area, and switch the original main screen to the secondary screen and adjust its size; Update the interface display status.

8. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 7, characterized in that, S6 includes: The mobile terminal preview or live broadcast interface features a marker toolbar containing various vector marker tools, which users can select by touch. Users touch and slide to draw and mark trajectories on the interface. The APP captures touch coordinate data in real time and generates a continuous sequence of trajectory points. Convert the trajectory point sequence into Bézier curve vector graphics data and remove redundant coordinate points; Provides options for customizing marker styles; The processed vector marker data is superimposed onto the corresponding coordinate positions of the video frame in real time; Provides mark undo and clear functions.

9. The method for synchronous acquisition and fusion communication of dual video streams between smart glasses and mobile phones according to claim 8, characterized in that, S7 includes: The mobile terminal re-encodes the fused video stream that has undergone synchronization calibration, layout adjustment, and marker overlay. Real-time monitoring of current network bandwidth status; dynamically adjusting the encoding bitrate based on bandwidth data. Select the corresponding transmission protocol; The encoded fused video stream is encapsulated according to the protocol format and transmitted in segments to the receiving end. Monitor network transmission status in real time and adjust video resolution according to bandwidth conditions; If the receiving end reports data loss, a partial retransmission mechanism is triggered to retransmit the lost fragments.

10. A dual-video stream synchronous acquisition and fusion communication system for smart glasses and mobile phones, characterized in that, include: Communication connection module, video acquisition and synchronization module, video fusion and processing module, user interaction module, encoding and transmission module; The communication connection module is used to establish a dedicated link between smart glasses and mobile terminals via Wi-Fi, complete device pairing and communication parameter adaptation, and provide a foundation for dual video stream transmission; The video acquisition synchronization module is electrically connected to the communication connection module and is used to receive the synchronization command from the mobile terminal, triggering both image acquisition modules to start simultaneously, acquire dual-view video respectively and achieve synchronous transmission. The video fusion processing module is electrically connected to the video acquisition synchronization module. It is used to receive two video streams, calibrate them using timestamp and frame-level alignment technology, and complete video fusion in combination with the selected layout mode. The user interaction module is electrically connected to the video fusion processing module, and is used to provide layout selection and main / sub screen switching functions, and supports users to add vector markers and overlay them in real time through touch operation; The encoding and transmission module is electrically connected to the user interaction module and is used to re-encode the fused video stream and send it to the other end of the call or the receiving end of the live streaming platform via the network.