A method and system for synchronized audio and video explanation based on a team explanation device
By using hybrid encoding and separate parsing technology, synchronized audio and video narration was achieved for tour guides in museums and scenic spots, solving the problem of single-voice narration and improving the visitor experience and management efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN HENGDA WENBO TECH CO LTD
- Filing Date
- 2026-06-26
- Publication Date
- 2026-07-31
AI Technical Summary
Existing tour guide systems in museums and scenic spots only support audio narration and cannot play videos simultaneously. They are also cumbersome to operate, have low management efficiency, and cannot provide complete information in crowded situations.
By collecting and preprocessing the transmitter signal, mixing and encoding audio data frames, and then separating and parsing them using data transmission paths and branch judgment logic, synchronous audio and video explanations can be achieved, and remote control is supported.
It enables real-time synchronized playback of audio guides and exhibit videos, simplifies equipment management operations, improves the visitor experience and management efficiency, and is suitable for large-scale guided tour scenarios.
Smart Images

Figure CN122496669A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of team explanation technology, and in particular to a method and system for synchronized audio and video explanation based on a team explanation device. Background Technology
[0002] Currently, the audio guides used in museums and scenic spots generally employ a single voice transmission mode. For example... Figure 1 As shown, the transmitter collects the guide's voice and digitally encodes it. The encoded data is then sent to the receiver via an RF front-end chip. The receiver's RF front-end chip decodes the data and reconstructs the sound for playback. However, current methods for group guided tours in museums or scenic areas still have significant shortcomings:
[0003] (1) Only supports audio explanation: The audience can only understand the characteristics of the scenic spots or museum artifacts through audio explanation, and cannot synchronize video screens, resulting in a single form of visit;
[0004] (2) The guided tour format is monotonous: When there are crowds in museums or scenic spots, the audience can hardly see the exhibits and can only hear the explanation. The information is not fully conveyed and the experience is generally poor.
[0005] (3) Unified control is not supported: the receiver needs to be turned on and off separately, the volume needs to be adjusted and the channel needs to be changed, which is cumbersome and the management efficiency is low.
[0006] Therefore, there is an urgent need for a method and system for synchronized audio and video explanation based on a team explanation device to address the shortcomings of existing technologies. Summary of the Invention
[0007] The purpose of this invention is to propose a method and system for synchronized audio and video explanation based on a team tour guide, which integrates voice explanation, synchronized video playback, and batch remote control to improve the visitor experience and management efficiency.
[0008] Firstly, to achieve the above objectives, the present invention provides a method for synchronized audio and video explanation based on a team explanation device, comprising:
[0009] S1. Collect transmitter signals and preprocess them to obtain an audio dataset for explanation. The transmitter signals include interface operation instructions, positioning signals and PCM audio data.
[0010] S2. Based on the aforementioned audio dataset, combine control and positioning information to perform hybrid encoding to obtain hybrid encoded audio data.
[0011] S3. Encapsulate the hybrid encoded audio data into frames to obtain hybrid encoded data frames;
[0012] S4. Based on the hybrid encoded data frame, the data transmission path and branch judgment logic are combined to perform separation and parsing to obtain the audio and video synchronized explanation results.
[0013] Optionally, S1, acquire the transmitter signal, perform preprocessing, and obtain the narration audio dataset, including:
[0014] The transmitter is powered on, and the UI operation acquisition, positioning signal acquisition, and microphone sound acquisition are started simultaneously.
[0015] Collect UI operations and detect whether a play operation is selected. If so, obtain the UI operation instruction; otherwise, return to UI operation collection.
[0016] The system collects positioning signals and checks whether a trigger signal has been received. If so, it acquires the positioning signal; otherwise, it returns to the positioning signal acquisition process.
[0017] Based on the microphone sound acquisition, obtain the microphone audio signal;
[0018] The interface operation instructions and the microphone audio signal are preprocessed to obtain the preprocessed interface operation instructions and PCM audio data;
[0019] Based on the location signal, it is detected whether location information is received. If so, the location signal is preprocessed to obtain the preprocessed location signal; otherwise, the location signal acquisition is performed repeatedly.
[0020] The preprocessed interface operation instructions, PCM audio data, and positioning signals are obtained as the explanatory audio dataset.
[0021] Optionally, S2, based on the narration audio dataset and control and positioning information, perform hybrid encoding to obtain hybrid encoded narration audio data, including:
[0022] The audio dataset used for explanation is sampled and quantized at fixed intervals to obtain raw PCM audio data.
[0023] The raw PCM audio data is encoded to obtain basic compressed audio data;
[0024] Based on the audio dataset described above, obtain control and positioning information;
[0025] Based on the data frame structure, the control and positioning information and the basic compressed audio data are fused and encoded to obtain a hybrid data frame;
[0026] Cyclic redundancy check is performed on the mixed data frame to obtain the mixed-encoded audio data for explanation.
[0027] Optionally, the data frame structure includes a frame header identifier, a frame length identifier, an audio data segment, a control information embedding segment, a positioning information embedding segment, a cyclic redundancy check bit, and a frame tail identifier.
[0028] Optionally, S3, performing frame encapsulation on the hybrid encoded explanatory audio data to obtain hybrid encoded data frames, including:
[0029] Based on the aforementioned hybrid encoding, the audio data is encapsulated in conjunction with the radio frequency data frame structure to obtain a standard radio frequency data frame.
[0030] The standard radio frequency data frame is transmitted via radio frequency to obtain a baseband radio frequency data frame;
[0031] Based on the baseband radio frequency data frame and the radio frequency data frame structure, a parsing and verification is performed to obtain a hybrid coded data frame.
[0032] Optionally, S4, based on the hybrid encoded data frame and the data transmission path and branch judgment logic, separate and parse the data to obtain the audio and video synchronized explanation result, including:
[0033] Based on the data frame structure, the hybrid encoded data frame is parsed according to the data transmission path to obtain separate hybrid encoded data frames, which include UI playback control data, positioning signal data and audio encoded data.
[0034] The UI playback control data is type-identified to determine whether it is a control command. If it is, the system operation is executed and the end of the path is obtained as the audio and video synchronization explanation result. Otherwise, the playback control data is obtained.
[0035] Based on the positioning signal data and the multi-source fusion priority rules, multi-source matching is performed to extract the positioning code and match the corresponding audio and video data to obtain the positioning code and audio and video timestamp information.
[0036] The audio encoded data is decoded to obtain decoded PCM audio data;
[0037] Based on the playback control data, the positioning code and audio / video timestamp information, video is played synchronously. The decoded PCM audio data is mixed into the current audio / video playback channel. The audio / video playback information is obtained by combining batch instruction priority rules and radio frequency anti-interference switching algorithm.
[0038] Based on the audio and video playback information, perform device offline detection and signal interruption detection to obtain the audio and video synchronized explanation results.
[0039] Optionally, based on the audio and video playback information, device offline detection and signal interruption detection are performed to obtain the audio and video synchronized explanation results, including:
[0040] Based on the audio and video playback information, determine whether to receive the response confirmation packet. If yes, perform the first operation; otherwise, immediately start the data retransmission mechanism, obtain the data retransmission information, and perform the second operation.
[0041] The first operation is as follows: monitor whether the signal is interrupted according to the audio and video playback information. If so, execute the signal interruption buffering strategy, obtain the target audio and control data, and execute the third operation. Otherwise, obtain the audio and video synchronized explanation result based on the audio and video playback information.
[0042] The second operation is: determine whether the data retransmission information is greater than the retransmission threshold. If so, obtain the device offline status as the audio and video synchronization explanation result; otherwise, execute the first operation.
[0043] The third operation is to supplement the transmission of the target audio and control data in accordance with the instruction encoding order to obtain the audio and video synchronized explanation result.
[0044] Secondly, to achieve the above objectives, the present invention provides an audio-visual synchronized explanation system based on a team explanation device, including a signal acquisition module, a hybrid encoding module, a radio frequency front-end data transmission module, and a separate decoding module;
[0045] The signal acquisition module is used to acquire transmitter signals for preprocessing and obtain an audio dataset for explanation. The transmitter signals include interface operation instructions, positioning signals and PCM audio data.
[0046] The hybrid encoding module is used to perform hybrid encoding based on the narration audio dataset and control and positioning information to obtain hybrid encoded narration audio data;
[0047] The radio frequency front-end data transmission module is used to encapsulate the hybrid encoded audio data into frames to obtain hybrid encoded data frames.
[0048] The separation decoding module is used to separate and parse the hybrid encoded data frame based on the data transmission path and branch judgment logic to obtain the audio and video synchronized explanation result.
[0049] Thirdly, to achieve the above objectives, the present invention provides an electronic device, comprising: one or more processors; and a storage device having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation of the first aspect.
[0050] Fourthly, to achieve the above objectives, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any implementation of the first aspect.
[0051] Compared with the closest existing technology, the present invention has the following advantages:
[0052] This invention breaks through the traditional one-way audio guide mode, supporting real-time listening to explanations during visits, two-way interaction with guides, simultaneous viewing of exhibit explanation videos, and remote unified control operations such as receiver shutdown, channel switching, and volume adjustment. It is suitable for various large-scale guided tour scenarios such as museums, scenic spots, and study tour groups. Specific effects are as follows:
[0053] (1) Real-time audio and video synchronization: The audio explanation and the exhibit video are accurately synchronized, changing the traditional explanation mode of only hearing the sound and not seeing the exhibit, enriching the form of information presentation, and significantly improving the immersive experience and viewing experience of the visitor;
[0054] (2) Centralized group control management: The transmitter can batch control the volume, channel, power on / off status and playback status of all receivers, greatly simplifying the operation and maintenance management process;
[0055] (3) Multi-source intelligent triggering: It integrates RFID, Bluetooth and infrared multi-source positioning, which can automatically match and trigger the product explanation content without manual switching;
[0056] (4) Stable and reliable communication transmission: The system adopts a dedicated Sub-1GHz frequency band and combines a channel RSSI and bit error rate joint monitoring algorithm. When the RSSI exceeds the standard continuously or the bit error rate exceeds 10% for three consecutive times, the system automatically switches frequencies. The system has three priority backup frequency bands, which are switched in order of priority. After a successful switch, the frequency band is locked for 5 minutes, which lays a solid anti-interference foundation for the synchronous transmission of audio and video throughout the process. At the same time, it is equipped with data loss retransmission, offline device detection, signal interruption buffering and orderly retransmission of instructions after link recovery, and deduplication and merging mechanism at the receiving end, which effectively avoids problems such as signal stuttering, disconnection and repeated playback in complex environments, and ensures continuous and reliable transmission.
[0057] (5) Comprehensive upgrade of application experience: It has been upgraded from one-way voice explanation to visual interactive explanation. A single transmitter can be equipped with multiple receivers to meet the needs of large-scale use by large groups such as study tours, museums, and scenic spots. Attached Figure Description
[0058] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0059] Figure 1 This is a schematic diagram of the structure of an existing team presentation device proposed in an embodiment of the present invention;
[0060] Figure 2 This is a flowchart illustrating a method for synchronized audio and video explanation based on a team explanation device, according to an embodiment of the present invention.
[0061] Figure 3 This is a flowchart of the system workflow proposed in an embodiment of the present invention;
[0062] Figure 4 This is a flowchart of the transmitter signal acquisition process proposed in an embodiment of the present invention;
[0063] Figure 5 This is a flowchart of the hybrid coding process proposed in an embodiment of the present invention;
[0064] Figure 6 This is a flowchart of the radio frequency front-end data transmission process proposed in an embodiment of the present invention;
[0065] Figure 7 This is a flowchart illustrating the audio and control data separation and decoding process proposed in an embodiment of the present invention;
[0066] Figure 8 This is a flowchart of an audio-visual synchronized explanation system based on a team explainer, according to an embodiment of the present invention.
[0067] Figure 9 Here are system configuration diagrams proposed in the embodiments of the present invention, (a) being a configuration diagram of the transmitting end and (b) being a configuration diagram of the receiving end;
[0068] Figure 10 This is a diagram illustrating the overall system architecture for a cultural tourism and study tour scenario proposed in this embodiment of the invention.
[0069] Figure 11 This is a schematic diagram of the teacher-side control function proposed in an embodiment of the present invention;
[0070] Figure 12 This is a schematic diagram of the structure of the electronic device proposed in an embodiment of the present invention. Detailed Implementation
[0071] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0072] The terminology used in the embodiments section of this invention is for the purpose of explaining specific embodiments of the invention only, and is not intended to limit the invention.
[0073] like Figure 2-3 As shown, this embodiment of the invention provides a method for synchronized audio and video explanation based on a team explanation device, including:
[0074] S1. Collect and preprocess the transmitter signal to obtain the audio dataset for explanation;
[0075] After the transmitter is powered on, it initiates the UI operation acquisition, positioning signal acquisition, and microphone sound acquisition processes in parallel. It can acquire interface operation commands, receive positioning signals, and collect microphone audio signals. Specifically, it samples the microphone audio signals using Pulse Code Modulation (PCM) to obtain PCM audio data. It then maps the corresponding exhibit codes using multi-source positioning signals from Radio Frequency Identification (RFID), Bluetooth, and infrared. Simultaneously, it acquires interface operation commands such as video playback, volume adjustment, channel switching, and power off. After preprocessing, it forms an explanatory audio dataset that can be used for encoding, providing a basic data source for subsequent mixed encoding.
[0076] S2. Based on the aforementioned audio dataset, combine control and positioning information to perform hybrid encoding to obtain hybrid encoded audio data.
[0077] The hybrid encoding process is triggered at a fixed period. First, the PCM audio data in the audio dataset is encoded using Opus to generate basic compressed audio frames. Then, control information and positioning information, including receiver one-click power-off, channel changing, and volume adjustment, are fused and encoded with the basic compressed audio frames according to a preset frame structure. Control information is packaged every 50 milliseconds (ms) and redundantly embedded into 5 consecutive audio frames to improve transmission fault tolerance. High-priority instructions are directly inserted into the next audio frame, finally resulting in hybrid encoded audio data that integrates audio and control information.
[0078] S3. Encapsulate the hybrid encoded audio data into frames to obtain hybrid encoded data frames;
[0079] The hybrid encoded audio data is encapsulated, and preamble bytes, synchronization words, frame length identifiers, hybrid encoded data, and cyclic redundancy check bits are added sequentially to form a radio frequency data frame that conforms to the wireless transmission standard. The data is then broadcast to each receiver via the radio frequency front-end chip to complete the wireless data transmission, enabling the receiver to obtain the hybrid encoded data frame that can be used for parsing.
[0080] S4. Based on the hybrid encoded data frame, the data transmission path and branch judgment logic are combined to perform separation and parsing to obtain the audio and video synchronized explanation results;
[0081] After receiving the mixed-encoded data frame, the receiver separates and parses the mixed-encoded data frame according to the data transmission path and branch judgment logic, splitting it into UI playback control data, positioning signal data, and audio encoded data; the audio encoded data is decoded and restored by Opus and played, and the corresponding exhibit video is played synchronously according to the video number and playback timestamp. Operations such as volume adjustment, channel switching, and batch shutdown are executed according to the command priority, and finally, the voice explanation, video display and remote control are output synchronously, completing the audio and video synchronous explanation.
[0082] In summary, steps S1 to S4 involve the transmitter acquiring and preprocessing the audio, positioning signals, and control commands in parallel. The Opus-encoded audio data and control information are embedded and fused in a frame structure with a 10ms period, and then redundantly packaged and encapsulated into radio frequency data frames for broadcast transmission. The receiver separates the audio, video, and control information through CRC verification and branch parsing. This achieves low-latency, high-reliability synchronous output of the narration voice, exhibit video, and remote commands, while also possessing high compression ratio audio transmission efficiency, control information fault tolerance, and multi-device adaptability, ensuring the audio-video synchronization quality and wireless transmission stability of the narration system.
[0083] like Figure 4 As shown, in one possible implementation, step S1 in the above embodiment may specifically include the following steps:
[0084] S1-1. Transmitter is powered on, and UI operation acquisition, positioning signal acquisition and microphone sound acquisition are started simultaneously.
[0085] Once the transmitter is powered on and the user enters the exhibit area and is triggered by the positioning signal, the process of collecting UI operation data, positioning signal data, and microphone sound data is started simultaneously and in parallel. The data is continuously collected and classified into the corresponding collection queues to provide a complete data source for subsequent mixed encoding.
[0086] S1-2. Collect UI operations and detect whether the play operation is selected. If yes, obtain the UI operation instruction; otherwise, return to UI operation collection.
[0087] In the interface operation acquisition process, when the controller is selected to perform a video playback operation, it collects information such as the currently playing audio file number, playback volume, and playback progress in real time, and sends the above information along with the playback action signal into the interface UI operation acquisition queue; if no playback operation is selected, the interface UI operation acquisition is continuously executed in a loop.
[0088] S1-3. Collect the positioning signal and detect whether a trigger signal is received. If so, acquire the positioning signal; otherwise, return to the positioning signal acquisition process.
[0089] In the positioning signal acquisition process, the transmitter receives three positioning signals in real time: RFID, Bluetooth, and infrared. When any valid trigger signal is received, the exhibit code and other information are mapped out through the positioning signal and sent to the positioning signal acquisition queue. If no trigger signal is received, the positioning signal acquisition is continuously performed in a loop.
[0090] S1-4. Obtain the microphone audio signal based on the microphone sound acquisition;
[0091] In the microphone sound acquisition process, the transmitter acquires audio data in real time at a sampling rate of 48kHz to obtain PCM audio data, encodes the PCM audio data, and puts the encoded data into the microphone sound acquisition queue.
[0092] S1-5. Preprocess the interface operation instructions and the microphone audio signal to obtain the preprocessed interface operation instructions and PCM audio data.
[0093] The acquired interface operation instructions and microphone audio signals are preprocessed by format regularization and data verification, and converted into a standard data format that meets the encoding requirements to obtain preprocessed interface operation instructions and preprocessed PCM audio data.
[0094] S1-6. Detect whether positioning information is received based on the positioning signal. If yes, preprocess the positioning signal to obtain the preprocessed positioning signal. Otherwise, repeatedly perform positioning signal acquisition.
[0095] Based on the acquired raw positioning signal, further detection is performed to determine whether valid positioning information such as exhibit codes has been successfully parsed. If valid positioning information is successfully received and parsed, the positioning signal is preprocessed by format mapping and data encapsulation to obtain a preprocessed positioning signal that can be used for encoding. If no valid positioning information is received or parsed, the positioning signal acquisition and detection process is executed repeatedly.
[0096] S1-7. Obtain the preprocessed interface operation instructions, PCM audio data and positioning signals as the explanatory audio dataset;
[0097] The preprocessed interface operation instructions, PCM audio data and positioning signals are uniformly encapsulated to form an explanatory audio dataset containing control information, audio data and positioning trigger information, providing complete input data for subsequent mixed encoding of audio, video and control data.
[0098] In summary, steps S1-1 to S1-7 simultaneously initiate three parallel processes after the transmitter is powered on: UI operation acquisition, positioning signal acquisition, and microphone sound acquisition. Interface operation commands are acquired through real-time detection of video playback operations. Positioning signals are acquired and mapped to exhibit codes via RFID, Bluetooth, or infrared trigger signals. Simultaneously, microphone audio is acquired in real-time at a 48kHz sampling rate, generating PCM audio data and completing audio encoding. The interface operation commands, microphone audio signals, and positioning signals are then preprocessed according to specifications, ultimately integrating them to form a complete audio dataset for explanation and sending it to the audio and control hybrid encoding module. This process can simultaneously acquire multi-source signals and standardize preprocessed data, providing a stable and reliable data source for subsequent hybrid encoding, RF transmission, and synchronized audio-video output, ensuring complete acquisition of explanation information and a solid foundation for transmission.
[0099] As one possible implementation, in the above embodiments, step S2 may specifically include the following steps:
[0100] S2-1. Sample and quantize the audio dataset at a fixed period to obtain the original PCM audio data.
[0101] Using an internal timer as the trigger source, the PCM audio data collected by the microphone is read from the microphone sound acquisition queue in the audio dataset every 10 milliseconds (ms). After sampling and quantization processing, the raw PCM audio data is generated to provide the basic audio source for subsequent encoding.
[0102] S2-2. Encode the original PCM audio data to obtain basic compressed audio data;
[0103] The obtained PCM raw audio data is sent to the Opus encoder for high compression ratio encoding to generate basic compressed audio data that meets the requirements of wireless transmission. This data will serve as the core frame data part of subsequent frame encapsulation.
[0104] S2-3. Based on the aforementioned audio dataset, obtain control and positioning information;
[0105] The UI operation acquisition queue in the audio dataset is used to extract control information generated by user operations. This control information includes functions such as one-click power off of the receiver, changing channels, and volume adjustment. At the same time, the location information of the exhibits, such as the location number, exhibit ID, and location timestamp, is read from the location signal acquisition queue. This information is then integrated with the above information to form complete control and location information, providing control-side data for subsequent frame encapsulation.
[0106] S2-4. Based on the data frame structure, the control and positioning information and the basic compressed audio data are fused and encoded to obtain a mixed data frame;
[0107] According to the preset data frame structure, the obtained basic compressed audio data is used as the main frame data. At the end of each 10ms audio frame, 5 bytes are reserved for control information marking. Control and positioning information are embedded into the frame structure as 2 bytes at the end of the frame. Following the embedding rules, the audio data and control / positioning information are finally fused and encoded to generate a hybrid data frame. The embedding rules are as follows:
[0108] (1) Control information is packaged every 50 milliseconds and embedded into 5 audio frames to achieve redundancy and fault tolerance;
[0109] (2) High-priority instructions are inserted directly into the next frame without waiting for the packaging cycle.
[0110] S2-5. Perform cyclic redundancy check on the mixed data frame to obtain the mixed encoded audio data for explanation;
[0111] Cyclic Redundancy Check (CRC) is performed on the generated hybrid data frame. The CRC result is written into the corresponding position in the data frame, and a frame header identifier and a frame tail identifier are added to complete the data integrity verification and final encapsulation. This results in hybrid coded audio data that can be directly used for radio frequency transmission and has data integrity verification capabilities, providing a reliable data frame for subsequent radio frequency transmission.
[0112] like Figure 5As shown, the hybrid encoding part consists of two parts: audio data Opus encoding and combined data frame encapsulation. Data is obtained from the UI operation acquisition queue, the positioning signal acquisition queue, and the microphone sound acquisition queue, respectively, and combined and encapsulated into data frames. Then, Opus encoding is performed. The specific data frame structure is shown in Table 1, including a frame header identifier (fixed value, used for frame synchronization), a frame length identifier (recording the total number of bytes in the frame), an audio data segment (baseline compressed audio data after Opus encoding), a control information embedding segment (5 bytes reserved at the end of the frame, used to carry various control commands), a positioning information embedding segment (reusing the reserved area at the end of the frame, combined with control commands, carrying device positioning-related information), a cyclic redundancy check bit (used for data integrity verification), and a frame tail identifier (fixed value, used for frame end marking).
[0113] Table 1
[0114]
[0115] The control information embedding segment and the positioning information embedding segment are both set in the reserved area at the end of the frame. The control information embedding segment is shown in Table 2-4, which includes a 1-byte instruction number, a 2-byte UI operator, and a 3-byte UI playback timestamp. The positioning information embedding segment is shown in Table 5, which includes a 1-byte signal source identifier and a 2-byte signal number. It is arranged in combination with the control information embedding segment in the reserved area at the end of the frame.
[0116] Table 2
[0117]
[0118] Table 3
[0119]
[0120] Table 4
[0121]
[0122] Table 5
[0123]
[0124] In summary, steps S2-1 to S2-5 first sample and quantize the audio at a fixed period to obtain the original PCM audio, then generate compressed audio data through Opus encoding. Simultaneously, device control and exhibit positioning information are collected. Then, according to a preset frame structure, the control and positioning information is embedded into the audio frame structure as a 2-byte tail to complete the data frame fusion and encapsulation. CRC verification is used to further refine the frame encapsulation. This process achieves synchronous transmission of control, positioning information, and audio service data within the same frame, without requiring additional separate channel occupancy. The frame structure is well-organized and highly compatible, the information embedding position is fixed, and parsing is convenient. It also improves the reliability and real-time performance of wireless transmission, while ensuring both audio transmission quality and accurate issuance and recognition of positioning and control commands.
[0125] like Figure 6 As shown, in one possible implementation, step S3 in the above embodiment may specifically include the following steps:
[0126] S3-1. Based on the hybrid encoding, the audio data is encapsulated in conjunction with the radio frequency data frame structure to obtain a standard radio frequency data frame.
[0127] The mixed-encoded data, containing explanatory audio (Opus-encoded data) and control information (such as positioning and operation commands), is used as the payload and encapsulated according to a preset RF data frame structure. As shown in Table 6, this frame structure consists of a preamble byte, a synchronization word, frame data, and a CRC checksum. The preamble byte and synchronization word are used for frame synchronization capture at the receiving end, the frame data carries the mixed-encoded explanatory audio data, and the CRC checksum is used for error detection during transmission. By writing the mixed-encoded explanatory audio data into the data payload area and completing frame header splicing and checksum calculation, an RF data frame conforming to the RF transmission protocol specification is generated, providing a complete and recognizable frame format for subsequent transmission.
[0128] Table 6
[0129]
[0130] S3-2. Perform radio frequency transmission on the standard radio frequency data frame to obtain the baseband radio frequency data frame;
[0131] The packaged standard RF data frame is sent to the RF transmitting circuit. After baseband modulation, carrier modulation and power amplification, it is converted into an air RF signal by the RF front-end circuit and transmitted to the wireless channel through the antenna. The receiving end captures the air-transmitted RF signal through the RF receiving circuit. After down-conversion, demodulation and synchronization processing, the baseband RF data frame is restored, and the signal transmission and reception of the wireless transmission link are completed.
[0132] S3-3. Perform parsing and verification based on the baseband radio frequency data frame and the radio frequency data frame structure to obtain the hybrid coded data frame;
[0133] The receiving end performs frame synchronization positioning and integrity verification on the acquired baseband radio frequency data frame according to the radio frequency data frame structure: the frame start position is locked by identifying the synchronization word, and the data integrity is verified based on the CRC check field in the frame; after the verification is passed, the preamble byte, synchronization word and check bit and other transmission additional fields are stripped, and only the effective payload part is retained to complete the parsing and verification of the radio frequency data frame, and finally restore the hybrid encoded data frame for subsequent decoding and playback.
[0134] In summary, steps S3-1 to S3-3 encapsulate the mixed encoded data containing narration audio and control information into a standard radio frequency data frame according to the radio frequency frame structure, transmit it through the radio frequency link, and parse, verify, and restore it at the receiving end. This achieves integrated and reliable transmission of narration audio and control data, combining transmission efficiency and anti-interference capability, and ensuring synchronous and low-distortion transmission of audio and control information.
[0135] like Figure 7 As shown, in one possible implementation, step S4 in the above embodiment may specifically include the following steps:
[0136] S4-1. Based on the data frame structure, parse the hybrid encoded data frame according to the data transmission path to obtain the separated hybrid encoded data frames;
[0137] After receiving the RF data packet and verifying the CRC32 checksum of the packet header, the receiver parses the hybrid coded data frame field by field according to the data frame structure. It sequentially extracts the frame header identifier, frame length, audio data segment, command number, UI operator, playback timestamp, positioning signal data, CRC16 checksum, and frame tail identifier. Following a complete data transmission path consisting of three independent branches—UI playback control data path, positioning signal data path, and audio encoded data path—the parsed data is separated into three parallel data streams: UI playback control data, positioning signal data, and audio encoded data, achieving physical separation of control information, positioning information, and audio data. Specifically, the UI playback control data includes the command number, command type, command parameters, and playback timestamp; the positioning signal data includes the signal source type (RFID, Bluetooth, infrared) and signal number; and the audio encoded data consists of Opus-compressed audio data segments within the frame.
[0138] S4-2. The UI playback control data is type-identified to determine whether it is a control command. If it is, the system operation is executed and the end of the path is obtained as the audio and video synchronization explanation result. Otherwise, the playback control data is obtained.
[0139] The UI playback control data is judged by instruction type to distinguish between system control data and playback control data. If it is determined to be a system control instruction (such as power off or change channel), the corresponding system operation such as power off or change channel is executed immediately. After the operation is completed, the path terminates and outputs the system operation completion as the audio and video synchronous explanation result. If it is determined to be playback control information, the current playback is controlled by fast forward, rewind, playback of a specified number, etc., and playback control data such as the playback audio file number, playback volume, and playback progress are extracted before proceeding to the subsequent synchronous playback process.
[0140] S4-3. Based on the positioning signal data and the multi-source fusion priority rules, perform multi-source matching, extract the positioning code and match the corresponding audio and video data, and obtain the positioning code and audio and video timestamp information.
[0141] For the separated positioning signal data, a multi-source fusion priority rule is used for matching, prioritizing RFID (high-precision) signals, followed by Bluetooth Beacon (medium-precision) signals, and finally infrared (low-precision) signals. High-precision positioning source data is used first, and a two-byte positioning code (0~65535) is extracted from it. This code corresponds to the exhibit's voice or video data. At the same time, the millisecond-level audio and video timestamps carried in the UI playback control data are extracted. These timestamps are the time progress value of the audio and video playback start time, in milliseconds, with a maximum value of 16,777,215 milliseconds. Finally, the positioning code and audio and video timestamp information used for precise synchronization are obtained, providing content and timing basis for synchronized video playback.
[0142] S4-4. Decode the audio encoded data to obtain the decoded PCM audio data;
[0143] The separated audio encoded data is decoded using Opus, and the compressed audio data is restored to PCM audio data with a sampling rate of 48kHz. This data is the original audio data of the narrator's microphone in real time. The decoding and restoration of the narration audio is completed, which is prepared for subsequent mixing into the playback channel.
[0144] S4-5. Based on the playback control data, the positioning code and audio / video timestamp information, perform synchronized video playback, mix the decoded PCM audio data into the current audio / video playback channel, and combine batch instruction priority rules and radio frequency anti-interference switching algorithm to obtain audio / video playback information;
[0145] Based on playback control data, positioning encoding, and audio / video timestamp information, synchronous video playback is initiated to achieve synchronized audio and video playback between the transmitting and receiving ends, ensuring strict consistency in playback progress. Decoded PCM audio data is mixed into the current audio / video playback channel to achieve synchronized output of the narrator's voice and the video audio track. During processing, the priority rule for batch commands is followed: system command > audio control > video control > positioning synchronization. Commands of the same priority are processed sequentially in the message queue, with each queue processing once every 10ms. Simultaneously, a radio frequency anti-interference switching algorithm is activated. This involves using available frequency bands less than 1GHz (sub-1GHz) and employing an anti-interference switching algorithm combining channel monitoring and bit error rate analysis. If the Received Signal Strength Indicator (RSSI) of the channel is consistently higher than the baseline value or the bit error rate of three consecutive packets is greater than 10%, frequency hopping is automatically triggered. Three backup frequency bands are preset and attempted according to priority. After a successful switch, the switching is maintained for 5 minutes, ultimately generating complete audio / video playback information.
[0146] S4-6. Perform device offline detection and signal interruption detection based on the audio and video playback information, and obtain the audio and video synchronized explanation results;
[0147] An anomaly protection mechanism is activated based on audio and video playback information. Through feedback from the receiving end and retransmission from the sending end, offline device detection and corresponding handling are completed. For scenarios of instantaneous communication interruption, local caching is used to achieve smooth retransmission after interruption recovery, avoiding duplicate execution of audio, video, and commands. At the same time, anti-interference strategies such as channel monitoring and automatic frequency hopping are adopted to maintain the stability and reliability of the transmission link, ultimately ensuring continuous audio and video playback, accurate timing, and a complete and synchronized narration effect.
[0148] In summary, after receiving the RF data packet that has passed CRC32 verification, the receiver in steps S4-1 to S4-6 parses the mixed encoded data according to the data frame structure and transmission path, separating the UI playback control data, positioning signal data, and audio encoded data. It distinguishes between system control commands and playback control information by command type identification, matches the exhibit code and audio / video timestamps to the positioning signal according to multi-source fusion priority rules, decodes the audio encoded data into PCM narration audio via Opus, and then combines the playback control data, positioning code, and timestamps to achieve synchronized audio and video playback. The narration audio is also mixed into the audio / video channel. Simultaneously, it processes commands according to batch priority rules, and utilizes sub-1GHz band anti-interference frequency hopping, ACK retransmission, offline device detection, and signal buffering mechanisms. Ultimately, it achieves integrated real-time restoration of narration voice, synchronized video playback, and remote group control operations such as volume / channel / power-off. This effectively solves the problems of traditional narration formats being monotonous, lacking interactivity, susceptible to transmission interference, and inconvenient equipment management, significantly improving the stability of narration and the visitor experience in cultural tourism and study scenarios.
[0149] As one possible implementation, in the above embodiments, step S4-6 may specifically include the following steps:
[0150] S4-6-1. Based on the audio and video playback information, determine whether to receive the response confirmation packet. If yes, proceed to step S4-6-2. Otherwise, immediately start the data retransmission mechanism, obtain the number of data retransmissions, and proceed to step S4-6-3.
[0151] The transmitting end monitors in real time whether it has received an acknowledgment (ACK) packet from the receiving end based on the current audio and video playback information. This ACK packet is actively sent by the receiving end every 500ms and carries the latest reception timestamp, which is used by the transmitting end to confirm that the data has been delivered normally. If it is determined that the ACK packet has been received normally, the signal interruption monitoring process is initiated. If the ACK packet is not received, the data loss retransmission mechanism is immediately activated, and the corresponding number of data retransmissions is generated to prepare for subsequent retransmissions.
[0152] S4-6-2. Monitor whether the signal is interrupted based on the audio and video playback information. If so, execute the signal interruption buffering strategy, obtain the target audio and control data, and execute step S4-6-4. Otherwise, obtain the audio and video synchronization explanation result based on the audio and video playback information.
[0153] After successfully receiving the ACK response packet, the transmitter continues to monitor whether the radio frequency signal is interrupted based on the audio and video playback information. If the signal is interrupted, the signal interruption buffering strategy is executed, and the audio and control data of the most recent 2 seconds cached locally by the transmitter is used as the target audio and control data to be retransmitted, and then the retransmission execution process is entered; if the signal is not interrupted, the normal audio and video synchronized explanation result is directly output based on the current stable audio and video playback information.
[0154] S4-6-3. Determine whether the data retransmission information is greater than the retransmission threshold. If so, obtain the device offline as the audio and video synchronization explanation result; otherwise, execute step S4-6-2.
[0155] If no ACK packet is received, the transmitter immediately initiates data retransmission according to the data loss retransmission mechanism. At the same time, it checks whether the current number of data retransmissions has reached the retransmission threshold of 3 times. If the number of retransmissions has exceeded 3 times, that is, no response for 3 consecutive ACK cycles, the device is determined to be offline. The transmitter stops sending data and records the offline time. The offline status of the device is used as the final audio and video synchronization explanation result. At the same time, after the receiver restarts, it will actively send a synchronization request packet to the transmitter to re-establish the communication connection. If the number of retransmissions has not reached 3 times, a signal interruption detection operation is performed.
[0156] S4-6-4. Based on the target audio and control data and the instruction encoding sequence, the audio and video synchronous explanation results are obtained.
[0157] After the signal is restored, the transmitting end transmits the target audio and control data obtained from the buffer strictly in the order of the instruction numbers to the receiving end. During the reception process, the receiving end performs deduplication and merging of duplicate data to avoid the problem of repeated playback of audio and video and repeated execution of control instructions. Finally, the data retransmission is completed and a stable and reliable audio and video synchronized explanation result is output.
[0158] In summary, steps S4-6-1 to S4-6-4 effectively ensure data transmission reliability through a triple anomaly handling mechanism of data loss retransmission, device offline detection, and signal interruption buffering. This avoids repeated playback of audio and video and erroneous execution of commands, improves the system's anti-interference capability and communication stability, and ensures continuous and smooth synchronous audio and video explanations and accurate and efficient group control of devices in cultural tourism and study tour scenarios.
[0159] Further reference Figure 8 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of an audio-visual synchronized explanation system based on a team explanation device. This system embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.
[0160] like Figure 8 As shown, an audio-visual synchronous explanation system based on a team explainer in this embodiment includes a signal acquisition module, a hybrid encoding module, a radio frequency front-end data transmission module, and a separate decoding module;
[0161] like Figure 9 (a)- Figure 9 As shown in (b), the system includes one transmitter and at least one receiver. The transmitter and receiver interact with each other via a wireless channel. The system as a whole consists of a signal acquisition module, a hybrid encoding module, a radio frequency front-end data transmission module, and a separation decoding module, which work together to achieve synchronized audio and video narration. Specifically, the signal acquisition module and the hybrid encoding module are deployed at the transmitter to acquire interface operation commands, positioning signals, and microphone audio data, and to perform hybrid encoding of audio, control, and positioning information. The radio frequency front-end data transmission module is deployed at both the transmitter and receiver, responsible for the wireless transmission and reception of hybrid encoded data frames. The separation decoding module is deployed at the receiver to parse, decode, and execute commands on the received hybrid encoded data frames. Thus, through the hardware division of labor between the transmitter and receiver and the software collaboration of the four modules, a synchronized narration service integrating voice narration, video synchronization, and remote group control is achieved. Detailed descriptions of each module are as follows:
[0162] The signal acquisition module is used to acquire transmitter signals for preprocessing and obtain an audio dataset for explanation. The transmitter signals include interface operation instructions, positioning signals and PCM audio data.
[0163] This module is used to complete the acquisition and preprocessing of all signals from the transmitter. It collects interface operation commands, multi-source positioning signals, and 48kHz sampling rate PCM audio data from the microphone input in real time. It organizes and processes the playback control, volume adjustment, channel switching commands generated by the interface operation, the exhibit coding information mapped by RFID / Bluetooth / infrared positioning signals, and the raw audio data of the guide's real-time explanation to form a standardized audio dataset for explanation, providing complete and reliable raw input for subsequent encoding and transmission.
[0164] The hybrid encoding module is used to perform hybrid encoding based on the narration audio dataset and control and positioning information to obtain hybrid encoded narration audio data;
[0165] This module is used to integrate and encode the audio dataset output by the signal acquisition module with control and positioning information. According to the preset frame structure and timing rules, the audio PCM data is compressed and encoded using Opus. Interface operation instructions, positioning signals, playback timestamps and system control instructions are embedded in the audio frames. The data encapsulation is completed using a high compression ratio and a redundancy fault tolerance mechanism to generate hybrid encoded audio data that can be directly used for wireless transmission, realizing the same-frame fusion of audio, video and control information.
[0166] The radio frequency front-end data transmission module is used to encapsulate the hybrid encoded audio data into frames to obtain hybrid encoded data frames.
[0167] This module is used to encapsulate the hybrid encoded audio data output by the hybrid encoding module into frames according to the sub-1GHz radio frequency communication standard, and add preamble, synchronization word, frame length and CRC check information. The hybrid encoded data frames are stably sent to the receiving end through the wireless channel. It also supports channel monitoring, RSSI judgment and bit error rate detection. It can automatically switch to the backup frequency band according to the interference situation to ensure that the data is delivered to the receiving end stably, reliably and with low latency in complex cultural and tourism scenarios.
[0168] The separation decoding module is used to perform separation and parsing based on the hybrid encoded data frame, combined with the data transmission path and branch judgment logic, to obtain the audio and video synchronized explanation result;
[0169] This module receives mixed-encoded data frames sent by the RF front-end data transmission module. After passing the CRC check, it parses the data according to the data transmission path and branch judgment logic, separating UI playback control data, positioning signal data, and audio encoded data. It completes Opus audio decoding, command type recognition, multi-source fusion matching of positioning signals, audio and video timestamp synchronization, and command priority processing. Combined with anomaly mechanisms such as data retransmission, offline detection, and signal buffering, it finally outputs a synchronized audio and video explanation result with real-time voice restoration, synchronized video playback, and remote group control execution.
[0170] The fused data is transmitted from the transmitter to the receiver via the radio frequency front end; this step involves over-the-air transmission and reception of the fused data. The receiver separates and decodes the audio and control data from the received data, reconstructing the audio from the narrator's microphone, synchronously playing the video, and simultaneously performing operations such as powering off, changing channels, and adjusting volume. This step reconstructs the narrator's audio and actions on the receiver.
[0171] In summary, the system comprises a transmitter, at least one receiver, a signal acquisition module, a hybrid encoding module, an RF front-end data transmission module, and a separation decoding module. The signal acquisition module and hybrid encoding module are located at the transmitter, the RF front-end data transmission module is deployed at both the transmitter and receiver, and the separation decoding module is located at the receiver. The system acquires interface operation commands, positioning signals, and microphone audio input signals through the signal acquisition module at the transmitter, obtaining the narrator's audio and operation information. The hybrid encoding module fuses and encodes the audio, control, and positioning information. The RF front-end data transmission module then wirelessly transmits the fused and encoded data from the transmitter to each receiver. The receiver uses the separation decoding module to parse and decode the received data, restoring the microphone audio, synchronously playing the corresponding video, and automatically performing control operations such as volume adjustment, channel changing, and power off. Overall, the system provides a unified narration service integrating real-time audio playback, synchronized video display, and batch remote control, effectively enriching narration formats, enhancing interactivity and scene adaptability, and significantly improving the narration experience and equipment management efficiency in museums, scenic spots, and cultural tourism study tours.
[0172] Taking cultural tourism and study tour scenarios as an example, this embodiment proposes an audio-visual synchronized narration system based on a team tour guide, such as... Figure 10 As shown, this system uses the teacher's end (controller) as the core and the student's end as the receiving node. A single teacher's end connects to an unlimited number of student ends, enabling unified management and synchronous explanation services from the teacher's end to all student ends. Figure 11 As shown, the teacher's end can uniformly execute control functions such as voice transmission, video playback, volume adjustment, triggering explanations, changing channels, and batch shutdown on all student ends. The implementation logic of each function is as follows:
[0173] Voice transmission: When the teacher needs to give an explanation, they can speak directly into the microphone of the controller. The controller will collect the explanation audio, mix and encode it, and then broadcast it directly to each receiver (student) through the voice transmission channel. The students can listen to the teacher's voice explanation about the exhibits in real time, achieving zero-delay voice transmission.
[0174] Video Playback: Teachers can select specific exhibit videos on the display interface and push them for playback. Students will receive the push notification and the corresponding video will play synchronously. Teachers can also adjust the playback progress, and students can adjust it in real time. When teachers explain exhibits, they need to provide a detailed explanation through a video. Students can watch the video synchronously with the teacher and listen to the explanation as if they were there.
[0175] Adjust volume: The controller (teacher's end) can adjust the volume of all receivers (student's end) in real time and push it to all receivers (student's end), which is suitable for teaching when the environment changes (noisy or quiet) to ensure better listening effect.
[0176] Triggered Explanation: The controller (teacher's end) can collect RFID / Bluetooth / infrared trigger signals in real time. Each trigger signal corresponds to a number, and the number corresponds to the exhibit or cultural relic information. When the controller (teacher's end) enters the trigger range, it can trigger the corresponding explanation video, which is simultaneously pushed to all receivers (student's end). This is suitable for teachers to automatically provide explanations when switching between different exhibits.
[0177] Channel switching: The controller (teacher's end) can monitor the frequency and strength of radio frequency signals in the environment in real time through the radio frequency module. When the link signal quality deteriorates, the channel can be switched in real time and the signal can be sent to the receiver (student's end) in real time to ensure better communication quality.
[0178] Batch shutdown: The controller (teacher's end) can send shutdown commands to all receivers (students' ends) in real time. This is suitable for shutting down all receivers (students' ends) in batches after the lecture is completed, without requiring manual operation by the receivers (students' ends).
[0179] The teacher's end encodes the collected voice, video, control, and location information and transmits it to each student's end via a radio frequency front-end. The student's end receives and separates the decoded data, restores the narration, plays the video synchronously, and executes commands such as volume adjustment, channel switching, and power off. Ultimately, it realizes an integrated synchronous service for voice narration, video display, location triggering, and device management during the study tour, which greatly improves the organizational efficiency and interactive experience of cultural tourism study tours.
[0180] In this embodiment, the specific processing of an audio-visual synchronized explanation system based on a team explainer and the resulting technical effects can be referred to separately. Figure 2The relevant descriptions of steps S1, S2, S3 and S4 in the corresponding embodiments will not be repeated here.
[0181] It should be noted that the implementation details and technical effects of each module and unit in the device provided in the embodiments of this disclosure can be referred to the description of other embodiments in this disclosure, and will not be repeated here.
[0182] The following is for reference. Figure 12 It shows a schematic diagram of the structure of a computer system 500 suitable for implementing the electronic device of the present disclosure. Figure 12 The computer system 500 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0183] like Figure 12 As shown, the computer system 500 may include a processing device 501 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM 502 or a program loaded from storage device 508 into random access RAM 503. RAM 503 also stores various programs and data required for the operation of the computer system 500. The processing device 501, ROM 502, and RAM 503 are interconnected via bus 504. I / O interface 505 is also connected to bus 504.
[0184] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows computer system 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12 A computer system 500 with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0185] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0186] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0187] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0188] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following functions: Figure 1 The illustrated embodiments and their alternative implementations demonstrate a method for synchronized audio and video explanation based on a team explainer.
[0189] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0191] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the unit itself; for example, an acquisition module can also be described as "acquiring preset prompts, including modality fusion prompts, attention mechanism prompts, and / or time-related prompts."
[0192] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A method for synchronized audio and video explanation based on a team explanation device, characterized in that, include: S1. Collect transmitter signals and preprocess them to obtain an audio dataset for explanation. The transmitter signals include interface operation instructions, positioning signals and PCM audio data. S2. Based on the aforementioned audio dataset, combine control and positioning information to perform hybrid encoding to obtain hybrid encoded audio data. S3. Encapsulate the hybrid encoded audio data into frames to obtain hybrid encoded data frames; S4. Based on the hybrid encoded data frame, combined with the data transmission path and branch judgment logic, separate and parse the data to obtain the audio and video synchronized explanation result, including: Based on the data frame structure, the hybrid encoded data frame is parsed according to the data transmission path to obtain separate hybrid encoded data frames, which include UI playback control data, positioning signal data and audio encoded data. The UI playback control data is type-identified to determine whether it is a control command. If it is, the system operation is executed and the end of the path is obtained as the audio and video synchronization explanation result. Otherwise, the playback control data is obtained. Based on the positioning signal data and the multi-source fusion priority rules, multi-source matching is performed to extract the positioning code and match the corresponding audio and video data to obtain the positioning code and audio and video timestamp information. The audio encoded data is decoded to obtain decoded PCM audio data; Based on the playback control data, the positioning code and audio / video timestamp information, video is played synchronously. The decoded PCM audio data is mixed into the current audio / video playback channel, and audio / video playback information is obtained by combining batch instruction priority rules and radio frequency anti-interference switching algorithm. Based on the audio and video playback information, perform device offline detection and signal interruption detection to obtain the audio and video synchronized explanation results.
2. The method for synchronized audio and video explanation based on a team explanation device according to claim 1, characterized in that, S1. Acquire the transmitter signal and preprocess it to obtain the audio dataset for explanation, including: The transmitter is powered on, and the UI operation acquisition, positioning signal acquisition, and microphone sound acquisition are started simultaneously. Collect UI operations and detect whether a play operation is selected. If so, obtain the UI operation instruction; otherwise, return to UI operation collection. The system collects positioning signals and checks whether a trigger signal is received. If so, it acquires the positioning signal; otherwise, it returns to the positioning signal acquisition process. Based on the microphone sound acquisition, obtain the microphone audio signal; The interface operation instructions and the microphone audio signal are preprocessed to obtain the preprocessed interface operation instructions and PCM audio data; Based on the location signal, it is detected whether location information is received. If so, the location signal is preprocessed to obtain the preprocessed location signal; otherwise, the location signal acquisition is performed repeatedly. The preprocessed interface operation instructions, PCM audio data, and positioning signals are obtained as the explanatory audio dataset.
3. The method for synchronized audio and video explanation based on a team explanation device according to claim 1, characterized in that, S2. Based on the aforementioned audio dataset, a hybrid encoding process is performed using control and positioning information to obtain hybrid encoded audio data, including: The audio dataset used for explanation is sampled and quantized at fixed intervals to obtain raw PCM audio data. The raw PCM audio data is encoded to obtain basic compressed audio data; Based on the audio dataset provided, obtain control and positioning information; Based on the data frame structure, the control and positioning information and the basic compressed audio data are fused and encoded to obtain a mixed data frame; Cyclic redundancy check is performed on the mixed data frame to obtain the mixed-encoded audio data for explanation.
4. The method for synchronized audio and video explanation based on a team explanation device according to claim 3, characterized in that, The data frame structure includes a frame header identifier, a frame length identifier, an audio data segment, a control information embedding segment, a location information embedding segment, a cyclic redundancy check bit, and a frame tail identifier.
5. The method for synchronized audio and video explanation based on a team explanation device according to claim 1, characterized in that, S3. Encapsulate the hybrid encoded audio data into frames to obtain hybrid encoded data frames, including: Based on the aforementioned hybrid encoding, the audio data is encapsulated in conjunction with the radio frequency data frame structure to obtain a standard radio frequency data frame. The standard radio frequency data frame is transmitted via radio frequency to obtain a baseband radio frequency data frame; Based on the baseband radio frequency data frame and the radio frequency data frame structure, a parsing and verification is performed to obtain a hybrid coded data frame.
6. The method for synchronized audio and video explanation based on a team explanation device according to claim 1, characterized in that, Based on the audio and video playback information, perform device offline detection and signal interruption detection to obtain the audio and video synchronized explanation results, including: Based on the audio and video playback information, determine whether to receive the response confirmation packet. If yes, perform the first operation; otherwise, immediately start the data retransmission mechanism, obtain the data retransmission information, and perform the second operation. The first operation is as follows: monitor whether the signal is interrupted according to the audio and video playback information. If so, execute the signal interruption buffering strategy, obtain the target audio and control data, and execute the third operation. Otherwise, obtain the audio and video synchronized explanation result based on the audio and video playback information. The second operation is: determine whether the data retransmission information is greater than the retransmission threshold. If so, obtain the device offline status as the audio and video synchronization explanation result; otherwise, execute the first operation. The third operation is to supplement the transmission of the target audio and control data in accordance with the instruction encoding order to obtain the audio and video synchronized explanation result.
7. A synchronized audio-visual explanation system based on a team explanation device, using the method described in any one of claims 1-6, characterized in that, It includes a signal acquisition module, a hybrid coding module, an RF front-end data transmission module, and a separate decoding module; The signal acquisition module is used to acquire transmitter signals for preprocessing and obtain an audio dataset for explanation. The transmitter signals include interface operation instructions, positioning signals and PCM audio data. The hybrid encoding module is used to perform hybrid encoding based on the narration audio dataset and control and positioning information to obtain hybrid encoded narration audio data; The radio frequency front-end data transmission module is used to encapsulate the hybrid encoded audio data into frames to obtain hybrid encoded data frames. The separation decoding module is used to separate and parse the hybrid encoded data frame based on the data transmission path and branch judgment logic to obtain the audio and video synchronized explanation result.
8. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any one of claims 1-6.