AI-based virtual anchor interactive broadcasting communication system

CN122372127BActive Publication Date: 2026-08-28XIAMEN YISHI TIAODONG ARTIFICIAL INTELLIGENCE TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610793856.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-28
Estimated Expiration
2046-06-03

AI Technical Summary

Technical Problem

由此引出的核心技术问题是,现有广播通信系统固定复用机制无法在信道波动时保障交互关键数据的传输可靠性,易导致虚拟主播交互中断

Benefits of technology

[0048]1.本发明通过人工智能语义提取模型解析虚拟主播多模态数据流,将面部微表情参数与交互指令标记为高优先级关键数据,将躯干动作参数与背景音频标记为常规数据,并根据优先级标记动态调整正交频分复用子载波分配策略。将高优先级关键数据分配至分配权重系数最大的第一子载波集合作为所述专属子载波群,将常规数据分配至剩余的第二子载波集合作为高吞吐量子载波群,在接收端依据动态信令提取高优先级关键数据优先送入渲染引擎。此种处理在信道发生波动时,确保关键的面部与指令数据占据最优信道资源,避免其与常规数据竞争导致的丢包,维持交互过程基础动作与指令的连续性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372127B_ABST
    Figure CN122372127B_ABST
Patent Text Reader

Abstract

The present application relates to the field of broadcast communication, and specifically to an AI-based virtual anchor interactive broadcast communication system. The sending end parses virtual anchor multi-modal data stream through an artificial intelligence semantic extraction model, marks facial micro-expression parameters and interaction instructions as high-priority key data, and marks trunk action parameters and background audio as regular data; according to the priority, dynamically adjusts the orthogonal frequency division multiplexing sub-carrier allocation strategy, allocates the high-priority key data to the first sub-carrier set with the largest allocation weight coefficient as the exclusive sub-carrier group, and allocates the regular data to the remaining second sub-carrier set as the high-throughput sub-carrier group; the receiving end extracts high-priority key data according to dynamic signaling and preferentially feeds it into the rendering engine. The present application maintains stable transmission of key facial parameters and interaction instructions when the channel fluctuates, reduces the key data packet loss rate and delay, and avoids interaction interruption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of broadcast communication, specifically to an AI-based virtual anchor interactive broadcast communication system. Background Technology

[0002] Existing broadcast communication systems typically use a uniform encapsulation format to combine multimodal data streams, such as facial expressions, interactive commands, body movements, and background audio, into a single data stream for transmission. In this approach, the sending end assigns the same transmission priority to all data types, arranging them according to a fixed data frame structure. The receiving end, upon receiving the data, decapsulates it in a fixed order and extracts various parameters before sending them to the rendering engine. This transmission mechanism fails to consider the varying impacts of different data types on virtual anchor interaction, placing all data under the same transmission guarantee level for channel coding and modulation. This results in channel resources being allocated to different data types in a fixed proportion, making flexible adjustments based on data content impossible.

[0003] In physical layer transmission of broadcast channels, orthogonal frequency division multiplexing (OFDM) is a commonly used modulation method. Existing technologies typically allocate subcarriers based on channel state information. Their allocation strategies primarily rely on the instantaneous signal-to-noise ratio (SNR) of each subcarrier to allocate bits and power, or they use a fixed subcarrier mapping table to evenly distribute data across all subcarriers. This subcarrier allocation method aims to maximize overall throughput or balance the bit error rate, treating the virtual broadcaster data stream to be transmitted as an indiscriminate bit stream. When the data stream enters the channel coding and interleaving stage, a uniform coding rate and interleaving depth are used to process all data, without distinguishing whether the data contains critical control information with extremely high real-time interaction requirements.

[0004] Because existing broadcast communication systems employ indiscriminate channel multiplexing and coding strategies when transmitting multimodal data from virtual anchors, critical data such as facial micro-expression parameters and interaction commands, along with regular data like torso movements and background audio, face the same risk of being discarded or erroneous when the broadcast channel experiences deep fading or bandwidth congestion. The absence or delay of facial micro-expressions and interaction commands directly leads to severe visual abnormalities in the virtual anchor, such as interaction interruptions and facial stiffness, while slight delays in regular data have a relatively smaller impact on visual presentation. The core technical problem arising from this is that the fixed multiplexing mechanism of existing broadcast communication systems cannot guarantee the reliability of transmission of critical interactive data during channel fluctuations, easily leading to virtual anchor interaction interruptions. Summary of the Invention

[0005] The purpose of this invention is to provide an AI-based virtual anchor interactive broadcast communication system, which can effectively solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] An AI-based virtual anchor interactive broadcast communication system includes:

[0008] The sending end semantic extraction device is used to receive the virtual anchor multimodal data stream, parse the virtual anchor multimodal data stream through the AI ​​semantic extraction model, mark the facial micro-expression parameters and interaction instructions as high-priority key data, and mark the torso movement parameters and background audio as regular data;

[0009] A transmitting-end channel coding mapping device, connected to the transmitting-end semantic extraction device, is used to receive the high-priority key data and the regular data, dynamically adjust the orthogonal frequency division multiplexing subcarrier allocation strategy according to the priority mark, allocate the high-priority key data to the first subcarrier set with the largest allocation weight coefficient as the dedicated subcarrier group, allocate the regular data to the remaining second subcarrier set as the high-throughput subcarrier group, and generate a composite broadcast signal;

[0010] The receiving end demultiplexing device is connected to the transmitting end channel coding mapping device through a broadcast channel, and is used to receive the composite broadcast signal and extract the high-priority key data from the composite broadcast signal according to dynamic signaling;

[0011] The receiving end rendering device is connected to the receiving end demultiplexing device and is used to send the high-priority key data into the rendering engine to execute the virtual anchor image rendering, thereby forming the AI-based virtual anchor interactive broadcast communication system.

[0012] Preferably, the semantic extraction device at the sending end includes:

[0013] A cross-modal feature extraction component is used to extract features from the facial video frame sequence and the interactive instruction text sequence in the multimodal data stream of the virtual anchor, respectively, to obtain facial micro-expression feature vectors and instruction semantic feature vectors;

[0014] A spatiotemporal alignment and fusion component, connected to the cross-modal feature extraction component, is used to calculate the cross-correlation coefficient between the facial micro-expression feature vector and the instruction semantic feature vector in the time dimension, and align the facial micro-expression feature vector and the instruction semantic feature vector based on the cross-correlation coefficient to obtain an aligned multimodal joint feature matrix;

[0015] A semantic importance classification component, connected to the spatiotemporal alignment fusion component, is used to input the multimodal joint feature matrix into the classifier, output a semantic importance score, and mark the original data corresponding to the facial micro-expression feature vectors and instruction semantic feature vectors whose semantic importance scores are greater than a preset threshold as the high-priority key data.

[0016] Preferably, the transmitting end channel coding mapping device includes:

[0017] The channel state prediction component is used to obtain historical fading statistical parameters of the broadcast channel, input the historical fading statistical parameters into the autoregressive prediction model, and output a sequence of subcarrier signal-to-noise ratio prediction values ​​for future time slots.

[0018] A subcarrier allocation weight calculation component, connected to the channel state prediction component, is used to construct a subcarrier allocation optimization objective function based on the subcarrier signal-to-noise ratio prediction value sequence and the priority flag received from the transmitting end semantic extraction device, and solve the subcarrier allocation optimization objective function to obtain the allocation weight coefficient of each subcarrier;

[0019] An adaptive resource mapping component, connected to the subcarrier allocation weight calculation component, is used to allocate the high-priority key data to the first subcarrier set with the largest allocation weight coefficient as the dedicated subcarrier group, and allocate the regular data to the remaining second subcarrier set as the high-throughput subcarrier group, based on the allocation weight coefficient.

[0020] Preferably, the transmitting end channel coding mapping device further includes:

[0021] A differentiated channel coding component is used to perform low-rate LDPC coding and orthogonal spreading sequence spread spectrum processing on the high-priority key data mapped to the dedicated subcarrier group, and to perform high-rate LDPC coding and no spread spectrum processing on the regular data mapped to the high-throughput subcarrier group.

[0022] An interleaving depth adjustment component, connected to the differentiated channel coding component, is used to configure a first interleaving depth for the high-priority key data after low-code-rate LDPC coding and spread spectrum processing, configure a second interleaving depth greater than the first interleaving depth for the regular data after high-code-rate LDPC coding, and multiplex the data after configuring the interleaving depth to generate the composite broadcast signal.

[0023] Preferably, the receiving end demultiplexing device includes:

[0024] A frame header parsing component is used to extract a physical frame header from the received composite broadcast signal, and to parse a subcarrier mapping table and an interleaving depth identifier from the physical frame header. The subcarrier mapping table records the frequency domain position information of the dedicated subcarrier group and the high-throughput subcarrier group.

[0025] The demapping matrix reconstruction component, connected to the frame header parsing component, is used to reconstruct the subcarrier demapping matrix and deinterleaving parameters based on the subcarrier mapping table and the interleaving depth identifier;

[0026] The data extraction component, connected to the demapping matrix reconstruction component, is used to separate the data stream corresponding to the dedicated subcarrier group from the frequency domain data of the composite broadcast signal using the subcarrier demapping matrix, perform deinterleaving processing on the data stream according to the deinterleaving parameters, and output the high-priority key data.

[0027] Preferably, the receiving rendering device includes:

[0028] A dual-buffering scheduling component is used to construct an active buffer and a background buffer. The high-priority key data received from the receiving end demultiplexing device is written into the active buffer, and the regular data extracted from the composite broadcast signal is written into the background buffer.

[0029] An asynchronous rendering execution component, connected to the dual-buffer scheduling component, is used to prioritize reading the high-priority key data from the active buffer to drive the rendering of the screen corresponding to the virtual anchor's facial micro-expressions and interaction commands in each rendering cycle, and then read the regular data from the background buffer to drive the rendering of the virtual anchor's torso movements and background audio.

[0030] A cache update component, connected to the asynchronous rendering execution component, is used to clear the active cache area after the rendering cycle ends, and wait for the next rendering cycle to write the new high-priority key data.

[0031] Preferably, the spatiotemporal alignment and fusion component includes:

[0032] A time-series sliding window construction sub-component is used to construct a first time-series sliding window for the facial micro-expression feature vector and a second time-series sliding window for the instruction semantic feature vector, wherein the first time-series sliding window and the second time-series sliding window have overlapping time intervals;

[0033] A local cross-correlation calculation sub-component, connected to the temporal sliding window construction sub-component, is used to calculate the cosine similarity between the facial micro-expression feature vector and the instruction semantic feature vector within the overlapping time interval, and to use the cosine similarity as the cross-correlation number.

[0034] The feature filtering and splicing sub-component, connected to the local cross-correlation calculation sub-component, is used to filter the facial micro-expression feature vectors and the instruction semantic feature vectors whose cross-correlation coefficient is greater than the similarity threshold, splice and fuse them, and use the spliced ​​and fused vector matrix as the multimodal joint feature matrix.

[0035] Preferably, the channel state prediction component includes:

[0036] The delay compensation sub-component is used to obtain the propagation delay parameters of downlink transmission in the broadcast channel, and to perform time axis alignment compensation on the historical fading statistics parameters based on the propagation delay parameters to obtain the compensated fading sequence.

[0037] The long short-term memory prediction sub-component, connected to the delay compensation sub-component, is used to input the compensated fading sequence into the long short-term memory network, extract the temporal dependency features of the compensated fading sequence, and output the initial signal-to-noise ratio prediction value of the future time slot.

[0038] The residual correction sub-component, connected to the long short-term memory prediction sub-component, is used to calculate the prediction error between the initial signal-to-noise ratio prediction value and the actual signal-to-noise ratio measurement value in the previous time slot, and to use the prediction error to perform additive correction on the initial signal-to-noise ratio prediction value in the current time slot, thereby outputting the subcarrier signal-to-noise ratio prediction value sequence.

[0039] Preferably, the differentiated channel coding component includes:

[0040] An urgency assessment subcomponent is used to parse the interaction instructions in the high-priority key data, extract the response timestamp and interaction type label of the interaction instructions, and calculate the semantic urgency index based on the response timestamp and the interaction type label.

[0041] The spreading code length mapping subcomponent, connected to the urgency evaluation subcomponent, is used to pre-establish the mapping relationship between the semantic urgency index and the spreading code length in the orthogonal spreading code set, and determine the target spreading code length corresponding to the current high-priority key data based on the mapping relationship.

[0042] A spreading sequence generation sub-component, connected to the spreading code length mapping sub-component, is used to select a target spreading sequence that conforms to the target spreading code length from the orthogonal spreading code set, and to use the target spreading sequence to perform frequency domain expansion on the high-priority key data after low-code-rate LDPC encoding.

[0043] Preferably, the asynchronous rendering execution component includes:

[0044] The deformation primitive extraction sub-component is used to read facial micro-expression parameters from the high-priority key data in the active cache area and convert the facial micro-expression parameters into facial deformation primitives composed of facial vertex offset vectors.

[0045] The skeleton matrix extraction sub-component is used to read the torso motion parameters from the background buffer in the regular data and convert the torso motion parameters into a joint rotation matrix.

[0046] An interpolation smoothing subcomponent, connected to the deformation primitive extraction subcomponent and the bone matrix extraction subcomponent respectively, is used to detect the update timing difference between the facial deformation primitive and the joint rotation matrix in the rendering frame sequence. When there is an update timing difference, intermediate transition state data is generated using a spherical linear interpolation algorithm in the unupdated dimension. The facial deformation primitive, the joint rotation matrix and the intermediate transition state data are then merged and output to the rendering pipeline.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] 1. This invention analyzes the multimodal data stream of a virtual anchor using an artificial intelligence semantic extraction model. Facial micro-expression parameters and interaction commands are marked as high-priority key data, while torso movement parameters and background audio are marked as regular data. The orthogonal frequency division multiplexing (OFDM) subcarrier allocation strategy is dynamically adjusted based on these priority markings. High-priority key data is allocated to the first subcarrier set with the highest allocation weight coefficient as its dedicated subcarrier group, while regular data is allocated to the remaining second subcarrier set as a high-throughput subcarrier group. At the receiving end, high-priority key data is extracted based on dynamic signaling and preferentially sent to the rendering engine. This processing ensures that critical facial and command data occupy optimal channel resources during channel fluctuations, avoiding packet loss caused by competition with regular data and maintaining the continuity of basic actions and commands during the interaction process.

[0049] 2. This invention predicts the subcarrier signal-to-noise ratio (SNR) of future time slots based on historical fading statistical parameters and constructs an optimization objective function using priority marking to calculate allocation weight coefficients, achieving dynamic adaptation of subcarrier resources. For high-priority critical data on dedicated subcarrier groups, low-rate channel coding combined with orthogonal spreading sequences is used for spreading processing, with a lower interleaving depth. For regular data, high-rate channel coding is used without spreading, with a higher interleaving depth. The spreading code length is dynamically adjusted based on semantic urgency indicators to further enhance the physical layer anti-fading capability of critical data. This differentiated channel coding and spreading processing mechanism enhances the error correction and anti-interference capabilities of critical data at the physical layer signal level, reducing the bit error rate and transmission delay of critical interactive data.

[0050] 3. This invention constructs an active buffer and a background buffer at the receiving end. High-priority key data extracted is written to the active buffer to prioritize the rendering of facial micro-expressions and interactive commands, while regular data is written to the background buffer to drive torso motion and background audio rendering. Facial deformation primitives are extracted from facial micro-expression parameters, and joint rotation matrices are extracted from torso motion parameters. When a difference in update timing is detected, a spherical linear interpolation algorithm is used to generate intermediate transition state data, which is then merged and output to the rendering pipeline. This dual-buffer scheduling and interpolation smoothing mechanism avoids blockages in the rendering pipeline caused by waiting for regular data, eliminates visual stuttering caused by asynchronous updates of facial and torso motions, and improves the smoothness and stability of the end-to-end rendering process. Attached Figure Description

[0051] Figure 1 This is a flowchart illustrating the overall workflow of an AI-based virtual anchor interactive broadcast communication system according to the present invention.

[0052] Figure 2 This is a flowchart of the data processing of the semantic extraction device at the sending end of the present invention;

[0053] Figure 3 This is a flowchart illustrating the resource allocation and coding process of the transmitting-end channel coding mapping device of the present invention.

[0054] Figure 4 This is a flowchart of the data extraction process of the receiving end demultiplexing device of the present invention;

[0055] Figure 5 This is a flowchart of the dual-buffered scheduling rendering process of the receiving-end rendering device of the present invention;

[0056] Figure 6 This is a flowchart of the interpolation smoothing process of the asynchronous rendering execution component of the present invention. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] Please refer to Figure 1This embodiment provides an AI-based virtual anchor interactive broadcast communication system, including a transmitting-end semantic extraction device, a transmitting-end channel coding mapping device, a receiving-end demultiplexing device, and a receiving-end rendering device. The transmitting-end semantic extraction device receives a multimodal data stream from the virtual anchor, which includes a facial video frame sequence, a torso motion parameter sequence, an interactive instruction text sequence, and a background audio sequence. The transmitting-end semantic extraction device parses the multimodal data stream using an AI semantic extraction model, marking facial micro-expression parameters and interactive instructions as high-priority key data, and marking torso motion parameters and background audio as regular data. The transmitting-end channel coding mapping device is connected to the transmitting-end semantic extraction device, receives the high-priority key data and regular data, and dynamically adjusts the orthogonal frequency division multiplexing subcarrier allocation strategy according to the priority markings. The high-priority key data is allocated to the first subcarrier set with the largest allocation weight coefficient as the dedicated subcarrier group, and the regular data is allocated to the remaining second subcarrier set as the high-throughput subcarrier group, generating a composite broadcast signal. The receiver demultiplexing device is connected to the transmitter channel coding and mapping device via a broadcast channel. It receives composite broadcast signals and extracts high-priority key data from the composite broadcast signals based on dynamic signaling. The receiver rendering device is connected to the receiver demultiplexing device and prioritizes the high-priority key data into the rendering engine for virtual anchor image rendering.

[0059] The virtual anchor multimodal data stream is output by the virtual anchor generation system. The facial video frame sequence is acquired at 30 frames per second, with each frame containing pixel data of the virtual anchor's facial area. The torso motion parameter sequence contains the rotation angles and position coordinates of the virtual anchor's 23 joints, sampled at 60 Hz. The interactive command text sequence is converted from user-inputted text commands, including command content, sending timestamp, and user identifier. The background audio sequence is stereo digital audio data with a sampling rate of 44.1 kHz. The sending-end semantic extraction device performs timestamp synchronization processing on the input multimodal data stream, aligning the data from different modalities along a unified timeline to ensure the accuracy of subsequent semantic analysis.

[0060] Before the start of each transmission time slot, the transmitting-end channel coding and mapping device acquires the current broadcast channel status information, including the signal-to-noise ratio, channel gain, and interference level of each subcarrier. Based on the data priority marker and channel status information, the transmitting-end channel coding and mapping device dynamically calculates the allocation weight of each subcarrier, forming a dedicated subcarrier group with the highest weights for transmitting high-priority critical data; the remaining subcarriers form a high-throughput subcarrier group for transmitting regular data. The number of subcarriers in the dedicated subcarrier group is dynamically adjusted according to the transmission requirements of high-priority critical data. When the amount of high-priority critical data increases, additional subcarriers are automatically allocated from the high-throughput subcarrier group to the dedicated subcarrier group.

[0061] After receiving the composite broadcast signal, the receiving demultiplexing device first performs synchronization processing, including carrier synchronization, symbol synchronization, and frame synchronization. After synchronization, the receiving demultiplexing device extracts the subcarrier mapping table and interleaving depth identifier from the physical frame header of the composite broadcast signal, and reconstructs the demapping matrix and deinterleaving parameters based on this information. Using the demapping matrix, the receiving demultiplexing device separates the data stream corresponding to the dedicated subcarrier group from the frequency domain data of the composite broadcast signal, performs deinterleaving and decoding processing on this data stream, and outputs high-priority key data. For the data stream corresponding to the high-throughput subcarrier group, the receiving demultiplexing device performs corresponding deinterleaving and decoding processing after extracting and decoding the high-priority key data, and outputs regular data.

[0062] The receiving-end rendering device constructs a dual-buffer structure, including an active buffer and a background buffer. High-priority critical data is written to the active buffer, while regular data is written to the background buffer. Within each rendering cycle, the receiving-end rendering device first reads data from the active buffer to drive the rendering of the virtual anchor's facial micro-expressions and interactive commands; then it reads data from the background buffer to drive the rendering of the virtual anchor's torso movements and background audio. The rendering cycle is consistent with the display device's refresh rate to ensure smooth image output. Once the data in the active buffer has been read, the buffer update component clears the active buffer, waiting for the high-priority critical data to be written in the next transmission slot. The transmission characteristics and priority division of different data types are shown in Table 1.

[0063] Table 1 Transmission characteristics and priority classification of different data types

[0064]

[0065] Table 1 illustrates the transmission characteristics and priority classification of different data types in the virtual anchor's multimodal data stream. Facial micro-expression parameters and interactive commands have high requirements for transmission latency and bit error rate, and are classified as Level 1 high-priority critical data; torso motion parameters and background audio have relatively lower requirements for transmission latency and bit error rate, and are classified as Level 2 regular data. Based on this priority classification, the system allocates different channel resources and transmission guarantee mechanisms to different types of data to ensure reliable transmission of critical data during channel fluctuations.

[0066] In this embodiment, the broadcast channel employs Orthogonal Frequency Division Multiplexing (OFDM) technology, with a subcarrier spacing of 15 kHz. Each transmission time slot contains 14 OFDM symbols, and the time slot length is 1 millisecond. The total system bandwidth is 20 MHz, containing 1200 available subcarriers. The number of subcarriers in the dedicated subcarrier group ranges from 100 to 300, dynamically adjusted according to the transmission requirements of high-priority critical data. When channel conditions are good, the number of subcarriers in the dedicated subcarrier group remains around 100; when channel conditions deteriorate, the number of subcarriers in the dedicated subcarrier group increases to 300 to improve the reliability of critical data transmission.

[0067] In a preferred embodiment, reference Figure 2 The semantic extraction device at the sending end includes a cross-modal feature extraction component, a spatiotemporal alignment and fusion component, and a semantic importance classification component. The cross-modal feature extraction component extracts features from the facial video frame sequence and the interactive command text sequence in the virtual anchor's multimodal data stream, obtaining facial micro-expression feature vectors and command semantic feature vectors. The spatiotemporal alignment and fusion component is connected to the cross-modal feature extraction component, calculates the cross-correlation coefficient between the facial micro-expression feature vectors and the command semantic feature vectors in the time dimension, and aligns the facial micro-expression feature vectors and command semantic feature vectors based on the cross-correlation coefficient, obtaining an aligned multimodal joint feature matrix. The semantic importance classification component is connected to the spatiotemporal alignment and fusion component, inputs the multimodal joint feature matrix into a classifier, outputs a semantic importance score, and marks the original data corresponding to facial micro-expression feature vectors and command semantic feature vectors with semantic importance scores greater than a preset threshold as high-priority key data.

[0068] The cross-modal feature extraction component comprises two parallel feature extraction branches, one for processing facial video frame sequences and the other for processing interactive command text sequences. The facial feature extraction branch employs a 3D convolutional neural network (CNN) structure, taking a sequence of 16 consecutive facial video frames as input and outputting a 512-dimensional facial micro-expression feature vector. The CNN consists of three 3D convolutional layers, two 3D pooling layers, and one fully connected layer. The first 3D convolutional layer uses 64 3×3×3 kernels with a stride of 1×1×1 and padding of 1×1×1; the second uses 128 3×3×3 kernels with a stride of 1×1×1 and padding of 1×1×1; and the third uses 256 3×3×3 kernels with a stride of 1×1×1 and padding of 1×1×1. Each 3D convolutional layer is followed by a batch normalization layer and a ReLU activation function layer. The 3D pooling layer uses max pooling with a kernel size of 2×2×2 and a stride of 2×2×2. The fully connected layer maps the feature maps output by the 3D convolutional layer into 512-dimensional facial micro-expression feature vectors.

[0069] The text feature extraction branch employs a bidirectional long short-term memory (BSSM) network structure. The input is the word embedding vector of the interactive instruction text sequence, and the output is a 256-dimensional instruction semantic feature vector. The word embedding vector has a dimension of 128 and is generated by a pre-trained word embedding model. The BSSM network contains two layers of bidirectional BSSM units, each containing 128 hidden units. The output of the BSSM network is passed through an average pooling layer and a fully connected layer, mapping to a 256-dimensional instruction semantic feature vector.

[0070] The spatiotemporal alignment and fusion component includes a temporal sliding window construction sub-component, a local cross-correlation calculation sub-component, and a feature selection and splicing sub-component. The temporal sliding window construction sub-component constructs a first temporal sliding window for facial micro-expression feature vectors and a second temporal sliding window for instruction semantic feature vectors, with overlapping time intervals between the first and second temporal sliding windows. The first temporal sliding window is 16 frames long, corresponding to a time length of 533 milliseconds; the second temporal sliding window is 8 instruction semantic feature vectors long, also corresponding to a time length of 533 milliseconds. Both windows have a sliding step size of 1 time unit to ensure overlap between adjacent windows.

[0071] The local cross-correlation calculation subcomponent calculates the cosine similarity between the facial micro-expression feature vector and the instruction semantic feature vector within the overlapping time interval, and uses the cosine similarity as the cross-correlation coefficient. The formula for calculating the cosine similarity is:

[0072]

[0073] in, This represents the i-th facial micro-expression feature vector. This represents the semantic feature vector of the j-th instruction. This represents the vector dot product operation. This represents the L2 norm of a vector.

[0074] The feature selection and splicing sub-component selects facial micro-expression feature vectors and command semantic feature vectors with a cross-correlation coefficient greater than a similarity threshold, and splices and fuses them. The spliced ​​and fused vector matrix is ​​used as the multimodal joint feature matrix. The similarity threshold is set to 0.7. When the cross-correlation coefficient is greater than 0.7, the corresponding facial micro-expression feature vector and command semantic feature vector are considered to have a strong correlation in time and semantics. The splicing and fusion operation concatenates the 512-dimensional facial micro-expression feature vector and the 256-dimensional command semantic feature vector into a 768-dimensional joint feature vector. Multiple joint feature vectors form the multimodal joint feature matrix.

[0075] The semantic importance classification component employs a multilayer perceptron classifier, taking a 768-dimensional multimodal joint feature matrix as input and outputting a semantic importance score between 0 and 1. The multilayer perceptron contains three fully connected layers: the first fully connected layer maps the 768-dimensional input to 512 dimensions, the second to 256 dimensions, and the third to 1 dimension. Each fully connected layer is followed by a batch normalization layer and a ReLU activation function layer, and the last fully connected layer is followed by a Sigmoid activation function layer, restricting the output value to between 0 and 1. A preset threshold of 0.6 is used; when the semantic importance score is greater than 0.6, the corresponding original data is marked as high-priority key data; otherwise, it is marked as regular data. The network structure parameters of the cross-modal feature extraction component are shown in Table 2.

[0076] Table 2 Network structure parameters of the cross-modal feature extraction component

[0077]

[0078] Table 2 shows the network structure parameters of the cross-modal feature extraction component. The facial feature extraction branch extracts spatiotemporal features from facial video frame sequences using a 3D convolutional neural network, while the text feature extraction branch extracts semantic features from interactive command text sequences using a bidirectional long short-term memory network. The outputs of the two branches are spatiotemporally aligned and fused before being input into the semantic importance classification component for priority labeling.

[0079] In this embodiment, the training process of the cross-modal feature extraction component adopts a supervised learning approach. The training dataset contains 10,000 sets of virtual anchor multimodal data samples, each sample containing a facial video frame sequence, an interactive command text sequence, and a corresponding semantic importance label. During training, the cross-entropy loss function is used, the Adam algorithm is employed as the optimizer, the learning rate is set to 0.001, the batch size is set to 32, and the number of training epochs is set to 50. After training, the parameters of the cross-modal feature extraction component and the semantic importance classification component are fixed for data processing in the actual system.

[0080] Further, refer to Figure 3The transmitting-end channel coding and mapping device includes a channel state prediction component, a subcarrier allocation weight calculation component, and an adaptive resource mapping component. The channel state prediction component acquires historical fading statistics of the broadcast channel, inputs these statistics into an autoregressive prediction model, and outputs a sequence of predicted subcarrier signal-to-noise ratio (SNR) values ​​for future time slots. The subcarrier allocation weight calculation component, connected to the channel state prediction component, constructs a subcarrier allocation optimization objective function based on the subcarrier SNR prediction sequence and the priority markers received from the transmitting-end semantic extraction device, and solves the subcarrier allocation optimization objective function to obtain the allocation weight coefficients for each subcarrier. The adaptive resource mapping component, also connected to the subcarrier allocation weight calculation component, allocates high-priority critical data to the first subcarrier set with the largest allocation weight coefficient as a dedicated subcarrier group, and allocates regular data to the remaining second subcarrier set as a high-throughput subcarrier group.

[0081] The channel state prediction component includes a delay compensation subcomponent, a long short-term memory prediction subcomponent, and a residual correction subcomponent. The delay compensation subcomponent acquires the propagation delay parameters of the downlink transmission in the broadcast channel and performs time-axis alignment compensation on the historical fading statistics parameters based on these parameters to obtain the compensated fading sequence. The propagation delay parameters are calculated from the distance between the transmitter and receiver and the electromagnetic wave propagation speed, using the following formula:

[0082]

[0083] in, Indicates delayed transmission. Indicates the distance between the sender and receiver. This represents the speed at which electromagnetic waves propagate in free space, approximately... meters per second.

[0084] The historical fading statistics parameters include the signal-to-noise ratio (SNR) measurements for each subcarrier over the past N time slots, where N is set to 20. The delay compensation subcomponent shifts the historical fading statistics parameters forward on the time axis. The time frame aligns the historical data timeline with the timeline of future time slots. The compensated fading sequence is 20 units long, with each element corresponding to a subcarrier's signal-to-noise ratio (SNR) measurement over the past 20 time slots.

[0085] The Long Short-Term Memory (LSTM) prediction subcomponent inputs the compensated fading sequence into the LSM network, extracts the temporal dependency features of the compensated fading sequence, and outputs the initial signal-to-noise ratio (SNR) prediction value for the future time slot. The LSM network consists of one LSM unit and one fully connected layer. The LSM unit contains 64 hidden units, with an input dimension of 1 and an output dimension of 64. The fully connected layer maps the 64-dimensional hidden states to a 1-dimensional initial SNR prediction value. The LSM network takes a compensated fading sequence of length 20 as input and outputs the initial SNR prediction value for the next time slot.

[0086] The residual correction sub-component calculates the prediction error between the initial SNR prediction value and the actual SNR measurement value of the previous time slot, and uses this prediction error to perform an additive correction on the initial SNR prediction value of the current time slot, outputting a sequence of subcarrier SNR prediction values. The formula for calculating the prediction error is:

[0087]

[0088] in, This represents the prediction error for the (k-1)th time slot. This represents the initial signal-to-noise ratio prediction value for the (k-1)th time slot. This represents the actual signal-to-noise ratio measurement for the (k-1)th time slot.

[0089] The formula for calculating the predicted subcarrier signal-to-noise ratio for the current time slot is:

[0090]

[0091] in, This represents the predicted signal-to-noise ratio (SNR) of the subcarrier in the k-th time slot. This represents the initial signal-to-noise ratio prediction value for the k-th time slot. This represents the correction factor, which ranges from 0 to 1. In this embodiment, it is set to 0.3.

[0092] The subcarrier allocation optimization objective function constructed by the subcarrier allocation weight calculation component is:

[0093]

[0094]

[0095] If subcarrier When allocated to high-priority critical data, then If subcarrier When assigned to regular data, then .

[0096] in, This indicates the total number of available subcarriers. Indicates the first The weighting coefficients for each subcarrier allocation. Indicates the first The transmit power of each subcarrier, Indicates the first Channel gain of each subcarrier Indicates noise power. Indicates the total transmission power. This represents the priority weight coefficient, which takes a value greater than 1; in this embodiment, it is set to 3. Indicates the first The predicted signal-to-noise ratio of each subcarrier.

[0097] A greedy algorithm is used to solve the subcarrier allocation optimization objective function. First, all subcarriers are sorted from highest to lowest according to their predicted signal-to-noise ratio (SNR). Then, the sorted subcarriers are allocated to high-priority critical data in sequence until their transmission requirements are met. The remaining subcarriers are allocated to regular data. During the allocation process, the transmit power of each subcarrier is allocated according to a water-filling algorithm to maximize the total system throughput.

[0098] The adaptive resource mapping component allocates high-priority critical data to the first subcarrier set with the largest allocation weight coefficient, forming a dedicated subcarrier group, based on allocation weight coefficients. Regular data is allocated to the remaining second subcarrier set, forming a high-throughput subcarrier group. The number of subcarriers in the dedicated subcarrier group is calculated based on the number of bits in the high-priority critical data and the number of bits each subcarrier can carry, using the following formula:

[0099]

[0100] in, This indicates the number of subcarriers in a dedicated subcarrier group. This indicates the number of bits representing high-priority critical data. This indicates the number of bits that each subcarrier can carry. This indicates the rounding up operation. Table 3 shows the subcarrier modulation schemes and bit carrying capacity under different signal-to-noise ratio conditions.

[0101] Table 3 Subcarrier modulation schemes and bit carrying capacity under different signal-to-noise ratio conditions

[0102]

[0103] Table 3 shows the subcarrier modulation schemes and bit carrying capacity under different signal-to-noise ratio (SNR) conditions. The adaptive resource mapping component selects an appropriate modulation scheme and coding rate based on the predicted SNR value of each subcarrier, determining the number of bits each subcarrier can carry. Then, based on the number of bits of high-priority key data, it calculates the required number of subcarriers and groups the subcarriers with the highest assigned weight coefficients into a dedicated subcarrier group.

[0104] In this embodiment, the transmitting-end channel coding and mapping device further includes a differentiated channel coding component and an interleaving depth adjustment component. The differentiated channel coding component performs low-rate LDPC coding on high-priority critical data mapped to dedicated subcarrier groups and spreads it using an orthogonal spreading sequence. It performs high-rate LDPC coding on regular data mapped to high-throughput subcarrier groups without spreading. The interleaving depth adjustment component is connected to the differentiated channel coding component. It configures a first interleaving depth for the high-priority critical data after low-rate LDPC coding and spreading, and configures a second interleaving depth greater than the first interleaving depth for the regular data after high-rate LDPC coding. The data with configured interleaving depths are then multiplexed to generate a composite broadcast signal.

[0105] The differentiated channel coding component includes an urgency assessment subcomponent, a spreading code length mapping subcomponent, and a spreading sequence generation subcomponent. The urgency assessment subcomponent parses the interaction instructions in high-priority key data, extracts the response timestamp and interaction type label of the interaction instructions, and calculates a semantic urgency index based on the response timestamp and interaction type label. The response timestamp represents the latest time that the interaction instruction needs to be responded to, and the interaction type label includes "real-time control," "information query," and "content recommendation," etc. The formula for calculating the semantic urgency index is:

[0106]

[0107] in, Indicates the semantic urgency index, Indicates the current time. Indicates the response timestamp. This indicates the weighting coefficient for the interaction type. The weighting coefficient for "Real-time Control" type interaction commands is set to 1.5, the weighting coefficient for "Information Query" type interaction commands is set to 1.0, and the weighting coefficient for "Content Recommendation" type interaction commands is set to 0.5.

[0108] The spreading code length mapping subcomponent pre-establishes a mapping relationship between the semantic urgency index and the length of spreading codes in the orthogonal spreading code set, and determines the target spreading code length corresponding to the current high-priority key data based on the mapping relationship. The mapping relationship is as follows: when the semantic urgency index is greater than 0.8, the target spreading code length is set to 64; when the semantic urgency index is between 0.5 and 0.8, the target spreading code length is set to 32; when the semantic urgency index is less than 0.5, the target spreading code length is set to 16.

[0109] The spreading sequence generation sub-component selects a target spreading sequence from the orthogonal spreading code set that conforms to the target spreading code length. This target spreading sequence is then used to perform frequency domain expansion on the high-priority critical data after low-rate LDPC encoding. The orthogonal spreading code set uses Walsh-Hadamard codes, which possess good orthogonality and autocorrelation. The calculation formula for the spreading process is:

[0110]

[0111] in, This represents the spread spectrum signal vector. This represents the encoded signal vector. Indicates the target spreading sequence. This represents the Kronecker product operation.

[0112] The interleaving depth adjustment component configures a first interleaving depth of 128 for high-priority critical data and a second interleaving depth of 512 for regular data. Interleaving processing uses a block interleaving method, writing the input data sequence into a two-dimensional matrix in row-major order and then reading it out in column-major order to achieve data interleaving. Interleaving depth represents the number of columns in the two-dimensional matrix; a larger interleaving depth provides stronger resistance to burst errors, but also increases transmission latency. High-priority critical data has higher requirements for transmission latency, so a smaller interleaving depth is configured; regular data has lower requirements for transmission latency, so a larger interleaving depth is configured to improve resistance to burst errors.

[0113] In this embodiment, quasi-cyclic LDPC encoding is used, with a code length of 1944 bits. The code rate of low-rate LDPC encoding is 1 / 3, and the code rate of high-rate LDPC encoding is 2 / 3. The generator matrix and parity check matrix of LDPC encoding are defined by standard specifications, and the encoding process adopts an encoding method based on the generator matrix. The decoding process uses the belief propagation algorithm, with the number of iterations set to 20.

[0114] refer to Figure 4The receiver demultiplexing device includes a frame header parsing component, a demapping matrix reconstruction component, and a data extraction component. The frame header parsing component extracts the physical frame header from the received composite broadcast signal, and parses the subcarrier mapping table and interleaving depth identifier from the physical frame header. The subcarrier mapping table records the frequency domain position information of dedicated subcarrier groups and high-throughput subcarrier groups. The physical frame header adopts a fixed-length structure, containing a synchronization sequence, frame number, subcarrier mapping table, interleaving depth identifier, and cyclic redundancy check (CRC) code. The synchronization sequence is used for carrier synchronization and symbol synchronization, and has a length of 64 bits; the frame number is used to identify different transmission frames, and has a length of 16 bits; the subcarrier mapping table records the allocation of each subcarrier, and has a length of 1200 bits; the interleaving depth identifier is used to indicate the interleaving depth of different data types, and has a length of 8 bits; the CRC code is used to detect transmission errors in the physical frame header, and has a length of 32 bits.

[0115] The demapping matrix reconstruction component is connected to the frame header parsing component. Based on the subcarrier mapping table and interleaving depth identifier, it reconstructs the subcarrier demapping matrix and deinterleaving parameters. The subcarrier demapping matrix is ​​a... The matrix represents the total number of available subcarriers, where M represents the total number of available subcarriers. The first column of the matrix represents the subcarrier index, and the second column represents the data type carried by that subcarrier, with "1" indicating high-priority critical data and "0" indicating regular data. Deinterleaving parameters include interleaving depth and block size, which are obtained from a predefined parameter table based on the interleaving depth identifier.

[0116] The data extraction component is connected to the demapping matrix reconstruction component. Using the subcarrier demapping matrix, it separates the data stream corresponding to the dedicated subcarrier group from the frequency domain data of the composite broadcast signal. Based on the deinterleaving parameters, it performs deinterleaving processing on the data stream, outputting high-priority key data. The frequency domain data is obtained through Fast Fourier Transform, with each subcarrier corresponding to a complex sample value. The data extraction component extracts the complex sample values ​​corresponding to all subcarriers marked "1" according to the subcarrier demapping matrix, forming the frequency domain data stream of the dedicated subcarrier group. Then, it performs demapping and decoding processing on this frequency domain data stream to obtain the encoded bit stream. Finally, it performs deinterleaving processing on the encoded bit stream according to the deinterleaving parameters, outputting high-priority key data. The physical frame header structure and the length of each field are shown in Table 4.

[0117] Table 4 Physical Frame Header Structure and Field Lengths

[0118]

[0119] Table 4 shows the structure of the physical frame header and the length of each field. The physical frame header is 1320 bits long and occupies one orthogonal frequency division multiplexing (OFDM) symbol. The receiver's demultiplexing device first detects the synchronization sequence to complete carrier synchronization and symbol synchronization; then it extracts the other fields of the physical frame header, parsing out the subcarrier mapping table and interleaving depth identifier; finally, it uses a cyclic redundancy check (CRC) code to verify the correctness of the physical frame header. If the CRC fails, the frame is discarded and a retransmission is requested.

[0120] refer to Figure 5 The receiver rendering unit includes a dual-buffer scheduling component, an asynchronous rendering execution component, and a buffer update component. The dual-buffer scheduling component constructs an active buffer and a background buffer. High-priority critical data received from the receiver demultiplexing unit is written to the active buffer, and regular data extracted from the composite broadcast signal is written to the background buffer. Both the active and background buffers use a circular queue structure, with each buffer being 10 megabytes in size. The active buffer stores high-priority critical data from the most recent transmission time slot, while the background buffer stores regular data from the most recent three transmission time slots.

[0121] The asynchronous rendering execution component is connected to the dual-buffer scheduling component. Within each rendering cycle, it prioritizes reading high-priority critical data from the active buffer to drive the rendering of the virtual anchor's facial micro-expressions and interaction commands. Subsequently, it reads regular data from the background buffer to drive the rendering of the virtual anchor's torso movements and background audio. The rendering cycle is consistent with the display device's refresh rate; in this embodiment, it is set to 16.67 milliseconds, corresponding to a 60Hz refresh rate. In the first 5 milliseconds of each rendering cycle, the asynchronous rendering execution component reads data from the active buffer to complete the rendering of the facial micro-expressions and interaction commands; in the last 11.67 milliseconds of the rendering cycle, it reads data from the background buffer to complete the rendering of the torso movements and background audio.

[0122] refer to Figure 6 The asynchronous rendering execution component includes a deformation primitive extraction subcomponent, a bone matrix extraction subcomponent, and an interpolation smoothing subcomponent. The deformation primitive extraction subcomponent reads facial micro-expression parameters from high-priority key data in the active cache and transforms these parameters into facial deformation primitives composed of facial vertex offset vectors. The virtual anchor's facial model contains 5000 vertices, each with 3 coordinate components. The facial micro-expression parameter is a 46-dimensional vector, with each component corresponding to the activation level of a facial action unit. The formula for calculating the facial deformation primitive is:

[0123]

[0124] in, This represents the coordinate matrix of the deformed facial vertices. The facial vertex coordinate matrix representing a neutral expression. This indicates the activation level of the k-th facial action unit. This represents the vertex offset matrix corresponding to the k-th facial motion unit.

[0125] The skeleton matrix extraction subcomponent reads torso motion parameters from the background buffer's regular data and transforms them into joint rotation matrices. The virtual anchor's skeletal system contains 23 joints, each with 3 rotational degrees of freedom. The torso motion parameters are a 69-dimensional vector, with each component corresponding to a joint's rotation angle. The joint rotation matrix is ​​represented using quaternions, and the conversion formula between quaternions and Euler angles is:

[0126]

[0127] in, Represents quaternions, The unit vector representing the axis of rotation.

[0128] The interpolation smoothing subcomponent is connected to the deformation primitive extraction subcomponent and the bone matrix extraction subcomponent, respectively. It detects the temporal differences in the update of facial deformation primitives and joint rotation matrices within the rendering frame sequence. When such differences exist, intermediate transition state data is generated using a spherical linear interpolation algorithm on the unupdated dimensions. The facial deformation primitives, joint rotation matrices, and intermediate transition state data are then merged and output to the rendering pipeline. The calculation formula for the spherical linear interpolation algorithm is as follows:

[0129]

[0130] in, This represents the interpolated quaternion. and A quaternion representing two keyframes. These represent interpolation coefficients, with values ​​ranging from 0 to 1. The angle between two quaternions is expressed by the following formula:

[0131]

[0132] The cache update component connects to the asynchronous rendering execution component. After the rendering cycle ends, it clears the active cache area, waiting for the next rendering cycle to write new high-priority critical data. The background cache area uses a first-in, first-out (FIFO) update strategy; when new regular data is written, the oldest data is automatically overwritten. The cache update component also monitors cache usage. When the data write delay in the active cache area exceeds 100 milliseconds, it automatically triggers the cache refresh mechanism, discarding expired data and requesting retransmission.

[0133] In a preferred embodiment, the receiving rendering device further includes an error-hiding component for handling transmission errors of high-priority critical data. When errors or loss of high-priority critical data are detected, the error-hiding component uses facial micro-expression parameters from the previous frame and interaction commands from the current frame to generate estimated facial micro-expression parameters through a linear interpolation algorithm, which are then used to drive the rendering of the virtual anchor's face. The presence of the error-hiding component can prevent obvious jumps or freezes in the virtual anchor's face when data transmission errors occur, thus improving the continuity of visual presentation.

[0134] In this embodiment, the rendering engine employs physically based rendering technology, supporting real-time global illumination and shadow effects. The rendering pipeline includes a vertex shader, a tessellation shader, a geometry shader, a fragment shader, and a computation shader. The vertex shader handles vertex position and normal transformations; the tessellation shader adds detail to the model; the geometry shader generates additional geometric primitives; the fragment shader calculates the color and lighting for each pixel; and the computation shader performs general computational tasks, such as particle system simulation and post-processing effects.

[0135] In a preferred embodiment, the transmitting channel coding mapping apparatus further includes a power control component for dynamically adjusting the transmit power of each subcarrier based on the predicted signal-to-noise ratio (SNR) of the subcarriers. The power control component employs a water-filling algorithm to allocate more transmit power to subcarriers with higher SNR, thereby maximizing the total system throughput. The calculation formula for the water-filling algorithm is as follows:

[0136]

[0137] in, Indicates the first The transmit power of each subcarrier, The Lagrange multiplier is determined by the total transmit power constraint. Indicates noise power. Indicates the first Channel gain of each subcarrier.

[0138] The power control component further improves the utilization efficiency of channel resources, increasing the total system throughput while keeping the total transmit power constant. Simultaneously, for subcarriers within a dedicated subcarrier group, the power control component prioritizes their transmit power, ensuring the reliability of high-priority critical data transmission.

[0139] In a preferred embodiment, the receiver demultiplexing device further includes a channel estimation component for estimating the channel state information of the broadcast channel. The channel estimation component uses pilot symbols inserted at the transmitter to estimate the channel gain of each subcarrier using a least squares algorithm or a minimum mean square error algorithm. The formula for the least squares algorithm is:

[0140]

[0141] in, This represents the least squares channel estimate. Indicates the received pilot symbol, Indicates the pilot symbol to be transmitted.

[0142] The formula for calculating the least mean square error algorithm is as follows:

[0143]

[0144] in, This represents the minimum mean square error channel estimate. The autocorrelation matrix represents the channel. Indicates noise power. Represents the identity matrix.

[0145] The channel state information output by the channel estimation component is fed back to the transmitter to update the historical fading statistics parameters of the channel state prediction component, thereby improving the accuracy of subcarrier signal-to-noise ratio prediction. The feedback channel uses a low-rate uplink control channel with a feedback period of 1 millisecond.

[0146] In this embodiment, the system's workflow is as follows: The virtual anchor generation system outputs a multimodal data stream. The sending-end semantic extraction device receives this data stream and performs semantic analysis, marking the data as high-priority key data and regular data. The sending-end channel coding and mapping device dynamically allocates subcarrier resources based on channel state prediction results and data priority markings, employing differentiated channel coding and interleaving processing for data of different priorities to generate a composite broadcast signal. The composite broadcast signal is transmitted to the receiving end through a broadcast channel. The receiving-end demultiplexing device extracts high-priority key data and regular data from the composite broadcast signal. The receiving-end rendering device uses a dual-buffering scheduling mechanism, prioritizing the rendering of facial micro-expressions and interactive command screens corresponding to high-priority key data, followed by rendering the torso movements and background audio corresponding to regular data, ultimately outputting a complete virtual anchor image and sound.

[0147] In a preferred embodiment, the system supports multiple users simultaneously receiving virtual anchor broadcast content. The transmitting-end channel coding and mapping device generates an independent subcarrier mapping table and interleaving depth identifier for each user, mapping high-priority key data of different users to different dedicated subcarrier groups. The receiving-end demultiplexing device extracts the corresponding high-priority key data and regular data from the composite broadcast signal based on its own user identifier. This multi-user support mechanism can provide personalized virtual anchor interactive services to multiple users on the same broadcast channel, improving the utilization efficiency of channel resources.

[0148] In this embodiment, the system's end-to-end latency is less than 100 milliseconds, which meets the real-time requirements of virtual anchor interaction services. When deep fading occurs in the broadcast channel, the bit error rate of high-priority critical data remains stable. Below, the bit error rate of conventional data remains at The system can dynamically adjust subcarrier allocation strategies and channel coding parameters when channel conditions change, ensuring the continuity and stability of the virtual anchor interaction process.

Claims

1. An AI-based virtual anchor interactive broadcast communication system, characterized in that, include: The sending end semantic extraction device is used to receive the virtual anchor multimodal data stream, parse the virtual anchor multimodal data stream through the AI ​​semantic extraction model, mark the facial micro-expression parameters and interaction instructions as high-priority key data, and mark the torso movement parameters and background audio as regular data; A transmitting end channel coding mapping device, connected to the transmitting end semantic extraction device, is used to receive the high-priority key data and the regular data, dynamically adjust the orthogonal frequency division multiplexing subcarrier allocation strategy according to the priority mark, allocate the high-priority key data to the first subcarrier set with the largest allocation weight coefficient as a dedicated subcarrier group, allocate the regular data to the remaining second subcarrier set as a high-throughput subcarrier group, and generate a composite broadcast signal. The receiving end demultiplexing device is connected to the transmitting end channel coding mapping device through a broadcast channel, and is used to receive the composite broadcast signal and extract the high-priority key data from the composite broadcast signal according to dynamic signaling; The receiving end rendering device is connected to the receiving end demultiplexing device and is used to send the high-priority key data into the rendering engine to execute the virtual anchor image rendering, thereby forming the AI-based virtual anchor interactive broadcast communication system.

2. The AI-based virtual anchor interactive broadcast communication system according to claim 1, characterized in that, The semantic extraction device at the sending end includes: A cross-modal feature extraction component is used to extract features from the facial video frame sequence and the interactive instruction text sequence in the multimodal data stream of the virtual anchor, respectively, to obtain facial micro-expression feature vectors and instruction semantic feature vectors; A spatiotemporal alignment and fusion component, connected to the cross-modal feature extraction component, is used to calculate the cross-correlation coefficient between the facial micro-expression feature vector and the instruction semantic feature vector in the time dimension, and align the facial micro-expression feature vector and the instruction semantic feature vector based on the cross-correlation coefficient to obtain an aligned multimodal joint feature matrix; A semantic importance classification component, connected to the spatiotemporal alignment fusion component, is used to input the multimodal joint feature matrix into the classifier, output a semantic importance score, and mark the original data corresponding to the facial micro-expression feature vectors and instruction semantic feature vectors whose semantic importance scores are greater than a preset threshold as the high-priority key data.

3. The AI-based virtual anchor interactive broadcast communication system according to claim 2, characterized in that, The transmitting end channel coding mapping device includes: The channel state prediction component is used to obtain historical fading statistical parameters of the broadcast channel, input the historical fading statistical parameters into the autoregressive prediction model, and output a sequence of subcarrier signal-to-noise ratio prediction values ​​for future time slots. A subcarrier allocation weight calculation component, connected to the channel state prediction component, is used to construct a subcarrier allocation optimization objective function based on the subcarrier signal-to-noise ratio prediction value sequence and the priority flag received from the transmitting end semantic extraction device, and solve the subcarrier allocation optimization objective function to obtain the allocation weight coefficient of each subcarrier; An adaptive resource mapping component, connected to the subcarrier allocation weight calculation component, is used to allocate the high-priority key data to the first subcarrier set with the largest allocation weight coefficient as the dedicated subcarrier group, and allocate the regular data to the remaining second subcarrier set as the high-throughput subcarrier group, based on the allocation weight coefficient.

4. The AI-based virtual anchor interactive broadcast communication system according to claim 1, characterized in that, The transmitting end channel coding mapping device further includes: A differentiated channel coding component is used to perform low-rate LDPC coding and orthogonal spreading sequence spread spectrum processing on the high-priority key data mapped to the dedicated subcarrier group, and to perform high-rate LDPC coding and no spread spectrum processing on the regular data mapped to the high-throughput subcarrier group. An interleaving depth adjustment component, connected to the differentiated channel coding component, is used to configure a first interleaving depth for the high-priority key data after low-code-rate LDPC coding and spread spectrum processing, configure a second interleaving depth greater than the first interleaving depth for the regular data after high-code-rate LDPC coding, and multiplex the data after configuring the interleaving depth to generate the composite broadcast signal.

5. The AI-based virtual anchor interactive broadcast communication system according to claim 1, characterized in that, The receiver demultiplexing device includes: A frame header parsing component is used to extract a physical frame header from the received composite broadcast signal, and to parse a subcarrier mapping table and an interleaving depth identifier from the physical frame header. The subcarrier mapping table records the frequency domain position information of the dedicated subcarrier group and the high-throughput subcarrier group. The demapping matrix reconstruction component, connected to the frame header parsing component, is used to reconstruct the subcarrier demapping matrix and deinterleaving parameters based on the subcarrier mapping table and the interleaving depth identifier; The data extraction component, connected to the demapping matrix reconstruction component, is used to separate the data stream corresponding to the dedicated subcarrier group from the frequency domain data of the composite broadcast signal using the subcarrier demapping matrix, perform deinterleaving processing on the data stream according to the deinterleaving parameters, and output the high-priority key data.

6. The AI-based virtual anchor interactive broadcast communication system according to claim 1, characterized in that, The receiving rendering device includes: A dual-buffering scheduling component is used to construct an active buffer and a background buffer. The high-priority key data received from the receiving end demultiplexing device is written into the active buffer, and the regular data extracted from the composite broadcast signal is written into the background buffer. An asynchronous rendering execution component, connected to the dual-buffer scheduling component, is used to prioritize reading the high-priority key data from the active buffer to drive the rendering of the screen corresponding to the virtual anchor's facial micro-expressions and interaction commands in each rendering cycle, and then read the regular data from the background buffer to drive the rendering of the virtual anchor's torso movements and background audio. A cache update component, connected to the asynchronous rendering execution component, is used to clear the active cache area after the rendering cycle ends, and wait for the next rendering cycle to write the new high-priority key data.

7. The AI-based virtual anchor interactive broadcast communication system according to claim 2, characterized in that, The spatiotemporal alignment and fusion component includes: A time-series sliding window construction sub-component is used to construct a first time-series sliding window for the facial micro-expression feature vector and a second time-series sliding window for the instruction semantic feature vector, wherein the first time-series sliding window and the second time-series sliding window have overlapping time intervals; A local cross-correlation calculation sub-component, connected to the temporal sliding window construction sub-component, is used to calculate the cosine similarity between the facial micro-expression feature vector and the instruction semantic feature vector within the overlapping time interval, and to use the cosine similarity as the cross-correlation number. The feature filtering and splicing sub-component, connected to the local cross-correlation calculation sub-component, is used to filter the facial micro-expression feature vectors and the instruction semantic feature vectors whose cross-correlation coefficient is greater than the similarity threshold, splice and fuse them, and use the spliced ​​and fused vector matrix as the multimodal joint feature matrix.

8. The AI-based virtual anchor interactive broadcast communication system according to claim 3, characterized in that, The channel state prediction component includes: The delay compensation sub-component is used to obtain the propagation delay parameters of downlink transmission in the broadcast channel, and to perform time axis alignment compensation on the historical fading statistics parameters based on the propagation delay parameters to obtain the compensated fading sequence. The long short-term memory prediction sub-component, connected to the delay compensation sub-component, is used to input the compensated fading sequence into the long short-term memory network, extract the temporal dependency features of the compensated fading sequence, and output the initial signal-to-noise ratio prediction value of the future time slot. The residual correction sub-component, connected to the long short-term memory prediction sub-component, is used to calculate the prediction error between the initial signal-to-noise ratio prediction value and the actual signal-to-noise ratio measurement value in the previous time slot, and to use the prediction error to perform additive correction on the initial signal-to-noise ratio prediction value in the current time slot, thereby outputting the subcarrier signal-to-noise ratio prediction value sequence.

9. The AI-based virtual anchor interactive broadcast communication system according to claim 4, characterized in that, The differentiated channel coding component includes: An urgency assessment subcomponent is used to parse the interaction instructions in the high-priority key data, extract the response timestamp and interaction type label of the interaction instructions, and calculate the semantic urgency index based on the response timestamp and the interaction type label. The spreading code length mapping subcomponent, connected to the urgency evaluation subcomponent, is used to pre-establish the mapping relationship between the semantic urgency index and the spreading code length in the orthogonal spreading code set, and determine the target spreading code length corresponding to the current high-priority key data based on the mapping relationship. A spreading sequence generation sub-component, connected to the spreading code length mapping sub-component, is used to select a target spreading sequence that conforms to the target spreading code length from the orthogonal spreading code set, and to use the target spreading sequence to perform frequency domain expansion on the high-priority key data after low-code-rate LDPC encoding.

10. The AI-based virtual anchor interactive broadcast communication system according to claim 6, characterized in that, The asynchronous rendering execution component includes: The deformation primitive extraction sub-component is used to read facial micro-expression parameters from the high-priority key data in the active cache area and convert the facial micro-expression parameters into facial deformation primitives composed of facial vertex offset vectors. The skeleton matrix extraction sub-component is used to read the torso motion parameters from the background buffer in the regular data and convert the torso motion parameters into a joint rotation matrix. An interpolation smoothing subcomponent, connected to the deformation primitive extraction subcomponent and the bone matrix extraction subcomponent respectively, is used to detect the update timing difference between the facial deformation primitive and the joint rotation matrix in the rendering frame sequence. When there is an update timing difference, intermediate transition state data is generated using a spherical linear interpolation algorithm in the unupdated dimension. The facial deformation primitive, the joint rotation matrix and the intermediate transition state data are then merged and output to the rendering pipeline.

Citation Information

Patent Citations

  • Audio and video player control method based on voice instruction

    CN121053987A

  • Dual-frequency signal adaptive switching transmission method

    CN121690437A