Speech recognition-based broadcasting host auxiliary system

By constructing a broadcasting and hosting assistance system based on a virtual potential energy field and the firefly algorithm, the problem of process gaps caused by the non-standardized expression of the host in live broadcasting was solved. The system achieved continuous tracking of the host's intentions and the coherence of the broadcast control process, thereby improving the automation level of live broadcasting.

CN122120478APending Publication Date: 2026-05-29MAOMING ZHENGJING INFORMATION TECHNOLOGY CONSULTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MAOMING ZHENGJING INFORMATION TECHNOLOGY CONSULTING CO LTD
Filing Date
2026-02-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing voice control methods are ill-equipped to handle impromptu speech, word order reversal, or semantically ambiguous expressions by presenters during live broadcasts. They also lack the ability to reason about ambiguous semantics, which can easily lead to interruptions in the live broadcast process.

Method used

A speech recognition-based broadcasting and hosting assistance system is adopted. The system acquires multimodal data through the acquisition module, constructs a virtual potential energy field through the feature extraction module, solves the sparse activation vector through the adjustment module, and obtains the centroid trajectory of the firefly population through the optimization module, thereby realizing material preloading and playback prompts.

Benefits of technology

It enables continuous tracking and precise positioning of the host's intentions, enhances the adaptability to non-standardized expressions, ensures the continuity and rapid response of the broadcast control process, and optimizes memory resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120478A_ABST
    Figure CN122120478A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on voice recognition's broadcasting host auxiliary system, it is related to voice recognition technical field, including acquisition module, feature extraction module, adjusting module and optimization module, the present application extracts program series single semantic feature in broadcasting preparation stage, and constructs virtual potential energy field, program is mapped as gravitational centroid, and potential energy guide groove and timing repulsive field are constructed based on timing, when live, the acoustic characteristics of host's speech are extracted in real time, construct target function based on sparse representation to solve activation vector, dynamically adjust the gravitational depth and effective action radius of potential energy field, optimization is carried out in potential energy field using firefly algorithm, real-time monitoring group barycenter trajectory, combined with physical operation signal executes material preloading or play control, the present application converts intention recognition into physical field optimization, solves the static text and dynamic voice dimension mismatch problem, realizes continuous tracking to non-standard expression using swarm intelligence, ensures the coherence of live process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and in particular to a broadcasting and hosting assistance system based on speech recognition. Background Technology

[0002] With the rapid development of the broadcasting, television, and online live streaming industries, the production process of live programs has become increasingly complex, placing extremely high demands on the comprehensive qualities of broadcasters and hosts. In traditional live streaming scenarios, hosts not only need to focus on content expression and control of the atmosphere, but also often need to handle tedious directing operations such as teleprompter page turning, background music switching, sound effects playback, and multimedia material scheduling. In order to reduce labor costs and improve broadcast security, voice recognition-based auxiliary systems have been gradually introduced into the field of broadcasting and hosting. These systems aim to assist hosts in completing material retrieval, process advancement, and equipment control through intelligent voice interaction technology, achieving a highly efficient live streaming mode with a single person and a single machine. This has significant application value for improving the automation level and response speed of program production.

[0003] Among existing voice control technologies, Chinese invention patent CN116939144A discloses a voice control method and device for online meetings. This technology directly obtains the host's control needs through voice recognition, eliminating the need for dialogue between the host and the back-end control personnel, thus saving operation time and improving control efficiency to a certain extent. However, the aforementioned existing technologies still have significant limitations when facing professional and complex live broadcasting scenarios. Existing voice control methods mostly adopt a discrete interaction mode of using wake words as commands and confirmations. This linear retrieval and matching mechanism is mainly for standardized and explicit commands, and it is difficult to deal with non-standardized expressions such as impromptu speech, word order reversal, or semantic ambiguity by the host during live broadcasts. It lacks the ability to reason about ambiguous semantics. Existing semantic matching technologies are usually based on isolated recognition of single-frame speech, ignoring the inherent temporal logic and contextual association of the program sequence on the timeline. This results in the system being unable to predict the next action based on the current broadcast progress during silent intervals or when there is a lack of explicit voice commands, which can easily cause process interruptions. Summary of the Invention

[0004] The technical problem solved by this invention is that the linear retrieval and matching mechanism of existing voice control methods is mainly unable to cope with the non-standardized expressions of hosts during live broadcasts, such as impromptu speech, word order reversal, or semantic ambiguity. It lacks the ability to reason about ambiguous semantics. Existing semantic matching technology ignores the inherent temporal logic and contextual association of the program sequence on the timeline, which can easily cause process interruptions.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a speech recognition-based broadcasting and hosting assistance system, comprising: a data acquisition module, a feature extraction module, an adjustment module, and an optimization module; The acquisition module is used to acquire multimodal data in live streaming scenarios; The feature extraction module is used to extract semantic features of the program sequence in multimodal data during the broadcast preparation stage, construct a virtual potential energy field, and map the semantic features to the gravitational center of mass in the virtual potential energy field. The adjustment module is used to extract the acoustic features of the host's audio stream from the multimodal data after the live broadcast starts, construct an objective function based on the acoustic features and semantic features, solve the objective function to obtain a sparse activation vector, and adjust the total potential value of the virtual potential field based on the sparse activation vector. The optimization module is used to map the total potential energy value of the virtual potential energy field to an absolute brightness value, combine the firefly algorithm to obtain the center-of-gravity trajectory of the firefly population, and execute material preloading instructions and generate a playback prompt signal.

[0006] Preferably, the multimodal data includes program playlists, host audio streams, and physical operation signals; The program playlist includes program entries arranged in chronological order. Each program entry is a structured text entry, and each entry contains a unique identifier, metadata tags, and media resource links. The host's audio stream includes digital audio signals acquired in real time by a microphone array; The physical operation signals include the travel value of the physical pusher, the on / off status of the physical button, and the status of the foot pedal.

[0007] Preferably, the process of extracting the semantic features includes: The pre-built synonym generation model is used to perform data augmentation on each structured text entry in the program sequence to generate a corresponding set of generalized descriptions. The structured text entries and the set of generalized descriptions are input into a pre-trained language encoder to obtain the baseline feature vector.

[0008] Preferably, the process of constructing the virtual potential energy field specifically includes: The reference feature vector corresponding to each structured text entry is mapped to a coordinate point, which serves as the gravitational center of mass of the virtual dynamic potential energy field. The effective radius of action of all gravitational centers of mass is the preset initial effective radius of action. Iterate through the coordinates of all programs on the program schedule and arrange all coordinates in the order of broadcast. Calculate the Euclidean distance between adjacent coordinate points. If the Euclidean distance between adjacent coordinate points is greater than the preset connectivity threshold, then construct a potential energy guiding trench connecting the adjacent coordinate points. The total potential energy at any point in the virtual potential energy field includes nodal gravitational potential energy, tubular guiding potential energy, and time-series repulsive potential energy. The mathematical expression for the total potential energy is: ; in, This represents the total potential energy. This refers to the total number of programs in the program schedule. For program number indexing, Let be any point in the virtual potential energy field. For the first The gravitational potential energy of the nodes corresponding to each program For the first The potential energy of the strip guides the tubular guiding potential energy corresponding to the groove. The number of grooves is determined by potential energy. It represents the temporal repulsive potential energy.

[0009] Preferably, the mathematical expression for the nodal gravity is: ; in, For the first The gravitational potential energy of the nodes corresponding to each program This is the real-time gravitational depth coefficient. Let be a point in the virtual potential energy field. For the first The core coordinates of the program The effective radius of action in real time (initial value is 5.0). Index the program by number; The mathematical expression for the tubular guided potential energy is: ; in, For the first The tubular guiding potential energy corresponding to the guiding groove. For the numbering index of potential energy guiding trenches, The preset depth of the trench is used to guide potential energy. Vertical distance represents the distance between points. to line segment The shortest vertical distance, The potential energy guides the width of the trench; The mathematical expression for the time-series repulsive potential energy is: ; in, For temporal repulsive potential energy, This represents a collection containing the indexes of all programs that have finished playing. The centroid coordinates of the already played programs, This is the repulsive force intensity coefficient.

[0010] Preferably, the process of extracting the acoustic features includes: The host's audio stream is time-sliced, the Mel frequency cepstral coefficients of the audio stream in each time slice are extracted, and an acoustic feature matrix is ​​constructed. Each row of the acoustic feature matrix represents the Mel frequency cepstral coefficients corresponding to the time slice.

[0011] Preferably, the process of adjusting the virtual potential energy field specifically includes: Construct a semantic feature matrix corresponding to the program sequence list, where each column of the semantic feature matrix corresponds to the baseline feature vector of a program item; Principal component analysis is used to reduce the dimensionality of the semantic feature matrix, and a dimensionality-reduced semantic feature mapping matrix is ​​constructed. The acoustic matrix is ​​averaged in the time domain to obtain the acoustic semantic mapping vector at the current time. Construct an intent-based objective function based on sparse representation to solve for the sparse activation vectors of the acoustic semantic mapping vector on the semantic feature basis matrix; The mathematical expression for the objective function to be solved is: ; in, Acoustic semantic mapping vector , Let be the sparse activation vector to be solved. For sparse regularization parameters; The objective function is solved using an iterative optimization algorithm to obtain a sparse activation vector, and the elements in the sparse activation vector are used as the real-time semantic matching weights of the program. The normalized information entropy of the sparse activation vector is used as the semantic entropy, and the effective radius of action of the total potential energy value of the virtual potential energy field is adjusted by the semantic entropy. The mathematical expression for adjusting the effective radius of action is: ; in, For the real-time effective radius of action, The initial effective radius of action, Normalized semantic entropy; The real-time gravitational depth coefficient of the total potential energy value of the virtual potential energy field is adjusted using a linear mapping formula and real-time semantic matching weights. The mathematical expression for adjusting the real-time gravitational depth coefficient is as follows: ; in, This is the real-time gravitational depth coefficient. Based on the fundamental gravitational depth coefficient, The acoustic gain coefficient. Weights for real-time semantic matching; The total potential energy value at any point in the virtual potential energy field is recalculated based on the real-time gravity depth coefficient and the effective radius of action.

[0012] Preferably, the process of obtaining the centroid trajectory of the firefly population specifically includes: A firefly population of a preset size is initialized and randomly distributed in a virtual potential energy field. The absolute brightness value of each firefly in the population is calculated. The absolute brightness value is negatively correlated with the total potential energy value of the location of each firefly in the population. Individuals in a firefly population with low absolute brightness values ​​update their positions to those within the population's perception range that have higher absolute brightness values. The mathematical expression for this position update is: ; in, For the first A firefly Location at any given moment In the fireflies The absolute brightness value within the perception range is better than firefly Location, To maximize attraction, The light intensity absorption coefficient is... Fireflies and The Euclidean distance between them for Random numbers within the interval This is the step size factor (in this embodiment, the value is 0.2). Preferably, the process of executing the material preloading instruction and generating the playback prompt signal specifically includes: Real-time calculation of the spatial distribution centroid of firefly swarms, average swarm movement speed, and continuous dwell time; Define a preload radius and a playback trigger radius that are concentrically set around each gravitational center of mass; Real-time monitoring of the center of gravity and physical triggering signals of firefly swarms; If no physical trigger signal is detected, and the center of gravity of the firefly swarm spatial distribution enters the preload radius of the gravitational center of mass, an anti-accidental touch judgment is executed. When all anti-accidental touch judgment conditions are met, a material preload instruction is generated, and the underlying interface is asynchronously called to load media resources into the memory buffer. The conditions for preventing accidental touches include: The average movement speed of the group is lower than the preset convergence threshold; The continuous dwell time of the center of gravity within the preload radius exceeds the preset time threshold. If no physical trigger signal is detected, and the center of gravity of the firefly swarm enters the playback trigger radius of the gravitational center of mass, and the swarm convergence density exceeds a preset threshold, a prompt signal will be generated. If a physical trigger signal is detected, and the centroid of the firefly swarm's spatial distribution does not enter the preload radius of any gravitational center of mass, the next program that has already been played is selected and played directly according to the timing of the program sequence. If a physical trigger signal is detected, and the centroid of the firefly swarm's spatial distribution enters the preload radius of the gravitational center of mass, the program corresponding to the current gravitational center of mass will be played directly.

[0013] Preferably, the optimization module further includes a timing monitoring unit for monitoring the playback status of the program, wherein the playback status is that playback has ended; If the current program has finished playing, then the program will be merged into a set containing all programs that have finished playing. The specific process of monitoring the playback status of a program includes: monitoring the digital playback progress and physical channel status of the current program; and determining that the current program has ended playback when any playback end condition is met. These playback end conditions include: The media playback kernel feedback file ends, including the media playback kernel feedback interruption signal and playback progress timestamp equal to the total duration of the media file; The physical fader travel value for controlling the program is lower than the preset mute threshold.

[0014] The beneficial effects of this invention are as follows: This invention introduces a spatial mapping mechanism based on a virtual potential energy field, which maps heterogeneous discrete program semantic features and continuous acoustic features into a unified topographic structure and energy distribution in the potential energy field. This solves the dimensional mismatch problem between static text and dynamic speech, transforming the complex intent recognition task into a physical optimization process in continuous space. This enables continuous tracking and precise positioning of the host's intent and enhances the adaptability to non-standardized expressions. Through the construction of potential energy guiding trenches and temporal repulsion fields, physical constraints on the live broadcast process logic are realized, effectively ensuring the continuity of the broadcast control process. By real-time monitoring of the spatial centroid, migration speed, and dwell time of the firefly swarm, long-distance search and local convergence states are accurately distinguished, effectively eliminating false triggers caused by signal jitter. Based on the design of concentric dual threshold regions, hierarchical control of asynchronous preloading of materials and instant playback triggering is realized, optimizing memory resource usage while ensuring rapid response. Attached Figure Description

[0015] Figure 1 This is a basic flowchart of a speech recognition-based broadcasting and hosting assistance system provided in one embodiment of the present invention. Detailed Implementation

[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0017] Example, refer to Figure 1 It provides a speech recognition-based broadcasting and hosting assistance system, including: a data acquisition module, a feature extraction module, an adjustment module, and an optimization module; The acquisition module is used to acquire multimodal data in live streaming scenarios; The feature extraction module is used to extract semantic features of the program sequence in multimodal data during the broadcast preparation stage, construct a virtual potential energy field, and map the semantic features to the gravitational center of mass in the virtual potential energy field. The adjustment module is used to extract the acoustic features of the host's audio stream from the multimodal data after the live broadcast starts, construct an objective function based on the acoustic and semantic features, solve the objective function to obtain sparse activation vectors, and adjust the total potential energy value of the virtual potential energy field based on the sparse activation vectors. The optimization module maps the total potential energy value of the virtual potential energy field to an absolute brightness value, combines the firefly algorithm to obtain the center-of-gravity trajectory of the firefly population, and executes material preloading instructions and generates a playback prompt signal.

[0018] This invention breaks away from the rigid model of traditional broadcasting and hosting assistance systems that rely on linear keyword matching. By utilizing the collaborative work of the acquisition module, feature extraction module, and dynamic field modulation module, the abstract host's intention is transformed into a quantifiable gravitational center of mass and energy distribution in a potential energy field. The optimization module introduces the firefly algorithm to replace traditional logical judgment and uses the convergence characteristics of swarm intelligence to smooth out random noise and instantaneous errors in speech recognition. This enables continuous tracking and robust inference of the host's intention in live broadcast scenarios, effectively avoiding process interruptions caused by single-frame speech recognition errors.

[0019] Multimodal data includes program playlists, host audio streams, and physical operation signals; The program schedule includes program entries arranged in chronological order. Each program entry is a structured text entry, and each entry contains a unique identifier, metadata tags, and media resource links. The presenter's audio stream consists of digital audio signals captured in real time by the microphone array; Physical operation signals include the travel value of the physical fader, the on / off state of the physical buttons, and the state of the foot pedal. This includes the travel value of the fader on the live streaming console, the trigger level of the physical buttons, and the on / off state of the foot pedal.

[0020] In one specific embodiment of the present invention, the broadcast control system retrieves program items from the program schedule and obtains the digital signal from the host's microphone from the digital mixing console; Metadata tags include keywords related to the program name, which are derived from the host's script and speech. For example, if the program is "Rice Fragrance", the keywords would be Jay Chou, pop, and healing.

[0021] The process of extracting semantic features includes: The pre-built synonym generation model is used to perform data augmentation on each structured text item in the program sequence, generating a corresponding set of generalized descriptions. The structured text entries and the set of generalized descriptions are input into a pre-trained language encoder to obtain the baseline feature vector.

[0022] In a specific embodiment of the present invention, the preset synonym generation model is a synonym generation model (BART) based on the Transformer architecture, which is used to generate generalized text with the same semantics as the original instruction but different expression (e.g., for the entry of playing the song "Rice Fragrance", generate the synonyms "play a song by Jay Chou" and "listen to Rice Fragrance"). In this embodiment, the number of generalized texts is greater than 5. The unique identifier, metadata tag, and generalized text from the program montage are input into the pre-trained language encoder (a lightweight DistilBERT model), and the corresponding feature vectors are output. The arithmetic mean of the feature vectors is calculated to obtain the baseline feature vector. In subsequent algorithms, the baseline feature vectors corresponding to structured text entries need to be mapped to coordinate points. In this embodiment, the program's unique identifier, metadata tag, and generalized text are used to ensure that a program does not need to appear twice, or that similar programs are mapped to the same coordinates. That is, each program has a unique mapped coordinate. This invention utilizes the BART model to augment program entries with data, generating a generalized description set containing various colloquial expressions. Combined with lightweight DistilBERT to extract baseline feature vectors, it can pre-construct centroids covering a broad semantic space during the broadcast preparation stage. Even if the host temporarily changes their wording or uses ambiguous pronouns during the live broadcast, it can still be mapped to the correct program node based on the similarity of the high-dimensional semantic space, significantly reducing the reliance on the host's standard language.

[0023] The process of constructing a virtual potential energy field specifically includes: The reference feature vector corresponding to each structured text entry is mapped to a coordinate point, which serves as the gravitational center of mass of the virtual dynamic potential energy field. The effective radius of action of all gravitational centers of mass is the preset initial effective radius of action. Iterate through the coordinates of all programs on the program schedule and arrange all coordinates in the order of broadcast. Calculate the Euclidean distance between adjacent coordinate points. If the Euclidean distance between adjacent coordinate points is greater than the preset connectivity threshold, then construct a potential energy guiding trench connecting the adjacent coordinate points. The total potential energy at any point in the virtual potential energy field includes nodal gravitational potential energy, tubular guiding potential energy, and time-series repulsive potential energy. The mathematical expression for the total potential energy is: ; in, This represents the total potential energy. This refers to the total number of programs in the program schedule. For program number indexing, Let be any point in the virtual potential energy field. For the first The gravitational potential energy of the nodes corresponding to each program For the first The potential energy of the strip guides the tubular guiding potential energy corresponding to the groove. The number of grooves is determined by potential energy. It represents the temporal repulsive potential energy.

[0024] This invention creates an inertial channel that conforms to the broadcast process by excavating low-potential-energy trenches between logically adjacent gravitational centers of mass. Even during transitional intervals with silence, no instructions, or weak voice features, the firefly swarm can naturally slide from the current node to the next logical node due to the physical constraints of the terrain gradient, thus ensuring the continuity of the live broadcast process and effectively solving the problem of intent loss that easily occurs during instruction gaps in traditional methods.

[0025] The mathematical expression for nodal gravity is: ; in, For the first The gravitational potential energy of the nodes corresponding to each program The real-time gravity depth coefficient (the basic gravity depth coefficient during the broadcast preparation stage is a preset constant) is used. (In this embodiment, the value is 10). Let be a point in the virtual potential energy field. For the first The core coordinates of the program The effective radius of action in real time (initial value is 5.0). Index the program by number; The mathematical expression for the tubular guided potential energy is: ; in, For the first The tubular guiding potential energy corresponding to the guiding groove. For the numbering index of potential energy guiding trenches, The preset depth of the potential energy guiding trench is set to 50 in this embodiment. Vertical distance represents the distance between points. to line segment The shortest vertical distance, To guide the trench width based on potential energy, this embodiment uses a value of 2.0; The mathematical expression for the temporal repulsive potential energy is: ; in, For temporal repulsive potential energy, This represents a collection containing the indexes of all programs that have finished playing. The centroid coordinates of the already played programs, The repulsion strength coefficient (500 in this embodiment) needs to be set very high to ensure that fireflies do not approach programs that have finished playing.

[0026] In a specific embodiment of the present invention, in response to the dimensional mismatch problem between discrete static semantic features and continuous dynamic sound features in live broadcast scenarios, traditional keyword matching methods perform linear retrieval in the text space, which is difficult to handle the non-standardized expressions of the host. Simple semantic matching is usually isolated single-frame recognition, ignoring the temporal logic of the program broadcast, which makes it impossible to predict the next action during silence or without instructions, and is prone to interruption or false triggering. In order to achieve continuous tracking and fuzzy reasoning of the host's intention, the present invention introduces a spatial mapping mechanism based on a virtual potential energy field, which transforms the host's intention recognition from a traditional classification problem into virtual potential energy field spatial optimization. Specifically, by constructing a dynamic potential energy field, heterogeneous multimodal features are mapped to the terrain structure and energy distribution in the dynamic potential energy field respectively. The virtual potential field exists within a bounded two-dimensional Euclidean space, where each coordinate point... Each corresponds to a total potential energy value; The structured text entries in the program schedule represent the logical constraints of the live broadcast process. Since the text entries are discrete on the time axis, the baseline feature vectors are mapped onto a virtual dynamic potential energy field through t-SNE dimensionality reduction. In the dynamic potential energy field, the structured text entries are mapped to fixed gravitational centers of mass (coordinate points). The coordinate distribution of these centroids is determined by the semantic features of the text. Programs with similar semantics are closer in Euclidean distance in space. This constitutes the static topological structure of the potential energy field and establishes the candidate solution space for intent search. The specific process of dimensionality reduction mapping via t-SNE includes: The original coordinates of all programs are calculated using t-SNE. The original coordinates are then normalized using Min-Max with a safety margin. Specifically, the minimum x-coordinate value, minimum y-coordinate value, maximum x-coordinate value, and maximum y-coordinate value in the original coordinates are identified. For each original coordinate, the original coordinate is mapped to the virtual potential energy field using a mapping formula. The size of the virtual potential energy field is... The effective mapping range of the virtual field is set as follows: In this embodiment, it is set as follows: That is, to ensure that the original coordinates will not fall into the boundary of the virtual potential field after mapping; The mathematical expression for the mapping formula is: ; ; in, The mapped coordinates, The coordinates before mapping The smallest x-coordinate value in the original coordinate system. The minimum ordinate value. The largest x-coordinate value, The largest ordinate value, This represents the effective mapping interval of the virtual field. In this embodiment, during the initialization phase of the virtual potential energy field, the initial effective radius of action of all gravitational centers of mass is set to a preset constant (5.0 in this embodiment). Calculate the current center of mass according to the program order. With the next gravitational center of mass Euclidean distance Set connectivity threshold It is 20.0; like Then in and Potential energy guiding grooves are superimposed on the line connecting the two centroids, which are geometric centers of the two centroids.

[0027] The temporal repulsive potential energy of this invention constructs a dynamic repulsive field of the broadcast program. Once the program ends, a high-intensity repulsive field is generated, which avoids the risk of firefly backflow causing repeated playback of the broadcast program from a physical mechanism perspective. At the same time, the gravity and groove model based on Gaussian distribution ensures the smoothness of the potential energy surface, making the firefly's optimal path continuous and stable, and avoiding trajectory oscillations caused by abrupt changes in potential energy.

[0028] The process of extracting acoustic features includes: The host's audio stream is time-sliced, the Mel frequency cepstral coefficients of the audio stream in each time slice are extracted, and an acoustic feature matrix is ​​constructed. Each row of the acoustic feature matrix represents the Mel frequency cepstral coefficients corresponding to the time slice.

[0029] In a specific embodiment of the present invention, a sliding time window is used to capture digital audio streams. In this embodiment, the length of the sliding window is set to 200ms. The duration of a complete syllable and a short instruction (such as play) is between 100 and 300ms. In this embodiment, 200ms is preferred to capture acoustic features with semantic recognizability, while the latency is low to meet the real-time requirements of live streaming. The sliding step size is set to 50ms. In this embodiment, the process of constructing the acoustic feature matrix specifically includes: The host's audio stream is segmented according to the preset time window length and preset sliding step size to obtain independent short audio segments; For each short audio frame, 13 basic MFCC coefficients are extracted, and the first and second differences of the basic MFCC coefficients are further calculated, resulting in a total of 39 feature values, which are then used to construct an acoustic vector. The acoustic vectors of each short audio frame are concatenated as row vectors to generate a dimension of... The acoustic feature matrix.

[0030] This invention achieves complete capture of the dynamic characteristics of speech signals through a sliding time window and acoustic feature matrix construction strategy. It not only obtains the static timbre features of speech, but also retains the dynamic evolution information in the time dimension. Compared with the simple mean vector, the matrix-based feature representation can provide richer acoustic details.

[0031] The process of adjusting the virtual potential field specifically includes: Construct a semantic feature matrix corresponding to the program sequence list, where each column of the semantic feature matrix corresponds to the baseline feature vector of a program item; Principal component analysis is used to reduce the dimensionality of the semantic feature matrix, and a dimensionality-reduced semantic feature mapping matrix is ​​constructed. The acoustic matrix is ​​averaged in the time domain to obtain the acoustic semantic mapping vector at the current time. Construct an intent-based objective function based on sparse representation to solve for the sparse activation vectors of the acoustic semantic mapping vector on the semantic feature basis matrix; The mathematical expression for the objective function intended to be solved is: ; in, For acoustic semantic mapping vectors, For the dimension-reduced semantic basis matrix, Let be the sparse activation vector to be solved. This is the sparsity regularization parameter (in this embodiment, the value is 0.01~0.1). For the Euclidean norm, for Norm; The objective function is solved using an iterative optimization algorithm (coordinate descent algorithm) to obtain a sparse activation vector, and the elements in the sparse activation vector are used as the real-time semantic matching weights of the program. The normalized information entropy of the sparse activation vector is used as the semantic entropy. The semantic entropy is used to adjust the effective radius of the total potential energy value of the virtual potential energy field. It represents the ambiguity of the current voice command. The larger the semantic entropy, the more ambiguous the directionality of the current voice command. The calculation process of semantic entropy specifically includes: for each element in the sparse activation vector Take absolute value Then calculate the probability distribution. , The mathematical expression is: ; in, for The probability distribution, , For elements in the sparse activation vector, This represents the total number of elements in the sparse activation vector; Using the normalized probability distribution The standard entropy value is calculated as semantic entropy. The mathematical expression for semantic entropy is: ; in, For semantic entropy, for The probability distribution, This represents the total number of elements in the sparse activation vector; The mathematical expression for adjusting the effective radius of action is: ; in, For the real-time effective radius of action, The initial effective radius of action (5.0 in this embodiment) is used. Normalized semantic entropy; The real-time gravitational depth coefficient of the total potential energy value of the virtual potential energy field is adjusted by using a linear mapping formula and real-time semantic matching weights. The real-time gravitational depth coefficient determines the peak brightness of the centroid in the firefly search space. The mathematical expression for adjusting the real-time gravitational depth coefficient is: ; in, This is the real-time gravitational depth coefficient. The basic gravity depth coefficient (set to 10 in this embodiment). This is the acoustic gain coefficient (set to 1000 in this embodiment). Weights for real-time semantic matching; The total potential energy value at any point in the virtual potential energy field is recalculated based on the real-time gravity depth coefficient and the effective radius of action.

[0032] In one specific embodiment of the present invention, the baseline feature vector of each program is arranged as a column vector to obtain a semantic feature matrix; Principal component analysis was used to reduce the dimensionality of the semantic feature matrix and extract the initial features. One principal component (in this embodiment) The value is consistent with the acoustic feature dimension, which is 39), constructing the dimension-reduced semantic feature mapping matrix (dimension). ); for the acoustic feature matrix at the current moment (dimensions) The acoustic semantic mapping vector (dimension) is obtained by performing an arithmetic mean over the time dimension. That is, taking the arithmetic mean of each column of the acoustic feature matrix.

[0033] This invention solves the problem of aligning high-dimensional semantics with low-dimensional acoustic features through PCA dimensionality reduction and sparse coding techniques, and endows it with adaptive fuzzy adjustment capabilities. By solving the sparse activation vector and calculating the semantic entropy, it can perceive the fuzziness of the command in real time. When the semantic entropy is high, it automatically uses the linear expansion formula to expand the effective radius of the gravitational center of mass, thereby realizing fuzzy search.

[0034] The process of obtaining the centroid trajectory of a firefly population specifically includes: Initialize a firefly population of a preset size and randomly distribute them in a virtual potential energy field. Calculate the absolute brightness value of each firefly in the population. The absolute brightness value is negatively correlated with the total potential energy value of the location of each firefly in the population; the absolute brightness value is equal to the negative total potential energy value. Individuals in a firefly population with low absolute brightness values ​​update their positions to individuals within the population with higher absolute brightness values ​​within their perception range. The mathematical expression for this position update is: ; in, For the first A firefly Location at any given moment For fireflies The absolute brightness value within the perception range is better than fireflies Location, The maximum attraction level is 1.0 in this embodiment. The light intensity absorption coefficient is 1.0 in this embodiment. Fireflies and The Euclidean distance between them for Random numbers within the interval This is the step size factor (in this embodiment, the value is 0.2). For relative attractiveness, The attractor term determines the direction and step size of the firefly's position update. The term is a random perturbation to prevent getting trapped in local optima; In one specific embodiment of the present invention, a firefly population of a preset size is initialized, in this embodiment 30-50, and the firefly population is randomly distributed in a virtual potential energy field (size...). Within the range of ), that is, the x and y coordinates of individual firefly populations follow... The random distribution; Traversing firefly populations, for fireflies and fireflies The absolute brightness values ​​of the two and Comparison is performed only when the judgment condition is met. When established, fireflies are triggered. To the fireflies The location update mechanism.

[0035] This invention establishes a negative correlation between absolute brightness and potential energy, and introduces random perturbation terms and step size factors to ensure that firefly swarms maintain a certain level of exploration activity in trenches or flat areas, preventing them from getting trapped in local optima. This allows the entire population to automatically converge to the target node with the highest probability, like a water flow, driven by both gravitational sources (voice commands) and terrain (temporal logic). Its group statistical characteristics are more robust than the judgment of a single individual.

[0036] The process of executing material preloading instructions and generating playback prompt signals specifically includes: Real-time calculation of the spatial distribution centroid of firefly swarms, average swarm movement speed, and continuous dwell time; Define a preload radius and a playback trigger radius that are concentrically set around each gravitational center of mass; Real-time monitoring of the center of gravity and physical triggering signals of firefly swarms; If no physical trigger signal is detected, and the center of gravity of the firefly swarm spatial distribution enters the preload radius of the gravitational center of mass, an anti-accidental touch judgment is executed. When all anti-accidental touch judgment conditions are met, a material preload instruction is generated, and the underlying interface is asynchronously called to load media resources into the memory buffer. The criteria for preventing accidental touches include: The average movement speed of the group is lower than the preset convergence threshold (5.0 units / second in this embodiment, where the unit refers to the coordinate scale value in the virtual potential energy field coordinate system). The continuous dwell time of the center of gravity position within the preload radius exceeds the preset time threshold (500ms in this embodiment). If no physical trigger signal is detected, and the center of gravity of the firefly swarm enters the playback trigger radius of the gravitational center of mass, and the swarm convergence density exceeds a preset threshold (90% in this embodiment), a prompt signal is generated; the swarm convergence density is defined as the ratio of the number of individual fireflies within the playback trigger radius of the target gravitational center of mass to the total number of firefly populations. If a physical trigger signal is detected, and the centroid of the firefly swarm's spatial distribution does not enter the preload radius of any gravitational center of mass, the next program that has already been played is selected and played directly according to the timing of the program sequence. If a physical trigger signal is detected, and the centroid of the firefly swarm's spatial distribution enters the preload radius of the gravitational center of mass, the program corresponding to the current gravitational center of mass will be played directly.

[0037] In a specific embodiment of the present invention, the spatial centroid coordinates of the firefly swarm are calculated ( ), To the scale of fireflies, Let be the coordinates of firefly i at time t; The average movement speed of the population is characterized by monitoring the displacement of the center of gravity in adjacent time steps, at each time step. average group movement speed The calculation is as follows: ; in, The average movement speed of the group Let be the center coordinates at time t. for The center coordinates at time t. The sampling interval is 50ms (i.e., 20 frames are refreshed per second). Specifically, the position update formula is executed once every 50ms, and the current centroid coordinates of the group are recorded once. A dwell timer is preset for each gravitational center of mass; the current center of mass... to the center of mass The distance is For each sample, the following logical judgment is performed: Determine if it satisfies If the condition is met (center of gravity is within the circle), then update the timer: If the condition is not met (center of gravity outside the circle), then reset the timer: ; Each center of mass There are concentric double-threshold regions surrounding it, with the outer layer being the preload radius. (In this embodiment, it is 20), the inner layer is the trigger radius. (This example is 5); Before the program began, individual firefly individuals updated their locations. Fireflies gradually fell into The outer layer sends an asynchronous preload instruction to the background to load the corresponding media resources from the disk into the memory buffer; when Furthermore, when the firefly swarm density exceeds 90%, a playback prompt signal is generated. If a silent frame indicating the end of the host's voice or a clear physical trigger signal is detected at this time, playback will be executed immediately. The physical trigger signal is the push-out of the live streaming console fader, which detects the trigger level of the physical button and the opening of the foot pedal.

[0038] This invention employs a multi-dimensional constraint-based anti-accidental touch judgment logic. By setting convergence speed thresholds and dwell time thresholds, it can accurately distinguish between the transit and lock states of fireflies, eliminating false triggers caused by signal jitter. Simultaneously, the design of concentric dual-threshold regions enables hierarchical response. Resources are loaded asynchronously first to eliminate IO latency, and then the system waits for the final trigger signal. This mechanism ensures that media resources are ready the moment playback is confirmed, achieving a zero-latency broadcast experience.

[0039] The optimization module also includes a timing monitoring unit, which monitors the playback status of the program, where the playback status is "playback ended". If the current program has finished playing, then the program will be merged into a set containing all programs that have finished playing. The specific process of monitoring the playback status of a program includes: monitoring the digital playback progress and physical channel status of the current program; and determining that the current program has ended playback when any playback end condition is met. These playback end conditions include: The media playback kernel feedback file ends, including the media playback kernel feedback interruption signal and playback progress timestamp equal to the total duration of the media file; The physical fader stroke value controlling the program is lower than the preset mute threshold (5% of the total fader stroke in this embodiment).

[0040] This invention introduces a spatial mapping mechanism based on a virtual potential energy field, which breaks through the limitations of traditional keyword matching in linear retrieval in discrete text space. By mapping heterogeneous discrete program semantic features and continuous acoustic features into terrain structure and energy distribution in a potential energy field, it solves the dimensional mismatch problem between static text and dynamic speech, transforming the complex intent recognition task into a physical optimization process in continuous space. This enables continuous tracking and accurate positioning of the host's intent and enhances the adaptability to non-standardized expressions. This invention constructs an intent calculation objective function based on sparse representation and introduces a semantic entropy evaluation index, which can perceive the ambiguity of the host's voice instructions in real time. When the voice instructions are unclear or ambiguous, the effective radius of action of the gravitational centroid is automatically expanded by dynamically expanding the semantic entropy to capture potential intentions. This dynamic field modulation mechanism can not only understand clear instructions, but also understand vague colloquial expressions. While ensuring a high recognition rate, it significantly reduces the dependence on the host's standard speech. This invention achieves physical constraints on the live broadcast process logic by constructing potential energy-guided trenches and a temporal repulsive field, effectively ensuring the continuity of the broadcast control process. Utilizing a Gaussian distribution to construct low-potential energy channels connecting adjacent program nodes, even when the host is silent or has weak voice characteristics, the firefly swarm can be guided by the physical inertia of the terrain gradient, migrating naturally and smoothly to the next logical node. Combined with the repulsive potential energy of already played nodes, this ensures that the live broadcast process strictly follows the preset temporal logic, solving the interruption problem easily encountered in traditional single-frame recognition technology. This invention accurately distinguishes between long-distance search and local convergence states by real-time monitoring of the firefly swarm's spatial centroid, migration speed, and dwell time, effectively eliminating false triggers caused by signal jitter. Based on a concentric dual-threshold region design, it achieves hierarchical control of asynchronous material preloading and instant playback triggering, optimizing memory resource usage while ensuring rapid response.

[0041] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0042] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the protection scope of the present invention.

Claims

1. A speech recognition-based broadcasting and hosting assistance system, characterized in that, include: Acquisition module, feature extraction module, adjustment module, optimization module; The acquisition module is used to acquire multimodal data in live streaming scenarios; The feature extraction module is used to extract semantic features of the program sequence in multimodal data during the broadcast preparation stage, construct a virtual potential energy field, and map the semantic features to the gravitational center of mass in the virtual potential energy field. The adjustment module is used to extract the acoustic features of the host's audio stream from the multimodal data after the live broadcast starts, construct an objective function based on the acoustic features and semantic features, solve the objective function to obtain a sparse activation vector, and adjust the total potential value of the virtual potential field based on the sparse activation vector. The optimization module is used to map the total potential energy value of the virtual potential energy field to an absolute brightness value, combine the firefly algorithm to obtain the center-of-gravity trajectory of the firefly population, and execute material preloading instructions and generate a playback prompt signal.

2. The speech recognition-based broadcasting and hosting assistance system as described in claim 1, characterized in that, The multimodal data includes program playlists, host audio streams, and physical operation signals; The program playlist includes program entries arranged in chronological order. Each program entry is a structured text entry, and each entry contains a unique identifier, metadata tags, and media resource links. The host's audio stream includes digital audio signals acquired in real time by a microphone array; The physical operation signals include the travel value of the physical pusher, the on / off status of the physical button, and the status of the foot pedal.

3. The speech recognition-based broadcasting and hosting assistance system as described in claim 2, characterized in that, The process of extracting the semantic features includes: The pre-built synonym generation model is used to perform data augmentation on each structured text entry in the program sequence to generate a corresponding set of generalized descriptions. The structured text entries and the set of generalized descriptions are input into a pre-trained language encoder to obtain the baseline feature vector.

4. The speech recognition-based broadcasting and hosting assistance system as described in claim 3, characterized in that, The process of constructing the virtual potential energy field specifically includes: The reference feature vector corresponding to each structured text entry is mapped to a coordinate point, which serves as the gravitational center of mass of the virtual dynamic potential energy field. The effective radius of action of all gravitational centers of mass is the preset initial effective radius of action. Iterate through the coordinates of all programs on the program schedule and arrange all coordinates in the order of broadcast. Calculate the Euclidean distance between adjacent coordinate points. If the Euclidean distance between adjacent coordinate points is greater than the preset connectivity threshold, then construct a potential energy guiding trench connecting the adjacent coordinate points. The total potential energy at any point in the virtual potential energy field includes nodal gravitational potential energy, tubular guiding potential energy, and time-series repulsive potential energy. The mathematical expression for the total potential energy is: ; in, This represents the total potential energy. This refers to the total number of programs in the program schedule. For program number indexing, Let be any point in the virtual potential energy field. For the first The gravitational potential energy of the nodes corresponding to each program For the first The potential energy of the strip guides the tubular guiding potential energy corresponding to the groove. The number of grooves is determined by potential energy. It represents the temporal repulsive potential energy.

5. The speech recognition-based broadcasting and hosting assistance system as described in claim 4, characterized in that, The mathematical expression for the nodal gravity is: ; in, For the first The gravitational potential energy of the nodes corresponding to each program This is the real-time gravitational depth coefficient. Let be a point in the virtual potential energy field. For the first The core coordinates of the program The effective radius of action in real time (initial value is 5.0). Index the program by number; The mathematical expression for the tubular guided potential energy is: ; in, For the first The tubular guiding potential energy corresponding to the guiding groove. For the numbering index of potential energy guiding trenches, The preset depth of the trench is used to guide potential energy. Vertical distance represents the distance between points. to line segment The shortest vertical distance, The potential energy guides the width of the trench; The mathematical expression for the time-series repulsive potential energy is: ; in, For temporal repulsive potential energy, This indicates an index containing all programs that have finished playing. The set, The centroid coordinates of the already played programs, This is the repulsive force intensity coefficient.

6. The speech recognition-based broadcasting and hosting assistance system as described in claim 5, characterized in that, The process of extracting the acoustic features includes: The host's audio stream is time-sliced, the Mel frequency cepstral coefficients of the audio stream in each time slice are extracted, and an acoustic feature matrix is ​​constructed. Each row of the acoustic feature matrix represents the Mel frequency cepstral coefficients corresponding to the time slice.

7. The speech recognition-based broadcasting and hosting assistance system as described in claim 6, characterized in that, The process of adjusting the virtual potential field specifically includes: Construct a semantic feature matrix corresponding to the program sequence list, where each column of the semantic feature matrix corresponds to the baseline feature vector of a program item; Principal component analysis is used to reduce the dimensionality of the semantic feature matrix, and a dimensionality-reduced semantic feature mapping matrix is ​​constructed. The acoustic matrix is ​​averaged in the time domain to obtain the acoustic semantic mapping vector at the current time. Construct an intent-based objective function based on sparse representation to solve for the sparse activation vectors of the acoustic semantic mapping vector on the semantic feature basis matrix; The mathematical expression for the objective function to be solved is: ; in, Acoustic semantic mapping vector For the dimension-reduced semantic basis matrix, Let be the sparse activation vector to be solved. For sparse regularization parameters; The objective function is solved using an iterative optimization algorithm to obtain a sparse activation vector, and the elements in the sparse activation vector are used as the real-time semantic matching weights of the program. The normalized information entropy of the sparse activation vector is used as the semantic entropy, and the effective radius of action of the total potential energy value of the virtual potential energy field is adjusted by the semantic entropy. The mathematical expression for adjusting the effective radius of action is: ; in, For the real-time effective radius of action, The initial effective radius of action, Normalized semantic entropy; The real-time gravitational depth coefficient of the total potential energy value of the virtual potential energy field is adjusted using a linear mapping formula and real-time semantic matching weights. The mathematical expression for adjusting the real-time gravitational depth coefficient is as follows: ; in, This is the real-time gravitational depth coefficient. Based on the fundamental gravitational depth coefficient, The acoustic gain coefficient. Weights for real-time semantic matching; The total potential energy value at any point in the virtual potential energy field is recalculated based on the real-time gravity depth coefficient and the effective radius of action.

8. The speech recognition-based broadcasting and hosting assistance system as described in claim 7, characterized in that, The process of obtaining the centroid trajectory of the firefly population specifically includes: A firefly population of a preset size is initialized and randomly distributed in a virtual potential energy field. The absolute brightness value of each firefly in the population is calculated. The absolute brightness value is negatively correlated with the total potential energy value of the location of each firefly in the population. Individuals in a firefly population with low absolute brightness values ​​update their positions to those of individuals within the population with higher absolute brightness values ​​within their perception range. The mathematical expression for this position update is: ; in, For the first A firefly Location at any given moment In the fireflies The absolute brightness value within the perception range is better than firefly Location, To maximize attraction, The light intensity absorption coefficient is... Fireflies and The Euclidean distance between them for Random numbers within the interval This is the step size factor.

9. The speech recognition-based broadcasting and hosting assistance system as described in claim 8, characterized in that, The process of executing the material preloading instruction and generating the playback prompt signal specifically includes: Real-time calculation of the spatial distribution centroid of firefly swarms, average swarm movement speed, and continuous dwell time; Define a preload radius and a playback trigger radius that are concentrically set around each gravitational center of mass; Real-time monitoring of the center of gravity and physical triggering signals of firefly swarms; If no physical trigger signal is detected, and the center of gravity of the firefly swarm spatial distribution enters the preload radius of the gravitational center of mass, an anti-accidental touch judgment is executed. When all anti-accidental touch judgment conditions are met, a material preload instruction is generated, and the underlying interface is asynchronously called to load media resources into the memory buffer. The conditions for preventing accidental touches include: The average movement speed of the group is lower than the preset convergence threshold; The continuous dwell time of the center of gravity within the preload radius exceeds the preset time threshold. If no physical trigger signal is detected, and the center of gravity of the firefly swarm enters the playback trigger radius of the gravitational center of mass, and the swarm convergence density exceeds a preset threshold, a prompt signal will be generated. If a physical trigger signal is detected, and the centroid of the firefly swarm's spatial distribution does not enter the preload radius of any gravitational center of mass, the next program that has already been played is selected and played directly according to the timing of the program sequence. If a physical trigger signal is detected, and the centroid of the firefly swarm's spatial distribution enters the preload radius of the gravitational center of mass, the program corresponding to the current gravitational center of mass will be played directly.

10. The speech recognition-based broadcasting and hosting assistance system as described in claim 9, characterized in that, The optimization module also includes a timing monitoring unit for monitoring the playback status of the program, wherein the playback status is that playback has ended. If the current program has finished playing, then the program will be merged into a set containing all programs that have finished playing. The specific process of monitoring the playback status of a program includes: monitoring the digital playback progress and physical channel status of the current program; and determining that the current program has ended playback when any playback end condition is met. These playback end conditions include: The media playback kernel feedback file ends, including the media playback kernel feedback interruption signal and playback progress timestamp equal to the total duration of the media file; The physical fader travel value for controlling the program is lower than the preset mute threshold.

Citation Information

Patent Citations

  • Voice control method and device for online conference

    CN116939144A