Real-time translation method for intelligent conference content based on speech recognition

Through distributed microphone arrays and hierarchical separation networks, multi-speaker voice is decoupled, combined with voiceprint fingerprint and bidirectional semantic bridging model, the semantic reconstruction problem of traditional translation systems in multi-speaker scenarios is solved, and high-quality real-time translation is achieved.

CN120164479BActive Publication Date: 2025-07-22广东公信智能会议股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510646085.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-22
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Traditional conference translation systems are difficult to achieve accurate semantic reconstruction in the mixed state of multiple spokesperson speech, resulting in semantic drift, sentence breakage, term mismatch and other problems in translation output, and lack effective analysis of spokesperson identity and language style.

Method used

A distributed microphone array and mutation noise suppression algorithm are used to generate enhanced speech streams, and multi-speaker speech decoupling is performed through a hierarchical separation network, combining voiceprint fingerprint maps and identity tags, and cross-language semantic consistency vectors are generated using a bidirectional semantic bridging model, and translated through a language adaptive neural compiler.

Benefits of technology

Improves the semantic fidelity and context consistency of the translated content, ensuring that the translation results match the spokesperson style, and are suitable for complex language communication scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164479B_ABST
    Figure CN120164479B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of speech processing, and particularly to a real-time translation method for intelligent conference content based on speech recognition, comprising the following steps: collecting an original speech signal, generating an enhanced speech stream by using a mutation noise suppression algorithm, and synchronously extracting a voiceprint fingerprint map; inputting the enhanced speech stream into a hierarchical separation network, decoupling multi-speaker speech based on the voiceprint fingerprint map, outputting speech segments with identity tags and triggering an incremental update of a term knowledge base; generating a cross-language semantic consistency vector, and simultaneously constructing a dynamically updated context memory pool; and converting the semantic consistency vector into a target language stream. The present invention improves the spatial accuracy and semantic independence of speech decoupling, provides a structured input basis for subsequent semantic modeling and translation, and is particularly applicable to conference scenarios with frequent cross-talk and overlapping speech streams.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and particularly to a real-time translation method for intelligent conference content based on speech recognition. Background Art

[0002] With the continuous growth of the demand for multilingual communication, especially in application scenarios such as international business negotiations, cross-border remote conferences, and multi-institutional collaborative decision-making, the real-time translation processing of conference content is becoming an important research direction of intelligent speech systems. Traditional conference translation systems usually rely on single speech stream input and linear semantic parsing processes, lacking the effective decoupling ability for the mixed state of multi-speaker voices, and it is difficult to achieve accurate semantic reconstruction in complex conversation environments.

[0003] In the prior art, the speech recognition and machine translation modules are mostly independent of each other, lacking a collaborative optimization mechanism based on semantic consistency, resulting in problems such as semantic drift, sentence pattern breakage, and term mismatch in the translation output. In addition, due to ignoring the style differences and term expression preferences of speaker identities, the generated results often lack consistency and context coherence, and it is difficult to meet the high standards of language quality and role differentiation in formal conference scenarios. Summary of the Invention

[0004] The present invention provides a real-time translation method for intelligent conference content based on speech recognition, a real-time translation method for conference content that completes speaker recognition, semantic depth modeling, and language style adaptive control under dynamic interference of multiple speech sources, realizes end-to-end closed-loop optimization from speech perception to language generation, improves the semantic fidelity, context consistency, and style matching degree of the translated content, and meets the intelligent translation requirements in complex language communication scenarios.

[0005] The real-time translation method for intelligent conference content based on speech recognition includes the following steps:

[0006] S1: Collect the original speech signal through a distributed microphone array, generate an enhanced speech stream using a sudden noise suppression algorithm, and synchronously extract the voiceprint fingerprint map;

[0007] S2: Input the enhanced speech stream into a hierarchical separation network, decouple the multi-speaker voices based on the voiceprint fingerprint map, output the speech segments with identity labels, and trigger the incremental update of the term knowledge base;

[0008] S3: Process the speech segments through a session-aware bidirectional semantic bridging model to generate a cross-lingual semantic consistency vector, and simultaneously construct a dynamically updated context memory pool;

[0009] S4: Use a language-adaptive neural compiler to convert the semantic consistency vector into a target language stream.

[0010] Optionally, S1 specifically includes:

[0011] S11, using a microphone array and performing spatial filtering on the original speech signal to generate an initial speech stream with enhanced directivity;

[0012] S12, performing time-frequency domain abrupt noise detection on the initial speech stream. When pulse noise or transient interference is detected, switch to the compensation mode, perform phase reconstruction on the damaged frequency band based on the time-frequency mask matrix, and generate an enhanced speech stream.

[0013] Optionally, S1 further includes parallel extraction of biometric parameters of the enhanced speech stream. The biometric parameters include fundamental frequency trajectory, formant distribution, and dynamic speech rate change amount, and generate a dimensionally compressed voiceprint fingerprint map through a deep feature encoder. Among them, the encoder uses a one-dimensional convolutional neural network with a temporal attention mechanism.

[0014] Optionally, S2 specifically includes:

[0015] S21, input the enhanced speech stream into a three-level separation network. The first level uses a complex domain beamforming network for spatial separation to generate a set of candidate speech segments; the second level constructs a voiceprint fingerprint contrastive learning network, generates a voiceprint attention mask based on the voiceprint fingerprint map, and decouples the features of the candidate speech segments; the third level fuses spatio-temporal features through a gated residual network and outputs a clean speech segment with an identity label;

[0016] S22, perform real-time term scanning on the clean speech segment. When an out-of-vocabulary term is detected, trigger an incremental update of the term knowledge base and generate a term vector:

[0017] S23, establish a mapping relationship table between the identity label and the term knowledge base, and bind the updated term vector to the voiceprint fingerprint map of the corresponding speaker.

[0018] Optionally, the generation of the term vector in S22 specifically includes:

[0019] S221, use a double-threshold N-gram model to identify candidate terms, and the first threshold filters speech segments with a term confidence > 0.85;

[0020] S222, through a cross-lingual term alignment engine, perform semantic similarity matching between the candidate terms and the multilingual knowledge graph;

[0021] S223, generate a term vector for the successfully matched terms and update it to the term knowledge base in an online clustering manner.

[0022] Optionally, the three-level separation network specifically includes:

[0023] The first stage, the spatial separation layer: perform multi-channel complex spectrum processing on the original speech signal through the frequency-domain beamforming method, construct an optimal spatial filtering weight matrix based on the estimated noise covariance structure and the target sound source steering vector, use the optimal spatial filtering weight matrix to weight and superimpose the spectrum signals of each channel, obtain a spatially directional speech output result, and output a set of candidate speech segments with speaker spatial characteristics as the input for voiceprint feature decoupling processing;

[0024] The second stage, the feature decoupling layer: perform frame-level feature alignment based on the Mel spectrum representation of the candidate speech segment set and the pre-extracted voiceprint fingerprint map, generate an attention mask matrix representing the speaker matching degree by calculating the similarity in the time-frequency two-dimensional space, and the attention mask matrix is used to enhance the feature components consistent with the target voiceprint in the speech signal and suppress the remaining non-target sound source information to achieve voiceprint-level decoupling of the candidate speech segment set;

[0025] The third stage, the spatio-temporal fusion layer: input the voiceprint-decoupled speech features into a bidirectional temporal modeling network to capture the dynamic dependencies between contexts, and introduce a frame-level regulation factor through a gated residual mechanism to achieve weighted fusion of the original speech information and the modeled features. The gating weight is modulated by the formant features in the voiceprint map to achieve speaker-adaptive feature enhancement, and finally output a clean speech segment bound with an identity label to provide input for semantic modeling and translation processes.

[0026] Optionally, the S3 specifically includes:

[0027] S31, perform semantic encoding on the speech segment with an identity label, generate a language-independent semantic unit sequence through a cross-lingual adversarial alignment network, and each semantic unit of the semantic unit sequence includes a part-of-speech tag, a logical role, and a cross-lingual shared semantic fingerprint;

[0028] S32, construct a bidirectional semantic bridging model. The forward branch uses a gated graph convolutional network to capture the dialogue logic flow, and the backward branch corrects the semantic deviation through a counterfactual reasoning module to output a cross-lingual semantic consistency vector;

[0029] S33, dynamically maintain a context memory pool, adopting a hierarchical storage architecture, including:

[0030] Short-term memory layer: cache the semantic units and their identity labels of the current speaker's last 3 rounds of conversations, and dynamically adjust the memory weight through a decay factor;

[0031] Long-term memory layer: store the conference global theme vector and the cross-speaker semantic association matrix, and perform memory distillation compression every 5 minutes.

[0032] Optionally, the forward branch is based on a gated graph convolution mechanism. The gated graph convolution mechanism includes, for each pair of adjacent nodes in the semantic graph, extracting their respective representation vectors, performing a concatenation operation in the vector space to construct a joint representation indicating the relationship between the nodes, linearly transforming the joint representation using a weight matrix, and compressing it to a standard interval through an activation function to generate a gated weight indicating the edge connection strength. The gated weight is used to control the degree of information transmission between nodes, adjust the semantic propagation path in the graph convolution process, and obtain a semantically enhanced representation result at the logical level by performing gated weight calculations on nodes and their adjacency structures across the entire graph;

[0033] The backward branch is based on counterfactual reasoning bias correction. Specifically, counterfactual reasoning bias correction includes constructing a set of virtual intervention graph structures based on the semantic graph. For each semantic unit, without changing the context conditions, the current node is removed and the semantic graph is re-inferred and calculated. By comparing the prediction results generated by the original graph and the intervention graph at the output layer, the difference is statistically used as a bias index to measure the degree of influence of the current semantic unit on the overall logical flow. The bias index is fed back to the forward gated path to adjust the representation or connection weights of specific nodes, realizing the reverse suppression and correction of semantic deviation phenomena.

[0034] Optionally, S4 includes inputting the semantic consistency vector into the neural compiler to generate a target language stream, specifically including:

[0035] S41, performing syntax tree projection: mapping the semantic unit sequence to the surface syntax structure of the target syntax;

[0036] S42, performing lexical instantiation: selecting target lexical items that match the context in combination with the term knowledge base;

[0037] S43, applying style constraints: applying the corresponding language style template according to the speaker identity label;

[0038] S44, outputting the target language stream.

[0039] Advantages of the present invention:

[0040] In the present invention, by constructing a three-level voice separation network, combining complex domain beamforming, a voiceprint fingerprint-driven attention mask mechanism, and a gated residual structure, it is possible to accurately separate the voice segments of multiple speakers under multiple interference backgrounds and automatically bind identity labels. This structure improves the spatial accuracy and semantic independence of voice decoupling, provides a structured input basis for subsequent semantic modeling and translation, and is particularly suitable for conference scenarios with frequent cross-talking and overlapping speech flows.

[0041] The present invention introduces a semantic encoding and cross - language adversarial alignment mechanism to ensure that semantic units form a language - independent shared representation between the source language and the target language. Through a bidirectional semantic bridging model, it jointly models the context logical structure and semantic deviation correction, constructs semantic vectors with cross - language consistency, provides stable and accurate semantic input for the translation system, and solves problems such as insufficient context awareness and serious semantic deviation in traditional translation systems.

[0042] In the process of translation generation, the present invention instantiates semantic vocabulary based on a term knowledge base updated in real - time, and at the same time introduces a speaker style template into the decoding path of the neural compiler, ensuring accurate conveyance of terms while keeping the translated sentences consistent with the speaker's style in terms of language style, tone logic, etc. Brief Description of the Drawings

[0043] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only for the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0044] Figure 1 It is a schematic flow chart of the translation method according to an embodiment of the present invention;

[0045] Figure 2 It is a schematic diagram of a three - level separation network according to an embodiment of the present invention. Detailed Embodiments

[0046] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well - known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawing part is only for more specifically describing the embodiments and is not intended to specifically limit the present invention.

[0047] It should be pointed out that in the specification, when referring to "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc., it indicates that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0048] Generally, terms can be understood at least in part from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or property in the singular sense, or can be used to describe a combination of features, structures, or properties in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but rather can alternatively, depending at least in part on the context, allow for the existence of other factors that are not necessarily explicitly described.

[0049] As Figure 1 - Figure 2 shown, a real-time translation method for intelligent conference content based on speech recognition includes the following steps:

[0050] S1: Collect the original speech signal through a distributed microphone array, generate an enhanced speech stream using a sudden noise suppression algorithm, and synchronously extract the voiceprint fingerprint map;

[0051] S2: Input the enhanced speech stream into a hierarchical separation network, decouple the multi-speaker speech based on the voiceprint fingerprint map, output the speech segments with identity tags, and trigger the update of the incremental term knowledge base;

[0052] S3: Process the speech segments through a session-aware bidirectional semantic bridging model to generate a cross-lingual semantic consistency vector, and simultaneously construct a dynamically updated context memory pool;

[0053] S4: Use a language-adaptive neural compiler to convert the semantic consistency vector into a target language stream.

[0054] S1 specifically includes:

[0055] S11, Directional speech acquisition uses an 8-channel microphone array with a circular distribution. The original speech signal is spatially filtered through an adaptive beamforming algorithm to generate an initially enhanced speech stream with directionality. Specifically, the minimum variance distortionless response (MVDR) algorithm is used to dynamically track the spatial position of the speaker, with an interference suppression angular range better than ±60° and a direction estimation accuracy of ±5°.

[0056] S12, Sudden noise suppression and phase reconstruction: Perform time-frequency domain sudden noise detection on the initial speech stream. When pulse noise or transient interference is detected, automatically switch to the dynamic compensation mode. Based on the time-frequency mask matrix, perform phase reconstruction on the damaged frequency band to generate an enhanced speech stream. The sudden noise detection algorithm uses the time-frequency energy mutation coefficient (TFEC) and is calculated as follows:

[0057] ;

[0058] where represents the complex amplitude of the speech spectrum at time at frequency , is the time window length (i.e., the number of historical reference frames, which is 5 to 20 frames) and is used to construct the historical average energy reference. represents the mutation ratio of the current moment to the historical energy (energy mutation coefficient). represents frequency in the past frame average power spectral density. If is determined to be burst noise, the phase reconstruction uses the Griffin-Lim algorithm, and a voiceprint similarity constraint term is introduced to improve the naturalness and speaker consistency of the reconstructed speech. The reconstruction target of the basic Griffin-Lim algorithm is: , where is the reconstructed time-domain speech signal. represents the inverse short-time Fourier transform. is the time-frequency amplitude matrix of the enhanced speech. is the phase estimate value after the

[0059] Based on the basic Griffin-Lim algorithm, a reconstruction target with voiceprint constraints is also introduced:

[0060] ;

[0061] where represents the short-time Fourier transform. is the voiceprint feature regularization weight coefficient; represents the distance function between voiceprint embeddings (Euclidean distance can be used). represents the voiceprint encoder output feature of the reconstructed speech . represents the reference voiceprint feature corresponding to the original speech segment (obtained from S13). is the time-domain speech signal generated in the current iteration. is the original enhanced speech stream sample.

[0062] S13. The voiceprint fingerprint map extracts the biometric parameters of the enhanced speech stream in parallel, including the fundamental frequency trajectory, formant distribution, and speech rate change amount, and generates a dimensionally compressed voiceprint fingerprint map through a deep feature encoder. The encoder is a one-dimensional convolutional neural network (1D-CNN) with a temporal attention mechanism, and the structure is as follows:

[0063] Input feature fusion: The following three-dimensional biometric features are fused after frame alignment: fundamental frequency trajectory, formant sequence, and speech rate change rate.

[0064] Network structure design:

[0065] Temporal attention layer: Importance weighting is performed on the time dimension.

[0066] Depthwise separable convolutional layer: Extract cross-band local dependence features;

[0067] Feature distillation (dimension compression) module: Compress the initial 512-dimensional speech representation into a 128-dimensional voiceprint fingerprint map.

[0068] The finally output voiceprint fingerprint map will be used as an auxiliary input for the hierarchical speech separation network in the subsequent S2 step, and the encoding weights will be shared during training to achieve end-to-end optimization.

[0069] The specific scheme for extracting the voiceprint fingerprint map is as follows:

[0070] 1. Input feature fusion expression: The original input features include the fundamental frequency trajectory , formant sequence , and speech rate change rate . The fused frame-level input feature vector is:

[0071] , where represents the input feature of the th frame, is the th normalized formant frequency, is the normalized fundamental frequency value, is the normalized speech rate change rate, forming the complete input matrix : , where is the number of frames.

[0072] 2. Temporal attention mechanism expression: Introduce the attention vector to weight the frame-level features:

[0073] , where , and construct the weighted feature representation:

[0074] ;

[0075] where is the attention projection matrix, is the bias term, is the attention vector, represents the attention weight, is the intermediate hidden layer dimension, represents the weighted input frame feature.

[0076] 3. Depthwise separable convolutional layer expression: Use depthwise separable convolution, that is, perform channel-wise convolution and then use pointwise convolution to merge: ;

[0077] Then, use pointwise convolution to integrate features: ;

[0078] Among them, represents the input of the -th channel within the time window , represents the convolutional kernel of the -th channel, is the channel bias term, represents the number of channels (corresponding to 5D input), is the convolutional kernel size, represents the pointwise convolution fusion weight, is the output vector of the depth convolution for the -th frame.

[0079] 4. Expression of the feature distillation module: Connect all frame features into a two-dimensional feature matrix , and compress it into a 128-dimensional voiceprint fingerprint map using linear projection: ; Among them,

[0080] represents taking the average of the features of all frames, represents the fully connected weight, is the bias vector, represents the final voiceprint fingerprint map (128-dimensional).

[0081] S2 specifically includes:

[0082] S21, three-level voice separation network:

[0083] The first level: spatial separation layer (complex domain beamforming network): Adopt the frequency domain MVDR algorithm to perform spatial separation on the spectral signal, and calculate the optimal beamforming weight vector for each frequency point, expressed as:

[0084] ;

[0085] Among them, represents the beamforming weight vector at frequency , represents the noise covariance matrix, is the number of microphones, is the steering vector of the target sound source, represents the conjugate transpose of the steering vector, and several candidate speech segment sets are output after separation;

[0086] The second level: feature decoupling layer (voiceprint fingerprint attention mask): Construct a voiceprint fingerprint contrast learning module, and calculate the Mel spectrogram for each candidate segment Cosine similarity of the frame frequency points with the voiceprint fingerprint map to generate a voiceprint attention mask:

[0087] ;

[0088] Among them, represents the time frame , frequency point voiceprint attention mask value, that is, the similarity weight, represents the Mel spectrogram feature vector of the voice segment at the time-frequency point , represents the frequency feature of the target speaker's voiceprint map, and the mask is used to retain the voice frequency band of the target speaker and suppress the interference of other sound sources;

[0089] The third level: spatio-temporal fusion layer (gated residual network): Adopt a gated residual structure and bidirectional LSTM for spatio-temporal feature fusion, and output a pure voice segment , with a speaker identity label; The basic structure of the gated residual network: , among them, represents the th pure voice segment separated by the speaker, is the input candidate voice segment, represents the bidirectional LSTM processing module, represents the time frame gating coefficient at, that is, the frame-level regulation factor, dynamically generated by combining the formant information of the voiceprint fingerprint, represents the element-wise product.

[0090] S22, Term Scanning and Knowledge Base Incremental Update:

[0091] S221, Term Recognition (N-gram Model Double Threshold Judgment): Use a sliding window to construct term candidates , calculate the confidence: , among them, is the N-gram order;

[0092] Judgment criteria:

[0093] If , it is considered a high-confidence term;

[0094] If , enter the secondary verification, is the first threshold, , is the second threshold, ;

[0095] S222, Term Alignment (Cross - language Atlas Matching): Match the term candidates with the term nodes in the multilingual knowledge graph for semantic similarity: , where is the semantic vector of the term. If , is the similarity threshold with a value of 0.82, then the match is determined to be successful.

[0096] S223, Dynamic Term Vector Clustering and Update: The terms with successful matches are encoded as term vectors , and an online clustering strategy is adopted to add them to the term knowledge base: , where is the clustering center vector, is the update rate; if the distances from all existing term clustering centers are greater than the threshold, a new term class is added.

[0097] S23, Binding of Term and Identity Mapping:

[0098] Establish a mapping table: ; Bind the term vector to the voiceprint fingerprint map of the speaker for subsequent semantic modeling and translation context use.

[0099] S3 specifically includes:

[0100] S31, Semantic Quantization Encoding: Perform semantic encoding on the voice segments with identity labels, and use a cross - language adversarial alignment network to generate language - independent semantic unit representations. Each semantic unit contains part - of - speech, logical role, and cross - language shared semantic fingerprints. The loss function of the adversarial alignment network is defined as:

[0101] ;

[0102] where is the semantic unit of the source language, is the semantic unit of the target language, is the semantic encoder (generating language - independent representations), represents the language discriminator (predicting which language the encoding result belongs to), is the adversarial training loss function. Perform multi - granularity hashing on the semantic units, including character - level, word - level, and sentence - level, and map the hashing results to the unit hypersphere to generate normalized vectors: ;

[0103] S32, Bidirectional Semantic Bridging Model: Construct a two - branch structure composed of a forward gated graph convolutional network and a backward counterfactual reasoning module for modeling the logical consistency and semantic deviation correction between semantic units.

[0104] Forward branch: Based on the gated graph convolutional mechanism, where the gated weight is calculated as: ; where represents the initial representation of node and in the semantic graph, is the gated weight matrix, represents the vector concatenation operation, represents the Sigmoid activation function, represents the gating strength that controls the node information transmission.

[0105] Backward branch: The counterfactual reasoning bias correction logic bias correction amount is defined as:

[0106] , where represents the predicted output of node in the original semantic graph, represents the counterfactual prediction after removing some semantic units, represents the number of semantic units within the current context, is an index used to judge the semantic logic stability and is used to backpropagate and correct the gating mechanism or node representation.

[0107] S33, Dynamic Context Memory Pool: Adopts a hierarchical architecture to manage short-term and long-term semantic information, supporting dynamic update, compression, and retrieval.

[0108] Memory weight decay in the short-term memory layer: ; where is the memory weight at time frame , is the decay rate constant, and the measured optimal value is 0.3 / second, is the current time, is the memory unit write time.

[0109] Long-term memory distillation compression:

[0110] Adopts the knowledge distillation method to process the global theme and correlation matrix. The steps are as follows: Project the memory matrix onto a subspace with rank (default ; Retain all vectors whose projection residuals in the low-rank subspace are greater than the specified threshold (default .

[0111] S4 specifically includes:

[0112] S41, Syntax tree projection: The input semantic unit sequence is represented as , Each semantic unit contains part of speech, logical role, and semantic fingerprint. During the projection stage of the syntax tree, it is mapped to the corresponding surface syntax structure in the target language. . The mapping function is expressed in the form of:

[0113] ; where represents the input sequence of semantic units, represents the syntax rule graph of the target language, represents the syntax tree construction function, represents the generated surface syntactic structure of the target language.

[0114] S42, Lexical instantiation: Based on the generated syntax structure, perform the operation of filling in lexical items, and select the most matching target language entries from the term knowledge base according to the term vector in the current context.

[0115] The target function for lexical selection is expressed as: ; where

[0116] is the target vocabulary corresponding to the -th semantic unit, is the set of target language vocabulary, represents the -th semantic fingerprint vector of the semantic unit, is the vector representation of the entry in the term database, represents the semantic weight in the current context (which can be determined by the speaker's identity or context);

[0117] S43, Application of style constraints: According to the identity label , apply personalized language style control during the generation process. The style template is embedded in the generation decoder in the form of a conditional vector , affecting word order, tone, conjunction, and sentence pattern selection strategies.

[0118] The neural compiler finally generates a target language fragment:

[0119] , where is the decoder module of the neural compiler, is the finally generated target language stream text, is the identity-based language style vector.

[0120] The present invention covers any alternatives, modifications, equivalent methods and solutions made to the essence and scope of the present invention. In order to enable the public to thoroughly understand the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0121] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A real-time translation method for intelligent conference content based on speech recognition, characterized in that The following steps are involved: S1: Collect the original speech signal through the distributed microphone array, generate the enhanced speech stream by using the mutation noise suppression algorithm, and extract the voiceprint fingerprint simultaneously; S2: Input the enhanced speech stream into the hierarchical separation network, decouple the multi-speaker speech based on the voiceprint fingerprint, output the speech fragments with identity tags and trigger the incremental terminology knowledge base update, including: S21, input the enhanced speech stream into a three-stage separation network. The first stage uses a complex domain beamforming network for spatial separation to generate a set of candidate speech segments; the second stage constructs a voiceprint fingerprint comparative learning network, generates a voiceprint attention mask based on the voiceprint fingerprint map, and performs feature decoupling on the candidate speech segments; the third stage fuses spatiotemporal features through a gated residual network to output clean speech segments with identity tags; S22 performs real-time term scanning on the clean speech clips. When unregistered terms are detected, the incremental update of the term knowledge base is triggered and a term vector is generated: S23, establishing a mapping relationship table between identity tags and terminology knowledge base, and binding the updated term vector to the voiceprint fingerprint of the corresponding speaker; S3: Processing the speech segment through a conversation-aware bidirectional semantic bridging model to generate a cross-language semantic consistency vector and construct a dynamically updated context memory pool; S4: Use a language-adaptive neural compiler to convert the semantic consistency vector into the target language stream.

2. The real-time translation method for intelligent conference content based on speech recognition according to claim 1, wherein, The S1 specifically includes: S11, uses a microphone array and performs spatial filtering on the original speech signal to generate an initial speech stream with enhanced directionality; S12, performing time-frequency domain mutation noise detection on the initial voice stream, when pulse noise or transient interference is detected, switching to compensation mode, performing phase reconstruction on the damaged frequency band based on the time-frequency mask matrix, and generating an enhanced voice stream.

3. The real-time translation method of intelligent conference content based on speech recognition according to claim 2, wherein The S1 also includes parallel extraction of biometric parameters of the enhanced speech stream, the biometric parameters including fundamental frequency trajectory, resonance peak distribution and dynamic speech rate change, and generates a dimensionally compressed voiceprint fingerprint map through a deep feature encoder, wherein the encoder adopts a one-dimensional convolutional neural network with a temporal attention mechanism.

4. The real-time translation method for intelligent conference content based on speech recognition according to claim 1, characterized in that, The generating term vector in S22 specifically includes: S221, using a dual-threshold N-gram model to identify candidate terms, the first threshold screens speech segments with term confidence greater than 0.85; S222, through the cross-language term alignment engine, semantic similarity matching is performed between candidate terms and multilingual knowledge graphs; S223, generating term vectors for successfully matched terms, and updating them to the term knowledge base in an online clustering manner.

5. The real-time translation method for intelligent conference content based on speech recognition according to claim 4, characterized in that The three-stage separation network specifically includes: The first stage, the spatial separation layer: the original speech signal is processed with multi-channel complex spectrum by the frequency domain beamforming method, and the optimal spatial filter weight matrix is constructed based on the estimated noise covariance structure and the target sound source steering vector. The optimal spatial filter weight matrix is used to perform weighted superposition on the spectrum signals of each channel to obtain a spatially directional speech output result, and a set of multiple candidate speech segments with speaker spatial characteristics are output as the input of the voiceprint feature decoupling process; In the second stage, the feature decoupling layer: Based on the Mel spectrogram representation of the candidate speech segment set and the pre-extracted voiceprint fingerprint map, frame-level feature alignment is performed. By calculating the similarity in the time-frequency two-dimensional space, an attention mask matrix representing the speaker matching degree is generated. The attention mask matrix is used to enhance the feature components in the speech signal that are consistent with the target voiceprint and suppress the remaining non-target sound source information, realizing the voiceprint-level decoupling of the candidate speech segment set; In the third stage, the spatio-temporal fusion layer: The speech features after voiceprint decoupling are input into a bidirectional temporal modeling network to capture the dynamic dependencies between contexts, and a frame-level regulation factor is introduced through a gated residual mechanism to achieve weighted fusion of the original speech information and the modeled features. The gating weights are modulated by the formant features in the voiceprint map to achieve speaker-adaptive feature enhancement. Finally, a pure speech segment with an associated identity label is output, providing input for the semantic modeling and translation processes.

6. The real-time translation method for intelligent conference content based on speech recognition according to claim 1, characterized in that The specific steps of S3 include: S31, Semantically encode the speech segments with identity labels, and generate a language-independent semantic unit sequence through a cross-lingual adversarial alignment network. Each semantic unit in the semantic unit sequence includes a part-of-speech tag, a logical role, and a cross-lingual shared semantic fingerprint; S32, Construct a bidirectional semantic bridging model. The forward branch uses a gated graph convolutional network to capture the dialogue logic flow, and the backward branch corrects the semantic deviation through a counterfactual reasoning module, and outputs a cross-lingual semantic consistency vector; S33, Dynamically maintain a context memory pool, adopting a hierarchical storage architecture, including: Short-term memory layer: Cache the semantic units and their identity labels of the current speaker's last 3 rounds of conversations, and dynamically adjust the memory weights through a decay factor; Long-term memory layer: Store the global conference theme vector and the cross-speaker semantic association matrix, and perform memory distillation compression every 5 minutes.

7. The real-time translation method of intelligent conference content based on speech recognition according to claim 6, wherein, The forward branch is based on the gated graph convolution mechanism. The gated graph convolution mechanism includes, for each pair of adjacent nodes in the semantic graph, extracting their respective representation vectors and performing a concatenation operation in the vector space to construct a joint representation representing the relationship between the nodes. A weight matrix is used to perform a linear transformation on the joint representation, and it is compressed to the standard interval through an activation function to generate a gating weight representing the edge connection strength. The gating weight is used to control the degree of information transmission between nodes, adjust the semantic propagation path in the graph convolution process, and obtain a semantically enhanced representation result at the logical level by performing gating weight calculations on the nodes and their adjacent structures within the entire graph; The backward branch is based on counterfactual reasoning deviation correction. Counterfactual reasoning deviation correction specifically includes, based on the semantic graph, constructing a set of virtual intervention graph structures. For each semantic unit, by removing the current node and re-performing semantic graph reasoning calculations without changing the context conditions, comparing the prediction results generated by the original graph and the intervention graph at the output layer, and statistically calculating the difference as a deviation index to measure the impact of the current semantic unit on the overall logic flow. The deviation index is fed back to the forward gating path to adjust the representation or connection weights of specific nodes, realizing the reverse suppression and correction of semantic deviation phenomena.

8. The real-time translation method for intelligent conference content based on speech recognition according to claim 1, characterized in that S4 includes inputting the semantic consistency vector into the neural compiler to generate a target language stream, specifically including: S41, performing syntax tree projection: mapping the semantic unit sequence to the surface syntax structure of the target grammar; S42, performing lexical instantiation: selecting target terms that match the context in combination with the term knowledge base; S43, applying style constraints: imposing corresponding language style templates according to the speaker identity label; S44, outputting the target language stream.

Citation Information

Patent Citations

  • Attribute word mining method and related product

    CN116484079A

  • Multi-person long speech semantic recognition and abstract generation method, system and device and medium

    CN118262705A