A voice processing method and a terminal device
By using a semantic coding model on terminal devices to identify and enhance target voice content in multi-person voice interaction scenarios, the problem of users having difficulty perceiving key information is solved, improving the flexibility of voice interaction and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LENOVO (BEIJING) LTD
- Filing Date
- 2026-05-28
- Publication Date
- 2026-07-24
Smart Images

Figure CN122454973A_ABST
Abstract
Description
Technical Field
[0001] This application mainly relates to the field of artificial intelligence technology, and more specifically to a speech processing method and terminal device. Background Technology
[0002] In multi-player voice interaction scenarios, such as multiplayer online games, players often communicate tactically through real-time voice. However, the mixing of multiple voices, environmental noise, and limitations of terminal device acquisition conditions make it difficult for users to perceive key voice information such as tactical instructions and status reminders in a timely manner, thereby affecting the game's collaborative effect and reducing the user experience. Summary of the Invention
[0003] In view of the above problems, this application provides the following solution:
[0004] The first aspect of this application provides a speech processing method, including:
[0005] Extract speech segments from a speech stream;
[0006] When the first speech segment in the speech segment has a first language expression content and corresponds to a first semantic, if the first semantic matches the target semantic, speech enhancement processing is performed on the first speech segment;
[0007] When a second speech segment in the speech segment has a second language expression content and corresponds to a second semantic, if the second semantic matches the target semantic, the same speech enhancement processing is performed on the second speech segment; wherein the first language expression content and the second language expression content are not completely the same.
[0008] In some implementations, the semantic matching with the target semantic includes:
[0009] Obtain the similarity between the semantic representation of the speech segment and at least one target semantic representation in the reference semantic set to determine whether the target speech segment meets the target semantic triggering condition;
[0010] If the target semantic triggering condition is met, speech enhancement processing is performed on the target speech segment to enhance the speech content corresponding to the target semantic representation.
[0011] In some implementations, the semantic representation is a semantic vector obtained by semantically encoding the corresponding speech segment using a semantic coding model, and the semantic vector is mapped in a unified semantic vector space;
[0012] The target semantic representation is used to define the speech content to be enhanced in the semantic vector space;
[0013] The semantic encoding model is a lightweight model deployed on terminal devices, and its parameters can be optimized based on model update data from the cloud.
[0014] In some implementations, the target semantic representation is generated based on user-defined speech samples and / or text samples processed by the semantic coding model;
[0015] Each of the target semantic representations is associated with at least one pre-configured speech enhancement strategy to perform speech enhancement processing on the target speech segment based on the speech enhancement strategy; the speech enhancement strategy includes at least one of the following:
[0016] Directional gain adjustment, dynamic range compression, specific frequency band enhancement, and noise suppression.
[0017] In some implementations, semantic vectors are obtained by semantically encoding the corresponding speech segments using a semantic coding model, including:
[0018] Feature representations are obtained by extracting features from the input using a semantic coding model; the input is the speech segment or text content converted from the speech segment.
[0019] The feature representations are semantically aggregated to generate aggregated semantic representations;
[0020] The aggregated semantic representation is encoded to output a semantic vector of a preset dimension.
[0021] In some implementations, prior to semantic encoding, the following is also included:
[0022] The speech segments are semantically summarized by a summary generation model to generate corresponding semantic summaries, which are then encoded by the semantic encoding model to obtain the semantic vector.
[0023] The number of parameters in the summary generation model is greater than that in the semantic coding model.
[0024] In some implementations, the target semantic triggering condition includes any of the following:
[0025] If the similarity of any speech segment exceeds the first similarity threshold;
[0026] The similarity of multiple consecutive speech segments exceeds the second similarity threshold;
[0027] If the smoothed value obtained after smoothing the similarity of multiple consecutive speech segments exceeds the third similarity threshold in N consecutive segments, then the N consecutive speech segments are identified as the target speech segment; where N≥2.
[0028] In some implementations, the voice stream is a real-time voice stream received by the terminal device during application operation, or an audio file obtained by recording the application operation process;
[0029] When the application is an online interactive application for multiple users, the voice stream contains mixed voices from multiple users;
[0030] The method further includes:
[0031] The enhanced speech segment is mixed with at least one audio signal and then output.
[0032] A second aspect of this application provides a terminal device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program in the memory to implement the voice processing method described in any of the preceding claims. Attached Figure Description
[0033] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0034] Figure 1 This is a schematic flowchart of the speech processing method proposed in Embodiment 1 of this application;
[0035] Figure 2 This is a schematic flowchart of the speech processing method proposed in Embodiment 2 of this application;
[0036] Figure 3 This is a schematic flowchart of the speech processing method proposed in Embodiment 3 of this application;
[0037] Figure 4 This is a schematic flowchart of the speech processing method proposed in Embodiment 4 of this application;
[0038] Figure 5 This is a schematic flowchart of the speech processing method proposed in Embodiment 5 of this application;
[0039] Figure 6 This is a flowchart illustrating the application of the speech processing method proposed in this application to an audio detection scenario;
[0040] Figure 7 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application;
[0041] Figure 8 This is a schematic diagram of the hardware structure of a terminal device proposed in an embodiment of this application. Detailed Implementation
[0042] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments and is not intended to limit the application. The embodiments of this application are described below with reference to the accompanying drawings. It will be understood by those skilled in the art that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0043] The terms “first,” “second,” etc., used throughout this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0044] It is understood that before using the technical solutions disclosed in the embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained. For example, in response to receiving a user's active request, a pop-up window may be used, and a textual prompt message may be presented in the pop-up window to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. The user can choose whether to provide personal information to the electronic device, application, server, or storage medium or other software or hardware that performs the operation of the technical solution of this application based on the prompt message. This application does not limit the prompt message and the method of user authorization implementation. The data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) shall comply with the requirements of relevant laws and regulations and related provisions.
[0045] Regarding the background technology description, in scenarios such as online meetings, intelligent voice assistants, and remote teaching, speech enhancement technologies such as noise reduction and echo cancellation are proposed to enhance the acquired speech stream and improve the overall sound quality. However, this method cannot distinguish the importance or semantic type of the speech content, making it difficult to accurately capture and enhance the target speech content, often resulting in key information being submerged in background noise or other irrelevant speech. To address this, speech recognition technology is proposed to convert speech into text and then perform semantic analysis using a large model. However, this increases processing latency and loses rich information such as emotion and speech patterns in the speech, leading to higher transmission latency and privacy risks, making it difficult to meet the low latency and data localization requirements of real-time interaction on the terminal side. Furthermore, matching through a pre-set vocabulary or rules of keywords / instructions is proposed, but this method lacks generalization ability and is difficult to adapt to the speech habits of different applications and users, resulting in limited generalization. To improve these problems, embodiments of this application provide a speech processing method and terminal device. The speech processing method of this application embodiment will be described in detail below with reference to the accompanying drawings.
[0046] Reference Figure 1 This is a flowchart illustrating the voice processing method proposed in Embodiment 1 of this application. This voice processing method can be applied to terminal devices such as smartphones, laptops, smart speakers, or conference terminals. Figure 1 As shown, the speech processing method proposed in this embodiment may include:
[0047] Step S11: Obtain speech segments from the speech stream.
[0048] In this application embodiment, the voice stream can be a continuous audio signal (audio file) obtained in real time by an audio acquisition device (such as a microphone) during the application's operation, or a real-time voice stream received by the terminal device during the application's operation (such as a voice stream acquired and sent by the peer of the interactive application), or an audio file read from a storage device, etc. This application does not limit the source and acquisition method of the voice stream.
[0049] After acquiring or identifying the audio stream, it can be divided into continuous audio segments according to a time window (e.g., a frame of preset duration 20ms-40ms) or based on voice activity detection. Each audio segment can include one or more phonemes, words, or phrases. For example, in an audio / video conferencing scenario, the terminal device can segment the received audio stream into the following audio segments: Audio segment A (speaker 1) is "I'm here to report on this week's project progress"; Audio segment B (speaker 2) is "The weather is nice today"; Audio segment C (speaker 3) is "The progress is currently 80% complete." In a game scenario, when a player says "Help me," the terminal device segments the audio stream into multiple audio segments, including the audio segment that fully expresses "Help me."
[0050] Step S12: When the first speech segment in the speech segment has a first language expression content and corresponds to a first semantic, if the first semantic matches the target semantic, perform speech enhancement processing on the first speech segment.
[0051] Step S13: When the second speech segment in the speech segment has a second language expression content and corresponds to a second semantic, if the second semantic matches the target semantic, perform speech enhancement processing on the second speech segment; the first language expression content and the second language expression content are not completely the same.
[0052] In this embodiment, one or more target semantics can be predefined. These can represent specific intentions, sound patterns, keywords, or sound features that require speech enhancement, either by default or user-defined. For example, if a user wants the system to enhance a speech expressing a "cry for help," the system can determine whether to perform speech enhancement on the current speech segment by judging whether the semantics corresponding to that segment are indeed a cry for help. Thus, regardless of differences in the linguistic content between different speech segments, if the semantics match the target semantics, speech enhancement will be applied to that speech segment. This application does not limit the content of the target semantics or its configuration method.
[0053] Therefore, this application breaks through the limitations of traditional speech enhancement that relies solely on acoustic features, introducing a semantic-level matching mechanism. Even if the first and second speech segments are not completely identical in terms of linguistic content (e.g., the speech segments "Please turn on the air conditioner" and "Raise the room temperature"), as long as their semantics match the same target semantic (e.g., adjusting the temperature), that is, different speech segments corresponding to different linguistic expressions within the same target semantic category, can be simultaneously identified and trigger speech enhancement to increase gain and suppress background noise. Thus, this application achieves normalized enhancement of diverse linguistic expressions under the same semantic meaning, significantly improving the coverage and intelligence of speech enhancement.
[0054] In one embodiment, based on the above analysis, for each acquired speech segment, speech recognition technology can be used to obtain the linguistic content of the speech segment. Semantic parsing is then used to obtain the corresponding semantics of the speech segment. It is determined whether the semantics match the target semantics. If they match, regardless of whether the linguistic content is consistent with the linguistic content of the target semantics, speech enhancement processing will be performed on the speech segment. This application does not limit the method of obtaining the linguistic content and semantics of each speech segment, or the method of determining whether the semantics of each speech segment matches the target semantics.
[0055] For example, taking the target semantics containing "project progress report" as an example, the first audio segment is audio segment A mentioned above, and the second audio segment is audio segment C. These two audio segments have completely different wording, but their corresponding semantics both match the target semantics. In this case, both audio segments are included in the enhancement scope. However, the semantics of audio segment B (casual conversation about the weather) do not match the target semantics, and no audio enhancement processing is performed on it. It can be seen that, compared to traditional keyword-based enhancement schemes that can only capture audio containing specific keywords such as "report" and "progress," this application, through semantic matching, allows for the capture of audio segments containing specific keywords such as "I'm here to report on this week's project progress" and "The progress is currently eighty percent complete," even though the wording is completely different. Figure 1 Expressions that correspond to the target semantics (each of which matches the target semantics) were all identified, significantly expanding the effective speech capture range.
[0056] If the target semantics in the game scenario include a cry for help (such as "help me"), the first voice segment obtained is "help me" and the second voice segment is "come help me quickly". Although the language content of these two voice segments is not exactly the same in words, their respective semantics match the target semantics of the cry for help. Both will trigger voice enhancement processing, can trigger the same voice enhancement strategy, and achieve the same voice enhancement effect. This solves the limitation of traditional methods that require users to say "help me" to trigger voice enhancement.
[0057] Therefore, the speech processing method proposed in this application, through semantic-level normalization enhancement, eliminates the dependence on fixed keywords, eliminating the need for users to memorize specific wake words. Furthermore, it ensures semantic matching with the target speech, enhancing the corresponding speech segment regardless of changes in the language content. This avoids missed detections of target speech due to wording differences, improves the completeness and reliability of speech enhancement in complex dialogue scenarios, and achieves intelligent processing of "different language content, semantic matching, and unified enhancement," reducing user memory costs and increasing recognition success rate (especially in mixed speech scenarios).
[0058] In practical applications, since the same target semantics can be expressed differently at different times and in different contexts, this application can continuously recognize and enhance speech segments that match different expressions of the target semantics, i.e., it supports multi-turn and variable expressions, and is particularly suitable for conversational interaction scenarios, such as any long-duration speech scenario such as in-vehicle, home, or headphones. Based on this, the linguistic expression content of each speech segment involved in this application can be a description of at least one feature at the lexical and grammatical level, such as words, phrases, sentence structures, and word choice. Therefore, the difference between the first speech segment and the second speech segment is at least a difference at the lexical and grammatical level, and may also include differences in at least one prosodic feature such as tone, speech rate, volume, and pauses, but it is not only a difference in prosodic features.
[0059] Furthermore, if the semantics of a speech segment do not match the target semantics, it indicates that the speech segment is an irrelevant speech segment (such as the casual conversation, noise, or background voices in the example above). It can maintain its original acoustic state (i.e., not process) or execute other preset processing strategies (which are different from the speech enhancement strategies executed on the first / second speech segments mentioned above). This avoids the problem of synchronous amplification of background noise caused by traditional global speech enhancement, improves noise resistance, and makes the output speech maintain a natural and comfortable listening experience while highlighting key content (its corresponding target semantics). It also improves the intelligibility of the enhanced speech, improves the speech recognition accuracy, reduces computing power consumption, reduces system latency, and improves the smoothness of interaction.
[0060] Therefore, the speech processing method proposed in this application achieves semantic-driven selective speech enhancement through a mechanism that allows different language expressions to be semantically consistent and matched with the target semantics. This not only accurately identifies the user's true intent but also avoids useless enhancement of irrelevant speech segments, significantly improving the flexibility, robustness, and noise resistance of the speech interaction system. It is especially suitable for real, variable, and non-standardized natural language scenarios, enhancing the naturalness of the interaction and the user experience.
[0061] Reference Figure 2 This is a flowchart illustrating the speech processing method proposed in Embodiment 2 of this application. This embodiment describes a possible refined implementation method of how to determine the semantic match between any speech segment and the target semantic in Embodiment 1 above, such as... Figure 2 As shown, this refined implementation method may include:
[0062] Step S21: Obtain the semantic representation of speech segments in the speech stream;
[0063] Step S22: Obtain the similarity between the semantic representation of the speech segment and at least one target semantic representation in the reference semantic set;
[0064] In this embodiment, the abstract semantic matching process can be transformed into a quantifiable semantic similarity calculation and condition-triggered judgment process. Therefore, for each acquired speech segment, its speech content can be processed into a semantic representation carrying semantic information (which can also be called a feature representation) for subsequent semantic similarity analysis.
[0065] In addition, the reference semantic set in this application may contain at least one target semantic or target semantic representation (which has the same form as the semantic representation of a speech segment), each target semantic / target semantic representation representing a speech content that a user expects to enhance, such as intent, a specific sound pattern, keywords or specific sound features, etc., without limitation.
[0066] In one possible implementation, the semantic representation (target speech representation and semantic representation corresponding to speech segments) mentioned above may include label representation. In this case, this application can use a speech recognition model or a classification model to perform semantic understanding on the speech segments to determine semantic labels containing semantic information such as intent and slots of the speech segments, so as to constitute the semantic representation of the speech segments. In this way, the similarity calculation in step S22 is transformed into label matching, which greatly reduces the computational overhead and is especially suitable for resource-constrained end devices.
[0067] Optionally, the semantic representation of this application can also be expressed using a probability distribution (i.e., the probability value of the input speech belonging to each semantic category), such as {play music: 0.92, pause music: 0.05, adjust volume: 0.03}. In this case, this application can directly use the probability values of each semantic label as similarity or triggering conditions, such as determining whether the semantic label corresponding to the maximum probability value is the target semantic label (target semantic label), or whether the semantic label corresponding to the probability value greater than the probability threshold in the probability distribution contains the target semantic label. Without additional similarity calculation, the speech segment requiring speech enhancement can be quickly identified. Preferably, if the maximum probability value in the probability distribution is between a critical value (e.g., 0.5), a secondary confirmation mechanism can be triggered, such as using other methods to determine whether the semantics of the current speech segment matches the target semantics, thereby improving system fault tolerance.
[0068] In another possible implementation, the semantic representation described above can also be a semantic graph structure (graph representation), that is, using a graph (nodes and edges) to represent the relationships between entities. For example, the semantic representation (graph representation) of the voice segment "Give Zhang San a phone call" can be: [Action: dial] → [Contact person: Zhang San] → [Device: mobile phone]. This facilitates structured reasoning, not only matching semantics but also reasoning logical relationships to distinguish between the voice segments "call Zhang San with a mobile phone" and "call Zhang San with a landline," and improves anti-interference capabilities. Even if a word is missing (such as not saying "phone"), as long as the core relationship of the semantic graph structure exists, it can still be determined as a match.
[0069] In another possible implementation, the semantic representation can also be semantic hash encoding, that is, the semantic information of the speech segment is mapped to binary encoding through a hash algorithm, such as the hash representation of a 64-bit or 128-bit 0 / 1 sequence (e.g., 01101001...), or vector representation (e.g., [0.12, 0.95, 0.34, 0.78] etc.). This transforms the similarity calculation process in step S22 into an XOR operation process, which greatly improves the calculation speed, is suitable for large-scale real-time retrieval, greatly improves performance, and since the hash value is irreversible, it is impossible to directly reverse the original text, which helps protect privacy information and improves the security of speech processing.
[0070] Furthermore, the aforementioned semantic representation can also be a multimodal joint representation, which combines the acoustic features of speech (such as timbre, intonation, and other prosodic features) with text semantics. For example, a semantic embedding model or coding model can be used to convert speech segments or their text into high-dimensional feature vectors, thereby distinguishing who is speaking in the speech segment and the emotional state of the speaker, achieving more precise personalized enhancement, such as specific channel enhancement for Group A speech. Therefore, the semantic representation of this application can include at least one of the following: feature vectors, semantic label sets, semantic graph construction, hash coding, probability distribution, and multimodal joint representation. This application does not limit the specific form of the semantic representation, as long as it can represent the semantic information of the speech segment.
[0071] Subsequently, for different forms of semantic representation, appropriate similarity calculation methods can be adopted, such as cosine similarity, Euclidean distance, or vector dot product, to calculate the similarity between the semantic representation of the current speech segment and each target semantic representation. This application does not restrict the similarity calculation method in step S22.
[0072] Step S23: Based on the similarity, determine whether the target speech segment meets the target semantic triggering condition. If yes, proceed to step S24; otherwise, do not perform speech enhancement processing.
[0073] Step S24: Perform speech enhancement processing on the target speech segment to enhance the speech content corresponding to the target semantic representation.
[0074] Based on the above analysis, the target speech segment in this embodiment may include the first speech segment and / or the second speech segment in the speech stream, and may also include other speech segments in the speech stream, depending on the situation. This application may also denote speech segments that meet the target semantic triggering conditions as target speech segments, and speech segments that do not meet the target semantic triggering conditions as non-target speech segments.
[0075] The target semantic triggering condition can include a similarity greater than a corresponding similarity threshold. Thus, by using the similarity between the semantic representation and any target semantic representation, it is possible to objectively and stably determine whether a corresponding speech segment needs enhancement processing. Speech segments with a similarity greater than the similarity threshold are identified as target speech segments that meet the target semantic triggering condition, i.e., target speech segments requiring enhancement. Speech enhancement processing can then be performed on these target speech segments to enhance (i.e., highlight) the speech content corresponding to the target semantic representation, making it clearer and more distinguishable in noisy backgrounds or among multiple speakers. For example, for a speech segment matching the target semantic "cry for help," its volume can be increased and its spatial localization effect enhanced, thereby avoiding the missed detection problems of synonyms, near-synonyms, and colloquial variations in keyword matching, and improving the accuracy and robustness of the speech recognition to be enhanced.
[0076] Therefore, this application embodiment achieves quantitative and adjustable matching of the semantics corresponding to each speech segment in the speech stream with the target semantics by explicitly calculating the similarity between semantic representations. Compared with the traditional keyword / speech content (text) matching method, this application's similarity matching allows for fuzzy semantics and near-synonyms. That is, different speech segments in the speech stream have different linguistic expressions, but their semantic representations are similar (similarity greater than the similarity threshold) and all match the target semantic representation (similarity greater than the similarity threshold). These different speech segments will all undergo speech enhancement processing, which significantly improves the recall rate of the speech segments to be enhanced. At the same time, the accuracy can be controlled by setting the threshold to avoid false triggering, that is, to avoid performing speech enhancement processing on speech segments that do not meet the semantic triggering conditions, causing interference and unnecessary resource consumption.
[0077] Reference Figure 3 This is a flowchart illustrating the speech processing method proposed in Embodiment 3 of this application. This embodiment uses the semantic representation of a semantic vector as an example to explain the refined implementation process of matching the semantics of a speech segment with the target semantics. Figure 3 As shown, the implementation process may include:
[0078] Step S31: Semantically encode the speech segment using a semantic coding model to obtain the corresponding semantic vector;
[0079] In the embodiments of the present application, the semantic encoding model can be a deep learning model, such as a pre-trained speech-semantic joint embedding model. This model can map speech segments of any length or the text content converted therefrom into a semantic vector of a fixed dimension (predetermined dimension), and ensure that the distances between speech segments with similar semantics in the unified semantic vector space are relatively close. For example, mapping "救救我" and "help me" into two semantic vectors in the semantic vector space, and their similarity is very high (such as greater than 0.9).
[0080] Based on this, in the model training stage, taking the game scenario as an example (but not limited to this scenario, it can also be other voice interaction scenarios, such as video conferencing scenarios, voice recording scenarios in complex environments, etc.), as Figure 4 shown, game voice data from different game types and different player voice styles can be collected in the cloud, a training data set can be constructed, and a speech-semantic model with long-context modeling capabilities can be used for training to learn the similarity representation of speech at the semantic level, that is, to learn the mapping relationship between speech segments and semantic representations, and the trained semantic encoding model can be deployed to the terminal device to map the speech segments of the speech stream into the unified semantic vector space to obtain the corresponding semantic vectors. The present application does not elaborate on the training implementation process of the semantic encoding model.
[0081] It should be understood that for terminal devices with limited computing power resources and memory resources, such as mobile phones, tablets, etc., the present application can perform lightweight processing on the semantic encoding model trained in the cloud, such as using at least one of methods such as knowledge distillation, model pruning (structured pruning or unstructured pruning), low-rank decomposition / matrix decomposition, replacing with depthwise separable convolution, parameter quantization (quantization-aware training or post-training quantization), weight sharing or early exit mechanism, etc., to compress the large cloud model into a lightweight model and directly deploy it as the semantic encoding model to the terminal device, so that the lightweight model used in the execution process of the method of the present application can significantly reduce the inference latency and memory occupancy while maintaining the speech enhancement accuracy and meet the real-time processing requirements. Optionally, the parameters of the lightweight model can be optimized based on the model update data from the cloud, that is, the lightweight model can periodically receive update data (such as new parameters or model files) from the cloud to continuously improve the encoding accuracy or adapt to new semantic categories.
[0082] Therefore, the end-to-end real-time semantic coding method proposed in this embodiment avoids the latency and privacy risks associated with uploading voice data to the cloud. The design of a unified semantic vector space makes semantic matching simple and efficient. The lightweight model, coupled with a cloud update mechanism, enables the terminal to operate offline or in weak network environments and to receive capability upgrades when connected to the network, balancing privacy, low latency, and high performance to better meet the speech recognition needs of real-time voice interaction scenarios.
[0083] Step S32: Obtain the similarity between each semantic vector and at least one target semantic vector in the reference semantic set;
[0084] The method for obtaining the target semantic vector contained in the reference semantic set is the same as the method for obtaining the semantic vector of the speech segment described above, and will not be repeated here. Therefore, the target semantic vector and the semantic vectors corresponding to each speech segment in the speech stream are both mapped to a unified semantic vector space. The geometric relationships between the semantic vectors in this semantic vector space (such as the distance and angle between vectors) reflect the similarity of these semantic vectors in semantics and instruction expression patterns (the functional role and behavioral paradigm of speech in interaction). Therefore, after obtaining the semantic vector of the speech segment, this application can measure semantic similarity through vector similarity calculation.
[0085] Each target semantic representation (target semantic vector) is used to define the speech content to be enhanced in the semantic vector space. In other words, it represents the anchor point of the target semantic that needs to be enhanced in the interactive scenario. In this semantic vector space, the hypersphere region centered on the anchor point and with a certain distance as the radius (or angle or other geometric relationship) defines the speech content to be enhanced. Any speech segments that fall into this region (regardless of whether their wording, speech style, etc. are the same) will be enhanced.
[0086] In this embodiment, the instruction expression pattern refers to the functional intent type and corresponding expression paradigm of a speech segment in an interactive scenario, that is, what the speech segment is doing and how it requests / triggers an action. It goes beyond simple lexical semantics (the literal meaning of the language content) and at least covers the pragmatic function of speech in a specific interactive context. Among them, the pragmatic function level can be the instruction role of the speech to reflect the type of intent that the speech segment focuses on. In the semantic vector space, the semantic encoding model not only clusters speech with similar literal semantics, but also maps speech with the same instruction role to adjacent regions and matches them with the target semantics (anchor point) of that region.
[0087] For example, in a scenario reflecting a distress call pattern, whether a user says "Help me," "Call an ambulance," or "Is anyone there? I've fallen," the semantics all point to needing emergency assistance, and the pragmatic function is to send a distress signal. In other words, these speech fragments share the expression pattern of "distress call," matching the pre-configured target semantic vector reflecting this pattern. Similarly, in a scenario reflecting a control command pattern, whether a user says "Turn on the air conditioner," "Turn the temperature down," or "It's a bit hot," the semantics are all related to temperature adjustment, and the pragmatic function is to issue a control command to the system. In other words, these speech fragments share the expression pattern of "control command," matching the pre-configured target semantic vector reflecting this pattern.
[0088] Based on the above description of the speech content to be enhanced, since different command patterns are often accompanied by stable acoustic behavioral features (prosodic features), the edge semantic coding model can directly perform semantic coding on speech segments to avoid losing other information besides the linguistic expression content during text transcription. Therefore, these prosodic features are also implicitly captured during semantic coding. In other words, the above command expression patterns can also cover acoustic behavioral features. For example, the distress command expression pattern can cover prosodic features such as rapid speech rate, pitch rise, sudden energy increase, and short sentence structure; the control command expression pattern can cover prosodic features such as steady speech rate, pitch drop, verb at the beginning of the sentence, and clear pauses; and the response / confirmation command expression pattern can cover prosodic features such as short and low voice, affirmative word stress (such as "yes" and "okay"), and rapid ending. Thus, even if two speech segments in a speech stream have significantly different linguistic content (such as text content), but exhibit similar acoustic behavioral paradigms—for example, both being urgent cries for help rather than calm statements—the semantic vectors of these two speech segments remain relatively close in the semantic vector space. This indicates that the semantic vectors of the two speech segments match, and also match the semantic vector of the target seeking help. Therefore, this approach demonstrates the accuracy and reliability of semantic coding models in identifying intent through direct speech encoding rather than text transcription, thus improving the accuracy of identifying the speech content to be enhanced.
[0089] Furthermore, the aforementioned command expression patterns can also reflect sentence structure tendencies, that is, the structural preferences of language organization in speech segments, such as direct imperatives beginning with the verb infinitive (e.g., "turn off the lights"); indirect requests with declarative sentences and needs hints (e.g., "it's too bright here"); and situational triggers describing the current state and calling for a response (e.g., "I can't find the exit"). During semantic encoding, by encoding this structured expression tendency, "turn off the lights" (direct control) and "it's too bright, it's blinding" (indirect request) can be recognized as the same command expression pattern (lighting control), both matching the target semantic vector reflecting lighting control. The corresponding speech segments are then enhanced, thus achieving robust target semantic matching on the device side.
[0090] In summary, the edge-side semantic encoding model semantically encodes speech segments, mapping them to a unified semantic vector space to obtain corresponding semantic vectors. This process encodes both the semantic dimension (thematic relevance of the language content, such as medical emergency or environmental adjustment) and the command expression pattern dimension (the functional type of the voice in the interaction, such as calling for help or control). This allows for the suggestion of corresponding anchor vectors (i.e., target semantic vectors) in the semantic vector space or its different regions. These target semantic vectors not only represent the semantics of the language content but also a command expression pattern. Therefore, during the recognition of speech segments to be enhanced from a speech stream, for different speech segments in the stream that are semantically similar but have different wording, or even different accents / speeds / tones (such as the first and second speech segments), as long as their semantic vectors are close to the anchor vectors (target semantic vectors) at the pattern level, they can be uniformly recognized and enhancement triggered. This makes the system particularly suitable for emergency scenarios (where users may shout or use non-standard vocabulary) and colloquial interactions (where the same thing can be expressed in multiple natural ways), significantly outperforming traditional text keyword matching schemes.
[0091] Step S33: Based on the similarity, determine that the target speech segment meets the target semantic triggering condition, and perform speech enhancement processing on the corresponding target speech segment according to the speech enhancement strategy corresponding to the target speech segment.
[0092] Based on the above analysis, this application represents the user's intent in the form of vectors, allowing different speech segments with different wording and speech styles but similar semantics to be uniformly identified in the subsequent target semantic vector matching process. Therefore, by directly analyzing the semantic similarity between the semantic vector of the speech segment and each target semantic vector, it is determined that in the unified semantic vector space, the speech segment corresponding to each semantic vector clustered in the region with a certain target semantic vector as the anchor point is identified as the target speech segment that needs speech enhancement. That is to say, these semantic vectors located in this region have a high similarity with the target semantic vector, such as the similarity being greater than the corresponding similarity threshold, and the corresponding speech segment is determined to meet the target semantic triggering condition, but is not limited to this.
[0093] In this approach, because various linguistic expressions can be mapped to the same instruction expression pattern, the common features of the deep instruction expression model are preserved during semantic encoding. This allows different linguistic expressions to cluster in the semantic vector space due to their similar patterns, and be identified as similar to the same target semantic vector. Therefore, the method proposed in this embodiment for determining the matching of the semantics of different speech segments with the target semantics identifies the shared instruction-level interaction intent behind the speech segments (i.e., the consistent instruction expression patterns reflected by each segment) rather than the consistency of the surface text (linguistic expression content). This improves the accuracy of speech recognition to be enhanced and solves the technical problems of traditional preset word lists or rule matching methods, as well as traditional global enhancement methods.
[0094] In practical applications of this application, for each target speech segment that meets the target semantic triggering conditions, speech enhancement processing can be performed directly according to a pre-configured speech enhancement strategy to enhance the speech content corresponding to the target semantic vector. The pre-configured speech enhancement strategy here can be a vendor-preset speech enhancement strategy, a user-defined speech enhancement strategy, or a speech enhancement strategy directly developed for third parties by a scenario-based SDK (Software Development Kit), etc. This application does not impose any restrictions on this.
[0095] Optionally, this application may also pre-configure a variety of speech enhancement strategies, such as at least one of directional gain adjustment, dynamic range compression, specific frequency band enhancement, and noise suppression. After determining at least one target speech segment that meets the target semantic triggering conditions, a speech enhancement strategy for each target speech segment can be selected from them according to a preset strategy selection method to achieve personalized speech enhancement for different target speech segments. This application does not impose any restrictions on this.
[0096] Among them, the Directional Gain Adjustment strategy can be based on a microphone array (or multi-channel headphones) to adjust the enhancement of each channel according to the direction of the sound source. It applies directional gain to the direction of the sound source of the target speech segment to enhance the speech in that direction and suppress the speech in other directions, thereby improving the perceived loudness (signal-to-noise ratio) of the target speech segment. It is especially suitable for in-vehicle, conference and headphone scenarios.
[0097] Dynamic Range Compression (DRC) is a technique that compresses the dynamic range of audio signals by avoiding excessive amplification of high sound pressure levels and appropriately boosting low sound pressure levels. This prevents popping sounds and improves the listening comfort of target speech segments, making it particularly suitable for media playback and broadcasting. Thus, when performing enhancement processing on a target speech segment, the compression ratio and threshold can be adjusted according to the speech amplitude of the target speech segment, resulting in a fuller, more resonant speech.
[0098] Band-specific enhancement is an adjustment strategy that enhances specific frequencies in the frequency domain, such as 300Hz-3400Hz for human voices and 20Hz-200Hz for bass frequencies in music. This can enhance the frequency bands related to speech intelligibility in the target speech segment, highlight the useful information in the target speech segment, suppress irrelevant frequency bands, and further improve its intelligibility.
[0099] Noise suppression strategies can estimate the noise spectrum of a speech segment and then use methods such as spectral subtraction or Wiener filtering to remove noise components from the mixed signal, reducing background interference (such as background chatter, wind noise, or keyboard noise) and improving speech recognition accuracy. Therefore, choosing this strategy to perform speech enhancement processing on a target speech segment can attenuate environmental noise, thereby improving the clarity of effective audio (such as the speaker's voice) and better meeting the speech processing needs of downstream tasks.
[0100] In summary, this application's embodiments, through a unified semantic vector space, directly encode speech segments with different linguistic expressions into semantic vectors using a semantic coding model. Furthermore, these semantic vectors can be directly compared within the same metric space (semantic vector space) (vector similarity comparison), achieving semantic-level normalization and enabling rapid and accurate identification of target speech segments requiring enhancement. Subsequently, differentiated speech enhancement processing can be performed on different types of target speech segments, thereby optimizing the auditory experience while maintaining speech intelligibility. Additionally, this application deploys the semantic coding model in a lightweight form on the terminal, satisfying the low-latency requirements of real-time processing while continuously optimizing parameters through model update data delivered from the cloud. This allows the terminal to operate offline or in weak network environments and receive capability upgrades when connected to the network, balancing real-time performance, privacy, and model evolution needs.
[0101] In some embodiments, this application can select corresponding speech enhancement strategies for different target speech segments that meet the target speech triggering conditions according to a target semantic-driven strategy selection method. These strategies may include, but are not limited to, at least one of directional gain adjustment, dynamic range compression, specific frequency band enhancement, and noise suppression. Therefore, this application can configure at least one speech enhancement strategy associated with each target semantic representation (such as a target semantic vector), i.e., binding differentiated speech enhancement strategies to different target semantics to achieve personalized speech enhancement. In this way, after determining the target semantic vector matching the semantic vector of the current speech segment, at least one of its associated pre-configured speech enhancement strategies can be selected to perform speech enhancement processing on the corresponding target speech segment based on the selected speech enhancement strategy, improving the flexibility and scene adaptability of speech enhancement.
[0102] For example, taking a multi-person remote conferencing application scenario on a terminal device as an example, the pre-configured reference semantic set includes two target semantics, "emergency instructions" and "background notifications" and / or their corresponding target semantic representations. The former target semantics are associated with two speech enhancement strategies, directional gain adjustment and specific frequency band enhancement, while the latter target semantics are associated with two speech enhancement strategies, noise suppression and light dynamic range compression. Speech segment A in the audio stream is speaker A urgently saying "Stop! There's a problem!" Semantic similarity analysis determines that the similarity (e.g., 0.91) between the semantic representation of speech segment A and the target semantic representation 1 corresponding to "emergency instruction" exceeds a similarity threshold, thus meeting the target semantic trigger condition. Therefore, a combination of speech enhancement strategies associated with the target semantic representation 1 (or any one of them) can be invoked to perform speech enhancement processing on speech segment A. For example, based on a directional gain adjustment strategy, the microphone array can be used to lock onto speaker A's location (sound source direction), suppressing sound sources from other directions and increasing the gain of speech segment A by 12dB. Furthermore, based on a specific frequency band enhancement strategy, the 1kHz-4kHz frequency band in speech segment A is enhanced, making the consonant "Stop!" more prominent and ensuring that remote participants can clearly hear the emergency instruction. Thus, in the hybrid output enhancement process, speech segment A significantly highlights the emergency instruction.
[0103] If voice segment B is the system notification tone "Five minutes until the end of the meeting," after semantic similarity analysis, if the similarity (e.g., 0.85) between the semantic representation of voice segment B and the target semantic representation 2 corresponding to "background notification" exceeds the similarity threshold, then the target semantic trigger condition is met. The voice enhancement strategy combination associated with the target semantic representation 2 (or any one of these strategies) can be invoked to perform voice enhancement processing on voice segment B. For example, based on a noise suppression strategy, the low-frequency hum of the air conditioner and the keyboard typing sound in voice segment B can be suppressed to reduce the background noise of the notification voice; and based on a light dynamic range compression strategy, the level fluctuation of the notification voice in voice segment B can be compressed within ±2dB to avoid volume jumps interrupting the meeting rhythm. This ensures that the notification voice in the enhanced voice segment B is clear but not abrupt, allowing participants to naturally perceive the time reminder without needing to deliberately increase the volume.
[0104] If speech segment C is speaker B saying "I think this solution is feasible", after semantic similarity analysis, it is determined that the similarity (e.g., 0.3 / 0.4) between the semantic vector of speech segment C and the target semantic representation 1 and the target semantic representation 2 is less than the similarity threshold. Therefore, speech segment C does not meet the target semantic triggering condition, and no speech enhancement strategy is triggered. Speech segment C is output in its original acoustic state.
[0105] Therefore, this embodiment achieves fine-grained semantic-level control over enhancement granularity by binding dedicated strategies (differentiated speech enhancement strategies) to different target semantics and selectively enhancing each speech segment contained in the speech stream. This solves the problem that global enhancement schemes cannot distinguish the importance or semantic type of speech content. Moreover, this application improves the targeting and auditory naturalness of enhancement processing by matching the combination of speech enhancement strategies with inherent semantic needs.
[0106] Preferably, the reference semantic set and each pre-configured voice enhancement strategy can be stored locally on the terminal device. Semantic matching and strategy invocation do not require cloud interaction. The latency from voice input to enhancement output is controlled within a short time (e.g., 50ms), which meets the low latency requirements of real-time conferencing scenarios and avoids uploading sensitive meeting content to the cloud.
[0107] In some embodiments, the strategy selection method may also include an on-demand resource allocation method. For example, for target speech segments reflecting simple instructions / tasks, a lightweight speech enhancement strategy, such as noise suppression, can be selected; for target speech segments reflecting complex instructions / tasks, a combined speech increment strategy can be selected, i.e., a combination of at least two speech enhancement strategies listed above. Optionally, the computing power and power consumption required for the execution of each speech enhancement strategy can be predetermined (e.g., determined experimentally or empirically). Based on this, and combined with the currently available resources of the device, at least one suitable speech enhancement strategy can be selected to achieve differentiated speech enhancement for different types of target speech segments. In summary, the embodiments of this application provide users with a high degree of personalization and customization capabilities. Different users and different scenarios can configure different speech enhancement strategies, greatly improving the applicability of the system and the user experience.
[0108] In some embodiments, such as Figure 4 As shown, in order to meet users' personalized configuration needs for target semantics, this application can use user-defined speech samples and / or text samples to perform semantic encoding through a semantic encoding model, and record the resulting semantic representation as the target semantic representation (such as the target semantic vector). That is, each target semantic representation (such as the target semantic vector) in the reference semantic set of this application can be generated by processing user-defined speech samples and / or text samples through a semantic encoding model, and each target semantic representation can be associated with at least one speech enhancement strategy.
[0109] Based on this, users can record one or more sample voice clips (with similar semantics) through the application interface as voice samples, such as voice segments corresponding to the target semantics of distress calls like "Help me," "Help," and "Help!" These clips are then input into a semantic encoding model to obtain corresponding semantic representations, which can be directly used as the target semantic representation. Alternatively, the semantic representations of these multiple voice samples can be averaged or clustered to obtain one or a set of target semantic representations. In this way, other target semantic representations can also be configured to form a reference semantic set. For example, in a multiplayer online tactical competitive game scenario, the environment is noisy (gunshots, explosions, footsteps), and players need to communicate tactics with teammates via voice. Users pre-configure three sets of target semantics / target semantic representations (such as target semantic vectors) to automatically capture and enhance key tactical instructions in noisy game audio, ensuring teammates can hear them clearly.
[0110] Based on this, users can customize personalized tactical commands based on voice samples. Players record several personalized tactical commands into their headset microphones, such as: "All assemble, prepare to engage" (low, fast, pre-battle mobilization tone) - voice sample 1; "Follow me" (short, whisper-like, stealth scenario) - voice sample 2; "Assemble at point A" (standard speaking speed, map marking coordination) - voice sample 3; "Assemble, assemble, don't be alone" (urgent, crisis scenario) - voice sample 4, etc. The semantic encoding model performs semantic encoding on the four voice samples respectively, generating corresponding semantic representations. These representations can be used directly as the target semantic representation, or the four semantic representations can be clustered, with the cluster center used as the target semantic representation. This captures the different linguistic expressions of players regarding assembly commands under various emotional states, using clustering to determine the target semantic meaning of assembly. The target semantic representation can then be stored in the headset's local reference semantic set instead of being uploaded to the game server or cloud, preventing third parties from obtaining team-specific tactical terminology and individual voice biometrics, thus ensuring competitive privacy and security. Each target semantic representation is associated with at least one speech enhancement strategy, such as directional gain adjustment (increasing the direction gain of human voices by 10dB and suppressing the direction of environmental gunshots) and specific frequency band enhancement (increasing the intelligibility of consonants in the 2kHz-4kHz range, allowing plosive sounds such as "gather" and "follow me" to penetrate the low-frequency noise of explosions).
[0111] Therefore, the first voice segment acquired during the game is "Come and assemble!" shouted by Player 1 in actual combat, and the second voice segment is "Come to me!" shouted by Player 2. Even though these two voice segments differ from the linguistic content of the aforementioned voice samples, their semantic representations satisfy the target semantic triggering condition if their similarity to the target semantic representation of "assembly" in the reference semantic set is satisfied. Both will trigger voice enhancement for their respective voice segments. If a third voice segment, such as a teammate casually saying "This game's equipment is good," is also acquired at this time, according to this semantic similarity analysis, its similarity to each target semantic representation in the reference semantic set does not satisfy the target semantic triggering condition, and voice enhancement will not be performed on the third voice segment. Thus, for tactical command-type target semantics, by semantically encoding their voice samples, the generated target voice vectors are highly adapted to the individual expression habits of players, avoiding the recognition bias of general models for personalized accents, intonations, and emotional expressions. This ensures that multiple voice segments with different linguistic content reflecting the target semantic of "assembly" are uniformly identified and enhanced, improving the accuracy and reliability of the voice recognition to be enhanced.
[0112] Optionally, this application can also directly input text samples (such as cries for help) in the application interface to convert them into semantic representations through a text semantic encoding model, which can then be used as a target semantic representation to form a reference semantic set. This application does not limit the configuration implementation method of each target semantic representation contained in the reference semantic set. Taking the game scenario described above as an example, players can input text entries in the game settings interface to construct a reference semantic set for the target semantic "cry for help". They can also construct reference semantic sets for other target semantics, and even form a large reference semantic set. For example, inputting text samples such as "need support", "save me", "I'm being targeted", "medic", and "SOS", after semantic encoding, obtains the corresponding semantic representations. Then, after averaging, the target semantic representation corresponding to "cry for help" is obtained. The associated speech enhancement strategies can include dynamic range compression (compression ratio 3:1, flattening the fluctuations in the level of the cry for help speech to avoid being masked by gunfire) and noise suppression (suppressing low-frequency noise below 500Hz by 12dB).
[0113] Thus, in actual gameplay, multiple voice segments from the acquired game voice stream, such as the first voice segment being "Come help me quickly," the second being "I'm about to fall," and the third being "An enemy has fallen," etc., according to the semantic similarity analysis described above, the semantics of the first and second voice segments both match the target semantic of "call for help," and both will undergo voice enhancement processing. However, the semantics of the third voice segment belong to battle reports rather than a call for help, meaning it does not meet the triggering condition of the target semantic of a call for help, and therefore no voice enhancement processing will be performed on the third voice segment. It is evident that this application rapidly constructs a standardized tactical terminology library (reference semantic set) through text samples, eliminating the need for players to record multiple rounds of voice, making it particularly suitable for rapid configuration before the start of the game. The semantic breadth provided by the text samples allows the system to cover diverse variations of the call for help phrasing used by players in actual gameplay. It should be noted that the configuration and implementation process for the reference semantic set of other voice content requiring enhancement processing in the game, such as status reminders, is similar and will not be detailed in detail here.
[0114] It should be understood that this application can also customize composite tactical commands based on a combination of voice and text samples. For example, for the target semantic of "retreat", a voice sample such as "retreat, retreat quickly" (in a panicked tone, accompanied by panting, simulating a real retreat scenario) can be recorded. Text samples such as "retreat", "survive", "stop fighting, run", and "strategic transfer" can also be input. After semantic encoding of each sample, the obtained semantic representations are fused to obtain the target semantic representation corresponding to "retreat", which is then stored and associated with voice enhancement strategies, such as directional gain adjustment (omnidirectional gain +8dB to ensure that teammates can hear the retreat command clearly regardless of their direction), dynamic range compression (compression ratio 2:1), and specific frequency band enhancement (boosting 1kHz-3kHz to highlight sibilants such as "retreat" and "run", creating auditory penetration in a chaotic battlefield).
[0115] Based on this, if the first voice segment in the game's voice stream is a player frantically shouting "Run!" during combat, the second voice segment is a player calmly saying "Strategic retreat first," and the third voice segment is a player tauntingly saying "You guys retreat first, I'll cover the rear," then, through the semantic similarity analysis above, we know that although the language content of the first and second voice segments is different, they both match the target semantic meaning of "retreat," and will trigger the voice enhancement strategy associated with this target semantic meaning, thus achieving voice enhancement processing for these two voice segments. While the language content of the third voice segment contains "retreat," its pragmatic function is to stay rather than retreat. Therefore, the semantics of the third voice segment do not match the target semantic meaning of "retreat," and will not trigger the voice enhancement strategy for it, avoiding false enhancement that could lead to misjudgments by teammates. Therefore, the implementation method of the joint custom target semantic representation in this embodiment can be seen that the voice samples contribute acoustic behavior features in the retreat scenario, and the text samples contribute the semantic coverage of the retreat concept (from professional language expressions to colloquial expressions). The fusion of the two enables the target semantic representation to not only recognize non-standard pronunciations in the player's real panic state, but also understand written tactical terms, and achieve high accuracy in retreat semantic capture in complex battlefield contexts.
[0116] It is understandable that for other terminal devices or other application scenarios, such as smart home control scenarios, online video conferencing scenarios, or live streaming scenarios, users can customize and configure voice samples and / or text samples through terminal devices, generate target semantic representations through corresponding semantic coding models, and form a reference semantic set stored locally on the terminal device, or even associated with the corresponding application in the terminal device. The implementation method of associating at least one pre-configured voice enhancement strategy with each target semantic representation is similar to the implementation method of the game scenario mentioned above, and this application will not provide detailed examples.
[0117] In the semantic similarity analysis process described in the embodiments above, this application places the predefined target semantic representation and the semantic representation of each speech segment in a unified semantic vector space. By calculating the similarity between the two in this semantic vector space (which can be represented by the distance or angle between vectors, etc.), semantic similarity-based matching is achieved to determine whether the target semantic triggering condition is met, i.e., the closer the distance / the smaller the angle, the better the semantic match. In conjunction with the relevant description of the semantic coding model above, the semantic vector space is obtained by training the model using large-scale speech and text (semantic description) pairing data in the cloud. The mapping relationship learned by the semantic coding model (which can be a model that directly encodes speech segments, or it can be called a speech semantic coding model) can be the output of the last layer or a specific layer of the model to constitute the semantic vector space. The implementation process is not detailed in this application.
[0118] The target semantic trigger condition can be flexibly configured according to the application scenario. In one possible implementation, it can be an enhancement trigger condition for a single speech segment. In this case, the target semantic trigger condition can include the similarity between any speech segment in the speech stream (i.e., the similarity between the semantic representation of that speech segment and at least one target semantic representation in the reference semantic set) exceeding a first similarity threshold (such as 0.85 or 0.8, which can be determined based on scenario requirements or experience, and its size is not limited). Speech segments that meet this target semantic trigger condition can be used as target speech segments to trigger speech enhancement. This type of target semantic trigger condition is suitable for scenarios with extremely high real-time requirements and a need for rapid response, such as enhancing emergency commands in games. This condition is simple and direct.
[0119] In another possible implementation, the target semantic triggering condition can be an enhancement triggering condition for multiple consecutive speech segments. This application can identify the target speech segment that satisfies the target semantic triggering condition by using the similarity scores of each of the multiple consecutive speech segments, thus avoiding false triggering due to transient noise. In this case, the target semantic triggering condition can include obtaining the similarity scores of each of the multiple consecutive speech segments (e.g., 3 or 5), all of which exceed a second similarity threshold (e.g., 0.75 or 0.8, which can be determined based on scenario requirements or experience, and its size is not limited). These speech segments are then used as the target speech segment to trigger speech enhancement. This target semantic triggering condition requires the semantic intent to remain consistent over a longer time window, effectively filtering out isolated noise or slips of the tongue, improving triggering accuracy, and is suitable for scenarios sensitive to false triggering.
[0120] In another possible implementation, where the target semantic trigger condition is an enhanced trigger condition for multiple consecutive speech segments, smoothing processing can be combined to identify the target speech segment. In this case, the target semantic trigger condition can include: after smoothing the similarity of the corresponding multiple consecutive speech segments (e.g., moving average, median filtering, or exponential weighted average), the smoothed value obtained exceeds the third similarity threshold (e.g., 0.8, which can be determined based on scenario requirements or experience, and its size is not limited) in N consecutive (N≥2) speech segments. At this time, these N consecutive speech segments can be identified as the target speech segment. This trigger condition, which combines temporal continuity constraints with similarity smoothing, can effectively avoid instantaneous false triggers, and is especially suitable for speech with relatively dispersed or slowly changing semantic expressions (e.g., hesitant or repetitive cries for help), and can more robustly identify the entire target semantic interval / region.
[0121] Therefore, this application provides configurable target semantic triggering conditions to adapt to different application scenarios with varying requirements for sensitivity, accuracy, and latency. Users or systems can select the most suitable target semantic triggering conditions based on actual usage to achieve optimal enhancement effects. The implementation process for performing speech enhancement processing on target speech segments can be referred to the relevant description of speech enhancement strategies above, and will not be repeated here.
[0122] In practical applications, the voice stream in this application can be a real-time voice stream received by the terminal device during application operation, or an audio file obtained by recording the application operation process. In scenarios where the application is an online interactive application for multiple users, such as games or video conferencing, the acquired voice stream contains mixed voices from multiple users. For example, in mobile games, the system simultaneously collects the voices of local players and the voices of remote teammates received via the network, mixing them into frame-by-frame voice streams. It is evident that this mixed voice contains specific voice content (such as key voice content closely related to game operations, such as tactical instructions or status reminders in the game scene) and non-specific voice content (such as voice content related to player emotions or casual conversation in the game scene). The voice content to be enhanced as defined by the aforementioned target semantic representation belongs to this specific voice content. This application does not limit its included content or acquisition method, and can refer to, but is not limited to, the relevant description of the user-defined configuration reference semantic set above.
[0123] Subsequently, the pre-configured speech content to be enhanced is used to generate at least one target semantic representation to define the speech content to be enhanced (specific speech content) through a semantic coding model. The similarity between the semantic representations of each segmented speech (such as a frame of speech or a speech segment within a preset duration / time window) and the target semantic representation can be used to filter out the target speech segments that need speech enhancement (i.e. meet the target semantic triggering conditions) from the mixed speech stream. These include key tactical commands issued by players ("retreat", "use ultimate", etc.). This avoids enhancing non-specific speech content such as chatter and background noise, which would obscure key tactical commands and make it difficult for users to perceive key information in a timely manner, thus affecting the game's collaborative effect.
[0124] Subsequently, this application only enhances the voice segments that meet the target semantic triggering conditions. For example, when a local player says "retreat" and it matches the target semantics, only that player's voice segment is enhanced, without enhancing the voices of other teammates. This significantly improves the audibility of specific voice content in the game and the overall player experience. Furthermore, the method of this application has high generalization and can adapt to different games and different players' language and voice habits. The enhanced voice segment is mixed with at least one audio signal (such as background music, game sound effects, or other audio signals) and then output. This output method can be determined based on the scenario, such as outputting through a speaker or headphones. Therefore, this embodiment achieves precise semantic selective enhancement in a multi-channel mixed voice environment. Users can clearly hear key instructions without being drowned out by irrelevant sounds, making it particularly suitable for team-based games and online meetings, significantly improving the auditory experience and information transmission efficiency.
[0125] Reference Figure 5 This is a flowchart illustrating the speech processing method proposed in Embodiment 5 of this application. This embodiment describes a possible refined implementation method for obtaining the semantic vector of a speech segment through a semantic coding model. This method is also applicable to obtaining the target semantic vector and can be further used to identify the target speech segment. Figure 5 As shown, this refined implementation method may include:
[0126] Step S51: Extract features from the input using a semantic coding model to obtain feature representations;
[0127] In this embodiment, the input to the semantic coding model can be a speech segment from a speech stream, or text content converted from that speech segment, such as text obtained through local lightweight ASR (Automatic Speech Recognition). There are no restrictions on this. It should be noted that if the semantic coding model learns a mapping relationship that directly maps speech to semantic representation during training, then the input is the original speech segment.
[0128] During feature extraction, the input can be processed by feature extraction layers in a semantic coding model, such as several convolutional layers or Transformer coding layers, to output a temporal feature sequence as the feature representation of the input. Optionally, when the input is text content, the input text content can be segmented into words, and each segmented word can be mapped into a vector, thereby obtaining a set of feature representations (temporal feature sequences) of the text content.
[0129] Step S52: Perform semantic aggregation on the feature representation to generate aggregated semantic representation;
[0130] In this embodiment, a self-attention mechanism or pooling operation can be used to aggregate discrete feature representations (temporal feature sequences) into a global semantic representation, which captures the overall semantic information of the entire speech segment. For example, average pooling or attention-weighted summation can be used. This application does not limit the implementation method of semantic aggregation.
[0131] Step S53: Encode the aggregated semantic representation and output a semantic vector of a preset dimension.
[0132] In this embodiment, the aggregated semantic representation can be mapped to (compressed to) a fixed dimension (such as a preset dimension of 256 or 512) through the fully connected layer of the semantic coding model. After normalization, the final semantic vector is obtained. This semantic vector can represent the semantic information of the input, and the semantic vectors obtained from semantically similar inputs have a higher similarity. It can be directly used for subsequent similarity calculation to identify the target speech segment that needs to be enhanced.
[0133] Therefore, the semantic encoding process proposed in this embodiment ensures end-to-end mapping from the original speech segment to the semantic vector, without the need for explicit speech recognition and text understanding, avoiding information loss and error accumulation. Furthermore, the fixed-dimensional (preset-dimensional) vector design facilitates storage and rapid comparison. This application does not limit the number of preset dimensions.
[0134] In some embodiments, if the acquired speech segment is too long and contains too much linguistic content, this application proposes to enhance the semantic summarization generation process before semantic encoding to further improve the encoding quality. Optionally, this application can use a summarization generation model to semantically summarize the speech segment, generate a corresponding semantic summary, and then use a semantic encoding model to encode the semantic summary to obtain the corresponding semantic vector. The summarization generation model has a larger number of parameters than the semantic encoding model; that is, the summarization generation model can be a large local model (a model with a large number of parameters, which can also be deployed in the cloud or other edge devices with strong computing power) to utilize its summarization capabilities to semantically process the speech segment, performing de-colloquialization, redundant information removal, and intent focusing, extracting the core semantics of the speech segment, resulting in a more compact semantic summary, which can be in text form or a structured semantic description. For example, for a speech segment containing background noise, "Uh…that…could you…help me…thank you," the summarization generation model outputs a semantic summary of a plea for help. However, this method of semantic summarization generation is not limited to this approach.
[0135] The summarization model can be a general or pre-trained AI (Artificial Intelligence) model, which can employ, but is not limited to, the Transformer or its architectural variants (such as encoder-only / decoder-only, encoder-decoder, or MoE (Mixture of Experts, a neural network architecture) or other infrastructures, learning the mapping relationship from raw speech data to semantic summaries through training on a large amount of data. In practical applications of this application, the AI model can include, but is not limited to, generative models and generative language models (GLMs). For example, one or more of the following: large language model (LLM), GPT (Generative Pre-trained Transformer) series models, T5 (Text to Text Transfer Transformer) models, large visual models, and multimodal large models.
[0136] Of course, to meet the local deployment requirements of terminal devices, the general AI model (pre-trained model) or the proprietary model trained on it (a model used only for generating semantic summaries) can be lightweighted before the resulting summarization generation model is deployed to the terminal device. During this process, it is necessary to ensure the integrity and accuracy of the generated semantic summaries and avoid excessive compression that could lead to the loss of semantic information represented by the subsequently generated semantic vectors. This application does not impose restrictions on the specific network structure, parameter size, training method, or deployment method of the model; these can be flexibly determined based on actual business needs.
[0137] Since semantic summarization removes redundant information and non-semantic noise, lightweight semantic coding models can output semantic vectors more accurately with less computational resources. This collaborative model of generating semantic summaries using a large model and semantic encoding using a small model allows for high-quality semantic summaries to be obtained by leveraging a large model (summarization generation model) occasionally called from the cloud or locally, even when only the small model (semantic coding model) is deployed on the terminal device. This improves the overall system's accuracy and robustness, solving the problem of insufficient encoding accuracy of purely lightweight models (semantic coding models) in complex noise scenarios. This application preprocesses speech segments using a summarization generation model, significantly reducing the difficulty of subsequent semantic encoding while maintaining the advantage of low latency on the edge. Optionally, semantic summarization can be completed offline asynchronously, allowing direct acquisition of the semantic summary corresponding to the speech segment for subsequent semantic encoding, but this is not limited to this.
[0138] Based on the semantic encoding implementation method described above, in its implementation process (which can be processed offline) for semantically encoding user-defined speech samples and / or text samples to generate corresponding target semantic representations, for at least one user-recorded speech sample, the summarizing ability of the AI model can be used to generate a structured semantic description as a semantic summary, and then the corresponding target semantic representation can be generated through the semantic encoding model. This ensures the accuracy and reliability of identifying the target speech segments to be enhanced in the speech stream, that is, reliably identifying multiple target speech segments with different linguistic expressions but semantically matching the target semantics, thereby triggering speech enhancement.
[0139] The AI model can also retrieve audio files semantically related to the input sample from an audio library to enrich the sample. It then summarizes and analyzes all semantically related audio files for the same sample to generate a corresponding semantic summary. This semantic summary can include structured semantic descriptions such as semantic tags, a content summary of the language expression, and keywords. For example, if the semantic tag (target semantics) is "English lesson," its content summary might be: "Explanation of English tenses…"; keywords might include "English," "classroom," and "tense." If the semantic tag is "conference," its content summary might be "patent review"; keywords might include "patent," "disclosure document," and "audio retrieval." The semantic encoding model can encode the semantic summary by encoding the content summary and / or keywords contained within it; the encoding process is not detailed in this application.
[0140] Based on this, in the process of identifying the target speech segments that need to be enhanced from the speech stream, this application calculates the similarity between the semantic representation corresponding to each speech segment and the target semantic representation to determine the target speech segments whose similarity is greater than the corresponding similarity threshold, i.e., those that match the target semantics, thereby triggering speech enhancement. For example, in actual classroom scenarios, teachers or students are not required to speak speech segments containing the linguistic expression "English class." They can directly speak speech segments containing other linguistic expressions that match the target semantics of "English class," such as "explaining English tenses" or "classroom," which can also trigger enhancement processing.
[0141] Optionally, during the training of the semantic coding model, this application may use different speech segments from the same audio file or their transcribed text as positive samples (such as English lessons, explanations of English tenses, classrooms, etc.) to generate similar semantic representations, and use speech segments from other audio files (usually semantically unrelated audio files) or their transcribed text as negative samples to generate semantic representations with greater discriminative power than the semantic representations of the positive samples. For example, the positive and negative samples "English lesson" and "meeting" are equivalent. By adjusting the model parameters, the similarity of the semantic representations corresponding to multiple samples matching the same target semantics is higher, while the similarity between the semantic representations of multiple samples matching different target semantics is lower. This application does not restrict the training implementation process.
[0142] It should be understood that the semantic encoding and semantic matching mechanism of this application can be used not only for real-time speech enhancement, but also for audio retrieval (such as...). Figure 6The processing described includes various scenarios such as: retrieving target speech segments whose semantic representation matches the target semantic representation from speech segments contained in at least one speech stream through semantic similarity calculation (the implementation process is not detailed here); content recommendation (i.e., filtering target recommended content from candidate content based on the similarity between the semantic representation of candidate content and the target semantic representation representing recommended content); and voice command recognition (the voice command is equivalent to the speech segment in the above embodiment, and the recognition process is the same as the implementation process of recognizing target speech segments that meet the target semantic triggering conditions described above, which is not detailed here). This expands the scope of protection. Therefore, this application maps different forms of input (speech, text, keywords) to a unified semantic vector space through a semantic coding model, and then judges whether the semantics of the input matches any target semantic in the reference semantic set through semantic vector similarity. This eliminates the dependence of traditional matching methods on specific words (such as keywords and other linguistic expressions), and improves the robustness and generalization of speech recognition.
[0143] Reference Figure 7 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application. Figure 7 As shown, the voice processing device may include:
[0144] The speech segment acquisition module 71 is used to acquire speech segments in the speech stream;
[0145] The speech enhancement processing module 72 is configured to perform speech enhancement processing on the first speech segment when the first speech segment in the speech segment has a first language expression content and corresponds to a first semantic, and if the first semantic matches the target semantic; and to perform speech enhancement processing on the second speech segment when the second speech segment in the speech segment has a second language expression content and corresponds to a second semantic, and if the second semantic matches the target semantic.
[0146] The content expressed in the first language is not entirely the same as the content expressed in the second language.
[0147] Optionally, the voice enhancement processing module 72 may include:
[0148] A similarity acquisition unit is used to acquire the similarity between the semantic representation of the speech segment and at least one target semantic representation in a reference semantic set;
[0149] A determining unit is configured to determine whether a target speech segment meets a target semantic triggering condition based on similarity; the target speech segment includes the first speech segment and / or the second speech segment.
[0150] The target semantic triggering conditions may include any of the following:
[0151] The similarity of any of the aforementioned speech segments exceeds a first similarity threshold;
[0152] The similarity of each of the consecutive speech segments exceeds the second similarity threshold;
[0153] After smoothing the similarity values corresponding to the consecutive multiple speech segments, if the smoothed values obtained exceed the third similarity threshold in all N consecutive speech segments, the consecutive N speech segments are identified as the target speech segment; where N≥2.
[0154] The speech enhancement processing unit is configured to perform speech enhancement processing on the target speech segment if the target semantic triggering condition is met, so as to enhance the speech content corresponding to the target semantic representation.
[0155] In some embodiments, the semantic representation is a semantic vector obtained by semantically encoding the corresponding speech segment using a semantic coding model, and the semantic vector is mapped in a unified semantic vector space; the target semantic representation is used to define the speech content to be enhanced in the semantic vector space.
[0156] The semantic encoding model is a lightweight model deployed on terminal devices, and the parameters of the lightweight model can be optimized based on model update data from the cloud.
[0157] In some embodiments, the target semantic representation is generated based on user-defined speech samples and / or text samples through the semantic coding model; wherein each target semantic representation is associated with at least one pre-configured speech enhancement strategy, and based on this, the speech enhancement processing module 72 can be specifically used to perform speech enhancement processing on the target speech segment based on the speech enhancement strategy; the speech enhancement strategy includes at least one of the following:
[0158] Directional gain adjustment, dynamic range compression, specific frequency band enhancement, and noise suppression.
[0159] In some embodiments, the semantic encoding module for semantically encoding corresponding speech segments using a semantic encoding model to obtain semantic vectors may include:
[0160] The feature extraction unit is used to extract features from the input through a semantic coding model to obtain a feature representation; the input is the speech segment or the text content converted from the speech segment.
[0161] A semantic aggregation unit is used to perform semantic aggregation on the feature representation to generate an aggregated semantic representation;
[0162] The encoding unit is used to encode the aggregated semantic representation and output a semantic vector of a preset dimension.
[0163] Optionally, it may also include:
[0164] The semantic summarization generation module is used to perform semantic summarization on the speech segment through a summarization generation model to generate a corresponding semantic summary, so that the semantic encoding module can encode the semantic summary through the semantic encoding model to obtain the semantic vector;
[0165] The number of parameters in the summary generation model is greater than the number of parameters in the semantic encoding model.
[0166] In some embodiments, the aforementioned voice stream is a real-time voice stream received by the terminal device during application operation, or an audio file obtained by recording the application operation process; wherein, when the application is an online interactive application for multiple users, the voice stream contains mixed voices of multiple users; based on this, the voice processing device may further include:
[0167] The output module is used to mix the speech segment obtained from the speech enhancement process with at least one audio signal and then output it.
[0168] Optionally, the mixed speech includes specific speech content and non-specific speech content; wherein the speech content to be enhanced as defined by the target semantic representation belongs to the specific speech content.
[0169] This application also provides a computer program product, including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the voice processing methods provided in this application.
[0170] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the voice processing methods provided in this application.
[0171] The computer-readable storage medium can be any available medium that an electronic device can store, or a data storage device such as a training device or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)). It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0172] Reference Figure 8 This is a schematic diagram of the hardware structure of a terminal device proposed in an embodiment of this application. The terminal device can be a user terminal such as a smartphone, tablet computer, game console, smart speaker, or smart wearable device. Figure 8 As shown, the terminal device may include at least one memory 81 and at least one processor 82, wherein: the memory 81 stores a computer program, and the processor 82 is used to load and execute the computer program in the memory 81 to implement the various steps of the voice processing method described in the method embodiments of this application. The implementation process can be referred to the description of the corresponding method embodiments above, and will not be repeated in this embodiment.
[0173] Optionally, the processor 82 can also call the intelligent program (such as an intelligent agent or other AI assistant) stored in the memory 81 to enter the working state. The intelligent program is configured to implement the speech processing method described in the above embodiments of this application, and output the enhanced speech segment or the signal mixed with at least one audio signal through at least one output component (such as a speaker or headphones).
[0174] Furthermore, the terminal device may also include a microphone array and a communication module. The microphone array is used to acquire local speech streams, and the communication module is used to receive speech streams from other users and send them to the processor 82. The processor 82 can run a lightweight semantic coding model and a similarity matching algorithm to obtain speech-enhanced speech segments.
[0175] It should be understood that, Figure 8 The structure of the terminal device shown does not constitute a limitation on the terminal device in the embodiments of this application. In practical applications, the terminal device may include more than Figure 8The more or fewer components shown, or combinations of certain components, such as gyroscopes, accelerometers, and gravity sensors used to obtain sensing parameters, power management modules, antennas, etc., are not described in detail in this application.
[0176] Finally, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0177] In the above embodiments, the implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware, or it can be implemented using dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memory, dedicated components, etc. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The various embodiments in this specification are described in a progressive or parallel manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to mutually. For the apparatus and terminal devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
Claims
1. A speech processing method, comprising: Extract speech segments from a speech stream; When the first speech segment in the speech segment has a first language expression content and corresponds to a first semantic, if the first semantic matches the target semantic, speech enhancement processing is performed on the first speech segment; When the second speech segment in the speech segment has a second language expression content and corresponds to a second semantic, if the second semantic matches the target semantic, speech enhancement processing is performed on the second speech segment; The content expressed in the first language is not entirely the same as the content expressed in the second language.
2. The method according to claim 1, wherein the semantic matching with the target semantic includes: The similarity between the semantic representation of the speech segment and at least one target semantic representation in the reference semantic set is obtained to determine whether the target speech segment meets the target semantic triggering condition; the target speech segment includes the first speech segment and / or the second speech segment; If the target semantic triggering condition is met, speech enhancement processing is performed on the target speech segment to enhance the speech content corresponding to the target semantic representation.
3. The method according to claim 2, wherein the semantic representation is a semantic vector obtained by semantically encoding the corresponding speech segment through a semantic coding model, and the semantic vector is mapped to a unified semantic vector space; The target semantic representation is used to define the speech content to be enhanced in the semantic vector space; in, The semantic encoding model is a lightweight model deployed on terminal devices, and the parameters of the lightweight model can be optimized based on model update data from the cloud.
4. The method according to claim 3, wherein the target semantic representation is generated based on user-defined speech samples and / or text samples processed by the semantic coding model; in, Each of the target semantic representations is associated with at least one pre-configured speech enhancement strategy to perform speech enhancement processing on the target speech segment based on the speech enhancement strategy; the speech enhancement strategy includes at least one of the following: Directional gain adjustment, dynamic range compression, specific frequency band enhancement, and noise suppression.
5. The method according to claim 3, wherein the step of semantically encoding the corresponding speech segment using a semantic coding model to obtain a semantic vector comprises: Feature representations are obtained by extracting features from the input using a semantic coding model. The input is the speech segment or the text content converted from the speech segment; The feature representations are semantically aggregated to generate aggregated semantic representations; The aggregated semantic representation is encoded to output a semantic vector of a preset dimension.
6. The method according to claim 3, further comprising, before the semantic encoding: The speech segment is semantically summarized by a summary generation model to generate a corresponding semantic summary, which is then encoded by the semantic encoding model to obtain the semantic vector. The number of parameters in the summary generation model is greater than the number of parameters in the semantic encoding model.
7. The method according to claim 2, wherein the target semantic triggering condition includes any one of the following: The similarity of any of the aforementioned speech segments exceeds a first similarity threshold; The similarity of each of multiple consecutive speech segments exceeds the second similarity threshold; After smoothing the similarity scores of multiple consecutive speech segments, if the smoothed values exceed a third similarity threshold in all N consecutive speech segments, then the N consecutive speech segments are identified as the target speech segment; wherein... N≥2。 8. The method according to any one of claims 3-7, wherein the voice stream is a real-time voice stream received by the terminal device during application operation, or an audio file obtained by recording the application operation process; in, When the application is a multi-user online interactive application, the voice stream contains mixed voices from multiple users; Also includes: The speech segment obtained from the speech enhancement process is mixed with at least one audio signal and then output.
9. The method according to claim 8, wherein the mixed speech includes specific speech content and non-specific speech content; in, The speech content to be enhanced, as defined by the target semantic representation, belongs to the specific speech content.
10. A terminal device, comprising a memory and a processor, wherein: The memory stores computer programs; The processor is configured to execute the computer program to implement the speech processing method as described in any one of claims 1-9.