Method and device for determining context hot words, computer device and storage medium

By performing temporal accumulation and local alignment processing on the boundary contributions of frame-level speech vectors in speech recognition technology, acoustic boundary sequences are determined, and contextual hot words are screened out. This solves the problems of alignment offset and local aggregation inaccuracy in existing technologies, and improves the accuracy and reliability of speech recognition.

CN122290574APending Publication Date: 2026-06-26BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BAIRONG ZHIXIN (BEIJING) TECH CO LTD
Filing Date
2026-03-26
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Under weak supervision without timestamps, existing speech recognition technologies suffer from local acoustic cues of short entities being submerged in background speech and irrelevant content, leading to diluted representations. Furthermore, the lack of soft matching methods with explicit monotonic time constraints makes it easy to cause alignment shifts and local aggregation inaccuracies.

Method used

By processing the speech to be recognized and the candidate word list, the frame-level speech vector and text semantic vector are determined. The acoustic boundary sequence is processed by time-series accumulation. Local alignment is performed by combining the text lexical length of the candidate hot words. Context hot words are selected, and the recognized text is generated using the speech recognition model.

Benefits of technology

It effectively suppresses alignment offset and local aggregation inaccuracy in weakly supervised scenarios, avoids interference from globally irrelevant information, improves the accuracy of context hot word recognition and the reliability of speech recognition, and has good versatility and transferability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290574A_ABST
    Figure CN122290574A_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, computer device, and storage medium for determining contextual hot words. The method includes: processing the speech to be recognized and candidate hot words in a candidate word list to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step; performing temporal accumulation processing of boundary contributions on the frame-level speech vectors to determine an acoustic boundary sequence, wherein the acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized; performing local alignment of the candidate hot words and the frame-level speech vectors on the acoustic boundary sequence according to the text lexical length of the candidate hot words to determine the acoustic vector group corresponding to the candidate hot words; and filtering each candidate hot word according to the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group to obtain the contextual hot words of the speech to be recognized, thereby improving the accuracy of contextual hot word recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining contextual hot words. Background Technology

[0002] With the development of large-scale speech models, open-domain speech recognition has achieved good results in general scenarios, but it still has significant shortcomings in the recognition of named entities, long-tail words, and low-frequency proper words. The effectiveness of existing context bias or retrieval enhancement methods largely depends on whether the front-end hot word retrieval module can stably and accurately locate target hot word segments in the entire sentence of speech.

[0003] Under weak supervision without timestamps, existing methods typically suffer from the following problems: On the one hand, the retrieval method using sentence-level global pooling can submerge the local acoustic cues corresponding to short entities in background speech and irrelevant content, leading to representation dilution; on the other hand, the lack of explicit monotonic temporal constraints in soft matching methods makes it easy to learn global semantic co-occurrence relationships rather than strict local temporal correspondences, thereby causing alignment shifts and inaccurate local aggregation. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for determining contextual hot words that can improve the accuracy of hot word determination, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for determining contextual hot words, including:

[0006] The speech to be recognized and the candidate hot words in the candidate word list are processed separately to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step.

[0007] The frame-level speech vectors undergo temporal accumulation processing for boundary contributions to determine an acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized.

[0008] Based on the text lexical length of the candidate hot words, the candidate hot words are locally aligned with the frame-level speech vectors on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words;

[0009] Based on the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group, the candidate hot words are filtered to obtain the context hot words of the speech to be identified.

[0010] In one embodiment, the step of locally aligning the candidate hot words with the frame-level speech vectors on the acoustic boundary sequence based on the text lexical length of the candidate hot words, and determining the acoustic vector group corresponding to the candidate hot words, includes:

[0011] The acoustic window of the candidate hot words is determined based on the text lexical length of the candidate hot words;

[0012] Based on the sliding of the acoustic window on the acoustic boundary sequence, the frame-level speech vectors corresponding to each acoustic boundary located within the acoustic window are determined as the acoustic vector group of the candidate hot words.

[0013] In one embodiment, the step of sliding the acoustic window on the acoustic boundary sequence and determining each frame-level speech vector corresponding to the acoustic boundary within the acoustic window each time as the acoustic vector group of the candidate hot words includes:

[0014] Based on the sliding of the acoustic window on the acoustic boundary sequence, the acoustic frame span corresponding to the acoustic boundary at multiple acoustic window positions is determined;

[0015] Each frame-level speech vector corresponding to the acoustic frame span within the acoustic window is determined as the acoustic vector group of the candidate hot words.

[0016] In one embodiment, the step of filtering each candidate hot word based on its text semantic vector and the frame-level speech vector of the acoustic vector group to obtain the context hot words of the speech to be recognized includes:

[0017] The candidate scores of the candidate hot words are determined based on the text semantic vectors of the candidate hot words and the frame-level speech vectors of the acoustic vector groups.

[0018] Among the candidate hot words, those whose candidate scores meet the screening criteria are determined as the context hot words of the speech to be recognized.

[0019] In one embodiment, determining the candidate score of the candidate hot word based on the text semantic vector of the candidate hot word and the frame-level speech vector of the acoustic vector group includes:

[0020] Determine the frame-level similarity between the text semantic vector of the candidate hot words and the corresponding acoustic vector group for each frame-level speech vector;

[0021] The frame-level similarity of each acoustic vector group is aggregated to determine the aggregated similarity of each acoustic vector group;

[0022] The highest aggregate similarity of the acoustic vector group is determined as the candidate score of the candidate hot word.

[0023] In one embodiment, determining the candidate hot words whose candidate scores meet the screening criteria as context hot words for the speech to be recognized from among the candidate hot words includes:

[0024] When the candidate hot words are arranged in descending order of candidate scores, the top target number of candidate hot words are determined as the context hot words of the speech to be recognized.

[0025] or,

[0026] Candidate hot words with scores greater than a preset score threshold are identified as context hot words of the speech to be recognized.

[0027] In one embodiment, the temporal accumulation processing of the frame-level speech vector for boundary contribution to determine the acoustic boundary sequence includes:

[0028] Predict the boundary contribution value for each frame-level speech vector;

[0029] The boundary contribution values ​​of each frame-level speech vector are accumulated frame by frame in chronological order, and the acoustic boundary sequence is determined based on the accumulated values.

[0030] In one embodiment, the boundary contribution value for each frame-level speech vector is accumulated frame by frame in chronological order, and the acoustic boundary sequence is determined based on the accumulated value, including:

[0031] The boundary contribution values ​​corresponding to the speech vectors at each frame level are accumulated in chronological order to obtain the cumulative value;

[0032] When the accumulated value reaches a preset threshold, a boundary event is triggered, which is used to indicate that the current frame that triggered the boundary event is determined as an acoustic boundary;

[0033] After the boundary event is triggered, based on the difference between the accumulated value and the preset threshold, the boundary contribution value of the subsequent frames of the current frame is continued to be accumulated; until the traversal is completed, an acoustic boundary sequence arranged in chronological order is obtained.

[0034] In one embodiment, the method further includes:

[0035] Generate hot word suggestions based on the context hot words;

[0036] Based on the hot word prompts, the speech to be recognized is recognized using a speech recognition model to generate recognized text.

[0037] Secondly, this application also provides a method for training a target network model, including:

[0038] Obtain a training dataset, which includes multiple sample data, each of which includes at least a whole sentence speech data, a whole sentence text data corresponding to the whole sentence speech data, and target hot word text contained in the whole sentence text data;

[0039] The whole-sentence speech data, the whole-sentence text data, and the target hot word text are processed respectively to determine the training frame-level speech vector, whole-sentence speech vector, whole-sentence text vector, and hot word text vector;

[0040] The trained frame-level speech vectors are input into the initial network model to obtain the training boundary contribution value;

[0041] The training boundary contribution value is subjected to time-series cumulative processing to obtain the training acoustic boundary;

[0042] Based on the training acoustic boundary, determine the hot word segment speech vector corresponding to the target hot word from the training frame-level speech vector;

[0043] Based on the whole sentence speech vector, the whole sentence text vector, the hot word text vector, the hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary, the initial network model is trained to generate the target network model.

[0044] In one embodiment, training the initial network model based on the whole-sentence speech vector, the whole-sentence text vector, the hot word text vector, the hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary to generate the target network model includes:

[0045] The data constraint loss is determined based on the difference between the number of training lexical units and the length of the lexical units in the entire sentence text;

[0046] The global contrast loss is determined based on the semantic difference between the whole-sentence speech vector and the whole-sentence text vector;

[0047] The local contrast loss is determined based on the semantic difference between the speech vector of the hot word segment and the text vector of the hot word.

[0048] The comprehensive loss value is determined based on the data constraint loss, the global comparison loss, and the local comparison loss.

[0049] The initial network model is trained based on the comprehensive loss value to generate the target network model.

[0050] Thirdly, this application also provides a voice recognition device, comprising:

[0051] The feature encoding module is used to process the speech to be recognized and the candidate hot words in the candidate word list respectively, and determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step.

[0052] The continuous integration-triggered alignment module is used to perform boundary contribution time-series accumulation processing on the frame-level speech vector to determine the acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vector in the speech to be recognized.

[0053] The candidate sorting module is used to perform local alignment between the candidate hot words and the frame-level speech vectors on the acoustic boundary sequence based on the text lexical length of the candidate hot words, and to determine the acoustic vector group corresponding to the candidate hot words;

[0054] The sorting and context recognition module is used to filter each candidate hot word according to the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group to obtain the context hot words of the speech to be recognized.

[0055] Fourthly, this application also provides a training device for a target model, comprising:

[0056] The first acquisition module is used to acquire a training dataset, which includes multiple sample data. Each sample data includes at least a whole sentence speech data, a whole sentence text data corresponding to the whole sentence speech data, and target hot word text contained in the whole sentence text data.

[0057] The first determining module is used to process the whole sentence speech data, the whole sentence text data and the target hot word text respectively to determine the training frame-level speech vector, whole sentence speech vector, whole sentence text vector and hot word text vector.

[0058] The second acquisition module is used to input the training frame-level speech vectors into the initial network model to obtain the training boundary contribution value.

[0059] The second determining module is used to perform time-series cumulative processing on the training boundary contribution value to obtain the training acoustic boundary.

[0060] The third determining module is used to determine the hot word segment speech vector corresponding to the target hot word from the training frame-level speech vector based on the training acoustic boundary.

[0061] The generation module is used to train the initial network model based on the whole sentence speech vector, the whole sentence text vector, the hot word text vector, the hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary, and generate the target network model.

[0062] Fifthly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0063] The speech to be recognized and the candidate hot words in the candidate word list are processed separately to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step.

[0064] The frame-level speech vectors undergo temporal accumulation processing for boundary contributions to determine an acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized.

[0065] Based on the text lexical length of the candidate hot words, the candidate hot words are locally aligned with the frame-level speech vectors on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words;

[0066] Based on the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group, the candidate hot words are filtered to obtain the context hot words of the speech to be identified.

[0067] Sixthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0068] The speech to be recognized and the candidate hot words in the candidate word list are processed separately to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step.

[0069] The frame-level speech vectors undergo temporal accumulation processing for boundary contributions to determine an acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized.

[0070] Based on the text lexical length of the candidate hot words, the candidate hot words are locally aligned with the frame-level speech vectors on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words;

[0071] Based on the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group, the candidate hot words are filtered to obtain the context hot words of the speech to be identified.

[0072] In a seventh aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0073] The speech to be recognized and the candidate hot words in the candidate word list are processed separately to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step.

[0074] The frame-level speech vectors undergo temporal accumulation processing for boundary contributions to determine an acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized.

[0075] Based on the text lexical length of the candidate hot words, the candidate hot words are locally aligned with the frame-level speech vectors on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words;

[0076] Based on the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group, the candidate hot words are filtered to obtain the context hot words of the speech to be identified.

[0077] The aforementioned method, apparatus, computer device, computer-readable storage medium, and computer program product for determining contextual hot words can first process the speech to be recognized and the candidate hot words in the candidate word list separately to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step. Then, the frame-level speech vector can be processed by temporal accumulation of boundary contributions to determine the acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized. Then, based on the text lexical length of the candidate hot words, the candidate hot words are locally aligned with the frame-level speech vectors on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words. Finally, based on the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group, each candidate hot word can be filtered to obtain the contextual hot words of the speech to be recognized. Therefore, by performing temporal accumulation processing on the boundary contribution of the frame-level speech vectors of the speech to be recognized, the corresponding acoustic boundary sequence can be determined. Then, based on the text word length of the candidate hot words, local alignment between the candidate hot words and the frame-level speech vectors is performed on the acoustic boundary sequence. This can segment the frame-level speech vectors of the speech to be recognized into acoustic vector groups matching the candidate hot words, thus ensuring the temporal correspondence between the frame-level speech vectors and the text words of the hot words in local segments. This effectively suppresses alignment offset and local aggregation inaccuracies in weakly supervised scenarios, avoids interference from globally irrelevant information, and also avoids... The local acoustic cues of short entities are submerged in background speech and irrelevant content, resulting in sparse representation. Then, based on the text semantic vector of the candidate hot words and the frame-level speech vector of the acoustic vector group, the candidate hot words are matched and filtered only within the acoustic vector group. This avoids noise interference when comparing the whole sentence speech with the global hot words. It does not rely on additional language models and complex finite graphs, thus improving the accuracy of context hot word recognition of the speech to be recognized. In turn, it improves the accuracy and reliability of speech recognition based on context hot words of the speech to be recognized, and has good versatility and transferability. Attached Figure Description

[0078] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0079] Figure 1 This is a flowchart illustrating a method for determining contextual hot words in one embodiment;

[0080] Figure 2 This is a flowchart illustrating the process of determining an acoustic boundary sequence in one embodiment;

[0081] Figure 3 This is a flowchart illustrating the process of determining contextual hot words for the speech to be recognized in one embodiment;

[0082] Figure 4 This is a flowchart illustrating the training method of the target network model in one embodiment;

[0083] Figure 5 This is a flowchart illustrating the process of determining contextual hot words for the speech to be recognized in another embodiment;

[0084] Figure 6 This is a schematic diagram of the speech recognition process in one embodiment;

[0085] Figure 7 This is a structural block diagram of a device for determining contextual hot words in one embodiment;

[0086] Figure 8 This is a structural block diagram of a training device for a target network model in one embodiment;

[0087] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0088] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0089] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0090] It should be understood that various forms of processes shown in this application can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and no limitation is imposed here.

[0091] The user data, data acquisition, and / or use involved in the embodiments of this application strictly comply with the laws, regulations, and industry standards of relevant countries and regions. The collection and acquisition of data involved in the embodiments of this application are all done in advance by actively prompting or prominently displaying information to inform users and obtaining authorization, or by obtaining full authorization from all parties. The processing, manipulation, forwarding, and use of data involved in the embodiments of this application are all carried out on the premise that the user or relevant party is fully informed and authorized. In implementing the various embodiments of this application, the types of data or information, scope of use, and usage scenarios that may be involved are informed to users or relevant parties and authorization is obtained through appropriate means. The specific methods of notification and authorization may vary according to the actual situation, and this application is not limited in this regard. The processing of personal information involved in the embodiments of this application is carried out under the premise of having a legal basis and only within the scope of regulations or agreements. Sensitive personal information such as biometric information, medical and health information, financial account information, and precise location information involved in the embodiments of this application are processed under the premise of having a specific purpose and sufficient necessity, and with the separate authorization and consent of the user or relevant party. In some embodiments of this application, if the user or relevant party refuses to process personal information other than the information necessary for basic functions, it will not affect the use of the basic functions of the embodiments of this application.

[0092] This application utilizes deep learning, neural network models, and other technologies to mine and analyze the inherent relationships in data that conform to natural laws, namely the inherent laws between speech acoustic features and text semantics. Based on this, it solves the technical problem of inaccurate hot word localization and achieves the technical effect of improving the accuracy of context hot word screening and the reliability of speech recognition.

[0093] In one embodiment, such as Figure 1As shown, a method for determining contextual hot words is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and can be implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0094] Step 102: Process the speech to be recognized and the candidate hot words in the candidate word list respectively to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step.

[0095] The speech to be recognized can be the original speech data that needs to be converted into corresponding text through speech recognition, such as speech data input by the user, audio data in audio files or video files, etc. This application does not limit this.

[0096] In some embodiments of this application, the candidate word list can be a general word list, which may include personal names, place names, organization names, business names, professional terms, or other context-related words. Alternatively, the candidate word list may also be a domain-specific word list related to a preset speech recognition task or scenario. For example, in scenarios such as e-commerce customer service or in-vehicle voice assistants, the candidate word list may be an e-commerce domain word list or a vehicle domain word list. The candidate word list may include one or more candidate hot words, etc., and this application does not limit this.

[0097] In some embodiments of this application, the frame-level speech vector can be the speech vector of each time frame obtained after processing the speech to be recognized, and the text semantic vector can be the semantic vector obtained after processing the candidate hot word text.

[0098] Understandably, the speech to be recognized can first be converted into an acoustic feature sequence, and then the acoustic feature sequence can be vectorized to obtain the frame-level speech vector at each time step. There are various ways to vectorize the acoustic feature sequence, such as encoding, linear transformation, or statistical feature mapping, etc., and this application does not limit the specific methods used.

[0099] The acoustic feature sequence can take many forms, such as Mel spectrum, Mel frequency cepstral, etc., and this application does not limit it.

[0100] In some embodiments, candidate hot words can be vectorized to obtain text semantic vectors corresponding to each candidate hot word. There are various ways to vectorize candidate hot words, such as encoding, lookup tables based on pre-trained word vectors, etc., and this application does not limit the specific methods used.

[0101] Optionally, the model can be used to process the speech to be recognized and the candidate hot words separately. For example, the acoustic feature sequence of the speech to be recognized can be encoded using a speech coding model to obtain the frame-level speech vector at each time step; the candidate hot words can be encoded using a text coding model to obtain the corresponding text semantic vector. The speech coding model and the text coding model can be two independent models, or they can be two branches of the same model, etc., and this application does not limit them in this way.

[0102] Step 104: Perform temporal accumulation processing on the boundary contribution of the frame-level speech vectors to determine the acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized.

[0103] Boundary contribution can be understood as the degree of contribution of each frame-level speech vector to the formation of the next word boundary. The boundary contribution value can be used to characterize the degree or contribution ratio of each frame-level speech vector to the formation of the next word boundary. Usually, the boundary contribution value is (0,1). Boundary contribution can be used for monotonic alignment of speech frames to word positions.

[0104] In this context, the acoustic boundary can be understood as the temporal boundary point used to divide the frame-level speech vector in the speech to be recognized. It indicates that the frame-level speech vector undergoes acoustic feature changes at that position, enabling the segmentation of the frame-level speech vector. It can be represented by the index of the speech frame. The acoustic boundary sequence can be a sequence composed of multiple acoustic boundaries arranged in chronological order. Since time increases gradually, the acoustic boundary sequence is a monotonically increasing sequence. For example, if there are 100 frames of frame-level speech vector, and the determined acoustic boundaries are frames 20, 50, 80, and 100, then the corresponding acoustic boundary sequence is [20, 50, 80, 100]. This application does not limit this.

[0105] In addition, time-series cumulative processing can be understood as performing step-by-step accumulation calculations on data according to the chronological order. In the embodiments of this application, it can be understood as the process of accumulating the boundary contribution value of each frame-level speech vector frame by frame according to the chronological order of each frame-level speech vector.

[0106] In some embodiments, there can be multiple ways to determine the acoustic boundary sequence based on frame-level speech vectors. For example, an acoustic boundary detection model can be used to process the frame-level speech vectors, predict the probability that each frame is an acoustic boundary, and determine the frames with probabilities greater than a certain probability threshold as acoustic boundaries, thus obtaining the acoustic boundary sequence in chronological order. For example, the frame-level speech vectors of the speech to be recognized can be input into the acoustic boundary detection model. After processing by the acoustic boundary detection model, the probability that each frame-level speech vector belongs to an acoustic boundary can be predicted. Then, the probability corresponding to each frame-level speech vector can be calibrated with a preset probability threshold. The frames with probabilities greater than the preset probability threshold can be determined as acoustic boundaries, and then arranged in chronological order to obtain the corresponding acoustic boundary sequence.

[0107] Alternatively, the acoustic boundary can be determined based on the similarity between the speech vectors of two adjacent frames. For example, if the similarity between adjacent frames is below a certain similarity threshold, it can be considered that the acoustic characteristics have changed significantly, and the preceding frame between adjacent frames can be determined as the acoustic boundary, thus obtaining the acoustic boundary sequence. Optionally, the following frame between adjacent frames can also be determined as the acoustic boundary, as long as the acoustic boundaries are both start or end frames; this application does not impose any limitations on this.

[0108] Alternatively, the acoustic boundary can be determined by accumulating the boundary contribution values ​​of the frame-level speech vectors. For example, the boundary contribution values ​​of the frame-level speech vectors can be determined first, and then time-series accumulating these boundary contribution values ​​can be performed to determine the acoustic boundary based on the accumulated value. Any acceptable method can be used to determine the boundary contribution values ​​of the frame-level speech vectors, and this application does not impose any limitations on this.

[0109] The boundary contributions of each frame-level speech vector can be predicted using a Continuous Integrate-and-Fire (CIF) predictor, yielding the boundary contribution values ​​for each frame-level speech vector. The CIF predictor processes the input frame-level speech vectors through one-dimensional convolution, layer normalization, activation functions, and linear transformations, outputting boundary contribution values ​​ranging from (0,1). A counter is then used to accumulate the boundary contribution values ​​of the frame-level speech vectors frame by frame in chronological order. When the accumulated value reaches a preset threshold, a boundary event is triggered, generating a word-level acoustic boundary. The index of the current frame triggering the boundary event can be recorded as an acoustic boundary. After each trigger, the portion of the accumulated value exceeding the preset threshold continues to participate in the accumulation process of subsequent frames until the boundary contribution values ​​of all frame-level speech vectors have been traversed, ultimately generating a monotonically increasing acoustic boundary sequence. By statistically analyzing the number of historical boundary triggers, each frame of speech can be mapped to its corresponding lexical index, constructing a non-decreasing frame-to-lexical mapping relationship. This provides a foundation for subsequent local window matching based on text lexical length awareness and speech-to-text alignment.

[0110] It should be noted that the above examples are merely illustrative and should not be construed as limiting the methods for determining acoustic boundary sequences based on frame-level speech vectors in the embodiments of this application.

[0111] Step 106: Based on the text lexical length of the candidate hot words, perform local alignment between the candidate hot words and the frame-level speech vectors on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words.

[0112] The length of the text lexicon of the candidate hot word can be the number of text lexicons contained in the candidate hot word. The acoustic vector group can be understood as a combination of frame-level speech vectors that are segmented by the acoustic boundary sequence and match the length of the text lexicons of the candidate hot word. The acoustic vector group may include one frame-level speech vector, or it may include multiple frame-level speech vectors, etc. This application does not limit this.

[0113] In addition, local alignment can be understood as matching local acoustic segments according to the text lexical length of candidate hot words based on the acoustic boundary, that is, matching the local frame-level speech vector corresponding to the acoustic boundary, thereby establishing a weakly supervised alignment relationship between speech frames and lexical positions.

[0114] Understandably, the acoustic boundary sequence can be divided according to the text lexical length of the candidate hot words to obtain one or more acoustic vector groups corresponding to each candidate hot word.

[0115] For example, if there are 6 acoustic boundaries in the acoustic boundary sequence, and the text byte length of a candidate hot word is 3, then the acoustic boundaries in the acoustic boundary sequence can be divided into an acoustic vector group of every 3. For example, acoustic boundary 1, acoustic boundary 2, and acoustic boundary 3 can be used as acoustic vector group 1; acoustic boundary 2, acoustic boundary 3, and acoustic boundary 4 can be used as acoustic vector group 2; acoustic boundary 3, acoustic boundary 4, and acoustic boundary 5 can be used as acoustic vector group 3; and acoustic boundary 4, acoustic boundary 5, and acoustic boundary 6 can be used as acoustic vector group 4.

[0116] Alternatively, a non-overlapping partitioning method can be used, which can avoid redundant calculations, reduce the number of acoustic vector groups, and lower subsequent computational costs. For example, acoustic boundary 1, acoustic boundary 2, and acoustic boundary 3 can be used as acoustic vector group 1, and acoustic boundary 4, acoustic boundary 5, and acoustic boundary 6 can be used as acoustic vector group 2.

[0117] It should be noted that the above examples are merely illustrative and should not be taken as limitations on the number of acoustic vector groups, etc., in the embodiments of this application.

[0118] Optionally, when determining the corresponding acoustic vector group based on the text byte length of the candidate hot words, if the number of remaining acoustic boundaries after grouping does not meet the text byte length, the remaining acoustic boundaries can be treated as a separate acoustic vector group. This allows for the full utilization of all acoustic vector information of the speech, avoids discarding the tail acoustic vectors, and ensures the integrity of the hot word retrieval coverage.

[0119] Therefore, in this embodiment, the length of the text lexical unit of the candidate hot words can be used as the basis for division to divide the acoustic boundary sequence, thereby dividing the frame-level speech vector into acoustic vector groups that match the candidate hot words. This achieves local temporal alignment between the frame-level speech vector and the text lexical unit of the hot words, effectively suppressing alignment offset and local aggregation inaccuracy in weakly supervised scenarios, avoiding interference from globally irrelevant information, providing a foundation for subsequent hot word similarity matching based on acoustic vector groups, and also providing support for improving the accuracy of hot word retrieval.

[0120] Optionally, the acoustic window of the candidate hot words can be determined first based on the length of the text lexical units of the candidate hot words. Then, the acoustic window can be slid across the acoustic boundary sequence, and the frame-level speech vectors corresponding to the acoustic boundaries within the acoustic window each time can be determined as the acoustic vector group of the candidate hot words.

[0121] In this context, an acoustic vector group can be understood as a combination of frame-level speech vectors corresponding to the acoustic boundaries corresponding to the text lexical length of a candidate hot word. Each acoustic boundary can correspond to one or more frame-level speech vectors, so an acoustic vector group can typically include multiple frame-level speech vectors. For example, if the text lexical length of candidate hot word 1 is 1, and the number of acoustic boundaries within the acoustic window is also 1, and the acoustic boundaries within the acoustic window in previous iterations are acoustic boundary 1, acoustic boundary 2, and acoustic boundary 3, then the frame-level speech vectors corresponding to acoustic boundary 1 can be defined as acoustic vector group 1, the frame-level speech vectors corresponding to acoustic boundary 2 can be defined as acoustic vector group 2, and the frame-level speech vectors corresponding to acoustic boundary 3 can be defined as acoustic vector group 3. When the text word length of the candidate hot words is 2 and the number of acoustic boundaries within the acoustic window is also 2, if the acoustic boundaries within the acoustic window in each iteration are: acoustic boundary 1 and acoustic boundary 2, acoustic boundary 2 and acoustic boundary 3, acoustic boundary 3 and acoustic boundary 4, then the frame-level speech vectors corresponding to acoustic boundary 1 and acoustic boundary 2 can be determined as acoustic vector group 1, the frame-level speech vectors corresponding to acoustic boundary 2 and acoustic boundary 3 can be determined as acoustic vector group 2, the frame-level speech vectors corresponding to acoustic boundary 3 and acoustic boundary 4 can be determined as acoustic vector group 3, and so on. This application does not limit this.

[0122] Understandably, the text element length of candidate hot words can be determined as the acoustic window width of the candidate hot words. For example, if the text element length of a candidate hot word is 4, then the acoustic window width of the candidate hot word is also 4. Alternatively, the text element length of the candidate hot words can be processed according to a certain transformation ratio and then determined as the acoustic window width of the candidate hot words. For example, if the text element length of a candidate hot word is 4, and the transformation ratio is 1.5, then the acoustic window width can be determined to be 6. Or, if the product of the transformation ratio and the text element length is not an integer, then the product can be rounded, such as by rounding to the nearest integer, rounding up, etc. The transformation ratio can be a pre-set value, or it can be adjusted as needed. Thus, by using an acoustic window width that is consistent with or proportionally corresponds to the text element length of the candidate hot words, both short and long hot words can obtain a local aggregation range that matches their own length.

[0123] For example, if there are 7 acoustic boundaries in the acoustic boundary sequence, corresponding to frames 3, 5, 9, 14, 19, 22, and 25, and if the text word length of candidate hot word 1 is 5, and the corresponding acoustic window width is also 5, then by sliding an acoustic window of width 5 across the acoustic boundary sequence, the determined acoustic vector groups for the candidate hot word can be acoustic vector group 1, acoustic vector group 2, and acoustic vector group 3, where:

[0124] Acoustic vector group 1 corresponds to the acoustic boundaries of frames 3, 5, 9, 14, and 19. That is, acoustic vector group 1 includes all frame-level speech vectors from frame 1 to frame 19.

[0125] Acoustic vector group 2 corresponds to the acoustic boundaries of frames 5, 9, 14, 19, and 22. That is, acoustic vector group 2 includes all frame-level speech vectors from frame 4 to frame 22.

[0126] Acoustic vector group 3 corresponds to the acoustic boundaries of frames 9, 14, 19, 22, and 25. That is, acoustic vector group 3 includes all frame-level speech vectors from frame 6 to frame 25.

[0127] It should be noted that the above examples are merely illustrative and should not be taken as limitations on the text lexical length, acoustic boundary sequence, acoustic vector group, etc. of candidate hot words in the embodiments of this application.

[0128] Therefore, in this embodiment, the acoustic window can be determined based on the length of the candidate hot word text lexical, and the acoustic boundary sequence can be divided based on the sliding of the window to obtain an acoustic vector group. This can comprehensively cover the possible matching positions of the candidate hot words in the speech, avoiding missed detections. At the same time, the width of the acoustic window is matched with the number of text vectors, which can effectively ensure the temporal correspondence between the frame-level speech vector and the hot text lexical in local segments. This effectively suppresses alignment offset and local aggregation inaccuracy in weakly supervised scenarios, avoids interference from globally irrelevant information, and improves the comprehensiveness and accuracy of hot word matching.

[0129] Optionally, the acoustic window can be slid across the acoustic boundary sequence to determine the acoustic frame span corresponding to the acoustic boundary at multiple acoustic window positions. Then, the frame-level speech vectors corresponding to each acoustic frame span within the acoustic window can be determined as the acoustic vector group of candidate hot words.

[0130] The acoustic frame span can be one or more frame-level speech vectors corresponding to the acoustic boundary.

[0131] For example, if there are 6 acoustic boundaries in the acoustic boundary sequence, the time frames corresponding to each acoustic boundary are frames 3, 5, 9, 14, and 20. If the width of the acoustic window is 3, based on the sliding of the acoustic window on the acoustic boundary sequence, the acoustic boundaries at multiple acoustic window positions can be determined as: frames 3, 5, and 9; frames 5, 9, and 14; and frames 9, 14, and 20. The corresponding acoustic frame spans are: frames 1 to 9; frames 4 to 14; and frames 6 to 20. Then, frames 1 to 9 corresponding to acoustic frame span 1 can be determined as acoustic vector group 1 of candidate hot words, frames 4 to 14 corresponding to acoustic frame span 2 can be determined as acoustic vector group 2 of candidate hot words, and frames 6 to 20 corresponding to acoustic frame span 3 can be determined as acoustic vector group 3 of candidate hot words.

[0132] It should be noted that the above examples are merely illustrative and should not be construed as limiting the acoustic frame span, acoustic vector group, etc., in the embodiments of this application.

[0133] Therefore, in this embodiment, unlike the method of compressing the entire speech into a single global vector and then matching it with candidate hot words, the corresponding acoustic window is determined based on the length of the text lexical unit of the candidate hot words, and then slid on the acoustic boundary sequence. The acoustic frame span corresponding to the acoustic boundary at each acoustic window position is determined, which can avoid matching deviation caused by length mismatch. The frame-level speech vector corresponding to the acoustic frame span within the acoustic window is determined as the acoustic vector group of the candidate hot words. This can focus on local speech features, eliminate globally irrelevant background information and noise interference, prevent short hot word acoustic cues from being submerged, and thus comprehensively cover the possible matching positions of candidate hot words in speech, avoiding missed detection. At the same time, the width of the acoustic window is matched with the number of text vectors, which can effectively ensure the temporal correspondence between the frame-level speech vector and the hot text lexical unit in the local segment, effectively suppressing alignment offset and local aggregation inaccuracy in weak supervision scenarios, avoiding interference from globally irrelevant information, and improving the comprehensiveness and accuracy of hot word matching.

[0134] Step 108: Based on the text semantic vector and frame-level speech vector of each candidate hot word, the candidate hot words are filtered to obtain the context hot words of the speech to be identified.

[0135] The frame-level speech vector of the acoustic vector group can be understood as the frame-level speech vector of all speech frames included in the acoustic vector group, but this application does not limit this.

[0136] In addition, the context hot words of the speech to be recognized can be understood as the hot words with a high degree of matching with the speech to be recognized determined from the candidate hot words after matching the text semantic vector with the frame-level speech vector. There can be one or more hot words, and this application does not limit them.

[0137] Understandably, the similarity between the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group can be calculated separately. Based on this similarity, the candidate hot words are filtered, and those meeting the filtering criteria are identified as context hot words for the speech to be recognized. This effectively locates the context hot words for the speech to be recognized, reduces the false recognition rate, and improves the accuracy of context hot words. Then, hot word prompts are generated based on these context hot words. Therefore, by performing local similarity calculations on acoustic vector groups segmented by acoustic boundary sequences, the method of global vector matching across the entire sentence is abandoned. The matching range is limited to local speech segments that match the hot words, effectively avoiding interference from irrelevant background information and preventing the dilution of local acoustic cues for short entities and long-tail hot words. This significantly improves the accuracy of hot word retrieval under weak supervision, ensuring the accuracy of context hot word selection from the source.

[0138] Optionally, there are various ways to calculate similarity, such as cosine similarity, Euclidean distance similarity, etc. This application does not limit the comparison.

[0139] Therefore, in this embodiment, based on the text lexical length of the candidate hot words, a corresponding acoustic window is determined. Then, the window slides along the acoustic boundary sequence, and the similarity between the frame-level speech vector and the text semantic vector of the candidate hot words within each acoustic window is locally aggregated. This allows the retrieval process to focus on the local speech segments where the hot words actually appear, effectively improving the accuracy and reliability of hot word matching and retrieval, and consequently improving the accuracy of speech recognition. In contrast, in related technologies, during the matching and filtering of entire sentences with candidate hot words, the entire sentence often contains a large amount of background information unrelated to the candidate hot words. This dilutes the local acoustic information of short hot words within the entire sentence, reducing the accuracy of hot word matching and consequently lowering the accuracy of subsequent speech recognition.

[0140] Optionally, hot word prompts can be generated first based on the context hot words of the speech to be recognized, and then the speech recognition model can be used to recognize the speech based on the hot word prompts to generate the recognized text.

[0141] The hot word prompts can be generated by concatenating the context hot words of the speech to be recognized, or they can be generated by combining the context hot words of the speech to be recognized with the prompt template. For example, the context hot words of the speech to be recognized can be filled into the prompt template, and the filled result can be used as the hot word prompts, etc. This application does not limit this.

[0142] Therefore, in this embodiment of the application, after generating hot word prompts, the context hot words of the speech to be recognized and the speech to be recognized can be input into the speech recognition model, so that the speech to be recognized can be processed by the speech recognition model to generate recognized text. Thus, the hot word prompts generated based on the highly accurate context hot words provide targeted context bias information for speech recognition, which makes the model more inclined to output content that matches the context hot words when decoding. This can effectively reduce the misrecognition rate of proper nouns, personal names, place names and other entities, and greatly improve the accuracy of speech recognition.

[0143] In the above method for determining contextual hot words, the speech to be recognized and the candidate hot words in the candidate word list can be processed separately to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step. Then, the frame-level speech vector can be processed by temporal accumulation of boundary contributions to determine the acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized. Then, according to the text lexical length of the candidate hot words, the candidate hot words are locally aligned with the acoustic boundaries on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words. After that, the contextual hot words of the speech to be recognized can be obtained by filtering the candidate hot words according to the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group. Therefore, by performing temporal accumulation processing on the boundary contributions of the frame-level speech vectors of the speech to be recognized, the corresponding acoustic boundary sequence can be determined. Then, based on the text word length of the candidate hot words, local alignment between the candidate hot words and the frame-level speech vectors is performed on the acoustic boundary sequence. This allows the frame-level speech vectors of the speech to be recognized to be segmented into acoustic vector groups matching the candidate hot words, thus ensuring the temporal correspondence between the frame-level speech vectors and the text words of the hot words in local segments. This effectively suppresses alignment offset and local aggregation inaccuracies in weakly supervised scenarios, avoids interference from globally irrelevant information, and also avoids... The local acoustic cues of short entities are submerged in background speech and irrelevant content, resulting in sparse representation. Then, based on the text semantic vector of the candidate hot words and the frame-level speech vector of the acoustic vector group, the candidate hot words are matched and filtered only within the acoustic vector group. This avoids noise interference when comparing the whole sentence speech with the global hot words. It does not rely on additional language models and complex finite graphs, thus improving the accuracy of context hot word recognition of the speech to be recognized. In turn, it improves the accuracy and reliability of speech recognition based on context hot words of the speech to be recognized, and has good versatility and transferability.

[0144] Compared with the prior art, the embodiments of this application have significant technical effects:

[0145] Compared with related technologies, in this embodiment, the determined contextual hot words do not contain background information unrelated to the candidate hot words. Therefore, the accuracy and matching degree of these contextual hot words with the speech to be recognized are both high, resulting in higher accuracy when performing speech recognition based on these contextual hot words. In contrast, the existing technologies that directly perform global matching and sorting of the entire speech sentence with candidate hot words often result in low accuracy of the selected candidate hot words because the entire speech sentence often contains a large amount of background information unrelated to the candidate hot words. Using these candidate hot words for speech recognition also leads to low accuracy in speech recognition.

[0146] Compared to another related technology, in this application's embodiment, contextual hot words are determined based on candidate hot words, and then hot word suggestions are generated. These hot word suggestions can be integrated into the large language model decoding process without relying on additional language models, finite state graphs, etc. When migrating and adapting to different large speech models, no significant changes to the model architecture are required, reducing cross-model adaptation costs. In contrast, related prior art relies on additional language models, finite state graphs, etc., resulting in complex system coupling and high adaptation costs when migrating to different large speech models.

[0147] In one exemplary embodiment, such as Figure 2 As shown, step 104 includes steps 202 to 204. Wherein:

[0148] Step 202: Predict the boundary contribution value of each frame-level speech vector.

[0149] Specifically, the boundary contribution value of each frame-level speech vector can be predicted using the target network model.

[0150] Optionally, each frame-level speech vector can be input into the target network model to obtain the boundary contribution value of each frame-level speech vector.

[0151] The target network model can be a trained neural network model. Each frame-level speech vector is input into the target network model, and after processing by the target network model, the boundary contribution value of each frame-level speech vector can be output.

[0152] Optionally, when processing the frame-level speech vectors in the input, the target network model can perform operations such as one-dimensional convolution, normalization, non-linear activation, and linear transformation on the frame-level speech vectors, and then determine the integral weight, i.e., the boundary contribution value, corresponding to each frame-level speech vector; or it can use other methods to process the frame-level speech vectors to determine the boundary contribution value, etc., which are not limited in this application.

[0153] In addition, the boundary contribution value can be understood as the integral weight corresponding to each frame-level speech vector, that is, the contribution ratio of each frame-level speech vector to the formation of the next word boundary. It can be used to accumulate frame by frame. For example, when the accumulation result reaches the threshold, an event is triggered, and the current frame can be marked as an acoustic boundary.

[0154] Step 204: The boundary contribution values ​​of each frame-level speech vector are accumulated frame by frame in chronological order, and the acoustic boundary sequence is determined based on the accumulated values.

[0155] In this process, the boundary contribution values ​​of each frame-level speech vector can be accumulated frame by frame in chronological order to obtain the corresponding cumulative value. Then, the acoustic boundary sequence can be determined based on the relationship between the cumulative value and the preset threshold.

[0156] The preset threshold serves as a condition for determining whether the cumulative value reaches the acoustic boundary. If the cumulative value is less than the threshold, the current frame is not considered an acoustic boundary; if the cumulative value is greater than or equal to the threshold, the current frame is considered an acoustic boundary. Therefore, the acoustic boundary can be determined based on the relationship between the cumulative value and the threshold, and an acoustic boundary sequence can be formed accordingly.

[0157] Therefore, in this embodiment, there is no need to rely on manual timestamp annotation. By accumulating the boundary contribution values ​​of frame-level speech vectors frame by frame, a boundary event is triggered when the accumulated value reaches a preset trigger threshold, forming a monotonically increasing acoustic boundary sequence along the time axis. This establishes a weakly supervised temporal correspondence between speech frames and word-level positions, significantly reducing annotation costs and effectively avoiding the drawbacks of weakly supervised schemes lacking explicit temporal constraints. Furthermore, by accumulating frame by frame in chronological order, the generated acoustic boundary sequence possesses a monotonically increasing temporal characteristic, providing stable explicit temporal constraints for subsequent local alignment of speech and hot words, effectively suppressing drift problems. In contrast, related technologies rely on global representation alignment, leading to information dilution, and the lack of an explicit monotonically constrained attention mechanism results in unstable localization, both of which reduce accuracy during subsequent speech and hot word alignment. The explicit monotonically increasing acoustic boundary sequence provided in this application can establish a weakly supervised temporal correspondence between speech frames and word-level positions, providing a foundation for subsequent hot word matching and alignment based on text word length.

[0158] Optionally, the boundary contribution values ​​corresponding to the speech vectors of each frame can be accumulated in chronological order to obtain a cumulative value. Then, when the cumulative value reaches a preset threshold, a boundary event is triggered. The boundary event is used to indicate that the current frame that triggers the boundary event is determined as the acoustic boundary. After the boundary event is triggered, the boundary contribution values ​​of subsequent frames of the current frame can be accumulated based on the difference between the cumulative value and the preset threshold until the traversal is completed, resulting in an acoustic boundary sequence arranged in chronological order.

[0159] The preset threshold can be a pre-set value, such as 1, or other values, or it can be adjusted according to actual needs, etc. This application does not limit it in this regard.

[0160] In addition, the boundary event can be the event that the cumulative value reaches a preset threshold, which can be used as the trigger signal for the acoustic boundary. It can be used to indicate that the current frame that triggers the boundary event is determined as the acoustic boundary. For example, if the cumulative value of the boundary contribution reaches the preset threshold in frame 10, then frame 10 can be determined as the acoustic boundary, etc. This application does not limit this.

[0161] In addition, the difference between the cumulative value and the preset threshold can be understood as the difference between the cumulative value and the preset threshold, or the difference can be numerically transformed, etc. This application does not limit this.

[0162] Understandably, the boundary contribution values ​​corresponding to the speech vectors of each frame can be accumulated sequentially in chronological order to obtain the accumulated value. Each time a boundary contribution value is added, it is checked whether the current accumulated value has reached a threshold. If the current accumulated value is less than the threshold, it can be considered that the acoustic boundary has not been reached, and accumulation can continue for the next frame. If the current accumulated value reaches the preset threshold, it can be considered that enough boundary contribution value has been accumulated to form the boundary of the next lexical unit, triggering a boundary event, and the current frame that triggered the boundary event is determined as the acoustic boundary. Then, the accumulated value can be subtracted from the preset threshold to obtain the updated accumulated value, and the boundary contribution values ​​of subsequent frames can be accumulated based on the updated accumulated value until the boundary contribution value of each frame-level speech vector has been accumulated, indicating that the traversal is complete. At this point, all acoustic boundaries in the cyclic traversal process can be arranged in chronological order to obtain the final acoustic boundary sequence.

[0163] For example, a speech to be recognized is segmented into frames 1 to 10 arranged in chronological order, and the boundary contribution values ​​of the frame-level speech vector of each frame can be 0.2, 0.3, 0.4, 0.1, 0.3, 0.3, 0.5, 0.1, 0.3, and 0.8, respectively.

[0164] If the threshold is 1, the boundary contribution values ​​of the frame-level speech vectors are accumulated starting from frame 1. At frame 4, the accumulated result equals the threshold 1, thus triggering the first boundary event and identifying frame 4 as the first acoustic boundary. Next, the accumulated value is subtracted from the threshold, and the boundary contribution value of frame 5 is accumulated to determine the new initial frame. The accumulation of the boundary contribution values ​​of the frame-level speech vectors is recalculated. At frame 7, the accumulated result is 1.1, which is greater than the threshold 1, thus triggering the second boundary event and identifying frame 7 as the second acoustic boundary. Then, the accumulated value is subtracted from the threshold, and the boundary contribution value of frame 8 is accumulated. At frame 10, the accumulated result is 1.3, which is greater than the threshold 1, thus triggering the third boundary event and identifying frame 10 as the third acoustic boundary. Arranging the acoustic boundaries in chronological order yields the acoustic boundary sequence [frame 4, frame 7, frame 10].

[0165] It should be noted that the above examples are merely illustrative and should not be construed as limiting the acoustic boundaries, thresholds, acoustic boundary sequences, etc., in the embodiments of this application.

[0166] Therefore, in this embodiment, by accumulating boundary contribution values ​​frame by frame and determining acoustic boundaries in combination with preset thresholds, an acoustic boundary sequence is generated. This establishes a temporal correlation between speech frames and word-level positions under weak supervision. At the same time, by using a time-ordered accumulation rule and resetting the accumulated value after triggering, fine-grained frame-level division of continuous speech can be achieved, capturing the natural acoustic boundaries in speech. This results in the generated acoustic boundary sequence having a monotonically increasing explicit temporal feature, providing strong temporal constraints for the subsequent local alignment of hot words and speech.

[0167] In this embodiment, no manual timestamp annotation is required. By accumulating frame-by-frame contribution values ​​at the frame level, an acoustic boundary sequence can be generated. Under weak supervision, a temporal correspondence between speech frames and word-level positions is established, significantly reducing annotation costs. At the same time, it effectively avoids the drawback of weak supervision schemes lacking explicit temporal constraints. Furthermore, by accumulating frame-by-frame in chronological order, the generated acoustic boundary sequence possesses a monotonically increasing temporal characteristic, providing stable explicit temporal constraints for subsequent local alignment of speech and hot words, effectively suppressing the drift problem.

[0168] Compared with existing technologies, the embodiments of this application have significant technical advantages: They do not rely on manual timestamp annotation; instead, by accumulating the boundary contribution values ​​of frame-level speech vectors frame by frame, a monotonically increasing acoustic boundary sequence along the time axis is formed. This establishes a weakly supervised temporal correspondence between speech frames and word-level positions, resulting in higher accuracy in hot word localization. In contrast, related technologies rely on global representation alignment, which leads to information dilution, and the lack of an explicit monotonic constraint attention mechanism results in unstable hot word localization, leading to lower accuracy in subsequent hot word selection.

[0169] In one exemplary embodiment, such as Figure 3 As shown, step 108 includes steps 302 to 304. Wherein:

[0170] Step 302: Determine the candidate score of the candidate hot word based on the text semantic vector of the candidate hot word and the frame-level speech vector of the acoustic vector group.

[0171] Each candidate hot word can correspond to one or more acoustic vector groups. The similarity between the text semantic vector of each candidate hot word and the frame-level speech vector of its corresponding acoustic vector group can be determined first. Then, a candidate score for the candidate hot word can be determined based on this similarity. For example, the similarity can be directly used as the candidate score, or the similarity can be numerically transformed, such as through normalization or scaling, and the result can be used as the candidate score. This application does not limit this approach.

[0172] Optionally, when calculating the similarity between text semantic vectors and frame-level speech vectors, they can be mapped to a unified vector space. For example, the text semantic vectors and frame-level speech vectors can be transformed and normalized separately through projection layers to perform cross-modal similarity calculations, etc. This application does not limit this.

[0173] Optionally, the frame-level similarity between the text semantic vector of the candidate hot word and the corresponding acoustic vector group of each frame-level speech vector can be determined first. Then, the frame-level similarity of each acoustic vector group can be aggregated to determine the aggregated similarity of each acoustic vector group. Finally, the highest aggregated similarity of each acoustic vector group can be determined as the candidate score of the corresponding candidate hot word.

[0174] The aggregation process can be averaging, max pooling, or other processing methods, and this application does not limit the specific processing methods.

[0175] In addition, an acoustic vector group can include multiple frame-level speech vectors. Frame-level similarity can be calculated based on the text semantic vector of the candidate hot word and each frame-level speech vector in each candidate vector group. Then, multiple frame-level similarities belonging to the same acoustic vector group can be aggregated, and the aggregation result can be determined as the aggregated similarity of the acoustic vector group. Then, for multiple acoustic vector groups corresponding to the same candidate vector, the highest aggregated similarity among the multiple acoustic vector groups can be determined as the candidate score of the candidate hot word.

[0176] For example, if acoustic vector group 1 for candidate hot word 1 includes frame-level speech vectors from frames 1 to 5, acoustic vector group 2 includes frame-level speech vectors from frames 6 to 15, and acoustic vector group 3 includes frame-level speech vectors from frames 16 to 20, then the frame-level similarity between the text semantic vector of candidate hot word 1 and each frame-level speech vector from frames 1 to 5 in acoustic vector group 1 can be calculated separately. Then, the frame-level similarities of these 5 frames can be summed and averaged to obtain the aggregate similarity 1 for acoustic vector group 1. Similarly, for acoustic vector group 2, the frame-level similarity between the text semantic vector of candidate hot word 1 and the frame-level speech vectors from frames 6 to 15 in acoustic vector group 2 can be calculated separately. Then, the frame-level similarities of these 10 frames can be summed and averaged to obtain the aggregate similarity 2 for acoustic vector group 2. For acoustic vector group 3, the text semantic vector of candidate hot word 1 and the frame-level speech vectors of frames 16 to 20 in acoustic vector group 3 can be used to calculate the frame-level similarity. Then, the frame-level similarities of the above 5 frames can be summed and averaged to obtain the aggregate similarity 3 of acoustic vector group 3. Then, aggregate similarity 1, aggregate similarity 2 and aggregate similarity 3 can be compared, and the highest aggregate similarity can be determined as the candidate score of candidate hot word 1. For example, if aggregate similarity 3 is greater than aggregate similarity 1 and aggregate similarity 2, then aggregate similarity 3 can be determined as the candidate score of candidate hot word.

[0177] It should be noted that the above examples are merely illustrative and should not be taken as limitations on the number of acoustic vector groups, the number of frame-level speech vectors, aggregated similarity, etc. in the embodiments of this application.

[0178] Therefore, in this embodiment, the frame-level similarity between the semantic vector of the candidate hot word text and the frame-level speech vector within the acoustic vector group can be calculated, and then aggregated to obtain the aggregated similarity of each acoustic vector group. The highest aggregated similarity is used as the candidate score of the corresponding candidate hot word. By accurately calculating the similarity between speech and text at the frame level, the method of global vector matching of the whole sentence is avoided, and the matching of hot words and speech is focused on fine-grained acoustic segments. This can effectively avoid the interference of irrelevant background frames and prevent the dilution of local acoustic cues of short entities and long-tail hot words, which greatly improves the accuracy of hot word matching under weak supervision. At the same time, by aggregating the frame-level similarity of each acoustic vector group separately and then selecting the highest value as the candidate score, the averaging drawback of global pooling is eliminated, which provides a foundation for improving the stability and accuracy of hot word localization in the future.

[0179] Optionally, to facilitate the calculation of aggregated similarity and improve efficiency, a similarity matrix between frame-level speech vectors and candidate hot words can be constructed. For example, for each candidate hot word, the similarity between the text semantic vector of the candidate hot word and each frame-level speech vector can be calculated separately, thereby obtaining a two-dimensional similarity matrix composed of the speech frame dimension and the candidate hot word dimension, etc. This application does not limit this.

[0180] Optionally, in determining the aggregation similarity of each acoustic vector group, the acoustic speech vector corresponding to the acoustic vector group can be calculated first. For example, all frame-level speech vectors contained in the acoustic vector group can be aggregated to obtain the acoustic speech vector corresponding to the acoustic vector. The aggregation process can include mean pooling, max pooling, etc., which are not limited in this application.

[0181] Then, the aggregation similarity between the text semantic vector and each acoustic speech vector can be calculated separately, and so on.

[0182] It should be noted that the above examples are merely illustrative and should not be construed as limiting the methods for determining the aggregation similarity of acoustic vector groups in the embodiments of this application.

[0183] Step 304: Among the candidate hot words, the candidate hot words whose candidate scores meet the screening criteria are determined as the context hot words of the speech to be recognized.

[0184] The filtering conditions can be varied, such as ranking based on merit or filtering based on score thresholds, and this application does not limit them.

[0185] Optionally, if the candidate hot words are sorted in descending order of candidate scores, the top target number of candidate hot words can be determined as the context hot words of the speech to be recognized.

[0186] The target number can be pre-set, such as 3, 5, 10, etc., or it can be adjusted according to actual needs, etc. This application does not limit this.

[0187] For example, when the target number is 6, the candidate hot words can be sorted in descending order of candidate scores from high to low. Then, the first 6 candidate hot words in the sequence can be determined as the context hot words of the speech to be recognized, etc. This application does not limit this.

[0188] It is understandable that the higher the candidate score, the higher the correlation between the candidate hot word and the frame-level speech vector of the acoustic vector group. Therefore, in this embodiment, candidate hot words can be screened based on a preset score threshold.

[0189] Optionally, candidate hot words with scores greater than a preset score threshold can be identified as context hot words for the speech to be recognized. This allows for the selection of highly relevant context hot words for the speech to be recognized from the candidate hot words using the preset score threshold. Therefore, in this embodiment, candidate hot words whose scores meet the selection criteria can be identified as context hot words for the speech to be recognized. This ensures that the selected context hot words for the speech to be recognized have a high degree of correlation with the frame-level speech vector and the speech to be recognized, thereby improving the accuracy and reliability of determining the context hot words for the speech to be recognized.

[0190] Therefore, in this embodiment of the application, since the candidate score of the candidate hot word can reflect the degree of matching between the candidate hot word and the local acoustic segment of the speech to be identified, the candidate hot word can be screened according to the candidate score, which can effectively filter out noisy candidate words with low matching degree, so that the final context hot words have a high correlation with the speech acoustic features, effectively improve the effectiveness of context prompts, and avoid introducing erroneous biases for subsequent recognition as much as possible, thus providing a foundation for improving the accuracy of speech recognition.

[0191] In this embodiment, the candidate scores of candidate hot words can be determined first based on the text semantic vectors and frame-level speech vectors of the acoustic vector groups. Then, among the candidate hot words, those whose candidate scores meet the screening criteria are determined as the context hot words of the speech to be recognized. This achieves efficient screening and accurate positioning of candidate hot words, effectively improving the accuracy of speech hot word matching while reducing the computational load of subsequent processing, thus ensuring the accuracy of subsequent speech recognition.

[0192] Compared with existing technologies, the embodiments of this application have significant technical advantages: since the determined contextual hot words do not contain background information unrelated to the candidate hot words, the accuracy and matching degree of the contextual hot words with the speech to be recognized are both high, resulting in high accuracy when performing speech recognition based on the contextual hot words. In contrast, the existing technologies that directly perform global matching and sorting of the entire sentence speech with the candidate hot words often result in low accuracy of the selected candidate hot words because the entire sentence speech often contains a large amount of background information unrelated to the candidate hot words. Therefore, when using these candidate hot words for speech recognition, the accuracy of speech recognition is also low.

[0193] In one embodiment, such as Figure 4 As shown, a method for training a target model is provided. This embodiment illustrates the method by applying it to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes steps 402 to 412. Wherein:

[0194] Step 402: Obtain the training dataset. The training dataset includes multiple sample data. Each sample data includes at least a whole sentence speech data, a whole sentence text data corresponding to the whole sentence speech data, and the target hot word text contained in the whole sentence text data.

[0195] The whole-sentence speech data can be the speech of a complete sentence containing the hot words, that is, the complete speech of the hot words in their context. The whole-sentence text data is the complete text corresponding to the whole-sentence speech data. The target hot word text can be the hot word text contained in the whole-sentence text data.

[0196] Step 404: Process the whole sentence speech data, whole sentence text data and target hot word text respectively to determine the training frame-level speech vector, whole sentence speech vector, whole sentence text vector and hot word text vector.

[0197] Among them, feature vector processing can be performed on the whole sentence speech data to obtain the whole sentence speech vector, which can be used to represent the overall speech features of the whole sentence speech. In addition, the whole sentence speech data can be segmented into frames and feature encoded to obtain multiple training frame-level speech vectors arranged in time order, which can be used to represent the fine-grained speech features of speech in single frame time sequence.

[0198] In addition, a text encoder can be used to perform semantic encoding on the whole sentence text data and the target hot word text respectively, to obtain the whole sentence text vector corresponding to the whole sentence text data and the hot word text vector corresponding to the target hot word text. The whole sentence text vector can be used to represent the complete semantic information of the text, and the hot word text vector can be used to represent the semantic information of the hot word.

[0199] Step 406: Input the training frame-level speech vectors into the initial network model to obtain the training boundary contribution value.

[0200] The initial network model can be an untrained neural network model. Each training frame-level speech vector is input into this initial network model, and after processing, it outputs the training boundary contribution value of each training frame-level speech vector. This initial network model may include convolutional layers, normalization layers, non-linear activation layers, and linear transformation layers connected end-to-end. Optionally, the convolutional layers can be one-dimensional convolutional layers. This embodiment does not limit the specific structure of the initial network model; the structure can be determined based on the actual application scenario.

[0201] Optionally, the process by which the terminal processes the training frame-level speech vector through the initial network model can be as follows: sequentially performing operations such as one-dimensional convolution, normalization, non-linear activation, and linear transformation on the training frame-level speech vector to obtain the output result of the initial network model, and determining the output result corresponding to the training frame-level speech vector as the training boundary contribution value corresponding to the training frame-level speech vector; or other methods can be used to process the training frame-level speech vector to determine the training boundary contribution value, etc., and this application does not limit this.

[0202] Step 408: Perform time-series accumulation processing on the training boundary contribution value to obtain the training acoustic boundary.

[0203] The training accumulation result can be the value obtained by accumulating the training boundary contribution values ​​frame by frame in chronological order. The training acoustic boundary can be understood as the temporal boundary point used to divide frame-level speech vectors. It can be used to indicate the boundary position of frame-level speech vectors and can be represented by the index of the speech frame.

[0204] Optionally, the training boundary contribution values ​​of the training frame-level speech vectors can be accumulated according to the time sequence of each training frame-level speech vector to obtain the training accumulation result. Then, when the training accumulation result reaches a threshold, a boundary event is triggered, and the current frame that triggers the boundary event is determined as the training acoustic boundary. After the boundary event is triggered, the accumulated result can be subtracted from the preset threshold, and the boundary contribution values ​​of subsequent frames can continue to be accumulated until the traversal is completed, resulting in multiple training acoustic boundaries.

[0205] Step 410: Determine the hot word segment speech vector corresponding to the target hot word from the training frame-level speech vector based on the training acoustic boundary.

[0206] This method involves dividing the training frame-level speech vectors into multiple training vector groups using training acoustic boundaries. Each training vector group is then compared to the target hot word's text vector for similarity calculation. Based on the similarity value, the corresponding hot word segment speech vector is determined. For example, the training vector group with the highest similarity can be identified as the hot word segment speech vector corresponding to the target hot word, or other methods can be used to determine the hot word segment speech vector. This application does not limit the specific methods used.

[0207] Step 412: Based on the whole sentence speech vector, whole sentence text vector, hot word text vector, hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary, train the initial network model to generate the target network model.

[0208] Among these methods, the training frame-level speech vectors can be segmented and aggregated based on the training acoustic boundaries to obtain the corresponding number of training lexical units.

[0209] Optionally, a data constraint loss can be determined based on the difference between the number of training lexical units and the length of the entire sentence text lexical units. A global contrast loss can be determined based on the semantic difference between the entire sentence speech vector and the entire sentence text vector. A local contrast loss can then be determined based on the semantic difference between the hot word segment speech vector and the hot word text vector. Finally, a comprehensive loss is determined based on the data constraint loss, global contrast loss, and local contrast loss. The initial network model can then be trained using this comprehensive loss to generate the target network model.

[0210] Among them, the data constraint loss can be used to constrain the difference between the number of training lexical units and the length of the whole sentence text lexical units; the global contrast loss can be used to constrain the consistency between the whole sentence speech vector and the whole sentence text vector in global semantics; the local contrast loss can be used to constrain the consistency between the hot word segment speech vector and the hot word text vector in local hot word semantics; and the comprehensive loss value can be used to optimize the acoustic boundary detection and speech-text semantic matching performance of the overall model.

[0211] In addition, there are several ways to calculate data constraint loss. For example, the L1 loss function or L2 loss function can be used to calculate the difference between the number of training words and the length of words in the whole sentence text to obtain the data constraint loss.

[0212] In addition, there are several ways to determine the global contrast loss. For example, you can use Information Noise Contrastive Estimation Loss (InfoNCE Loss) or cross-entropy loss to calculate the semantic difference between the whole sentence speech vector and the whole sentence text vector to obtain the global contrast loss.

[0213] In addition, there are several ways to determine the local contrast loss. For example, InfoNCE Loss or cross-entropy loss can be used to calculate the semantic difference between the speech vector of hot word segments and the text vector of hot words to obtain the local contrast loss.

[0214] The data constraint loss, global contrast loss, and local contrast loss can then be fused to obtain a comprehensive loss value. For example, the data constraint loss, global contrast loss, and local contrast loss can be directly summed to obtain the comprehensive loss value; or the data constraint loss, global contrast loss, and local contrast loss can be weighted and fused to obtain the comprehensive loss value; or any other acceptable method can be used for fusion to obtain the comprehensive loss value, etc. This application does not limit this approach.

[0215] Therefore, in this embodiment, by fusing global contrast loss, local contrast loss, and quantity constraint loss, a comprehensive loss value is obtained. The model can be trained in stages or by joint training. The model can first be allowed to stably learn local alignment capabilities, and then the overall performance of the model can be improved by joint optimization of multiple losses. Through the synergistic optimization of the three types of losses, the model bias problem caused by training with a single loss can be effectively avoided, and the model can have stable boundary generation, complete semantic understanding of whole sentences, and accurate local hot word matching capabilities, thereby significantly improving the generalization and robustness of the model in weakly supervised scenarios.

[0216] Understandably, after obtaining the comprehensive loss value, the parameters of the initial network model can be iteratively updated using the backpropagation algorithm. Optimization strategies such as gradient descent can be employed to adjust the model parameters, gradually reducing the loss value. This allows the model's predictions to continuously approach the true labels, and training continues until the model performance stabilizes or a preset iteration stopping condition is met, thus terminating the training process. This ensures that the final trained target network model possesses reliable performance, accurately predicts the boundary contribution values ​​of speech frames, improves the accuracy of acoustic boundary segmentation, and provides a solid foundation for subsequent hot word retrieval speech recognition.

[0217] It is understandable that in the aforementioned method for determining contextual hot words, the target network model generated in this application can be used to predict the boundary contribution value of each frame-level speech vector during the prediction process. This allows the boundary contribution value of each frame-level speech vector to be obtained. The accuracy of this boundary contribution value is relatively high, and the acoustic boundary determined based on the boundary contribution value of each frame-level speech vector is also more reliable, thus providing a foundation for improving the accuracy of hot word retrieval and speech recognition in the future.

[0218] In this embodiment, a training dataset is first acquired, comprising multiple sample data points. Each sample data point includes at least a complete sentence of speech data, the corresponding complete sentence of text data, and the target hot word text contained within the complete sentence of text data. Then, the training frame-level speech vectors are input into the initial network model to obtain training boundary contribution values. These contribution values ​​are then time-series accumulated to obtain the training acoustic boundary. Based on the training acoustic boundary, the hot word segment speech vectors corresponding to the target hot words are determined from the training frame-level speech vectors. Finally, based on the complete sentence speech vectors, complete sentence text vectors, hot word text vectors, hot word segment speech vectors, and the number of training lexical units corresponding to the training acoustic boundary, the initial network model is trained to generate the target network model. Thus, through a training method that combines multi-feature collaborative processing and multi-dimensional loss joint optimization, the target network model can effectively learn the correlation between speech frame-level features and text semantic features, effectively improving the accuracy and reliability of acoustic boundary detection and providing a solid foundation for subsequent hot word retrieval speech recognition.

[0219] The method for determining contextual hot words provided in this application can be applied to any speech recognition scenario. The following section, in conjunction with a retrieval-enhanced contextual speech recognition system, further explains the process of determining contextual hot words provided in this application.

[0220] The retrieval-enhanced contextual speech recognition system may include a retrieval model and a large speech model. The retrieval model may include a feature encoding module, a continuous integral-triggered alignment module, a local matching module, a candidate ranking module, and a contextual cue recognition module. The feature encoding module may include a feature extraction module, a speech encoding module, and a text encoding module.

[0221] The system comprises several modules: a feature extraction module to extract acoustic features of the speech to be recognized; a speech coding module to encode these features, resulting in frame-level speech vectors arranged chronologically; a text coding module to encode candidate hot words, generating text semantic vectors corresponding to each candidate hot word; and a continuous integration-triggered alignment module to predict the integral weight (i.e., boundary contribution value) of each frame-level speech vector based on the output of the speech coding module, and to accumulate these boundary contribution values ​​chronologically. When the accumulated value reaches a preset trigger threshold, a trigger event is output, and the current frame corresponding to the trigger event is defined as an acoustic boundary. After the trigger event, the accumulated value is subtracted from the preset trigger threshold to retain the remaining accumulated value, and the boundary contribution values ​​of subsequent frames are accumulated, thus forming a monotonically increasing acoustic boundary sequence.

[0222] The local matching module can be used to determine the candidate acoustic segments, also known as acoustic vector groups, corresponding to candidate hot words based on the text lexical length and acoustic boundary sequence of the candidate hot words. Specifically, the acoustic window length corresponding to the candidate hot words can be determined first based on the text lexical length of the candidate hot words. Then, the acoustic window length is slid across the acoustic boundary sequence to determine candidate acoustic segments at multiple window positions. After that, based on the frame-level similarity between each frame-level speech vector and the text semantic vector of the candidate hot words, the frame-level similarity within the acoustic frame span corresponding to each window position is averaged to obtain a local matching branch. The maximum local matching score in each window position is determined as the candidate score of the candidate hot word, which can also be called the retrieval score.

[0223] For example, for the candidate hot word "XX Village", the text encoding module can output its corresponding text semantic vector. The local matching module can slide an acoustic window across the obtained acoustic boundary sequence based on the length of the text tokens of the candidate hot word to determine the acoustic vector group covered by each acoustic window. Then, the similarity between the text semantic vector of the candidate hot word and the frame-level speech vector of the acoustic vector group can be calculated and averaged. The maximum mean of the acoustic vector group can be used as the candidate score of the candidate hot word.

[0224] The candidate ranking module sorts candidate hot words according to their scores and outputs the top K candidate hot words as context hot words for the speech to be recognized. The context cue recognition module concatenates the context hot words of the speech to be recognized, i.e., the top K candidate hot words, to form a hot word cue text. This hot word cue text is then used as a natural language context cue input to the speech model to guide it in completing the enhanced contextual speech recognition.

[0225] The large-scale speech model can be used to process received hot word prompts and the speech to be recognized, and can output the final recognized text through prompt-based decoding. This large-scale speech model does not require changes to the network structure; it only needs to receive hot word text prompts generated from the search results to complete context-enhanced recognition. It has good modularity, pluggability, and cross-model transferability.

[0226] For example, the candidate ranking module can identify the top K candidate hot words with the highest scores as context hot words for the speech to be recognized. The context cue recognition module can then concatenate these K candidate hot words into a hot word cue text, which, along with the input speech to be recognized, is fed into the speech model. After receiving the hot word cue text, the speech model can, during the decoding process, be more inclined to output hot words that are acoustically consistent with the input speech, thereby reducing the probability of misidentifying low-frequency proper nouns as high-frequency general words and improving the accuracy of speech recognition.

[0227] It is understandable that in speech scenarios containing news broadcasts, government announcements, financial business names, or personal and place names, if the large speech model has difficulty recognizing certain low-frequency entity words, it can be combined with the contextual hot words of the speech to be recognized obtained through front-end hot word retrieval in this application, and used as contextual prompts to participate in decoding, thereby significantly improving the entity recognition effect.

[0228] Understandably, during the training of the retrieval model, the training data can include multiple sample data sets. Each sample data set includes at least a complete sentence of speech data, the corresponding complete sentence of text data, and the target hot word text contained within the complete sentence of text data. Training can employ joint optimization using multi-granularity objective functions. The training loss of the target network model can include global contrastive loss, local contrastive loss, and quantity constraint loss. The global contrastive loss constrains the consistency between the complete sentence-level speech vector and the complete sentence of text vector; the local contrastive loss constrains the consistency between the hot word fragment-level speech vector and the corresponding hot word text; and the quantity constraint loss constrains the difference between the number of training units corresponding to the boundary contribution value and the length of the complete sentence of text units, thus stabilizing the continuous integral to trigger the boundary learning process. The comprehensive loss can be obtained by weighted summation of the global contrastive loss, local contrastive loss, and quantity constraint loss. Through multi-granularity joint training, both sentence semantic consistency and fragment-level hot word localization capabilities can be considered, improving the stability and generalization of the retrieval module.

[0229] Optionally, the speech encoder or speech coding module in this application embodiment may adopt a parallel converter speech coding structure, the text encoder or text coding module may adopt a pre-trained Chinese text coding structure, and the speech large model may adopt any end-to-end speech recognition model that supports prompt input, and is not limited to a specific network backbone, etc., and this application does not limit it.

[0230] Therefore, in this embodiment of the application, hot word matching is performed by using local window acoustic vector groups instead of global pooling of the whole sentence, which can effectively alleviate the problem of dilution of acoustic information of short hot words and improve the accuracy of hot word retrieval. At the same time, by introducing a local aggregation strategy that is aware of the length of hot word text units, candidate hot words of different lengths can be matched with a more reasonable acoustic range, thereby reducing attention drift.

[0231] The following is combined with Figure 5 The process for determining the contextual keywords provided in this application will be further explained.

[0232] like Figure 5As shown, the original speech signal, i.e., the speech to be recognized, can be input into an audio encoder to extract frame-level speech vectors. These vectors are then subjected to nonlinear transformation and projection using a multilayer perceptron (MLP). Simultaneously, a continuous integral-and-fire (CIF) predictor is used to predict the boundary contribution values ​​of the frame-level speech vectors. The boundary contribution values ​​of each frame-level speech vector are then accumulated frame by frame sequentially to obtain the cumulative value c for the t-th frame. t When the accumulated value reaches a preset threshold, a boundary event is triggered, generating token-level acoustic boundaries. Among these, b k e represents the starting boundary of the acoustic frame span. k For the end boundary of the acoustic frame span, (b k e k () can represent the start and end frame positions of the acoustic frame span.

[0233] Additionally, a candidate hot word vocabulary can be input into a text encoder and transformed using a multilayer perceptron (MLP) to align its feature space with the transformed speech feature space. Then, a similarity matrix between frame-level speech vectors and text semantic vectors can be calculated. Based on the text lexical length of each candidate hot word, an acoustic window for the candidate hot word is determined. This acoustic window is then slid across the acoustic boundary sequence, performing length-aware localized matching. The frame-level speech vectors corresponding to the acoustic boundaries within the acoustic window are identified as the acoustic vector group for the candidate hot word. Local similarity can then be calculated using localized mean-slice between the text semantic vectors of the candidate hot words and the frame-level speech vectors of the acoustic vector groups. The final retrieval score for each candidate hot word is then output based on this local similarity. Finally, based on the final retrieval scores of each candidate hot word, contextual hot words matching the speech to be recognized are determined, achieving accurate hot word localization and retrieval under weak supervision.

[0234] The following is combined with Figure 6 The speech recognition process based on the context hot word determination method provided in this application is described.

[0235] The process involves inputting the speech signal to be recognized into an audio encoder, which then converts it into a speech embedding (i.e., a speech vector) that can be received by a Large Language Model (LLM). Additionally, candidate hot words from a hot word list, such as h0 "XX newspaper", h1 "Deng XX", etc., can be used. n Terms such as "XX Village" are input into the CLAR module, which uses the context hot words determination method provided in this application, i.e. Figure 5 The logic for determining context hot words outputs the top K candidate scores with high matching scores to the speech, and identifies the candidate hot words corresponding to these top K scores as context hot words. After processing by the Retrieval-Augmented Generation (RAG) module, context hot word biases are generated, providing hot word priority information for the large language model (LLM) decoder. The LLM decoder, after low-rank adaptation (LoRA) fine-tuning, can receive multi-source inputs: instructions, speech embeddings, hot word prompts, and transcription results. It performs autoregressive decoding and finally outputs a complete speech transcription result containing the target hot words, thus achieving end-to-end speech recognition with hot word enhancement.

[0236] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0237] Based on the same inventive concept, this application also provides a context hot word determination device for implementing the aforementioned method for determining context hot words. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more context hot word determination device embodiments provided below can be found in the limitations of the context hot word determination method described above, and will not be repeated here.

[0238] In one exemplary embodiment, such as Figure 7 As shown, a context hot word determination device 500 is provided, including: a feature encoding module 510, a continuous integration triggered alignment module 520, a candidate ranking module 530, and a ranking and context recognition module 540, wherein:

[0239] The feature encoding module 510 is used to process the speech to be recognized and the candidate hot words in the candidate word list respectively, and determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time.

[0240] The continuous integration-triggered alignment module 520 is used to perform temporal cumulative processing of boundary contributions on the frame-level speech vectors to determine the acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the acoustic vectors in the speech to be recognized.

[0241] The candidate sorting module 530 is used to perform local alignment between the candidate hot words and the frame-level speech vectors on the acoustic boundary sequence based on the text lexical length of the candidate hot words, and to determine the acoustic vector group corresponding to the candidate hot words.

[0242] The sorting and context recognition module 540 is used to filter each candidate hot word according to the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group to obtain the context hot words of the speech to be recognized.

[0243] In one embodiment, the candidate sorting module 530 includes:

[0244] The first candidate sorting unit is used to determine the acoustic window of the candidate hot words based on the text lexical length of the candidate hot words;

[0245] The second candidate sorting unit is used to slide the acoustic window on the acoustic boundary sequence and determine the frame-level speech vectors corresponding to the acoustic boundaries located within the acoustic window each time as the acoustic vector group of the candidate hot words.

[0246] In one embodiment, the second candidate sorting unit is specifically used for:

[0247] Based on the sliding of the acoustic window on the acoustic boundary sequence, the acoustic frame span corresponding to the acoustic boundary at multiple acoustic window positions is determined;

[0248] Each frame-level speech vector corresponding to the acoustic frame span within the acoustic window is determined as the acoustic vector group of the candidate hot words.

[0249] In one embodiment, the sorting and context recognition module 540 includes:

[0250] The first determining unit is used to determine the candidate score of the candidate hot word based on the text semantic vector of the candidate hot word and the frame-level speech vector of the acoustic vector group;

[0251] The second determining unit is used to determine the candidate hot words whose candidate scores meet the screening conditions as the context hot words of the speech to be recognized from among the candidate hot words.

[0252] In one embodiment, the first determining unit is specifically used for:

[0253] Determine the frame-level similarity between the text semantic vector of the candidate hot words and the corresponding acoustic vector group for each frame-level speech vector;

[0254] The frame-level similarity of each acoustic vector group is aggregated to determine the aggregated similarity of each acoustic vector group;

[0255] The highest aggregate similarity of the acoustic vector group is determined as the candidate score of the candidate hot word.

[0256] In one embodiment, the second determining unit is specifically used for:

[0257] When the candidate hot words are arranged in descending order of candidate scores, the top target number of candidate hot words are determined as the context hot words of the speech to be recognized.

[0258] or,

[0259] Candidate hot words with scores greater than a preset score threshold are identified as context hot words of the speech to be recognized.

[0260] In one embodiment, the continuous integration-triggered alignment module 520 includes:

[0261] A prediction unit is used to predict the boundary contribution value of each frame-level speech vector;

[0262] The third determining unit is used to accumulate the boundary contribution value of each frame-level speech vector frame by frame in chronological order, and determine the acoustic boundary sequence based on the accumulated value.

[0263] In one embodiment, the third determining unit is specifically used for:

[0264] The boundary contribution values ​​corresponding to the speech vectors at each frame level are accumulated in chronological order to obtain the cumulative value;

[0265] When the accumulated value reaches a preset threshold, a boundary event is triggered, which is used to indicate that the current frame that triggered the boundary event is determined as an acoustic boundary;

[0266] After the boundary event is triggered, based on the difference between the accumulated value and the preset threshold, the boundary contribution value of the subsequent frames of the current frame is continued to be accumulated until the traversal is completed, resulting in an acoustic boundary sequence arranged in chronological order.

[0267] In one embodiment, the apparatus further includes a sorting and contextual cue recognition module for:

[0268] Generate hot word suggestions based on the context hot words;

[0269] Based on the hot word prompts, the speech to be recognized is recognized using a speech recognition model to generate recognized text.

[0270] Each module in the aforementioned speech recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0271] Based on the same inventive concept, this application also provides a training apparatus for a target model to implement the training method for the target model described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more speech recognition apparatus embodiments provided below can be found in the limitations of the target model training method described above, and will not be repeated here.

[0272] In one exemplary embodiment, such as Figure 8 As shown, a training device 600 for a target model is provided, including: a first acquisition module 610, a first determination module 620, a second acquisition module 630, a second determination module 640, a third determination module 650, and a generation module 660.

[0273] The first acquisition module 610 is used to acquire a training dataset, which includes multiple sample data. Each sample data includes at least a whole sentence speech data, a whole sentence text data corresponding to the whole sentence speech data, and target hot word text contained in the whole sentence text data.

[0274] The first determining module 620 is used to process the whole sentence speech data, the whole sentence text data and the target hot word text respectively to determine the training frame-level speech vector, whole sentence speech vector, whole sentence text vector and hot word text vector.

[0275] The second acquisition module 630 is used to input the training frame-level speech vector into the initial network model to obtain the training boundary contribution value.

[0276] The second determining module 640 is used to perform time-series cumulative processing on the training boundary contribution value to obtain the training acoustic boundary.

[0277] The third determining module 650 is used to determine the hot word segment speech vector corresponding to the target hot word from the training frame-level speech vector based on the training acoustic boundary.

[0278] The generation module 660 is used to train the initial network model based on the whole sentence speech vector, the whole sentence text vector, the hot word text vector, the hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary, and generate the target network model.

[0279] In one embodiment, the generation module 660 is specifically used for:

[0280] The data constraint loss is determined based on the difference between the number of training lexical units and the length of the lexical units in the entire sentence text;

[0281] The global contrast loss is determined based on the semantic difference between the whole-sentence speech vector and the whole-sentence text vector;

[0282] The local contrast loss is determined based on the semantic difference between the speech vector of the hot word segment and the text vector of the hot word.

[0283] The comprehensive loss value is determined based on the data constraint loss, the global comparison loss, and the local comparison loss.

[0284] The initial network model is trained based on the comprehensive loss value to generate the target network model.

[0285] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a method for determining contextual hot words. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0286] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0287] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0288] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0289] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0290] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0291] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0292] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0293] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for determining contextual hot words, characterized in that, The method includes: The speech to be recognized and the candidate hot words in the candidate word list are processed separately to determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step. The frame-level speech vectors are subjected to temporal cumulative processing of boundary contributions to determine an acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the frame-level speech vectors in the speech to be recognized. Based on the text lexical length of the candidate hot words, the candidate hot words are locally aligned with the frame-level speech vectors on the acoustic boundary sequence to determine the acoustic vector group corresponding to the candidate hot words; Based on the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group, the candidate hot words are filtered to obtain the context hot words of the speech to be identified.

2. The method according to claim 1, characterized in that, The step of determining the acoustic vector group corresponding to the candidate hot words by locally aligning the candidate hot words with the frame-level speech vectors on the acoustic boundary sequence based on the text lexical length of the candidate hot words includes: The acoustic window of the candidate hot words is determined based on the text lexical length of the candidate hot words; Based on the sliding of the acoustic window on the acoustic boundary sequence, the frame-level speech vectors corresponding to each acoustic boundary located within the acoustic window are determined as the acoustic vector group of the candidate hot words.

3. The method according to claim 2, characterized in that, The step of sliding the acoustic window on the acoustic boundary sequence and determining the frame-level speech vectors corresponding to each acoustic boundary within the acoustic window as the acoustic vector group of the candidate hot words includes: Based on the sliding of the acoustic window on the acoustic boundary sequence, the acoustic frame span corresponding to the acoustic boundary at multiple acoustic window positions is determined; Each frame-level speech vector corresponding to the acoustic frame span within the acoustic window is determined as the acoustic vector group of the candidate hot words.

4. The method according to claim 1, characterized in that, The step of filtering each candidate hot word based on its text semantic vector and the frame-level speech vector of the acoustic vector group to obtain the context hot words of the speech to be recognized includes: The candidate scores of the candidate hot words are determined based on the text semantic vectors of the candidate hot words and the frame-level speech vectors of the acoustic vector groups. Among the candidate hot words, those whose candidate scores meet the screening criteria are determined as the context hot words of the speech to be recognized.

5. The method according to claim 4, characterized in that, The step of determining the candidate score of the candidate hot word based on the text semantic vector of the candidate hot word and the frame-level speech vector of the acoustic vector group includes: Determine the frame-level similarity between the text semantic vector of the candidate hot words and the corresponding acoustic vector group for each frame-level speech vector; The frame-level similarity of each acoustic vector group is aggregated to determine the aggregated similarity of each acoustic vector group; The highest aggregate similarity of the acoustic vector group is determined as the candidate score of the candidate hot word.

6. The method according to claim 4, characterized in that, The step of determining the candidate hot words whose candidate scores meet the screening criteria as context hot words for the speech to be recognized from among the candidate hot words includes: When the candidate hot words are arranged in descending order of candidate scores, the top target number of candidate hot words are determined as the context hot words of the speech to be recognized. or, Candidate hot words with scores greater than a preset score threshold are identified as context hot words of the speech to be recognized.

7. The method according to claim 1, characterized in that, The temporal cumulative processing of boundary contributions to the frame-level speech vectors to determine the acoustic boundary sequence includes: Predict the boundary contribution value for each frame-level speech vector; The boundary contribution values ​​of each frame-level speech vector are accumulated frame by frame in chronological order, and the acoustic boundary sequence is determined based on the accumulated values.

8. The method according to claim 7, characterized in that, The boundary contribution value for each frame-level speech vector is accumulated frame by frame in chronological order, and the acoustic boundary sequence is determined based on the accumulated value, including: The boundary contribution values ​​corresponding to the speech vectors at each frame level are accumulated in chronological order to obtain the cumulative value; When the accumulated value reaches a preset threshold, a boundary event is triggered, which is used to indicate that the current frame that triggered the boundary event is determined as an acoustic boundary; After the boundary event is triggered, based on the difference between the accumulated value and the preset threshold, the boundary contribution value of the subsequent frames of the current frame is continued to be accumulated until the traversal is completed, resulting in an acoustic boundary sequence arranged in chronological order.

9. The method according to claim 1, characterized in that, The method further includes: Generate hot word suggestions based on the context hot words; Based on the hot word prompts, the speech to be recognized is recognized using a speech recognition model to generate recognized text.

10. A method for training a target network model, characterized in that, The method includes: Obtain a training dataset, which includes multiple sample data, each of which includes at least a whole sentence speech data, a whole sentence text data corresponding to the whole sentence speech data, and target hot word text contained in the whole sentence text data; The whole-sentence speech data, the whole-sentence text data, and the target hot word text are processed respectively to determine the training frame-level speech vector, whole-sentence speech vector, whole-sentence text vector, and hot word text vector; The trained frame-level speech vectors are input into the initial network model to obtain the training boundary contribution value; The training boundary contribution value is subjected to time-series cumulative processing to obtain the training acoustic boundary; Based on the training acoustic boundary, determine the hot word segment speech vector corresponding to the target hot word from the training frame-level speech vector; Based on the whole sentence speech vector, the whole sentence text vector, the hot word text vector, the hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary, the initial network model is trained to generate the target network model.

11. The method according to claim 10, characterized in that, The process of training the initial network model based on the whole-sentence speech vector, the whole-sentence text vector, the hot word text vector, the hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary to generate the target network model includes: The data constraint loss is determined based on the difference between the number of training lexical units and the length of the lexical units in the entire sentence text; The global contrast loss is determined based on the semantic difference between the whole-sentence speech vector and the whole-sentence text vector; The local contrast loss is determined based on the semantic difference between the speech vector of the hot word segment and the text vector of the hot word. The comprehensive loss value is determined based on the data constraint loss, the global comparison loss, and the local comparison loss. The initial network model is trained based on the comprehensive loss value to generate the target network model.

12. A device for determining contextual hot words, characterized in that, The device includes: The feature encoding module is used to process the speech to be recognized and the candidate hot words in the candidate word list respectively, and determine the frame-level speech vector and the text semantic vector of each candidate hot word at each time step. The continuous integration-triggered alignment module is used to perform temporal cumulative processing of the boundary contribution of the frame-level speech vector to determine the acoustic boundary sequence. The acoustic boundaries in the acoustic boundary sequence are used to indicate the boundary positions of the acoustic vectors in the speech to be recognized. The candidate sorting module is used to perform local alignment between the candidate hot words and the frame-level speech vectors on the acoustic boundary sequence based on the text lexical length of the candidate hot words, and to determine the acoustic vector group corresponding to the candidate hot words; The sorting and context recognition module is used to filter each candidate hot word according to the text semantic vector of each candidate hot word and the frame-level speech vector of the acoustic vector group to obtain the context hot words of the speech to be recognized.

13. A training device for a target network model, characterized in that, The device includes: The first acquisition module is used to acquire a training dataset, which includes multiple sample data, each of which includes at least whole sentence speech data, whole sentence text data corresponding to the whole sentence speech data, and target hot word text contained in the whole sentence text data; The first determining module is used to process the whole sentence speech data, the whole sentence text data and the target hot word text respectively to determine the training frame-level speech vector, whole sentence speech vector, whole sentence text vector and hot word text vector; The second acquisition module is used to input the training frame-level speech vectors into the initial network model to obtain the training boundary contribution value. The second determining module is used to perform time-series cumulative processing on the training boundary contribution value to obtain the training acoustic boundary; The third determining module is used to determine the hot word segment speech vector corresponding to the target hot word from the training frame-level speech vector based on the training acoustic boundary; The generation module is used to train the initial network model based on the whole sentence speech vector, the whole sentence text vector, the hot word text vector, the hot word segment speech vector, and the number of training lexical units corresponding to the training acoustic boundary, and generate the target network model.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9, or the steps of the method according to any one of claims 10 to 11.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9, or the steps of the method according to any one of claims 10 to 11.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9, or the steps of the method according to any one of claims 10 to 11.