Wake-up method, device, equipment, medium and product

By constructing a structured audio library and a noise simulation training set, a personalized wake-up model is generated, which solves the personalized needs of in-vehicle wake-up models, improves the recognition rate and robustness of custom wake-up words, and realizes a personalized interactive experience.

CN121884785APending Publication Date: 2026-04-17IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2025-12-02
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing in-vehicle wake-up models are unable to meet the personalized needs of all users, especially since custom wake-up words have low wake-up rates and are complex to train, and fail to effectively utilize user interaction history information.

Method used

By collecting users' in-vehicle voice data, a structured and labeled audio library is built. Personalized training audio is generated based on users' acoustic characteristics and combined with noise environment simulation to form a training set. Transfer learning is then used to train the target wake-up model.

Benefits of technology

It significantly improves the recognition rate and robustness of custom wake words, realizes a personalized interactive experience with zero user intervention, and enhances the practicality and user satisfaction of the in-vehicle voice wake-up system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884785A_ABST
    Figure CN121884785A_ABST
Patent Text Reader

Abstract

The invention discloses a wake-up method and device, equipment, a medium and a product, and the method comprises the steps: obtaining a target audio clip matched with a wake-up word from an audio library in response to a triggering operation of a target user for the wake-up word; splicing the target audio clips to generate training audio with acoustic features of the target user; forming a training set by the noise adding audio and the training audio; and training a preset wake-up model by using the training set, obtaining a target wake-up model, and executing wake-up of the wake-up word through the target wake-up model. According to the method, the special awakening model is provided for the target user, so that the awakening effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a wake-up method, apparatus, device, medium, and product. Background Technology

[0002] With the application and popularization of intelligent voice technology in vehicles, voice wake-up function has now become an indispensable basic function in intelligent cars.

[0003] Currently, the main approach is to configure a universal wake-up model in the in-vehicle wake-up system to provide wake-up functionality for all users in the vehicle.

[0004] However, the above wake-up model is difficult to meet the personalized needs of all users. Summary of the Invention

[0005] The main purpose of this application is to provide a wake-up method, device, equipment, medium and product, providing users with a dedicated wake-up model to improve the wake-up effect.

[0006] To achieve the above objectives, firstly, this application provides a wake-up method, comprising: In response to the target user's trigger action on the wake word, retrieve the target audio segment that matches the wake word from the audio library; The target audio segments are spliced ​​together to generate training audio with the acoustic characteristics of the target user; The training set consists of noise-added frequencies and training audio. The preset wake-up model is trained using the training set to obtain the target wake-up model, which is then used to execute the wake word.

[0007] In one embodiment, in response to a target user's triggering operation on a wake word, retrieving a target audio segment matching the wake word from an audio library includes: The wake word is segmented to obtain the target text unit corresponding to the wake word; Obtain the audio library; The initial audio segment is obtained by retrieving and matching each text unit in the target text unit from the audio library; Select the audio segment with a signal-to-noise ratio greater than the preset value from the initial audio segments as the target audio segment.

[0008] In one embodiment, obtaining an audio library includes: Collect the user's raw audio; The original audio is converted to obtain the corresponding audio text. The audio text is segmented and filtered to obtain the target sub-audio text; Obtain the user ID of the target sub-audio text, the audio segment in the original audio that matches the target sub-audio text, and the audio feature tags of the audio segment; The audio library is constructed by matching the target sub-audio text, the user ID of the target sub-audio text, the audio segments in the original audio that match the target sub-audio text, and the audio feature tags of the audio segments.

[0009] In one embodiment, the audio text is segmented and filtered to obtain the target sub-audio text, including: The audio file is segmented to obtain the initial sub-audio text; Obtain the time information corresponding to the initial sub-audio text, where the time information includes the start time and end time; Based on time information, obtain the intermediate sub-audio text; The target sub-audio text is obtained by filtering the intermediate sub-audio text using preset conditions.

[0010] In one embodiment, obtaining intermediate sub-audio text based on time information includes: The time interval between adjacent audio segments in the initial sub-audio text is calculated by using the start and end times of each sub-audio text in the initial sub-audio text. Based on the time interval, determine how to segment adjacent audio files; The adjacent audio segments are segmented using a method that allows for the extraction of the intermediate sub-audio text.

[0011] In one embodiment, determining the segmentation method for adjacent audio based on time intervals includes: If the time interval is greater than or equal to the preset time interval, the adjacent audio segments will be segmented by word. If the time interval is less than the preset time interval, the adjacent audio segments will be segmented by whole words.

[0012] In one embodiment, the intermediate sub-audio text is filtered according to preset conditions to obtain the target sub-audio text, including: Determine whether the clarity, truncation, and overlap of the audio segment corresponding to the intermediate sub-audio text meet the preset thresholds; Select the sub-audio texts corresponding to audio segments that meet the preset threshold from the intermediate sub-audio texts as the target sub-audio texts.

[0013] In one embodiment, splicing target audio segments to generate training audio with the acoustic features of the target user includes: Retrieve the audio segments in the target audio segment that match each text unit; The audio segments matched by each text unit are concatenated according to the user ID and audio feature tags to obtain the combined audio corresponding to each text unit; Training audio with the acoustic characteristics of the target user is formed by combining the audio corresponding to each text unit.

[0014] In one embodiment, the training set comprises noise-added frequencies and training audio, including: Acquire noise audio; The training audio is noise-added by adding noise to the noisy audio, resulting in a noise-added audio. The noise-added audio is fused with the training audio to obtain the training set.

[0015] Secondly, embodiments of this application provide a wake-up device, including: The audio acquisition module is used to retrieve the target audio segment that matches the wake word from the audio library in response to the target user's trigger operation on the wake word; The audio splicing module is used to splice target audio segments to generate training audio with the acoustic characteristics of the target user; The training set acquisition module is used to construct a training set from noise-added frequencies and training audio. The wake-up module is used to train a preset wake-up model using a training set to obtain a target wake-up model, which is then used to execute the wake-up word.

[0016] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the methods described above.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0018] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0019] This application provides a wake-up method, apparatus, device, medium, and product, comprising: firstly, in response to a target user's trigger operation on a wake-up word, obtaining a target audio segment matching the wake-up word from an audio library; then, splicing the target audio segment to generate training audio with the target user's acoustic characteristics; subsequently, constructing a training set from the noise-added audio and the training audio; and using the training set to train a preset wake-up model to obtain a target wake-up model, thereby executing the wake-up word activation through the target wake-up model. This application provides a dedicated wake-up model for the target user to improve the wake-up effect. Attached Figure Description

[0020] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings: Figure 1 This is a flowchart illustrating a wake-up method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a wake-up device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0022] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein.

[0023] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0024] It should be understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0025] It should be understood that in this application, "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product or device.

[0026] It should be understood that in this application, "multiple" refers to two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, "and / or B" can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "Contains A, B, and C", "Contains A, B, and C" means that all three A, B, and C are contained; "Contains A, B, or C" means that one of A, B, and C is contained; "Contains A, B, and / or C" means that any one, two, or three of A, B, and C are contained.

[0027] It should be understood that in this application, "B corresponding to A", "B corresponding to A", "A corresponds to B", or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. Matching A and B is defined as a similarity between A and B that is greater than or equal to a preset threshold.

[0028] Depending on the context, "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection."

[0029] The data involved in this application may be data authorized by the tester or fully authorized by all parties. The collection, dissemination, and use of the data shall comply with the relevant laws, regulations and standards of the relevant countries and regions. The implementation methods / executives of this application may be combined with each other.

[0030] The technical solutions of this application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0031] The present application will now be described in conjunction with the accompanying drawings and specific embodiments.

[0032] With the application and popularization of intelligent voice technology in vehicles, voice wake-up has become an indispensable basic function in smart cars. However, due to limitations in model size and training conditions, the wake-up model in the car is basically pre-trained before the car leaves the factory. But because different users have different acoustic characteristics, it does not mean that all users can be successfully woken up.

[0033] In recent years, the "custom wake word" function has become a trend. However, since the text is user-defined, it is highly random and therefore it is impossible to collect a large amount of training audio in advance to train different audio. As a result, the wake-up rate of custom wake words in cars has not been well guaranteed or effectively improved so far.

[0034] In conclusion, to optimize the actual experience of in-vehicle voice control, it is highly necessary to conduct customized wake-up training based on the acoustic characteristics of different users to improve the wake-up effect.

[0035] Currently, a common method for training a primary wake-up word involves pre-collecting a large batch of audio recordings of the primary wake-up word (potentially covering different age groups, genders, speech rates, noise levels, etc.) to extract features and learn model parameters, forming a corresponding primary wake-up model. This model is then pre-installed in the vehicle's infotainment system for user use. While this model can cover a wide range of pronunciation features and achieve a high wake-up accuracy, it doesn't guarantee successful wake-up for all users due to their unique acoustic characteristics.

[0036] Custom wake words, being randomly defined by the user, are inherently uncertain and cannot be predicted or trained in advance. Therefore, only a basic, general wake-up model can be provided, containing pronunciation rules for common words. When a user speaks the corresponding word, a match to that pronunciation rule triggers the wake-up. However, due to the lack of targeted pre-training, its accuracy is significantly lower than that of a standard wake-up word. Some products on the market incorporate interactive elements that allow users to record a certain number of audio clips at different speaking speeds after setting a custom wake word, enabling targeted training to improve its wake-up rate. However, this technology requires active user participation, presenting a degree of complexity and time commitment, potentially leading to a poor user experience.

[0037] Furthermore, under the current technical solution, the historical voice information of users' daily interactions with the vehicle's infotainment system cannot be well reused for voice wake-up, and is not actively converted into personalized wake-up word training resources, which is a waste of resources.

[0038] To address the aforementioned problems, this application proposes a wake-up method.

[0039] Please see Figure 1 , Figure 1 This is a flowchart illustrating a wake-up method provided in an embodiment of this application. Figure 1 As shown, it includes the following steps: Step S101: In response to the target user's trigger operation on the wake word, retrieve the target audio segment that matches the wake word from the audio library.

[0040] The target users can be those in the vehicle, such as the driver and passengers.

[0041] Triggering operations for wake words include actions such as when a driver enters a custom wake word in the car via voice or text input.

[0042] In response to a user's trigger action on a wake word, the system retrieves a target audio segment matching the wake word from the audio library. First, the wake word is segmented to obtain the target text unit corresponding to the wake word. Then, the audio library is retrieved, and each text unit in the target text unit is matched with it to obtain an initial audio segment. Finally, an audio segment with a signal-to-noise ratio greater than the preset signal-to-noise ratio is selected from the initial audio segments as the target audio segment.

[0043] Specifically, the user inputs a custom wake word via voice or text, such as "Hello China". Then, the wake word text is segmented into its smallest semantic units or character units. For example, "Hello China" is segmented into: ["you", "good", "China"] or ["you", "good", "zhong", "guo"] (depending on the minimum granularity set by the system).

[0044] Assign a unique text ID to each segmented unit. For example: text ID: text01 → text: "you", text ID: text02 → text: "good", text ID: text03 → text: "China". The final target text unit is: {text01: "you", text02: "good", text03: "China"}.

[0045] The process of acquiring an audio library includes: collecting the user's original audio; converting the original audio to obtain the corresponding audio text; segmenting and filtering the audio text to obtain target sub-audio text; acquiring the user ID of the target sub-audio text, the audio segments in the original audio that match the target sub-audio text, and the audio feature tags of the audio segments; and matching the target sub-audio text, the user ID of the target sub-audio text, the audio segments in the original audio that match the target sub-audio text, and the audio feature tags of the audio segments to form an audio library.

[0046] The process of segmenting and filtering audio text to obtain target sub-audio text includes: segmenting the audio file to obtain initial sub-audio text; obtaining the time information corresponding to the initial sub-audio text, including start and end times; obtaining intermediate sub-audio text based on the time information; and filtering the intermediate sub-audio text according to preset conditions to obtain the target sub-audio text.

[0047] The process of obtaining intermediate sub-audio text based on time information includes: calculating the time interval between adjacent audios in the initial sub-audio text using the start and end times of each sub-audio text in the initial sub-audio text; determining the segmentation method of adjacent audios based on the time interval; and segmenting adjacent audios using the segmentation method of adjacent audios to obtain intermediate sub-audio text.

[0048] Among them, determining the segmentation method of adjacent audio based on time intervals includes: if the time interval is greater than or equal to a preset time interval, the segmentation method of adjacent audio is segmentation by character; if the time interval is less than the preset time interval, the segmentation method of adjacent audio is segmentation by whole word.

[0049] Among them, the intermediate sub-audio text is screened and processed according to preset conditions to obtain the target sub-audio text, including: judging whether the clarity, truncation, and overlap of the audio segment corresponding to the intermediate sub-audio text meet the preset thresholds; screening out the sub-audio text corresponding to the audio segment that meets the preset thresholds from the intermediate sub-audio text as the target sub-audio text.

[0050] Specifically, during the daily driving of the vehicle, the original audio of the user is collected in real time through the in-vehicle microphone, including the interactive voice between the user and the vehicle-mounted system, the conversation between passengers, etc. Then, the collected continuous speech stream is real-time transcribed into text by using the vehicle-mounted speech recognition module to form the audio text.

[0051] After the text conversion is completed, it enters the audio segmentation and screening stage.

[0052] First, the speech recognition model performs a preliminary segmentation of the audio based on the time stamp information of words, and outputs the initial sub-audio text and its corresponding time information (including the start time and the end time). For example, when the user says the sentence "I am very beautiful", the recognition model will output the content shown in Table 1.

[0053] Table 1 Table of initial sub-audio text and its corresponding time information

[0054] Then, based on the time interval between adjacent words, the segmentation granularity is further judged: if the interval ≥ 150 ms, it is segmented by character; if the interval ≤ 50 ms, it is merged and processed as a whole word. For example, if the interval between "very" and "beautiful" is only 10 ms, it is merged into "very beautiful" as a whole word unit to form the intermediate sub-audio text.

[0055] After obtaining the intermediate sub-audio text, the quality screening step is executed, and it is judged whether the audio segment is qualified according to multiple preset conditions: Clarity judgment: Through signal-to-noise ratio evaluation, only the audio with SNR ≥ 15 dB is retained; Truncation detection: Exclude the segments with the silence at the beginning and end of the audio being 0 or having broken words; Overlap detection: Reject the audio with multiple voiceprint overlaps within the same time period.

[0056] After the above screening, the sub-audio text corresponding to the qualified audio segment is finally retained as the target sub-audio text.

[0057] Further feature extraction and tagging processing of the target sub-audio text: distinguish the speaker through voiceprint recognition technology, assign user ID, calculate the speech rate according to the audio duration and content, and label it as "fast, medium, slow" level, and label the tone tags such as "high, medium, low" according to the tone analysis results. At the same time, record and label the in-vehicle noise segments, and classify them into "low noise, medium noise, high noise" levels according to intensity.

[0058] Finally, the target sub-audio text, corresponding user ID, audio segment, speech rate tag, intonation tag, and other metadata are structured and stored by text ID, forming a continuously updated audio library as shown in Table 2. This library supports multi-dimensional retrieval by text ID, user ID, speech rate, intonation, etc., providing a high-quality audio material foundation for subsequent wake word training.

[0059] Table 2 Audio Library

[0060] This application embodiment collects user in-vehicle voice data in real time without user intervention. Based on a precise audio segmentation algorithm (combined with time interval thresholds to achieve intelligent segmentation at the character / word level) and a multi-dimensional quality screening mechanism (ensuring data quality from multiple angles such as signal-to-noise ratio, audio integrity, and voiceprint overlap), it constructs a structured and tagged user-specific audio library. Then, based on the principle of high similarity combination, it quickly retrieves and splices high-quality training audio that matches the user's acoustic characteristics from the audio library. This effectively solves the core pain points of insufficient training samples, low user cooperation, and limited personalization in traditional solutions. Ultimately, it significantly improves the recognition rate and robustness of custom wake-up words without zero user intervention, enabling the in-vehicle voice wake-up system to have continuous self-evolution capabilities and achieve a precise and personalized interactive experience.

[0061] Step S102: Segment the target audio segments to generate training audio with the acoustic features of the target user.

[0062] The process involves splicing together target audio segments to generate training audio with the acoustic features of the target user. This includes: acquiring audio segments that match each text unit in the target audio segment; splicing the audio segments that match each text unit according to the user ID and audio feature labels to obtain the combined audio corresponding to each text unit; and constructing the training audio with the acoustic features of the target user from the combined audio corresponding to each text unit.

[0063] Specifically, upon receiving a user-defined wake word (e.g., "Hello China"), the wake word is first segmented into the smallest text units: "you", "hello", and "China", which correspond to text IDs: text01, text02, and text03, respectively.

[0064] Then, all target audio segments matching each text unit are retrieved from the audio library. Taking user 001 as an example, the text unit "you" (text01) matches 4 audio segments with the following feature combinations: (medium speaking speed, medium intonation) × 2, (fast speaking speed, high intonation) × 1, and (slow speaking speed, low intonation) × 1. The text unit "good" (text02) matches 4 audio segments with the same feature combination as text01. The text unit "China" (text03) matches 3 audio segments with the following feature combinations: (medium speaking speed, medium intonation) × 2 and (fast speaking speed, high intonation) × 1.

[0065] Next, audio segments are concatenated according to the principle of "high similarity combination". Specifically, the audio segments are first classified according to user ID to ensure that each training audio contains only the pronunciation features of the same user. Then, under the same user ID, the audio segments are further grouped according to audio feature labels (speech rate label and intonation label), and only audio segments with completely identical feature labels are selected for concatenation.

[0066] For user 001, two sets of perfectly matching feature combinations were found: Group 1: Speech rate = medium, intonation = medium text01 contains two audio clips that meet the criteria (001.wav, 002.wav). text02 contains two audio clips that meet the criteria (001.wav, 002.wav). text03 contains two audio clips that meet the criteria (001.wav, 002.wav). By arranging and combining (2×2×2), 8 complete wake word audios can be generated.

[0067] Group 2: Speech speed = fast, tone = high text01 contains one audio clip (003.wav) that meets the criteria. text02 contains one audio clip (003.wav) that meets the criteria. text03 contains one matching audio clip (003.wav). By arranging and combining (1×1×1), a complete wake word audio can be generated.

[0068] As shown in Table 3, the selected audio segments were seamlessly spliced ​​according to the text unit order ("you" → "good" → "China"), ultimately generating 9 training audio clips with unique acoustic features for user 001. These audio clips not only maintained the user's specific pronunciation habits but also maintained a high degree of consistency in speech rate and intonation, ensuring the effectiveness of subsequent model training.

[0069] In contrast, user 002 only has one matching audio segment (005.wav) in text01, and cannot find feature-matching audio in other text units. Therefore, it is determined that user 002 is insufficient to form a splicing subunit and cannot generate effective training audio.

[0070] Table 3 Training Audio

[0071] This application's embodiments employ a dual screening mechanism based on user ID and acoustic feature tags (speech rate, intonation) to ensure that the generated training audio maintains a high degree of consistency in acoustic characteristics while preserving the user's unique pronunciation features, effectively solving the feature mismatch problem present in traditional audio splicing. Secondly, by arranging and combining audio segments with completely consistent feature tags, a considerable number of training samples can be generated from limited basic audio resources (e.g., generating 9 complete training audios from 9 basic segments), greatly improving the diversity and coverage of training data. Finally, this intelligent splicing method not only fully utilizes historical speech data collected without the user's awareness but also ensures the naturalness and usability of the generated audio through the consistency of acoustic features, providing a high-quality and highly targeted training foundation for subsequent model training, thereby significantly improving the accuracy and robustness of the personalized wake-up model.

[0072] Step S103: The training set is composed of noise-added frequencies and training audio.

[0073] The training set consists of noise-added audio and training audio, including: collecting noise-added audio; adding noise to the training audio to obtain noise-added audio; and fusing the noise-added audio with the training audio to obtain the training set.

[0074] Specifically, after successfully generating nine clean, personalized training audio tracks (such as "Hello China") for user 001, the enhancement steps for training set construction are executed, aiming to improve the robustness of the wake-up model by simulating a real in-car noise environment.

[0075] First, the pre-collected in-vehicle noise audio library is accessed. This library contains noise samples recorded under different driving conditions, as shown in Table 4, and has been categorized and labeled according to noise intensity: low noise environment: 30-45dB (such as when the vehicle is stationary or the electric motor is idling), medium noise environment: 50-65dB (such as when driving at a constant speed on urban roads), and high noise environment: 70-85dB (such as when driving on highways or in the rain).

[0076] Table 4 Noise Scene Type Table

[0077] Then, a stratified sampling strategy was adopted to randomly select representative samples from each noise category. For example, three different noise segments were selected from each of the low, medium, and high noise levels to ensure coverage of various typical driving scenarios.

[0078] Next, the noise addition process is performed. Based on nine clean training audio clips from user 001, each clean audio clip is fused with three noise samples of different levels. The noise addition process employs digital audio synthesis technology, adjusting the noise gain to achieve the target signal-to-noise ratio (e.g., 15dB, 10dB, 5dB) to simulate voice communication environments ranging from good to poor. During noise addition, it is necessary to ensure a natural match between the noise characteristics and the speech content. For audio with a slower speech rate and lower pitch, low-noise samples are prioritized; for audio with a faster speech rate and higher pitch, the proportion of medium-to-high noise samples is appropriately increased. The length of the noise samples is kept consistent with the training audio, and loop filling or intelligent truncation techniques are used to handle duration mismatches.

[0079] After noise reduction, 27 noisy audio tracks were obtained (9 clean audio tracks × 3 noise levels). Finally, the training set consisted of the original 9 clean training audio tracks and the 27 noisy audio tracks, totaling 36 training samples.

[0080] This application employs a stratified sampling strategy to select representative noise samples from low, medium, and high noise levels, ensuring that the training set comprehensively covers typical driving scenarios such as stationary vehicles, urban road driving, and highway driving, greatly improving the model's environmental adaptability. Secondly, it uses digital audio synthesis technology to intelligently fuse clean training audio with noise of different levels, and achieves reasonable matching between noise and speech based on speech rate and intonation features, effectively simulating voice communication conditions in a real in-vehicle environment and significantly enhancing the model's anti-interference capability. Finally, by expanding the 9 clean training audio samples to 36 training samples containing different noise conditions, the diversity and coverage of training data are greatly increased while maintaining the user's acoustic characteristics, laying a solid foundation for training a robust wake-up model that maintains high recognition rate in complex noise environments, ultimately enabling the voice wake-up system to achieve excellent performance in real driving environments.

[0081] Step S104: Train the preset wake-up model using the training set to obtain the target wake-up model, so as to execute the wake-up word through the target wake-up model.

[0082] Specifically, after completing the construction of the training set (containing 36 "Hello China" training samples from user 001, including 9 clean audio tracks and 27 noisy audio tracks), the system starts the training process for the personalized wake-up model.

[0083] First, a pre-set general wake-up model is loaded as the base model. This base model is pre-trained on a massive corpus of multiple speakers before leaving the factory and has basic voice wake-up capabilities, but its adaptability to the acoustic features of specific users is limited. Transfer learning technology is used to retain the feature extraction layer and most of the network parameters of the base model, replacing only the final classification layer and fine-tuning it accordingly.

[0084] During training, the 36 training samples were divided into a training set, a validation set, and a test set in a 7:2:1 ratio. The training set (25 samples) was used for updating model parameters, the validation set (7 samples) was used for hyperparameter tuning and monitoring the training process, and the test set (4 samples) was used for final model performance evaluation.

[0085] The model training adopts an end-to-end deep learning architecture, and the specific process is as follows: First, acoustic features such as MFCC (Mel-frequency cepstral coefficients) and FBank (filter bank features) are extracted from the input audio. The frame length is 25ms and the frame shift is 10ms. Then, the feature sequence is passed through multiple convolutional layers and LSTM layers to capture the spatiotemporal features of the audio. Next, the weights of key speech segments (such as the first word of the wake word) are enhanced to improve the model's sensitivity to important segments. The cross-entropy loss function is then used to calculate the difference between the model's prediction and the true label. Finally, the Adam optimizer is used to update the model parameters with a learning rate of 0.001, with a focus on adjusting the parameters of the last two layers.

[0086] During training, performance metrics on the validation set are monitored in real time. For example, when the wake-up accuracy on the validation set reaches above 98%, the false wake-up rate drops below 0.5%, and the loss function value no longer decreases significantly for five consecutive epochs, training automatically stops, the optimal model parameters are saved, and a target wake-up model specific to user 001 is formed.

[0087] The trained target wake-up model is encrypted and distributed to user 001's in-vehicle system via the vehicle-to-everything (V2X) network, replacing the original general wake-up model. Subsequently, when user 001 says "Hello China" in the car: the audio acquisition module captures the voice signal in real time, the target wake-up model performs rapid inference on the input audio (inference time <100ms), and when the confidence level exceeds the 0.95 threshold, a wake-up response is immediately triggered. Simultaneously, the audio data and recognition results of each wake-up are recorded for subsequent model iteration and optimization.

[0088] In addition, the model retraining process is automatically triggered every month or when the user adds a sufficient number of high-quality audio clips, ensuring that the wake-up model can continuously adapt to changes in the user's acoustic characteristics and achieve a personalized experience that becomes more accurate the more it is used.

[0089] This application uses a pre-trained general wake-up model as a foundation. By retaining the feature extraction layer and fine-tuning the classification layer through transfer learning strategies, it significantly improves training efficiency while ensuring the model's basic recognition capabilities, enabling personalized model training to be completed quickly with limited computing resources and training samples. Secondly, by establishing a multi-dimensional monitoring system that includes accuracy, false wake-up rate, and loss function value, it ensures that the model automatically stops training when it reaches optimal performance, effectively preventing overfitting and ensuring the model's generalization ability. Finally, by establishing a dynamic model update mechanism, a virtuous cycle of continuous optimization of the wake-up model based on user usage is achieved, enabling the voice wake-up system to have self-evolution capabilities. Ultimately, this achieves a precise personalized interactive experience, greatly improving the practicality and user satisfaction of the in-vehicle voice system.

[0090] This application provides a wake-up method, comprising: first, responding to a target user's trigger operation on a wake-up word, obtaining a target audio segment matching the wake-up word from an audio library; then, concatenating the target audio segment to generate training audio with the target user's acoustic characteristics; subsequently, constructing a training set from the noise-added audio and the training audio; and using the training set to train a preset wake-up model to obtain a target wake-up model, which is then used to execute the wake-up word activation. This application provides a dedicated wake-up model for the target user to improve the wake-up effect.

[0091] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0092] The following are device embodiments of this application. For details not described in detail, please refer to the corresponding method embodiments described above.

[0093] Figure 2 This illustration shows a schematic diagram of a wake-up device according to an embodiment of this application. For ease of explanation, only the parts related to the embodiment of this application are shown. The wake-up device includes: The audio acquisition module 201 is used to retrieve a target audio segment that matches the wake word from the audio library in response to the target user's trigger operation on the wake word. The audio splicing module 202 is used to splice target audio segments to generate training audio with the acoustic characteristics of the target user; The training set acquisition module 203 is used to construct a training set from noise-added frequencies and training audio. The wake-up module 204 is used to train a preset wake-up model using a training set to obtain a target wake-up model, so as to execute the wake-up word through the target wake-up model.

[0094] In one embodiment, the audio acquisition module 201 is further configured to segment the wake word to obtain the target text unit corresponding to the wake word; Obtain the audio library; The initial audio segment is obtained by retrieving and matching each text unit in the target text unit from the audio library; Select the audio segment with a signal-to-noise ratio greater than the preset value from the initial audio segments as the target audio segment.

[0095] In one embodiment, the audio acquisition module 201 is further configured to acquire the user's original audio. The original audio is converted to obtain the corresponding audio text. The audio text is segmented and filtered to obtain the target sub-audio text; Obtain the user ID of the target sub-audio text, the audio segment in the original audio that matches the target sub-audio text, and the audio feature tags of the audio segment; The audio library is constructed by matching the target sub-audio text, the user ID of the target sub-audio text, the audio segments in the original audio that match the target sub-audio text, and the audio feature tags of the audio segments.

[0096] In one embodiment, the audio acquisition module 201 is further configured to perform segmentation processing on the audio file to obtain initial sub-audio text; Obtain the time information corresponding to the initial sub-audio text, where the time information includes the start time and end time; Based on time information, obtain the intermediate sub-audio text; The target sub-audio text is obtained by filtering the intermediate sub-audio text using preset conditions.

[0097] In one embodiment, the audio acquisition module 201 is further configured to calculate the time interval between adjacent audios in the initial sub-audio text by using the start time and end time of each sub-audio text in the initial sub-audio text; Based on the time interval, determine how to segment adjacent audio files; The adjacent audio segments are segmented using a method that allows for the extraction of the intermediate sub-audio text.

[0098] In one embodiment, the audio acquisition module 201 is further configured to segment adjacent audio segments by word if the time interval is greater than or equal to a preset time interval. If the time interval is less than the preset time interval, the adjacent audio segments will be segmented by whole words.

[0099] In one embodiment, the audio acquisition module 201 is further configured to determine whether the clarity, truncation, and overlap of the audio segment corresponding to the intermediate sub-audio text meet a preset threshold. Select the sub-audio texts corresponding to audio segments that meet the preset threshold from the intermediate sub-audio texts as the target sub-audio texts.

[0100] In one embodiment, the audio splicing module 202 is further configured to acquire audio segments in the target audio segment that match each text unit; The audio segments matched by each text unit are concatenated according to the user ID and audio feature tags to obtain the combined audio corresponding to each text unit; Training audio with the acoustic characteristics of the target user is formed by combining the audio corresponding to each text unit.

[0101] In one embodiment, the training set acquisition module 203 is also used to acquire noisy audio; The training audio is noise-added by adding noise to the noisy audio, resulting in a noise-added audio. The noise-added audio is fused with the training audio to obtain the training set.

[0102] This application provides a wake-up device, specifically used for: firstly, responding to a target user's trigger operation on a wake-up word, retrieving a target audio segment matching the wake-up word from an audio library; then, splicing the target audio segment to generate training audio with the target user's acoustic characteristics; subsequently, constructing a training set from the noise-added audio and the training audio; and using the training set to train a preset wake-up model to obtain a target wake-up model, which is then used to execute the wake-up word activation. This application provides a dedicated wake-up model for the target user to improve the wake-up effect.

[0103] This application Figure 3 A schematic diagram of a computer device is provided. (Example) Figure 3 As shown, the computer device 3 in this embodiment includes a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various wake-up method embodiments described above, for example... Figure 1 Steps 101 to 104 are shown. Alternatively, when processor 301 executes computer program 303, it implements the functions of each module / unit in the above-described wake-up device embodiments, for example... Figure 2 The functions of modules / units 201 to 204 shown.

[0104] This application also provides a readable storage medium storing a computer program, which, when executed by a processor, is used to implement the wake-up methods provided in the various embodiments described above.

[0105] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0106] This application also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the execution of the executable instructions by the at least one processor causes the device to implement the wake-up methods provided in the various embodiments described above.

[0107] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0108] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A wake-up method, characterized in that, include: In response to a target user's trigger action on a wake word, a target audio segment matching the wake word is retrieved from an audio library; The target audio segments are spliced ​​together to generate training audio with the acoustic characteristics of the target user; The training set consists of the noise-added frequencies and the training audio. The preset wake-up model is trained using the training set to obtain a target wake-up model, which is then used to activate the wake-up word.

2. The wake-up method as described in claim 1, characterized in that, The step of responding to a target user's trigger operation on a wake word by retrieving a target audio segment matching the wake word from an audio library includes: The wake word is segmented to obtain the target text unit corresponding to the wake word; Obtain the audio library; The initial audio segment is obtained by retrieving and matching each text unit in the target text unit from the audio library; The audio segment with a signal-to-noise ratio greater than the preset value is selected from the initial audio segments as the target audio segment.

3. The wake-up method as described in claim 2, characterized in that, The process of obtaining the audio library includes: Collect the user's raw audio; The original audio is converted to obtain the corresponding audio text. The audio text is segmented and filtered to obtain the target sub-audio text; Obtain the user ID of the target sub-audio text, the audio segment in the original audio that matches the target sub-audio text, and the audio feature tags of the audio segment; The audio library is constructed by matching the target sub-audio text, the user ID of the target sub-audio text, the audio segments in the original audio that match the target sub-audio text, and the audio feature tags of the audio segments.

4. The wake-up method as described in claim 3, characterized in that, The process of segmenting and filtering the audio text to obtain the target sub-audio text includes: The audio file is segmented to obtain initial sub-audio text; Obtain the time information corresponding to the initial sub-audio text, wherein the time information includes the start time and the end time; Based on the time information, obtain the intermediate sub-audio text; The intermediate sub-audio text is filtered and processed according to preset conditions to obtain the target sub-audio text.

5. The wake-up method as described in claim 4, characterized in that, The step of obtaining the intermediate sub-audio text based on the time information includes: The time interval between adjacent audio segments in the initial sub-audio text is calculated using the start and end times of each sub-audio text in the initial sub-audio text. Based on the time interval, determine the segmentation method of the adjacent audio; The adjacent audio segments are segmented using the adjacent audio segmentation method to obtain the intermediate sub-audio text.

6. The wake-up method as described in claim 5, characterized in that, Determining the segmentation method of adjacent audio based on the time interval includes: If the time interval is greater than or equal to the preset time interval, then the adjacent audio segments are segmented by word. If the time interval is less than the preset time interval, the adjacent audio segments are segmented by whole words.

7. The wake-up method as described in claim 4, characterized in that, The step of filtering the intermediate sub-audio text according to preset conditions to obtain the target sub-audio text includes: Determine whether the clarity, truncation, and overlap of the audio segment corresponding to the intermediate sub-audio text meet preset thresholds; The target sub-audio text is selected from the intermediate sub-audio texts and the sub-audio texts corresponding to the audio segments that meet the preset threshold.

8. The wake-up method as described in claim 2, characterized in that, The step of splicing the target audio segments to generate training audio with the acoustic features of the target user includes: Obtain the audio segments in the target audio segment that match each text unit; The audio segments matched by each text unit are concatenated according to the user ID and audio feature tags to obtain the combined audio corresponding to each text unit; The training audio with the acoustic features of the target user is composed of the combined audio corresponding to each text unit.

9. The wake-up method as described in claim 1, characterized in that, The training set, consisting of noise-added frequencies and the training audio, includes: Collect noise audio; The training audio is noise-added using the noise audio to obtain the noise-added audio; The noise-added audio is fused with the training audio to obtain the training set.

10. A wake-up device, characterized in that, include: The audio acquisition module is used to retrieve a target audio segment that matches the wake word from the audio library in response to the target user's trigger operation on the wake word; An audio splicing module is used to splice the target audio segments to generate training audio with the acoustic characteristics of the target user; The training set acquisition module is used to construct a training set from the noise-added frequency and the training audio; The wake-up module is used to train a preset wake-up model using the training set to obtain a target wake-up model, so as to execute the wake-up word through the target wake-up model.

11. A computer device, characterized in that, Includes a memory, and one or more processors communicatively connected to the memory; The memory stores instructions that can be executed by the one or more processors to cause the one or more processors to implement the wake-up method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, Includes a program or instructions that, when run on a computer, implement the wake-up method according to any one of claims 1 to 9.

13. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the wake-up method according to any one of claims 1 to 9.