Voice assistant wake-up method, apparatus, device, and storage medium

By comprehensively processing vibration signals and cockpit image data, low-power and accurate wake-up of the intelligent voice assistant in new energy vehicles has been achieved, solving the problems of high power consumption and false wake-up caused by microphone monitoring, and improving user experience and safety.

CN122363486APending Publication Date: 2026-07-10DONGFENG LIUZHOU MOTOR
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGFENG LIUZHOU MOTOR
Filing Date
2026-04-09
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

The current method of waking up intelligent voice assistants in new energy vehicles relies on continuous microphone listening, which leads to high power consumption and a high false wake-up rate, affecting the user experience.

Method used

By extracting vibration features from vibration signals and matching them with a preset action library, combined with face matching and interactive environment scoring from cockpit image data, low-power and accurate wake-up can be achieved.

Benefits of technology

It reduces the standby power consumption of the voice assistant, reduces false wake-ups, improves security and user experience, and adapts to different vehicle usage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122363486A_ABST
    Figure CN122363486A_ABST
Patent Text Reader

Abstract

This application discloses a voice assistant wake-up method, device, equipment, and storage medium, relating to the field of voice control technology. The disclosed voice assistant wake-up method includes: extracting vibration features based on vibration signals; matching the vibration features with a preset action library to obtain action matching degree and action matching timing information; determining a face matching score and an interaction environment score based on cockpit image data when the action matching degree is greater than or equal to a preset action matching threshold and the action matching timing information meets the corresponding preset action timing conditions; determining a voice assistant wake-up score based on the action matching degree, face matching score, and interaction environment score; and generating a voice assistant wake-up control command when the voice assistant wake-up score is greater than a preset threshold to wake up the voice assistant. This solution can achieve low-power and accurate wake-up of intelligent voice assistants in new energy vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice control technology, and in particular to voice assistant wake-up methods, devices, equipment and storage media. Background Technology

[0002] With the accelerated development of intelligent new energy vehicles, the frequency of use of intelligent voice assistants and the requirements for user experience are constantly increasing. The performance of their wake-up control has become an important consideration in the design of intelligent interactive systems for new energy vehicles.

[0003] The current method of waking up intelligent voice assistants in new energy vehicles mainly relies on continuous microphone monitoring combined with keyword recognition technology. The microphone is always in standby monitoring mode, continuously collecting voice signals in the car. When a preset wake-up keyword appears in the collected voice signal, the keyword recognition technology is used to complete the matching and judgment, thereby waking up the voice assistant.

[0004] However, traditional wake-up methods based on continuous microphone monitoring and keyword recognition keep the microphone in standby mode for extended periods, resulting in excessive power consumption. Furthermore, ambient noise such as video playback and conversations can interfere with keyword recognition, leading to a high false wake-up rate for voice assistants and severely impacting the user experience. Therefore, achieving low-power, accurate wake-up of intelligent voice assistants in new energy vehicles remains a problem to be solved.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a voice assistant wake-up method, device, equipment and storage medium, aiming to solve the technical problem of how to achieve low-power and accurate wake-up of intelligent voice assistants in new energy vehicles.

[0007] To achieve the above objectives, this application proposes a voice assistant wake-up method, which includes: Vibration features are extracted based on vibration signals; The vibration characteristics are matched with a preset action library to obtain the action matching degree and action matching timing information. When the action matching degree is greater than or equal to the preset action matching threshold, and the action matching timing information meets the corresponding preset action timing conditions, the face matching score and the interaction environment score are determined based on the cockpit image data. When the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold, the voice assistant wake-up score is determined based on the action matching degree, the face matching score and the interaction environment score. When the voice assistant wake-up score is greater than a preset threshold, a voice assistant wake-up control command is generated to wake up the voice assistant.

[0008] In one embodiment, the step of determining the face matching score and the interaction environment score based on the cockpit image data includes: The cockpit image data is preprocessed to obtain a preprocessed image; The face matching score and the interaction environment score are determined based on the preprocessed image.

[0009] In one embodiment, the step of determining the face matching score and the interaction environment score based on the preprocessed image includes: Facial features are extracted based on the preprocessed image; The facial features are compared with a preset authorized user feature database to obtain a facial matching score; The user's head posture and gaze direction are determined based on the preprocessed image; The interaction environment score is determined based on the user's head posture and the user's gaze direction.

[0010] In one embodiment, the step of performing image preprocessing on the cockpit image data to obtain a preprocessed image includes: The cockpit image data is subjected to illumination correction processing to obtain a corrected image; The corrected image is then subjected to face alignment processing to obtain an aligned image; The aligned image is subjected to occlusion compensation processing to obtain an occlusion compensation image; When the ambient light intensity is detected to be lower than the preset light threshold, the infrared mode or the fill light is activated to perform illumination compensation on the occlusion compensation image to obtain a pre-processed image.

[0011] In one embodiment, the step of determining the voice assistant wake-up score based on the action matching degree, the face matching score, and the interaction environment score includes: The vibration intention weight coefficient, identity confidence weight coefficient, and environmental suitability confidence weight coefficient are determined based on the vehicle scenario. The voice assistant wake-up score is determined based on the vibration intention weight coefficient, the identity confidence weight coefficient, the environmental suitability confidence weight coefficient, the action matching degree, the face matching score, and the interaction environment score.

[0012] In one embodiment, the step of extracting vibration features based on vibration signals includes: Based on the vibration signal, the vehicle body vibration signal is filtered out to obtain the filtered vibration signal; The filtered vibration signal is denoised and gravity compensated to obtain a preprocessed vibration signal; Based on the preprocessed vibration signal, the peak acceleration, energy integral, and time-domain waveform matching degree are determined to obtain vibration characteristic parameters. Based on the vibration characteristic parameters, the user's action pattern and action duration information are determined to obtain vibration characteristics.

[0013] In one embodiment, the method further includes: When the action matching degree is less than the preset action matching threshold, return to the step of extracting vibration features based on vibration signals; If the face matching score is less than a preset face matching threshold, switch to a wake-up rejection state, record the wake-up event information, and push the wake-up event information to the authorized user terminal; When the interactive environment score is less than a preset environment score threshold, a voice prompt message is generated to prompt the authorized user to face the preset voice interaction interface.

[0014] Furthermore, to achieve the above objectives, this application also proposes a voice assistant wake-up device, which includes: The data acquisition module is used to extract vibration features based on vibration signals; The data processing module is used to match the vibration characteristics with a preset motion library to obtain motion matching degree and motion matching timing information; The data processing module is further configured to determine the face matching score and the interactive environment score based on the cockpit image data when the action matching degree is greater than or equal to the preset action matching threshold and the action matching timing information meets the corresponding preset action timing conditions. The wake-up scoring module is used to determine the voice assistant wake-up score based on the action matching degree, the face matching score, and the interaction environment score when the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold. The voice wake-up module is used to generate a voice assistant wake-up control command to wake up the voice assistant when the voice assistant wake-up score is greater than a preset threshold.

[0015] In addition, to achieve the above objectives, this application also proposes a voice assistant wake-up device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voice assistant wake-up method as described above.

[0016] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the voice assistant wake-up method described above.

[0017] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the voice assistant wake-up method described above.

[0018] One or more technical solutions proposed in this application have at least the following technical effects: By extracting vibration signals and matching them with a preset action library to determine the physical operation intent, the system effectively filters out false triggers caused by irrelevant vibrations such as road bumps, making the initial wake-up judgment more accurate. Based on low-power monitoring by vibration sensors, it replaces the continuous operation of microphones and cameras, reducing the standby power consumption of the voice assistant. By using cockpit image data to complete face matching and interaction environment scoring, it not only prevents unauthorized use by unauthorized users through identity verification, improving the safety of vehicle use, but also avoids activating the voice assistant in inappropriate scenarios such as distracted driving or making phone calls, reducing interference during driving. By weighted fusion calculation of three indicators to obtain a comprehensive wake-up score, the wake-up mechanism can adapt to different vehicle usage scenarios, achieving low-power and accurate wake-up of intelligent voice assistants in new energy vehicles, solving the problem of high false wake-up rate of traditional voice assistants due to environmental interference. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating an embodiment of the voice assistant wake-up method of this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the voice assistant wake-up method of this application; Figure 3 A simplified flowchart illustrating the voice assistant wake-up method provided in Embodiment 2 of this application; Figure 4This is a schematic diagram of the module structure of the voice assistant wake-up device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the voice assistant wake-up method in this application embodiment.

[0022] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0024] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0025] The main solution of this application embodiment is as follows: extract vibration features based on vibration signals; match the vibration features with a preset action library to obtain action matching degree and action matching timing information; when the action matching degree is greater than or equal to a preset action matching threshold and the action matching timing information meets the corresponding preset action timing conditions, determine the face matching score and the interaction environment score based on the cockpit image data; when the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold, determine the voice assistant wake-up score based on the action matching degree, face matching score, and interaction environment score; when the voice assistant wake-up score is greater than a preset threshold, generate a voice assistant wake-up control command to wake up the voice assistant.

[0026] The current method of waking up intelligent voice assistants in new energy vehicles mainly relies on continuous microphone monitoring combined with keyword recognition technology. The microphone is always in standby monitoring mode, continuously collecting voice signals in the car. When the preset wake-up keyword appears in the collected voice signal, the keyword recognition technology is used to complete the matching and judgment, thereby waking up the voice assistant.

[0027] However, traditional wake-up methods based on continuous microphone monitoring and keyword recognition keep the microphone in standby mode for extended periods, resulting in excessive power consumption. Furthermore, ambient noise such as video playback and conversations can interfere with keyword recognition, leading to a high false wake-up rate for voice assistants and severely impacting the user experience. Therefore, achieving low-power, accurate wake-up of intelligent voice assistants in new energy vehicles remains a problem to be solved.

[0028] This application provides a solution that uses vibration signal extraction and a preset action library matching to perform a dual determination of physical operation intent, effectively filtering out false triggers caused by irrelevant vibrations such as road bumps, making the initial wake-up judgment more accurate. Based on low-power monitoring by vibration sensors, it replaces the continuous operation of microphones and cameras, reducing the standby power consumption of the voice assistant. By using cockpit image data to complete face matching and interaction environment scoring, it not only prevents unauthorized use by unauthorized users through identity verification, improving the safety of vehicle use, but also avoids activating the voice assistant in inappropriate scenarios such as distracted driving or making phone calls, reducing interference during driving. By weighted fusion calculation of three indicators to obtain a comprehensive wake-up score, the wake-up mechanism can adapt to different vehicle usage scenarios, achieving low-power and accurate wake-up of intelligent voice assistants in new energy vehicles, solving the problem of high false wake-up rate of traditional voice assistants due to environmental interference.

[0029] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or voice assistant wake-up device capable of the above functions. The following description uses a voice assistant wake-up device as an example to illustrate this embodiment and the subsequent embodiments.

[0030] Based on this, the embodiments of this application provide a method for waking up a voice assistant, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the voice assistant wake-up method of this application.

[0031] In this embodiment, the voice assistant wake-up method includes steps S10 to S50: Step S10: Extract vibration features based on vibration signals; It should be noted that vibration signals are acceleration-related data generated by physical operations performed by the user inside the vehicle, as well as by the vehicle's own movement and component operation. This data is raw data collected by sensors deployed inside the vehicle, including signals generated by effective operations and various irrelevant background interference signals. Vibration characteristics are various parameters extracted from the vibration signals that characterize the user's physical operations.

[0032] It should be understood that the process involves first acquiring the raw vibration signals from easily touched locations inside the vehicle, filtering out low-frequency background noise such as vehicle body vibrations, retaining the high-frequency components generated by physical operations, and then calculating relevant parameters through a sliding window to extract the corresponding vibration features from the processed vibration signals.

[0033] In addition, before extracting vibration features based on vibration signals, micro-electro-mechanical system (MEMS) accelerometers can be deployed in easily touched locations inside the vehicle, such as steering wheel spokes, center console surfaces, and door armrests. These sensors operate continuously at a low preset sampling rate, such as 50Hz, to continuously collect vibration signals inside the vehicle, providing raw data for subsequent feature extraction.

[0034] Step S20: Match the vibration features with a preset motion library to obtain motion matching degree and motion matching timing information; It should be noted that the preset action library is a pre-defined collection of vibration feature models and timing rules corresponding to various wake-up actions. Each wake-up action model in the library corresponds to a preset physical operation's vibration and timing features, such as tapping the steering wheel twice with two fingers, lightly patting the side of the center console, or specific gesture trajectories—common wake-up operation models used by drivers and passengers. Action matching degree is a quantified value of the similarity between the vibration features and the models in the preset action library, reflecting the degree of fit between the physical action and the wake-up action.

[0035] Additionally, action matching timing information refers to the time-related characteristics of the user's actual physical operation process, including the number of times the operation is executed and the time interval between each operation. Different preset wake-up actions correspond to different timing requirements.

[0036] It should be understood that the actual vibration features extracted from the vibration signal are input into a preset action library, and compared one by one with the vibration features corresponding to all preset wake-up action models in the library. A feature comparison algorithm is used to calculate the similarity of each set of features. Simultaneously with feature comparison, time-related information of the actual physical operation is extracted, recording the number of times the operation is completed, the time interval between two adjacent operations, and other temporal information, forming corresponding action matching temporal information. This allows the matching judgment to cover both feature and temporal dimensions. Based on the feature comparison results, the similarity value between the actual vibration features and each preset wake-up action model is obtained, i.e., the action matching degree. At the same time, the corresponding action matching temporal information is output, completing the overall matching process with the preset action library.

[0037] Step S30: When the action matching degree is greater than or equal to the preset action matching threshold and the action matching timing information meets the corresponding preset action timing conditions, determine the face matching score and the interaction environment score based on the cockpit image data. It should be noted that the preset action matching threshold is a pre-set critical value used to determine whether the vibration characteristics match the preset action library model effectively. It serves as the standard for distinguishing between suspected wake-up intentions and invalid actions. The preset action timing conditions are pre-set rules used to determine whether the physical operation time-related data meets the wake-up action requirements, including requirements such as action interval and execution duration. For example, a wake-up operation with a tapping time interval of 0.2-0.5 seconds corresponds to a preset action timing condition; only operations meeting this condition are considered to have valid timing.

[0038] Additionally, the cockpit image data consists of high-definition images of the vehicle's interior cockpit area captured by driver status monitoring cameras or cockpit cameras deployed near the rearview mirror. The capture area primarily covers the driver and passenger areas, prioritizing image capture of the corresponding area based on the location of vibrations. The face matching score refers to the degree of similarity between facial features extracted from the cockpit image data and features in the authorized user feature library stored in the vehicle's infotainment system, expressed numerically as the face recognition matching result. Furthermore, the interaction environment score is a quantitative score obtained by analyzing the user's head posture, gaze direction, etc., based on the cockpit image data. It is used to determine whether the current in-vehicle interaction environment is suitable for activating the voice assistant; a higher score indicates a more suitable interaction environment.

[0039] It should be understood that the motion matching degree is first numerically judged to determine whether the actual obtained motion matching degree is greater than or equal to the preset motion matching threshold. If it is not reached, the subsequent process is terminated directly; if it is reached, the timing condition judgment continues. Then, the motion matching timing information is judged to determine whether the actual number of operations, operation intervals, and other timing information meet the preset motion timing conditions corresponding to the operation. If they do not meet the conditions, the process is terminated; if both the matching degree threshold and the timing conditions are met, a wake-up signal is sent to the camera.

[0040] Furthermore, upon receiving the wake-up signal, the camera immediately acquires high-definition images of the cabin, obtaining cabin image data. Simultaneously, it preprocesses the acquired cabin image data. Then, a lightweight deep learning model detects faces and extracts facial features from the preprocessed cabin image data. This facial feature is compared with the authorized user feature database to obtain a face matching score. At the same time, the user's head posture, gaze direction, and other information are analyzed to comprehensively determine the corresponding interaction environment score.

[0041] In one feasible implementation, step S30, which involves determining the face matching score and the interaction environment score based on the cockpit image data, may include steps S31-S32: Step S31: Perform image preprocessing on the cockpit image data to obtain a preprocessed image; It should be noted that cockpit image data refers to high-definition image information collected from the vehicle's cockpit area, covering various visual content such as the people in the cockpit, their facial expressions, and the surrounding environment.

[0042] In addition, image preprocessing is an optimization operation performed on the raw acquired cockpit image data. The preprocessed image is an optimized image obtained after image preprocessing of the cockpit image data, with better visual quality, which can meet the recognition requirements of subsequent face detection and environmental analysis.

[0043] It should be understood that the process involves first acquiring cockpit image data, then performing illumination correction for complex scenarios such as backlighting, nighttime, and wearing sunglasses, followed by face alignment of the face areas in the image, and occlusion detection. In low-light environments, the system can automatically switch to infrared mode or activate supplementary lighting. After all processing operations are completed, a pre-processed image is obtained.

[0044] In one feasible implementation, step S31 may include steps S311 to S314: Step S311: Perform illumination correction processing on the cockpit image data to obtain a corrected image; It should be noted that illumination correction processing is an image brightness and color adjustment operation performed on cockpit image data to address issues such as uneven lighting and backlighting, thereby improving the visual presentation of the image. The corrected image is the image obtained after illumination correction processing of the cockpit image data, with more uniform lighting distribution and clearer image details.

[0045] It should be understood that, in order to identify lighting problems such as backlighting and excessive contrast between light and dark in the cockpit image data, corresponding algorithms are used to perform brightness equalization and color restoration. After completing the lighting correction process, the corresponding corrected image is obtained.

[0046] Step S312: Perform face alignment processing on the corrected image to obtain an aligned image; It should be noted that face alignment processing addresses the position and angle issues of face regions in the corrected image. This optimization operation, achieved through facial landmark detection and image geometric transformation, adjusts the detected faces to preset standard positions and angles, unifying the visual recognition benchmark and avoiding subsequent feature extraction errors caused by deviations in face position and angle. The aligned image is the image obtained after face alignment processing of the corrected image, with facial feature points such as eyes, nose, and mouth positioned in preset standard locations.

[0047] It should be understood that the process begins by using a lightweight face detection algorithm to accurately locate the face region in the calibration image, simultaneously identifying the specific coordinates of key feature points such as the eyes, eyebrows, nose, mouth, and chin. Then, based on a pre-set standard face recognition template in the vehicle's infotainment system, the positional offset, angular rotation, and scaling values ​​between the face region in the calibration image and the standard template are calculated. Through geometric transformations such as image translation, rotation, and scaling, the detected face region is gradually adjusted to the preset standard position and angle, ensuring accurate matching between the key feature points of the face and the feature points of the standard template. After completing the geometric transformation, edge pixel completion and detail optimization are performed on the transformed image to ensure image integrity and feature clarity, avoiding image loss due to the transformation operation, ultimately resulting in an aligned image.

[0048] Step S313: Perform occlusion compensation processing on the aligned image to obtain an occlusion compensation image; It's important to note that occlusion compensation processing addresses the occlusion of key facial regions in aligned images. This image optimization operation utilizes image inpainting algorithms and feature completion techniques. It combines facial symmetry features with information from surrounding unoccluded areas to perform pixel-level repair and feature restoration of the occluded region. This compensates for the loss of facial features caused by occlusion, ensuring the completeness of subsequent facial feature extraction and adapting to common usage scenarios such as users wearing sunglasses or masks inside vehicles. The occlusion-compensated image is the result of occlusion compensation processing on the aligned image, where effective features of the occluded facial region are preserved or completed.

[0049] It should be understood that the process involves identifying the occlusion area and type of the face in the aligned image, using an appropriate algorithm to complete the feature of the occlusion area, and then obtaining the corresponding occlusion compensation image after the occlusion compensation process is completed.

[0050] Step S314: When the ambient light intensity is detected to be lower than the preset light threshold, the infrared mode is activated or the fill light is activated to perform illumination compensation on the occlusion compensation image to obtain a preprocessed image.

[0051] It should be noted that ambient light intensity refers to the actual light intensity inside the cockpit, reflecting the current brightness of the cockpit. The preset light threshold is a pre-set critical value for determining whether the cockpit lighting meets the requirements for image recognition.

[0052] Additionally, illumination compensation is a brightness enhancement and detail restoration operation performed on occlusion-compensated images in low-light environments, improving image quality in low light. The preprocessed image is the final optimized image obtained after illumination compensation of the occlusion-compensated image.

[0053] It should be understood that the actual ambient light intensity in the area captured by the camera image inside the cabin is first detected. This can be quantified by extracting the average brightness of all pixels in the occlusion compensation image, or by obtaining real-time light values ​​from the vehicle's light sensor. The detected ambient light intensity is then compared with a pre-set light threshold. If the detected ambient light intensity is higher than or equal to the pre-set light threshold, the current lighting environment is deemed suitable for visual recognition, and no further operation is required; the occlusion compensation image is directly used as the final pre-processed image. If the detected ambient light intensity is lower than the pre-set light threshold, it is determined to be a low-light environment, and the corresponding light compensation operation is immediately initiated. Two methods can be selected depending on the actual scenario: one is to activate the camera's infrared mode, using infrared imaging technology to re-capture images of the cabin area, and then fusing the re-captured clear infrared image with the facial features of the occlusion compensation image; the other is to activate the supplementary lighting near the camera to increase the visible light brightness of the acquisition area, performing professional brightness enhancement processing on the occlusion compensation image. Through the above light compensation operations, the final pre-processed image is obtained.

[0054] Step S32: Determine the face matching score and the interaction environment score based on the preprocessed image.

[0055] It should be noted that the face matching score ranges from 0 to 1. The higher the score, the higher the degree of matching between the actual facial features and the authorized user's features. If the score does not reach the preset threshold, the user will be directly identified as an unauthorized user.

[0056] In addition, the interactive environment score ranges from 0 to 1. The higher the score, the more the user's current head posture and gaze direction meet the requirements of the voice interaction scenario. The score takes into account multiple actual usage dimensions such as whether the user is facing the central control screen, whether they are driving with distraction, and whether they are making or receiving a phone call.

[0057] It should be understood that a lightweight deep learning model is first used to perform face detection and feature extraction on the preprocessed image. The extracted face features are compared one by one with the authorized user feature library to output the corresponding face matching score. Then, the head posture and gaze direction of the person in the preprocessed image are analyzed. Combined with the requirements of the voice interaction scenario, the environmental suitability is determined, and finally, the corresponding interaction environment score is given.

[0058] In one feasible implementation, step S32 may include steps S321 to S324: Step S321: Extract facial features based on the preprocessed image; It should be noted that the preprocessed image is an image obtained after a series of optimization processes on the cockpit image data. It has good visual quality and can meet various recognition requirements for face detection and feature extraction.

[0059] In addition, facial features are feature data extracted from the facial region that can characterize the unique attributes of the face. They cover specialized information such as facial key points and textures, and are the core basis for distinguishing the identities of different people.

[0060] It should be understood that lightweight deep learning models such as MobileNet and FaceNet are used to detect and locate face regions in preprocessed images, and then feature extraction algorithms are used to extract unique facial features that identify a person from the located face regions.

[0061] Step S322: Compare the facial features with a preset authorized user feature database to obtain a facial matching score; It should be noted that the preset authorized user feature library is a database of facial features of authorized users that is pre-stored in the vehicle's infotainment system to verify the voice assistant's usage rights. The library contains a standardized set of facial feature parameters of users with voice assistant usage rights, such as drivers and frequent passengers.

[0062] It should be understood that the extracted facial features are compared with the features of authorized users in the preset authorized user feature library one by one to calculate the similarity. The highest value is selected from all the calculated similarity values ​​and is determined as the final facial matching score. If there is no facial feature of any authorized user in the preset authorized user feature library that has a valid similarity with the extracted facial features, the facial matching score is directly determined to be 0.

[0063] Step S323: Determine the user's head pose and gaze direction based on the preprocessed image; It should be noted that user head posture refers to the physical posture characteristics of a user's head in the cabin of a new energy vehicle, such as the angle of head placement and spatial position. It includes multiple dimensions such as the left and right rotation angle, the up and down tilt angle, and the forward and backward tilt angle. It is an important basis for judging whether the user is in a suitable state for voice interaction. For example, when the driver turns his head to talk to the rear passenger, it can be judged as an unsuitable head posture for voice interaction.

[0064] Additionally, the user's gaze direction refers to the spatial direction in which the user's eyes are directed. When analyzing the user's gaze direction, it is determined whether the gaze is directed towards the central control screen, whether the user is observing the road conditions, or whether the user is looking towards the area where they are making or receiving calls, such as a mobile phone.

[0065] It should be understood that key point detection is performed on the user's head region in the preprocessed image, the user's head pose is calculated based on the coordinates of the detected key points, and then the user's eye state is analyzed through an eye-tracking algorithm to determine the corresponding user's gaze direction.

[0066] Step S324: Determine the interaction environment score based on the user's head posture and the user's gaze direction.

[0067] It should be noted that the quantitative scoring rules are a pre-defined system that assigns corresponding base scores to different user head postures and gaze directions. These rules are developed based on the actual needs of voice interaction in new energy vehicles, assigning higher base scores to suitable interaction states and lower base scores to unsuitable ones. For example, the interaction environment score is 0.9 when the user is facing the central control screen and paying attention, and 0.5 when the user is observing the rearview mirror.

[0068] It should be understood that, based on the requirements of the voice interaction scenario, a comprehensive analysis of the user's head posture and gaze direction is performed. According to preset state classification standards, the results of these two indicators are divided into different state levels. For example, the user's head posture is classified as normal interaction posture or distracted driving posture, and the user's gaze direction is classified as facing the central control screen, facing the driving road, or facing other areas. Then, according to pre-set quantitative scoring rules, corresponding base scores are assigned to different state levels of the user's head posture and gaze direction. A gaze direction facing the central control screen and a normal interaction head posture are assigned higher base scores, while a head posture indicating distracted driving and a gaze direction not facing the central control screen are assigned lower base scores. Next, a weighted fusion method is used to comprehensively calculate the base scores corresponding to the user's head posture and gaze direction according to preset weights, resulting in a comprehensive quantitative value. This value is determined as the final interaction environment score. If the user is detected to be in a state completely unsuitable for interaction, such as making or receiving a phone call or being severely distracted while driving, the interaction environment score is directly judged as low.

[0069] Step S40: When the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold, determine the voice assistant wake-up score based on the action matching degree, the face matching score and the interaction environment score. It should be noted that the preset face matching threshold is a pre-set threshold used to determine whether a person is an authorized user, and it is the core standard for successful identity verification. The preset environment scoring threshold is a pre-set threshold used to determine whether the current cabin environment is suitable for voice interaction, and it is the key standard for successful environment verification.

[0070] In addition, the voice assistant wake-up score is a comprehensive quantitative value obtained by fusing action matching degree, face matching score and interaction environment score, and is used to evaluate the effectiveness and matching degree of voice assistant wake-up.

[0071] It should be understood that the face matching score is first verified to meet the preset face matching threshold, and the interaction environment score is confirmed to meet the preset environment score threshold. When both conditions are met, the action matching degree, face matching score and interaction environment score are spatiotemporally aligned and weighted and fused. The voice assistant wake-up score is calculated and determined through a reasonable fusion method.

[0072] In one feasible implementation, step S40, which determines the voice assistant wake-up score based on the action matching degree, the face matching score, and the interaction environment score, may include steps S41-S42: Step S41: Determine the vibration intent weight coefficient, identity confidence weight coefficient, and environmental suitability confidence weight coefficient based on the vehicle scenario. It should be noted that vehicle scenarios refer to the scenario types corresponding to the actual driving and usage states of new energy vehicles. These scenarios are specific to the vehicle based on real-time parameters such as vehicle speed, parking brake status, gear information, and road conditions. They mainly include parking scenarios, driving scenarios, and low-speed driving scenarios. Different vehicle scenarios have different focuses in determining the wake-up of the voice assistant, which is the basis for the allocation of weight coefficients. The determination results will determine the degree of influence of each dimension index in the wake-up score calculation.

[0073] In addition, the vibration intention weighting coefficient is a pre-set parameter used to weight the action matching degree, reflecting the importance of the action matching degree in the calculation of the voice assistant wake-up score.

[0074] In addition, the identity confidence weighting coefficient is a pre-set parameter used to weight the face matching score, reflecting the importance of the face matching score in the calculation of the voice assistant wake-up score.

[0075] In addition, the environmental suitability confidence weighting coefficient is a pre-set parameter used to weight the interaction environment score, reflecting the importance of the interaction environment score in the calculation of the voice assistant wake-up score.

[0076] It should be understood that various operating parameters of new energy vehicles are collected in real time, including vehicle speed, parking brake activation / deactivation status, gear information, and road condition data. These parameters are used to determine the current vehicle scenario, identifying whether the vehicle is in a parking, driving, or low-speed driving scenario. Then, a pre-defined rule for matching scenarios and weights is retrieved. This rule is formulated based on the usage needs and safety requirements of different vehicle scenarios, and an initial weight coefficient allocation scheme is configured for each scenario. According to this rule, corresponding vibration intent weight coefficients, identity confidence weight coefficients, and environmental suitability confidence weight coefficients are matched to the currently determined vehicle scenario. Furthermore, these weight coefficients can be dynamically adjusted dynamically based on actual usage needs and safety standards. For example, the vibration intent weight coefficient is increased in parking scenarios, the environmental suitability confidence weight coefficient is increased in driving scenarios, and the identity confidence weight coefficient maintains a high proportion in all scenarios, thus completing the determination of the three weight coefficients.

[0077] Step S42: Determine the voice assistant wake-up score based on the vibration intention weight coefficient, the identity confidence weight coefficient, the environmental suitability confidence weight coefficient, the action matching degree, the face matching score, and the interaction environment score.

[0078] It should be understood that the weighted score for action matching is obtained by multiplying the action matching degree by the vibration intent weight coefficient; the weighted score for face matching is obtained by multiplying the face matching score by the identity confidence weight coefficient; and the weighted score for interaction environment is obtained by multiplying the interaction environment score by the environment suitability confidence weight coefficient. This method quantifies the influence of each indicator in different vehicle scenarios. Next, either a weighted summation or Bayesian fusion calculation method is chosen to combine the three weighted scores. If a weighted summation method is used, the three weighted scores are directly added together; if a Bayesian fusion method is used, a probability model is incorporated for comprehensive calculation. The final comprehensive quantitative value is the voice assistant wake-up score.

[0079] For example, the formula for calculating the voice assistant wake-up score using a weighted summation method is as follows:

[0080] In the formula, This represents the voice assistant wake-up score, used to quantitatively evaluate the wake-up matching degree and wake-up effectiveness of intelligent voice assistants; This represents a pre-defined vibration intent weighting coefficient, used for weighting the vibration intent confidence value. This represents a pre-defined identity confidence weighting coefficient, used for weighting identity confidence values. This represents the pre-defined confidence level weighting coefficient for environmental suitability, used for weighting the environmental suitability confidence values. This indicates the accuracy of the action matching, representing the degree to which the wake-up action of the intelligent voice assistant in a new energy vehicle is matched. This represents the face matching score, specifically the face recognition matching score during the wake-up process of the intelligent voice assistant in new energy vehicles. This represents the interaction environment score, a comprehensive result of posture and gaze scores in the context of waking up the intelligent voice assistant in a new energy vehicle. , , It can be dynamically adjusted according to the vehicle usage scenario.

[0081] Step S50: When the voice assistant wake-up score is greater than a preset threshold, a voice assistant wake-up control command is generated to wake up the voice assistant.

[0082] It should be noted that the preset threshold is a pre-set critical value used to determine whether the conditions for voice assistant wake-up are met; it is the standard for triggering the voice assistant wake-up operation. The voice assistant wake-up control command refers to the control signal generated after all wake-up conditions are met, used to instructively activate the voice assistant. This signal is transmitted to the voice assistant's control module. The voice assistant is an intelligent functional module in new energy vehicles that enables voice interaction. It can receive user voice commands and complete various vehicle operations such as navigation settings, media playback, air conditioning control, and window adjustment. It is an important component of vehicle intelligent interaction, and its wake-up state represents entering voice command listening mode.

[0083] It should be understood that the calculated voice assistant wake-up score is compared with a preset threshold. If the voice assistant wake-up score exceeds the preset threshold, a corresponding voice assistant wake-up control command is generated. This command is transmitted to the voice assistant's control module in the vehicle via the in-vehicle control link. Upon receiving the wake-up control command, the control module immediately activates the microphone array deployed on the roof or A-pillar. The microphone array enables high-power voice acquisition and processing, and the vehicle's infotainment screen displays a wake-up animation or issues a voice prompt, thus completing the formal wake-up of the voice assistant and putting it into a listening state to receive the user's voice commands. If the voice assistant wake-up score does not exceed the preset threshold, it is determined that the wake-up conditions are not met, the wake-up process is terminated, and no wake-up operation is performed, ensuring that the voice assistant is not ineffectively triggered.

[0084] In addition, after the voice assistant is activated and the user's original voice command is collected, the original voice signal will first undergo noise reduction and echo cancellation preprocessing to filter out environmental noise and echo interference in the car and obtain a clear voice signal. Then, the voice signal is sent to the voice recognition engine for automatic voice recognition to convert the voice signal into corresponding text information. Subsequently, the semantics of the text information are analyzed through natural language understanding to identify the user's actual operation intention and execute the corresponding vehicle operation according to the identified intention, including various functions such as navigation settings, media playback, air conditioning control, and window adjustment.

[0085] Furthermore, after executing the voice command, the system records relevant data such as vibration characteristics, user identity, and environmental status for each successful wake-up. This data is then uploaded to the cloud in real time via the vehicle network, allowing the cloud to perform professional analysis and mining of the uploaded data. Based on the analysis results, the system continuously optimizes relevant models, including vibration motion models, facial recognition models, and fusion weight parameter models that can be dynamically adjusted according to different scenarios. This continuously improves the accuracy and adaptability of the entire voice assistant wake-up control method, enabling personalized interactive experiences.

[0086] In one feasible implementation, the method may further include steps A10 to A30: Step A10: When the motion matching degree is less than the preset motion matching threshold, return to the step of extracting vibration features based on vibration signals; It should be understood that the obtained action matching degree is compared with the preset action matching threshold. If the action matching degree is lower than the threshold, the current matching result is abandoned and the process returns to the step of extracting vibration features based on vibration signals to perform a new round of vibration feature extraction and subsequent matching operations.

[0087] Step A20: When the face matching score is less than the preset face matching threshold, switch to the wake-up rejection state, record the wake-up event information, and push the wake-up event information to the authorized user terminal; It should be noted that the wake-up denial state refers to a working state where the voice assistant is prohibited from being activated. In this state, the voice assistant will not respond to any wake-up-related trigger operations. Wake-up event information is a collection of information formed by recording the relevant details of this unauthorized wake-up operation, including the specific time the operation occurred, the vehicle's current geographical location, the physical operation type corresponding to the vibration characteristics that triggered the operation, and visual information captured by the cockpit image. Furthermore, the authorized user terminal refers to the mobile terminal or dedicated in-vehicle terminal used by the authorized user of the vehicle, which is the carrier for receiving wake-up event push information.

[0088] It should be understood that the face matching score is compared with the preset face matching threshold. If the face matching score is lower than the threshold, the voice assistant is switched to a wake-up rejection state. At the same time, various relevant data of this wake-up attempt are recorded to form wake-up event information, and then the wake-up event information is pushed to the authorized user terminal used by the vehicle's authorized user.

[0089] In step A30, if the interactive environment score is less than a preset environment score threshold, a voice prompt message is generated to prompt the authorized user to face the preset voice interaction interface.

[0090] It should be noted that voice prompts are audio messages used to provide voice reminders to authorized users, directly conveying the environmental adaptation requirements for voice interaction. The preset voice interaction interface is a pre-defined interface used for voice interaction within the vehicle, which can be the vehicle's central control screen. Furthermore, authorized users refer to individuals who have completed information entry into the vehicle's infotainment system and obtained permission to use the voice assistant, including drivers and frequent passengers.

[0091] It should be understood that, provided the face matching score reaches the preset face matching threshold, if the value of the interaction environment score is lower than the preset environment score threshold, a corresponding voice prompt will be generated and played through the vehicle's audio playback device to remind the authorized user to adjust their head and gaze to face the preset voice interaction interface.

[0092] This embodiment provides a voice assistant wake-up method. It achieves dual determination of physical operation intent through vibration signal extraction and matching with a preset action library, effectively filtering out false triggers caused by irrelevant vibrations such as road bumps, making the initial wake-up judgment more accurate. Based on low-power monitoring with vibration sensors, it replaces the continuous operation of microphones and cameras, reducing the standby power consumption of the voice assistant. It uses cockpit image data to complete face matching and interaction environment scoring, preventing unauthorized use by users through identity verification, improving vehicle safety, and avoiding activation of the voice assistant in inappropriate scenarios such as distracted driving or making phone calls, reducing interference during driving. By weighted fusion calculation of three indicators to obtain a comprehensive wake-up score, the wake-up mechanism can adapt to different vehicle usage scenarios, achieving low-power and accurate wake-up of intelligent voice assistants in new energy vehicles, solving the problem of high false wake-up rates of traditional voice assistants due to environmental interference.

[0093] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S10 may include steps S11 to S14: Step S11: Filter out the vehicle body vibration signal based on the vibration signal to obtain the filtered vibration signal; It should be noted that the filtered vibration signal is the optimized vibration data obtained after the original vibration signal has undergone the vehicle body vibration signal filtering operation. This data retains only the high-frequency vibration components generated by the user's physical operation and eliminates the interference of low-frequency vehicle body vibration.

[0094] It should be understood that the process begins by acquiring the raw vibration signal, which is then filtered in the frequency domain using a high-pass filter. The high-pass filter allows high-frequency signals to pass through while blocking low-frequency signals, thus filtering out the high-frequency components generated by the user's physical operations. This removes low-frequency body vibration signals caused by vehicle movement, engine vibration, etc. During the filtering process, algorithms are used to maintain the original characteristics of the high-frequency vibration signal without distortion, ensuring that key information such as the vibration amplitude and waveform of the user's physical operations are fully preserved. After the filtering operation is completed, a filtered vibration signal containing only the high-frequency vibration information of the user's effective operations is obtained.

[0095] Step S12: The filtered vibration signal is denoised and gravity compensated to obtain a preprocessed vibration signal; It should be noted that denoising is an operation that eliminates stray noise remaining in the filtered vibration signal, which can further improve the purity of the vibration data and reduce noise interference.

[0096] Additionally, gravity compensation processing is a calibration operation performed to address the static deviation caused by gravity acceleration when the accelerometer collects data. Accelerometers deployed in different locations within the vehicle are affected by varying degrees of gravity acceleration due to different installation angles, resulting in fixed numerical deviations in the collected vibration data. This processing accurately eliminates these deviations, ensuring the vibration data accurately reflects the actual vibration conditions during user physical operations. Furthermore, the pre-processed vibration signal is the vibration data obtained after noise reduction and gravity compensation.

[0097] It should be understood that the filtered vibration signal is first denoised. A signal denoising algorithm is used to identify random noise in the signal, eliminating noise while preserving the true vibration waveform, amplitude, and other dynamic characteristics to ensure the purity of the vibration signal. After denoising, gravity compensation processing is performed on the signal. Based on the actual deployment angle and installation position of the accelerometer inside the vehicle, the algorithm calculates the specific impact of gravitational acceleration on the sensor's data acquisition. This impact value is then removed from the vibration signal to eliminate the interference of static deviation, resulting in a high-precision, interference-free pre-processed vibration signal.

[0098] Step S13: Determine the peak acceleration, energy integral, and time-domain waveform matching degree based on the preprocessed vibration signal to obtain vibration characteristic parameters; It should be noted that peak acceleration refers to the maximum acceleration of the pre-processed vibration signal within a fixed time range. It is used to characterize the vibration intensity of the user's physical operation. Different physical operations, such as tapping the steering wheel with two fingers or lightly patting the side of the center console, will produce different peak accelerations due to different operating forces. This indicator can accurately distinguish the differences in the force of the user's operation.

[0099] Additionally, energy integral refers to the cumulative energy value of the preprocessed vibration signal over time, used to characterize the vibration energy of the user's physical operation. It is obtained by integrating the amplitude of the vibration signal over time. Different physical operations will have different energy integral values ​​due to differences in the amplitude of the action and the duration of a single operation. This indicator can help distinguish different types of physical operations. Furthermore, time-domain waveform matching degree refers to the similarity between the time-domain waveform of the preprocessed vibration signal and the preset basic vibration waveform, used to characterize the vibration pattern of the user's physical operation. The time-domain waveform can intuitively reflect the law of vibration signal change over time. Different physical operations will form unique time-domain waveforms. This indicator is the basis for distinguishing the specific type of user operation.

[0100] It should be understood that using a sliding window approach for real-time data analysis and parameter calculation of preprocessed vibration signals allows the sliding window to continuously capture dynamic changes in the signal within a fixed time interval, ensuring the real-time and continuous nature of parameter calculation and adapting to the real-time monitoring needs of vehicle-mounted scenarios. For each sliding window of the preprocessed vibration signal, the algorithm first calculates the maximum acceleration value of the signal within that time interval, defining this value as the peak acceleration. Then, the amplitude of the vibration signal within the window is integrated over time to obtain the cumulative energy value, i.e., the energy integral, within that time interval. Next, a comprehensive comparison and analysis is performed between the time-domain waveform of the preprocessed vibration signal within the window and a preset baseline vibration waveform, calculating the similarity between the two, i.e., the time-domain waveform matching degree. Finally, the calculated peak acceleration, energy integral, and time-domain waveform matching degree are integrated to form complete vibration characteristic parameters.

[0101] Step S14: Determine the user's action pattern and action duration information based on the vibration characteristic parameters to obtain vibration characteristics.

[0102] It should be noted that the action mode refers to the specific physical operation type performed by the user based on the vibration characteristic parameters. For example, the preset voice assistant wake-up action types include tapping the steering wheel twice with two fingers, lightly tapping the side of the center console, specific gesture trajectories such as drawing circles, and horizontal swiping. Different action modes correspond to different combinations of vibration characteristic parameters and are identifiers that distinguish different wake-up operation types of users.

[0103] Additionally, the action duration information refers to the time consumed by a user to complete a preset physical operation, including time dimension features such as the duration of a single operation and the time interval between multiple consecutive operations.

[0104] It should be understood that the vibration characteristic parameters are first comprehensively compared with a preset operation type feature library. This library stores the combination patterns of vibration characteristic parameters corresponding to various preset wake-up actions, such as two-finger taps, light taps on the center console, and specific gesture trajectories. Through feature comparison algorithms, the specific physical operation type performed by the current user is determined, and this type is identified as the action mode. If the vibration characteristic parameters do not match the combination pattern of any preset operation type sufficiently, it is determined that there is no valid action mode. Subsequently, the time-series data of the vibration characteristic parameters is analyzed to calculate the duration of a single physical operation performed by the user. If it is a series of consecutive physical operations, the time interval between two adjacent operations is also calculated. These time-dimensional features are integrated to determine the action duration. The determined action mode and the calculated action duration are integrated to form a vibration characteristic that uniquely represents the user's physical operation. This characteristic fully encompasses the type and time pattern of the user's operation.

[0105] This embodiment provides a voice assistant wake-up method that sequentially filters out vehicle vibration signals, performs noise reduction processing, and gravity compensation processing on vibration signals, effectively eliminating the interference of background noise and gravity factors on vibration acquisition. By extracting core vibration feature parameters such as peak acceleration, energy integral, and time-domain waveform matching degree, and combining these parameters to determine action patterns and action duration information and integrating them into vibration features, it can completely and accurately represent the user's specific physical operations, providing a data foundation for subsequent matching with a preset action library. This effectively reduces the occurrence of mismatched actions due to feature extraction errors, and improves the accuracy and reliability of recognizing the user's physical wake-up intent.

[0106] For example, to help understand the implementation flow of the voice assistant wake-up method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 3 , Figure 3 A simplified flowchart of a voice assistant wake-up method is provided, specifically: The process is divided into four stages: continuous vibration monitoring at the perception layer, visual wake-up and acquisition at the perception layer, visual recognition and verification at the decision layer, fusion decision layer, execution and application layer, and personalized model update at the cloud-based vehicle network. Continuous vibration monitoring at the perception layer involves accelerometers or inertial measurement units (IMUs) deployed on the steering wheel or center console to collect vibration signals. These signals are preprocessed, including filtering, noise reduction, and gravity compensation, followed by vibration feature extraction, including tapping patterns, intensity, and duration. It then determines if the signal matches a preset physical action, such as a two-finger tap or a specific gesture. If not, it returns to continue collecting vibration signals; if so, it wakes up the in-vehicle camera to trigger an event. The visual wake-up and acquisition stage at the perception layer captures images of the cabin, including the driver and passenger areas. The images are preprocessed, including lighting correction and face alignment. Then, the visual recognition and verification stage at the decision layer performs face detection and recognition to determine if the user is authorized. If not, the wake-up attempt is rejected and recorded. If authorized, it determines whether the user's gaze is directed towards the screen or whether their posture indicates focused driving. If the user's gaze is not directed towards the screen or they are not focused on driving, the interaction environment is deemed unsuitable. A delayed wake-up or voice prompt to speak towards the screen is implemented, and user characteristics are updated. If the user's gaze is directed towards the screen and they are focused on driving, the process proceeds to the fusion decision layer. Multimodal fusion decision-making is performed by determining environmental suitability (interaction environment score), identity confidence (face matching score), and vibration intent confidence (action matching degree), and a weighted confidence threshold is applied. The fusion decision determines whether the result constitutes a formal wake-up. If not, the system remains in standby or low-power mode, recording events for semi-learning. If yes, the process proceeds to the execution and application layer. The voice assistant is activated, entering listening mode. The microphone array starts, receiving voice commands, followed by speech recognition and natural language understanding. In-vehicle control tasks are executed, and interactive feedback, including voice prompts or screen animations, is provided. Updated user characteristics and interaction-related data are uploaded to the cloud or vehicle network for personalized model updates and optimization. The optimized model is then reused in various stages to improve the accuracy and adaptability of the process.

[0107] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the voice assistant wake-up method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0108] This application also provides a voice assistant wake-up device, please refer to... Figure 4 The voice assistant wake-up device includes: Data acquisition module 10 is used to extract vibration features based on vibration signals; Data processing module 20 is used to match the vibration characteristics with a preset motion library to obtain motion matching degree and motion matching timing information; The data processing module 20 is further configured to determine the face matching score and the interactive environment score based on the cockpit image data when the action matching degree is greater than or equal to the preset action matching threshold and the action matching timing information meets the corresponding preset action timing conditions. The wake-up scoring module 30 is used to determine the voice assistant wake-up score based on the action matching degree, the face matching score, and the interaction environment score when the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold. The voice wake-up module 40 is used to generate a voice assistant wake-up control command to wake up the voice assistant when the voice assistant wake-up score is greater than a preset threshold.

[0109] In one embodiment, the data processing module 20 is further configured to perform image preprocessing on the cockpit image data to obtain a preprocessed image; The face matching score and the interaction environment score are determined based on the preprocessed image.

[0110] In one embodiment, the data processing module 20 is further configured to extract facial features based on the preprocessed image; The facial features are compared with a preset authorized user feature database to obtain a facial matching score; The user's head posture and gaze direction are determined based on the preprocessed image; The interaction environment score is determined based on the user's head posture and the user's gaze direction.

[0111] In one embodiment, the data processing module 20 is further configured to perform illumination correction processing on the cockpit image data to obtain a corrected image; The corrected image is then subjected to face alignment processing to obtain an aligned image; The aligned image is subjected to occlusion compensation processing to obtain an occlusion compensation image; When the ambient light intensity is detected to be lower than the preset light threshold, the infrared mode or the fill light is activated to perform illumination compensation on the occlusion compensation image to obtain a pre-processed image.

[0112] In one embodiment, the wake-up scoring module 30 is further configured to determine the vibration intention weight coefficient, the identity confidence weight coefficient, and the environmental suitability confidence weight coefficient based on the vehicle scenario. The voice assistant wake-up score is determined based on the vibration intention weight coefficient, the identity confidence weight coefficient, the environmental suitability confidence weight coefficient, the action matching degree, the face matching score, and the interaction environment score.

[0113] In one embodiment, the data acquisition module 10 is further configured to filter out the vehicle body vibration signal based on the vibration signal to obtain a filtered vibration signal; The filtered vibration signal is denoised and gravity compensated to obtain a preprocessed vibration signal; Based on the preprocessed vibration signal, the peak acceleration, energy integral, and time-domain waveform matching degree are determined to obtain vibration characteristic parameters. Based on the vibration characteristic parameters, the user's action pattern and action duration information are determined to obtain vibration characteristics.

[0114] In one embodiment, the data processing module 20 is further configured to return to the step of extracting vibration features based on vibration signals when the action matching degree is less than a preset action matching threshold; If the face matching score is less than a preset face matching threshold, switch to a wake-up rejection state, record the wake-up event information, and push the wake-up event information to the authorized user terminal; When the interactive environment score is less than a preset environment score threshold, a voice prompt message is generated to prompt the authorized user to face the preset voice interaction interface.

[0115] The voice assistant wake-up device provided in this application, employing the voice assistant wake-up method in the above embodiments, can solve the technical problem of how to achieve low-power and accurate wake-up of intelligent voice assistants in new energy vehicles. Compared with the prior art, the beneficial effects of the voice assistant wake-up device provided in this application are the same as those of the voice assistant wake-up method provided in the above embodiments, and other technical features in the voice assistant wake-up device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0116] This application provides a voice assistant wake-up device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the voice assistant wake-up method in the first embodiment described above.

[0117] The following is for reference. Figure 5The diagram illustrates a structural schematic suitable for implementing a voice assistant wake-up device according to embodiments of this application. The voice assistant wake-up device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The voice assistant wake-up device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0118] like Figure 5 As shown, the voice assistant wake-up device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the voice assistant wake-up device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the voice assistant wake-up device to communicate wirelessly or wiredly with other devices to exchange data. While the figures show voice assistant wake-up devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0119] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0120] The voice assistant wake-up device provided in this application, employing the voice assistant wake-up method described in the above embodiments, can solve the technical problem of how to achieve low-power and accurate wake-up of intelligent voice assistants in new energy vehicles. Compared with the prior art, the beneficial effects of the voice assistant wake-up device provided in this application are the same as those of the voice assistant wake-up method provided in the above embodiments, and other technical features of this voice assistant wake-up device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0121] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0122] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0123] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the voice assistant wake-up method in the above embodiments.

[0124] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or flash memory, optical fiber, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0125] The aforementioned computer-readable storage medium may be included in the voice assistant wake-up device; or it may exist independently and not assembled into the voice assistant wake-up device.

[0126] The aforementioned computer-readable storage medium carries one or more programs. When one or more of these programs are executed by a voice assistant wake-up device, the voice assistant wakes up the device by: extracting vibration features based on vibration signals; matching the vibration features with a preset action library to obtain action matching degree and action matching timing information; determining a face matching score and an interaction environment score based on cockpit image data when the action matching degree is greater than or equal to a preset action matching threshold and the action matching timing information meets the corresponding preset action timing conditions; determining a voice assistant wake-up score based on action matching degree, face matching score, and interaction environment score when the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold; and generating a voice assistant wake-up control command to wake up the voice assistant when the voice assistant wake-up score is greater than a preset threshold.

[0127] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0129] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0130] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described voice assistant wake-up method, thereby solving the technical problem of how to achieve low-power and accurate wake-up of intelligent voice assistants in new energy vehicles. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the voice assistant wake-up method provided in the above embodiments, and will not be repeated here.

[0131] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the voice assistant wake-up method described above.

[0132] The computer program product provided in this application solves the technical problem of how to achieve low-power and accurate wake-up of intelligent voice assistants in new energy vehicles. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the voice assistant wake-up method provided in the above embodiments, and will not be repeated here.

[0133] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for waking up a voice assistant, characterized in that, The voice assistant wake-up method includes: Vibration features are extracted based on vibration signals; The vibration characteristics are matched with a preset action library to obtain the action matching degree and action matching timing information. When the action matching degree is greater than or equal to the preset action matching threshold, and the action matching timing information meets the corresponding preset action timing conditions, the face matching score and the interaction environment score are determined based on the cockpit image data. When the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold, the voice assistant wake-up score is determined based on the action matching degree, the face matching score and the interaction environment score. When the voice assistant wake-up score is greater than a preset threshold, a voice assistant wake-up control command is generated to wake up the voice assistant.

2. The method as described in claim 1, characterized in that, The steps for determining the face matching score and interaction environment score based on the cockpit image data include: The cockpit image data is preprocessed to obtain a preprocessed image; The face matching score and the interaction environment score are determined based on the preprocessed image.

3. The method as described in claim 2, characterized in that, The steps of determining the face matching score and the interaction environment score based on the preprocessed image include: Facial features are extracted based on the preprocessed image; The facial features are compared with a preset authorized user feature database to obtain a facial matching score; The user's head posture and gaze direction are determined based on the preprocessed image; The interaction environment score is determined based on the user's head posture and the user's gaze direction.

4. The method as described in claim 2, characterized in that, The step of performing image preprocessing on the cockpit image data to obtain a preprocessed image includes: The cockpit image data is subjected to illumination correction processing to obtain a corrected image; The corrected image is then subjected to face alignment processing to obtain an aligned image; The aligned image is subjected to occlusion compensation processing to obtain an occlusion compensation image; When the ambient light intensity is detected to be lower than the preset light threshold, the infrared mode or the fill light is activated to perform illumination compensation on the occlusion compensation image to obtain a pre-processed image.

5. The method according to any one of claims 1 to 4, characterized in that, The steps for determining the voice assistant wake-up score based on the action matching degree, the face matching score, and the interaction environment score include: The vibration intention weight coefficient, identity confidence weight coefficient, and environmental suitability confidence weight coefficient are determined based on the vehicle scenario. The voice assistant wake-up score is determined based on the vibration intention weight coefficient, the identity confidence weight coefficient, the environmental suitability confidence weight coefficient, the action matching degree, the face matching score, and the interaction environment score.

6. The method according to any one of claims 1 to 4, characterized in that, The steps for extracting vibration features based on vibration signals include: Based on the vibration signal, the vehicle body vibration signal is filtered out to obtain the filtered vibration signal; The filtered vibration signal is denoised and gravity compensated to obtain a preprocessed vibration signal; Based on the preprocessed vibration signal, the peak acceleration, energy integral, and time-domain waveform matching degree are determined to obtain vibration characteristic parameters. Based on the vibration characteristic parameters, the user's action pattern and action duration information are determined to obtain vibration characteristics.

7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: When the action matching degree is less than the preset action matching threshold, return to the step of extracting vibration features based on vibration signals; If the face matching score is less than a preset face matching threshold, switch to a wake-up rejection state, record the wake-up event information, and push the wake-up event information to the authorized user terminal; When the interactive environment score is less than a preset environment score threshold, a voice prompt message is generated to prompt the authorized user to face the preset voice interaction interface.

8. A voice assistant wake-up device, characterized in that, The device includes: The data acquisition module is used to extract vibration features based on vibration signals; The data processing module is used to match the vibration characteristics with a preset motion library to obtain motion matching degree and motion matching timing information; The data processing module is further configured to determine the face matching score and the interactive environment score based on the cockpit image data when the action matching degree is greater than or equal to the preset action matching threshold and the action matching timing information meets the corresponding preset action timing conditions. The wake-up scoring module is used to determine the voice assistant wake-up score based on the action matching degree, the face matching score, and the interaction environment score when the face matching score is greater than or equal to a preset face matching threshold and the interaction environment score is greater than or equal to a preset environment score threshold. The voice wake-up module is used to generate a voice assistant wake-up control command to wake up the voice assistant when the voice assistant wake-up score is greater than a preset threshold.

9. A voice assistant wake-up device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voice assistant wake-up method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the voice assistant wake-up method as described in any one of claims 1 to 7.