Voice filtering method, device, terminal and medium independent of in-vehicle infotainment system functions
By monitoring the user's behavior status when speaking at the vehicle terminal, determining whether the voice is irrelevant to the vehicle terminal function, and stop uploading when it is judged to be irrelevant, the problem of low interaction accuracy of the vehicle voice function is solved, and more efficient voice filtering and recognition effects are achieved.
Patent Information
- Application Number
- CN202211053112.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The vehicle's voice function interaction accuracy is low, especially when a user makes a phone call or talks to a person in the car. The voice recognition system will mistakenly recognize non-functional voices, resulting in functionally unrelated texts on the screen.
During the process of uploading the user's voice to the cloud server in real time, by monitoring the user's behavior status when speaking, it is determined whether the user's voice has nothing to do with the function of the vehicle computer. If it is determined that it is irrelevant, stop uploading the user's voice. The specific implementation includes obtaining the image to be identified through the imaging device, inputting it into the behavior recognition model, extracting behavior characteristics to determine the behavior state, and judging the speech correlation based on the preset angle threshold.
Effectively filter out user voices that are not related to the vehicle-machine function, improve the accuracy of voice function interaction, and ensure that the cloud server only processes voices related to the vehicle-machine function.
Smart Images

Figure CN115527548B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent vehicles, and in particular, to a voice filtering method, device, terminal, and medium that are independent of the functions of the in-vehicle computer. Background Art
[0002] With the development of intelligent vehicles, the cloud server of the current vehicle networking system is configured with the ASR (Automatic Speech Recognition) function, which can perform online cloud speech recognition on the continuous speech of vehicle users, complete the interaction with users, and execute specific functions. However, when the user is making a call or talking to someone in the vehicle, etc., in many cases, these non-functional-purpose voices will be recognized, and there will be situations such as text unrelated to the function being displayed on the screen, and the accuracy of the voice function interaction of the vehicle is relatively low. Summary of the Invention
[0003] In view of this, embodiments of the present application provide a voice filtering method, device, terminal, and medium that are independent of the functions of the in-vehicle computer, so as to solve the problem of relatively low accuracy of the voice function interaction of the vehicle.
[0004] In a first aspect, embodiments of the present application provide a voice filtering method that is independent of the functions of the in-vehicle computer, including:
[0005] Connect to the cloud server according to the wake-up instruction input by the user;
[0006] During the process of uploading the collected user voice to the cloud server in real time, monitor the behavior state of the user when speaking, and determine whether the user voice is unrelated to the functions of the in-vehicle computer according to the behavior state;
[0007] When it is determined that the user voice is unrelated to the functions of the in-vehicle computer, stop uploading the user voice to the cloud server.
[0008] As described above in the aspect and any possible implementation manner, a further implementation manner is provided, where the monitoring of the behavior state of the user when speaking includes:
[0009] When the user is speaking, obtain an image to be recognized through a camera device;
[0010] Input the image to be recognized into a behavior recognition model, where the behavior recognition model is trained using a training set including user behavior actions;
[0011] Extract behavior features from the image to be recognized through the behavior recognition model, and output and determine the behavior state of the user when speaking according to the extracted behavior features.
[0012] For the aspects and any possible implementation manners described above, a further implementation manner is provided. Monitoring the behavioral state of the user when speaking includes:
[0013] When the user is speaking, obtaining an image of the driver's head position through a camera device;
[0014] Judging whether the user's voice is irrelevant to the in-vehicle device function according to the behavioral state includes:
[0015] According to the image of the driver's head position, determining the deviation angle of the driver's head relative to looking straight ahead;
[0016] Judging whether the user's voice is irrelevant to the in-vehicle device function according to the deviation angle and a preset first angle threshold. Wherein, if the deviation angle is greater than the first angle threshold, it is determined that the user's voice is irrelevant to the in-vehicle device function.
[0017] For the aspects and any possible implementation manners described above, a further implementation manner is provided. Before monitoring the behavioral state of the user when speaking, the method further includes:
[0018] Obtaining the identification information of the driver, where the identification information is used to identify the driver's identity;
[0019] Obtaining the individual difference adjustment information of the driver according to the identification information. The individual difference adjustment information includes individual difference characteristic information associated with determining the deviation angle and adjustment parameters. The individual difference characteristic information records different values of the same individual characteristics of drivers with different identities. The adjustment parameters are used to adjust the first angle threshold. The individual difference characteristic information and the adjustment parameters have a mapping relationship to set the first angle threshold for drivers with different identities according to the mapping relationship;
[0020] Setting the first angle threshold according to the individual difference adjustment information of the driver.
[0021] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The method further includes:
[0022] Obtaining the individual difference characteristic information of the driver, where the individual difference characteristic information records different values of the same individual characteristics of drivers with different identities;
[0023] Uploading the individual difference characteristic information of the driver and the behavioral state to the cloud server to obtain determination adjustment information, where the determination adjustment information includes an adjustment value for adjusting the determination threshold corresponding to the behavioral state;
[0024] Adjust the determination of the behavior state according to the determined adjustment information.
[0025] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. When it is determined that the user voice is irrelevant to the in-vehicle device function, stopping uploading the user voice to the cloud server includes:
[0026] When the user voice is irrelevant to the in-vehicle device function, send a truncation instruction to the voice module;
[0027] According to the truncation feedback signal returned by the voice module, determine that the user voice is truncated in the voice module, and stop sending the user voice to the cloud.
[0028] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. After stopping uploading the user voice to the cloud server, the method further includes:
[0029] According to the wake-up instruction input again by the user, or after an interval of a preset time period, send a voice upload resumption instruction to the voice module;
[0030] According to the resumption feedback signal returned by the voice module, send the current user voice to the cloud.
[0031] In a second aspect, an embodiment of the present application provides a voice filtering device for an in-vehicle device function irrelevance, including:
[0032] A cloud connection module, configured to connect to a cloud server according to a wake-up instruction input by a user;
[0033] A voice relevance determination module, configured to monitor the behavior state when the user is speaking during the process of real-time uploading the collected user voice to the cloud server, and determine whether the user voice is irrelevant to the in-vehicle device function according to the behavior state;
[0034] A voice truncation module, configured to stop uploading the user voice to the cloud server when it is determined that the user voice is irrelevant to the in-vehicle device function.
[0035] Further, the voice relevance determination module includes:
[0036] An image acquisition unit, configured to obtain an image to be recognized through a camera device when the user is speaking;
[0037] An image input unit, configured to input the image to be recognized into a behavior recognition model, where the behavior recognition model is trained using a training set including user behavior actions;
[0038] A feature extraction unit, configured to extract behavioral features of the image to be recognized through the behavior recognition model, and output and determine the behavior state of the user when speaking according to the extracted behavioral features.
[0039] Further, the speech relevance determination module further includes:
[0040] A head image acquisition unit, configured to obtain an image of the driver's head position through a camera device when the user is speaking;
[0041] A deviation angle determination unit, configured to determine the deviation angle of the driver's head relative to looking straight ahead according to the driver's head position image;
[0042] A speech relevance determination unit, configured to determine whether the user's speech is irrelevant to the in-vehicle device functions according to the deviation angle and a preset first angle threshold. Wherein, if the deviation angle is greater than the first angle threshold, it is determined that the user's speech is irrelevant to the in-vehicle device functions.
[0043] Further, the speech filtering device irrelevant to the in-vehicle device functions is further specifically configured to:
[0044] Obtain the identification information of the driver, where the identification information is used to identify the driver's identity;
[0045] Obtain the individual difference adjustment information of the driver according to the identification information. The individual difference adjustment information includes individual difference feature information associated with determining the deviation angle and adjustment parameters. The individual difference feature information records different values of different drivers with the same identity on the same individual characteristics. The adjustment parameters are used to adjust the first angle threshold, and the individual difference feature information and the adjustment parameters have a mapping relationship to set the first angle threshold for different drivers with different identities according to the mapping relationship;
[0046] Set the first angle threshold according to the individual difference adjustment information of the driver.
[0047] Further, the speech filtering device irrelevant to the in-vehicle device functions is further specifically configured to:
[0048] Obtain the individual difference feature information of the driver, where the individual difference feature information records different values of different drivers with the same identity on the same individual characteristics;
[0049] Upload the individual difference feature information of the driver and the behavior state to the cloud server, and obtain determination adjustment information. The determination adjustment information includes adjustment values for adjusting the determination threshold corresponding to the behavior state;
[0050] Adjust the determination of the behavior state according to the determined adjustment information.
[0051] Further, the voice truncation module includes:
[0052] A truncation instruction sending unit, configured to send a truncation instruction to the voice module when the user voice is irrelevant to the in-vehicle device function;
[0053] A voice truncation unit, which determines that the user voice is truncated in the voice module according to the truncation feedback signal returned by the voice module, and stops sending the user voice to the cloud.
[0054] Further, the voice filtering device irrelevant to the in-vehicle device function is further specifically configured to:
[0055] According to the wake-up instruction re-entered by the user, or after a preset time interval, send a resume voice upload instruction to the voice module;
[0056] According to the resume feedback signal returned by the voice module, send the current user voice to the cloud.
[0057] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it performs the steps of the voice filtering method for functions irrelevant to the in-vehicle device described in the first aspect.
[0058] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of the voice filtering method for functions irrelevant to the in-vehicle device described in the first aspect are implemented.
[0059] In the embodiment of the present application, first, connect to the cloud server according to the wake-up instruction input by the user, so that the cloud server can enter the state of user voice recognition, and timely recognize the user voice and perform text conversion; then, during the process of real-time uploading the collected user voice to the cloud server, monitor the behavior state of the user when speaking, and determine whether the user voice is irrelevant to the in-vehicle device function according to the behavior state, and can filter out the user voice irrelevant to the in-vehicle device function, and judge the validity of the user voice through the behavior state of the user; finally, when it is determined that the user voice is irrelevant to the in-vehicle device function, stop uploading the user voice to the cloud server in time, so that these meaningless voices will not be sent to the cloud server for voice analysis, and the user voices recognized by the cloud server are all voices related to the in-vehicle device function. The present application can effectively improve the accuracy of the voice function interaction of the vehicle. Description of the Drawings
[0060] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0061] Figure 1 is a flowchart of a voice filtering method independent of the functions of the in-vehicle device terminal in the embodiments of the present application;
[0062] Figure 2 is a schematic diagram of a principle block of a device corresponding one-to-one to the voice filtering method independent of the functions of the in-vehicle device terminal in the embodiments of the present application;
[0063] Figure 3 is a schematic diagram of a computer device in the embodiments of the present application. Detailed implementation manners
[0064] To better understand the technical solutions of the present application, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0065] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0066] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "an", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0067] It should be understood that the term " / and / " used herein is only a description of the same field of related objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the related objects before and after.
[0068] It should be understood that although terms such as first, second, and third may be used in the embodiments of the present application to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present application, the first preset range can also be referred to as the second preset range, and similarly, the second preset range can also be referred to as the first preset range.
[0069] Depending on the context, as used herein, the term "if" may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0070] This application provides a voice filtering method independent of in-vehicle unit functions. Figure 1 It is a flowchart of a voice filtering method independent of in-vehicle unit functions in an embodiment of this application. This method can be applied to an in-vehicle terminal and a corresponding cloud server. The cloud server can have computing power services for speech recognition and speech-to-text conversion. When a user issues a voice, the cloud server can recognize and convert the voice issued by the user into text and send it to be displayed on the interface of the in-vehicle terminal. As Figure 1 shown, the voice filtering method independent of in-vehicle unit functions includes the following steps executed by the in-vehicle terminal:
[0071] S10: Connect to the cloud server according to the wake-up instruction input by the user.
[0072] Among them, the wake-up instruction refers to an instruction to start the voice recognition function of the vehicle, which can enable the in-vehicle terminal to connect to the cloud server to perform voice recognition on the user voice related to the in-vehicle terminal issued, and send the recognized text to be displayed on the screen of the in-vehicle terminal. It should be noted that in this application, the in-vehicle unit side and the in-vehicle terminal are two concepts. The in-vehicle unit side includes the in-vehicle terminal. The in-vehicle unit side is a more general concept and can be understood as the vehicle side. The user voice related to the in-vehicle unit side refers to the user voice related to realizing vehicle functions. Through the user voice related to the in-vehicle unit side, the vehicle can be controlled to realize specific functions, such as navigation, controlling the window switch, controlling the vehicle air conditioner temperature, etc. The in-vehicle terminal is a terminal device carried on the vehicle, which can perform information interaction with Internet devices such as the user's handheld terminal device and the server. The in-vehicle terminal may further include a display screen, on which objects, controls, etc. for interacting with the user are displayed. The user can directly input function instructions related to the in-vehicle unit side through the in-vehicle terminal and execute corresponding vehicle functions.
[0073] S20: During the process of real-time uploading the collected user voice to the cloud server, monitor the behavior state of the user when speaking, and determine whether the user voice is irrelevant to the in-vehicle unit functions according to the behavior state.
[0074] Among them, the cloud server can allocate speech recognition computing power, and itself or the computing power server communicating with it has the ability of speech recognition. It should be noted that in this application, the terminal device is not used to implement speech recognition, but the cloud server is used for speech recognition. The reason is that the speech recognition accuracy of the terminal device is limited, and the speech recognition related to the in-vehicle terminal has relatively high requirements for recognition accuracy. If the terminal device is used to implement speech recognition, it may occur that the speech recognition is inaccurate, resulting in the inability to generate instructions related to vehicle functions. More seriously, incorrect instructions may be generated due to recognition errors, which may even bring safety problems to users. On the other hand, using the cloud server to implement speech recognition is more secure than the terminal device, and can effectively prevent other users from inputting some instructions that are not expected by the user through the terminal device. It should also be noted that the speech recognition of the cloud server involves computing power allocation. Note that the speech recognition service of the cloud server requires GPU (graphics processing unit) computing power support, and the cost is generally relatively high. Some ASR manufacturers charge according to the number of concurrencies or even the number of words for accessing the recognition engine. In view of this premise, this application emphasizes that recognizing user speech unrelated to the in-vehicle terminal can reduce the meaningless recognition of the cloud server's ASR recognition service, make the voice function interaction of the vehicle more accurate, and at the same time can save more costs.
[0075] In one embodiment, after the user speech is collected, it is not directly uploaded to the cloud server. Instead, when the user speech is collected, the behavior state of the user when speaking is synchronously detected. Then, according to the user's behavior state, such as when talking to passengers, singing, or on the phone with others, it is judged whether the user speech is unrelated to the in-vehicle terminal function through these behavior states. In the embodiment of this application, by judging whether the user speech is related to the in-vehicle terminal function, it helps to filter out the meaningless speech of the user that is unrelated to the in-vehicle terminal function.
[0076] S30: When it is determined that the user speech is unrelated to the in-vehicle terminal function, stop uploading the user speech to the cloud server.
[0077] In one embodiment, if it is determined that the user speech is unrelated to the in-vehicle terminal function, the user speech is immediately terminated. The cloud server will not receive these user speeches that are unrelated to the in-vehicle terminal function, and only performs speech recognition on the successfully uploaded user speeches. In this way, the accuracy of the voice function interaction of the vehicle can be effectively improved.
[0078] In the embodiment of the present application, first, connect to the cloud server according to the wake-up instruction input by the user, so that the cloud server can enter the state of user speech recognition, timely recognize the user speech and perform text conversion; then, during the process of real-time uploading the collected user speech to the cloud server, monitor the behavior state of the user when speaking, and judge whether the user speech is irrelevant to the functions of the in-vehicle device according to the behavior state, so as to filter out the user speech that is irrelevant to the functions of the in-vehicle device, and judge the validity of the user speech through the behavior state of the user; finally, when it is determined that the user speech is irrelevant to the functions of the in-vehicle device, stop uploading the user speech to the cloud server in time, so that these meaningless speeches will not be sent to the cloud server for speech analysis, and the user speech recognized by the cloud server is all speech related to the functions of the in-vehicle device. The present application can effectively improve the accuracy of the voice function interaction of the vehicle.
[0079] Further, in step S20, that is, in the step of monitoring the behavior state of the user when speaking, it specifically includes the following steps:
[0080] S211: When the user is speaking, obtain the image to be recognized through the imaging device.
[0081] Wherein, the image to be recognized is the user image collected by the imaging device to be used for behavior state recognition.
[0082] In one embodiment, at least one imaging device is provided above the vehicle steering wheel, facing the user direction. After the user inputs the wake-up instruction, the imaging device can be started synchronously and photograph the user to collect the image to be recognized related to the user when the user is speaking.
[0083] S212: Input the image to be recognized into the behavior recognition model, wherein the behavior recognition model is trained by using a training set including user behavior actions.
[0084] In one embodiment, the behavior recognition model is used to recognize the user's behavior actions to determine the user's behavior state. It should be noted that the behavior recognition model is trained using a training set of the user's behavior actions. The training images in this training set may specifically include various user behavior action images that have nothing to do with the functions of the in-vehicle device. These various user behavior images that have nothing to do with the functions of the in-vehicle device may include pre-shot user behavior action images, or some existing obtainable user behavior action images. In addition, control user behavior action images should also be set. These user behavior action images are input into a convolutional neural network, etc. for training according to a preset ratio. Through multiple iterations, after the network parameters are updated backward according to the loss function, when the training times are reached or the update is within the preset change threshold range, the behavior recognition model is trained. In the embodiment of the present application, the image to be recognized is input into the behavior recognition model, and it can be determined by the behavior recognition model whether the user's current behavior action is related to the in-vehicle device.
[0085] S213: Extract the behavior features of the image to be recognized through the behavior recognition model, and output and determine the behavior state when the user is speaking according to the extracted behavior features.
[0086] In one embodiment, the recognition model is generally trained using a convolutional neural network or a neural network improved based on convolution. During the training process, image feature extraction is performed through convolution to extract the deep features in the training images and associate them with the results output by the behavior recognition model, so as to achieve an accurate recognition effect. It can be understood that the image to be recognized is an image obtained by photographing the user in real time. This image to be recognized is unique, but the deep features contained in this image to be recognized can reflect the behavior state of the image. Specifically, the trained behavior recognition model is used to extract the behavior features of the image to be recognized in terms of the behavior state and output the model results. According to the results output by this model, the behavior state when the user is speaking can be determined.
[0087] In steps S211 - S213, when the user is speaking, the user is synchronously imaged, and the captured image to be recognized is input into the behavior recognition model, so as to determine the behavior state when the user is speaking according to the pre-trained behavior recognition model. In this embodiment, by detecting the state of the user's accompanying behavior when the user issues a voice, it is possible to well distinguish whether the user's voice is related to the in-vehicle device.
[0088] Further, in step S20, that is, the step of monitoring the behavior state when the user is speaking, specifically, it further includes: when the user is speaking, obtaining an image of the driver's head position through a camera device.
[0089] Specifically, the user can be the driver of a vehicle. In one embodiment, the behavior state of the user when speaking is determined mainly based on the position of the driver's head. If there is a large deviation in the position of the user's head, it is considered that the user is doing other things related to the functions of the in-vehicle device.
[0090] Further, it is determined whether the user's speech is unrelated to the functions of the in-vehicle device according to the behavior state, including:
[0091] S221: Determine the deviation angle of the driver's head relative to looking straight ahead according to the image of the driver's head position.
[0092] In one embodiment, the imaging device can be set directly in front of the driver. When the driver is driving, the driver's eyes are looking straight ahead. When the driver issues the user speech, the imaging device synchronously acquires the image of the driver's head position and compares it with the preset angle of the driver's eyes looking straight ahead, so as to determine the deviation angle of the driver's head relative to looking straight ahead through angle comparison.
[0093] S222: Determine whether the user's speech is unrelated to the functions of the in-vehicle device according to the deviation angle and the preset first angle threshold. Among them, if the deviation angle is greater than the first angle threshold, it is determined that the user's speech is unrelated to the functions of the in-vehicle device.
[0094] Among them, the first angle threshold can be specifically set to 60°. When the deviation angle of the driver's head is within 60°, it can be considered that the driver is looking straight ahead or at the display interface of the in-vehicle terminal. When the deviation angle of the driver's head is greater than 60°, the driver may be talking to other people in the vehicle or looking for something, etc., which are some things unrelated to the functions of the in-vehicle device. In one embodiment, when the deviation angle of the driver's head is greater than the first angle threshold, it can be determined that the user's speech is unrelated to the functions of the in-vehicle device.
[0095] Further, if it is determined that the user's speech is unrelated to the functions of the in-vehicle device, the in-vehicle terminal can let the user adjust the head position and then issue speech related to the functions of the in-vehicle device by means of voice playback reminder or screen text reminder. Further, the reminder method of the in-vehicle terminal can be given according to the deviation angle of the driver's head. Specifically, if the deviation angle of the driver's head is not less than 60° and not greater than 90°, the in-vehicle terminal can use the method of displaying text on the screen to remind the user. If the deviation angle of the driver's head is not less than 90°, the in-vehicle terminal can use the method of voice playback to remind the user.
[0096] In steps S221 - S222, the behavior state of the user is determined based on the head position of the driver. Compared with the determination of the overall image of the user, this determination is more targeted and also conforms to most determinations that have nothing to do with the functions of the in - vehicle device terminal. Compared with the need for strong computing power for the determination of the overall user image, using the head position of the driver to judge the behavior state of the user will have lower costs and is also more suitable for promotion.
[0097] Furthermore, before step S20, that is, before the step of monitoring the behavior state of the user when speaking, the following steps are specifically included:
[0098] S231: Obtain the identification information of the driver, where the identification information is used to identify the driver's identity.
[0099] Among them, the identification information of the driver refers to the information that can uniquely identify the driver's identity. For example, biometric features such as the driver's voiceprint, fingerprint, face, etc. can all be used as the identification information of the driver. In one embodiment, face recognition or voiceprint recognition can be used to obtain the identification information of the driver. In this way, after uniquely determining the identification information of the driver, the relevance determination of the user's voice and the functions of the in - vehicle device terminal can be better made according to the physical differences between the driver and others.
[0100] S232: Obtain the individual - difference adjustment information of the driver according to the identification information, where the individual - difference adjustment information includes individual - difference characteristic information associated with determining the deviation angle and adjustment parameters. The individual - difference characteristic information records different values of different - identity drivers on the same individual characteristics, and the adjustment parameters are used to adjust the first - angle threshold. The individual - difference characteristic information and the adjustment parameters have a mapping relationship to set the first - angle threshold for different - identity drivers according to the mapping relationship.
[0101] S233: Set the first - angle threshold according to the individual - difference adjustment information of the driver.
[0102] In one embodiment, the head shapes and neck thicknesses of different drivers will cause certain errors in the recognition of the deviation angle. It can be understood that theoretically, the deviation angle will not be affected by the head shape and neck thickness of the driver in terms of accuracy, but this requires the use of very high - precision equipment to achieve. While using a general image - recognition model, it will be affected to a certain extent by the individual - difference adjustment information of these drivers. Therefore, in view of the feasibility and popularization of practical applications, the present application introduces the individual differences between drivers to make the relevance determination of the user's voice and the functions of the in - vehicle device terminal more accurate.
[0103] Among them, the individual difference characteristic information can be the width and height of the user's head, the width and length of the neck, and even the shoulder width can also be used as individual difference characteristic information. The adjustment parameter is used to adjust the first angle threshold. For example, if the adjustment parameter corresponding to the individual difference characteristic information of user A is +10°, then when obtaining the angle threshold, 10° should be added to the standard angle threshold. For example, it is set from 60° to 70°. In this way, starting from the individual differences of the driver, the user's behavior state can be judged more accurately.
[0104] In steps S231 - S233, the angle offset is precisely adjusted starting from the individual differences among drivers, making it possible to judge the user's behavior state more accurately.
[0105] Furthermore, the voice filtering method unrelated to the vehicle machine terminal function further includes the following steps:
[0106] S41: Obtain the individual difference characteristic information of the driver. Among them, the individual difference characteristic information records different values of different drivers with the same individual characteristics.
[0107] In one embodiment, the same vehicle may correspond to different drivers, and there are individual difference characteristic information among different drivers, and these individual difference characteristic information can be collected in advance. For example, the driver can manually enter or automatically enter different values of different drivers with the same individual characteristics through image recognition.
[0108] S42: Upload the individual difference characteristic information and behavior state of the driver to the cloud server, and obtain the determination adjustment information. Among them, the determination adjustment information includes an adjustment value, which is used to adjust the determination threshold corresponding to the behavior state.
[0109] In one embodiment, after different drivers enter the individual difference characteristic information, it will be uploaded to the cloud, and the determination adjustment information corresponding to different individual difference characteristic information will be calculated through a preset algorithm, that is, the adjustment value specifically required to adjust the determination threshold corresponding to the behavior state, such as the angle threshold of the driver's head.
[0110] S43: Adjust the determination of the behavior state according to the determination adjustment information.
[0111] In one embodiment, if the deviation angle of the driver's head is judged, the angle threshold is adjusted according to the judgment adjustment information; in addition to judging the deviation angle of the driver's head, other determination thresholds related to individual differences can also be judged, and these judgments can be used to determine the driver's behavior state.
[0112] In steps S41 - S43, the individual difference characteristic information of different drivers can be collected in advance, and the relevant determination thresholds can be adjusted through the calculated determination adjustment information, so as to more accurately determine the driver's behavior state from the individual differences.
[0113] Further, in step S30, that is, when it is determined that the user voice is irrelevant to the in - vehicle device function, in the step of stopping uploading the user voice to the cloud server, it specifically includes the following steps:
[0114] S311: When the user voice is irrelevant to the in - vehicle device function, send a truncation instruction to the voice module.
[0115] Among them, the voice module can be used for the collection and upload of user voices. When the user voice is meaningless for the implementation of the in - vehicle device function, the in - vehicle terminal sends a truncation instruction to the built - in or external voice module to stop uploading the user voice irrelevant to the in - vehicle device function according to this truncation instruction.
[0116] S312: According to the truncation feedback signal returned by the voice module, determine that the user voice is truncated in the voice module and stop sending the user voice to the cloud.
[0117] In an embodiment, the in - vehicle terminal will receive the feedback signal of the user voice truncation. By receiving this feedback signal, a reminder message for reminding the user to re - enter the user voice can be triggered.
[0118] In steps S311 - S312, the truncation of the user voice irrelevant to the in - vehicle device function is realized through the voice module, so that the user voice that really needs to be speech - recognized can be uploaded to the cloud, which can effectively filter out the meaningless user voices and effectively improve the accuracy of the voice function interaction of the vehicle.
[0119] Further, after step S30, that is, after the step of stopping uploading the user voice to the cloud server, the method further includes:
[0120] S321: According to the wake - up instruction input again by the user, or, after an interval of a preset time period, send a voice upload resume instruction to the voice module.
[0121] Among them, the wake - up instruction is an instruction to start the voice recognition function of the vehicle. The voice upload resume instruction is an instruction to let the voice module resume uploading the user voice to the cloud server.
[0122] In an embodiment, the user can trigger the wake - up instruction again by uttering a preset voice, so that the in - vehicle terminal can stop the voice truncation in time. Or, if the user does not wake up the in - vehicle terminal again to implement voice recognition, after an interval of a preset time period such as 5s, 10s, a voice upload resume instruction can be sent to the voice module.
[0123] S322: Send the current user voice to the cloud according to the recovery feedback signal returned by the voice module.
[0124] In one embodiment, after the recovery voice is uploaded, the voice module collects the current user voice and sends the user voice to the cloud.
[0125] In steps S321 - S322, a mechanism for re - uploading user voice is provided. After the user voice upload is truncated, the user can still quickly resume the voice upload state, enabling the vehicle to promptly return to the normal voice recognition state and effectively improving the efficiency of user voice recognition.
[0126] In the embodiment of the present application, first, connect to the cloud server according to the wake - up instruction input by the user, so that the cloud server can enter the user voice recognition state, promptly recognize the user voice and perform text conversion; then, during the process of real - time uploading the collected user voice to the cloud server, monitor the behavior state of the user when speaking, and judge whether the user voice is irrelevant to the in - vehicle device function according to the behavior state, and can filter out the user voice that is irrelevant to the in - vehicle device function, and judge the validity of the user voice through the behavior state of the user; finally, when it is determined that the user voice is irrelevant to the in - vehicle device function, promptly stop uploading the user voice to the cloud server, so that these meaningless voices will not be sent to the cloud server for voice analysis, and the user voices recognized by the cloud server are all voices related to the in - vehicle device function. The present application can effectively improve the accuracy of the voice function interaction of the vehicle.
[0127] Furthermore, the present application also synchronously captures an image of the user when the user is speaking, and inputs the captured image to be recognized into the behavior recognition model, so as to determine the behavior state of the user when speaking according to the pre - trained behavior recognition model. Through the detection of the state of the accompanying behavior when the user emits voice, it is possible to well distinguish whether the user voice is relevant to the in - vehicle device.
[0128] Furthermore, the present application also determines the behavior state of the user based on the head position of the driver. Compared with the determination of the overall user image, this determination is more targeted and also conforms to most determinations of irrelevance to the in - vehicle device function. Compared with the premise that a relatively strong computing power is required for the determination of the overall user image, using this method of determining the behavior state based on the head position of the driver has a lower cost and is also more suitable for promotion.
[0129] Furthermore, the present application also makes a precise adjustment to the angle offset starting from the individual differences between drivers, so that the judgment of the user behavior state can be more accurate.
[0130] Further, the present application can also pre-collect the individual difference characteristic information of different drivers, and adjust the relevant determination thresholds through the calculated determination adjustment information, so as to more accurately determine the driver's behavior state from the individual differences.
[0131] Further, the present application also truncates the user voice unrelated to the functions of the in-vehicle device through the voice module, so that the user voice that really needs to be recognized can be uploaded to the cloud. In this way, meaningless user voices can be effectively filtered out, and the accuracy of the voice function interaction of the vehicle can be effectively improved.
[0132] Further, the present application also provides a mechanism for re-uploading user voices. After the upload of the user voice is truncated, the user can still quickly resume the voice upload state, so that the vehicle can return to the normal voice recognition state in time, and the efficiency of user voice recognition can be effectively improved.
[0133] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0134] Figure 2 It is the principle block diagram of a device corresponding to a voice filtering method unrelated to the functions of the in-vehicle device in the embodiments of the present application. As Figure 2 shown, the voice filtering device unrelated to the functions of the in-vehicle device includes a cloud connection module 10, a voice relevance determination module 20, and a voice truncation module 30.
[0135] The cloud connection module 10 is used to connect to the cloud server according to the wake-up instruction input by the user.
[0136] The voice relevance determination module 20 is used to monitor the behavior state of the user when speaking during the process of uploading the collected user voice to the cloud server in real time, and determine whether the user voice is unrelated to the functions of the in-vehicle device according to the behavior state.
[0137] The voice truncation module 30 is used to stop uploading the user voice to the cloud server when it is determined that the user voice is unrelated to the functions of the in-vehicle device.
[0138] Further, the voice relevance determination module 20 includes:
[0139] The image acquisition unit is used to obtain the image to be recognized through the imaging device when the user is speaking
[0140] The image input unit is used to input the image to be recognized into the behavior recognition model, where the behavior recognition model is trained by using a training set including user behavior actions.
[0141] A feature extraction unit, configured to extract behavioral features from an image to be recognized through a behavior recognition model, and output and determine the behavioral state of the user when speaking according to the extracted behavioral features.
[0142] Further, the voice relevance determination module 20 further includes:
[0143] A head image acquisition unit, configured to obtain an image of the driver's head position through a camera device when the user is speaking.
[0144] An offset angle determination unit, configured to determine the offset angle of the driver's head relative to looking straight ahead according to the image of the driver's head position.
[0145] A voice relevance determination unit, configured to determine whether the user's voice is irrelevant to the in-vehicle device function according to the offset angle and a preset first angle threshold, wherein if the offset angle is greater than the first angle threshold, it is determined that the user's voice is irrelevant to the in-vehicle device function.
[0146] Further, the voice filtering device irrelevant to the in-vehicle device function is further specifically configured to:
[0147] Obtain the identification information of the driver, where the identification information is used to identify the driver's identity.
[0148] Obtain the individual difference adjustment information of the driver according to the identification information, where the individual difference adjustment information includes individual difference feature information associated with determining the offset angle and adjustment parameters. The individual difference feature information records different values of different drivers with the same identity on the same individual feature, and the adjustment parameters are used to adjust the first angle threshold. The individual difference feature information and the adjustment parameters have a mapping relationship to set the first angle threshold for different drivers according to the mapping relationship.
[0149] Set the first angle threshold according to the individual difference adjustment information of the driver.
[0150] Further, the voice filtering device irrelevant to the in-vehicle device function is further specifically configured to:
[0151] Obtain the individual difference feature information of the driver, where the individual difference feature information records different values of different drivers with the same identity on the same individual feature.
[0152] Upload the individual difference feature information and the behavioral state of the driver to the cloud server to obtain determination adjustment information, where the determination adjustment information includes adjustment values for adjusting the determination threshold corresponding to the behavioral state.
[0153] Adjust the determination of the behavioral state according to the determination adjustment information.
[0154] Further, the voice truncation module 30 includes:
[0155] A truncation instruction sending unit, configured to send a truncation instruction to the voice module when the user voice is not related to the in-vehicle device function.
[0156] A voice truncation unit, determines that the user voice is truncated in the voice module according to the truncation feedback signal returned by the voice module, and stops sending the user voice to the cloud.
[0157] Furthermore, the voice filtering device unrelated to the in-vehicle device function is further specifically configured to:
[0158] Send a voice upload resumption instruction to the voice module according to the wake-up instruction re-entered by the user, or after a preset time interval.
[0159] Send the current user voice to the cloud according to the resumption feedback signal returned by the voice module.
[0160] In the embodiment of the present application, first connect to the cloud server according to the wake-up instruction input by the user, so that the cloud server can enter the state of user voice recognition, and timely recognize the user voice and perform text conversion; then during the process of real-time uploading the collected user voice to the cloud server, monitor the behavior state of the user when speaking, and judge whether the user voice is related to the in-vehicle device function according to the behavior state, and can filter out the user voice unrelated to the in-vehicle device function, and judge the validity of the user voice through the behavior state of the user; finally, when it is determined that the user voice is unrelated to the in-vehicle device function, stop uploading the user voice to the cloud server in time, so that these meaningless voices will not be sent to the cloud server for voice analysis, and the user voices recognized by the cloud server are all voices related to the in-vehicle device function. The present application can effectively improve the accuracy of the voice function interaction of the vehicle.
[0161] Further, the present application also synchronously captures an image of the user while the user is speaking, and inputs the captured image to be recognized into the behavior recognition model, so as to determine the behavior state of the user when speaking according to the behavior recognition model trained in advance. Through the detection of the state of the accompanying behavior of the user when issuing the voice, the user voice can be well distinguished from whether it is related to the in-vehicle device. Further, the present application also determines the behavior state of the user based on the head position of the driver. Compared with the determination of the overall image of the user, this determination is more targeted and also conforms to most determinations that have nothing to do with the functions of the in-vehicle device. Compared with the premise that a relatively strong computing power is required for the determination of the overall image of the user, using the head position of the driver to judge the behavior state of the user will have a lower cost and is also more suitable for popularization. Further, the present application also makes a precise adjustment to the angle offset starting from the individual differences between drivers, so that the behavior state of the user can be judged more accurately when judging. Further, the present application can also pre-collect the individual difference characteristic information of different drivers, and adjust the relevant determination threshold through the calculated determination adjustment information, so as to more accurately judge the behavior state of the driver starting from the individual differences. Further, the present application also truncates the user voice that has nothing to do with the functions of the in-vehicle device through the voice module, so that the user voice that really needs to be recognized can be uploaded to the cloud. In this way, meaningless user voices can be effectively filtered out, and the accuracy of the voice function interaction of the vehicle can be effectively improved. Further, the present application also provides a mechanism for re-uploading the user voice. After the upload of the user voice is truncated, the user can still quickly return to the state of voice upload, so that the vehicle can return to the normal voice recognition state in time, and the efficiency of user voice recognition can be effectively improved.
[0162] The present application also provides a computer-readable storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the voice filtering method unrelated to the functions of the in-vehicle device as described in the embodiment is implemented.
[0163] Figure 3 It is a schematic diagram of a computer device in an embodiment of the present application.
[0164] As Figure 3 shown, the computer device 110 includes a processor 111, a memory 112, and computer-readable instructions 113 stored in the memory 112 and executable on the processor 111. When the processor 111 executes the computer-readable instructions 113, each step of the voice filtering method unrelated to the functions of the in-vehicle device is implemented.
[0165] Exemplarily, the computer-readable instructions 113 can be divided into one or more modules / units, which are stored in the memory 112 and executed by the processor 111 to complete this application. One or more modules / units can be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer-readable instructions 113 in the computer device 110.
[0166] The computer device 110 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor 111 and a memory 112. Those skilled in the art can understand that Figure 3 merely examples of the computer device 110, which do not constitute a limitation on the computer device 110, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, a bus, etc.
[0167] The so-called processor 111 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0168] The memory 112 may be an internal storage unit of the computer device 110, such as the hard disk or memory of the computer device 110. The memory 112 may also be an external storage device of the computer device 110, such as a plug-in hard disk equipped on the computer device 110, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 112 may also include both the internal storage unit and the external storage device of the computer device 110. The memory 112 is used to store computer-readable instructions and other programs and data required by the computer device. The memory 112 may also be used to temporarily store data that has been output or will be output.
[0169] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0170] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0171] In the embodiments of the present application, the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0172] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0173] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer-readable instructions include computer-readable instruction codes, and the computer-readable instruction codes can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device that can carry the computer-readable instruction code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0174] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0175] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A voice filtering method independent of in-vehicle device functions, characterized in that it includes connecting to a cloud server according to a wake-up instruction input by a user; during the process of real-time uploading of the collected user voice to the cloud server, monitoring the behavior state of the user when speaking, and judging whether the user voice is independent of in-vehicle device functions according to the behavior state; when it is determined that the user voice is independent of in-vehicle device functions, stopping uploading the user voice to the cloud server; wherein, the monitoring of the behavior state of the user when speaking includes when the user is speaking, obtaining an image of the driver's head position through a camera device; the judging whether the user voice is independent of in-vehicle device functions according to the behavior state includes determining the deviation angle of the driver's head relative to looking straight ahead according to the driver's head position image; judging whether the user voice is independent of in-vehicle device functions according to the deviation angle and a preset first angle threshold, wherein if the deviation angle is greater than the first angle threshold, it is determined that the user voice is independent of in-vehicle device functions; before the monitoring of the behavior state of the user when speaking, the method further includes obtaining the identification information of the driver, wherein the identification information is used to identify the driver's identity; obtaining the individual difference adjustment information of the driver according to the identification information, wherein the individual difference adjustment information includes individual difference characteristic information associated with determining the deviation angle and adjustment parameters, the individual difference characteristic information records different values of different-identity drivers on the same individual characteristics, the adjustment parameters are used to adjust the first angle threshold, and the individual difference characteristic information and the adjustment parameters have a mapping relationship to set the first angle threshold for different-identity drivers according to the mapping relationship; setting the first angle threshold according to the individual difference adjustment information of the driver.
2. The method according to claim 1, characterized in that the monitoring of the behavior state of the user when speaking includes when the user is speaking, obtaining an image to be recognized through a camera device; inputting the image to be recognized into a behavior recognition model, wherein the behavior recognition model is trained using a training set including user behavior actions; performing behavior feature extraction on the image to be recognized through the behavior recognition model, and outputting and determining the behavior state of the user when speaking according to the extracted behavior features.
3. The method according to claim 1, characterized in that the method further includes obtaining the individual difference characteristic information of the driver, wherein the individual difference characteristic information records different values of different-identity drivers on the same individual characteristics; uploading the individual difference characteristic information of the driver and the behavior state to the cloud server, and obtaining determination adjustment information, wherein the determination adjustment information includes an adjustment value for adjusting the determination threshold corresponding to the behavior state; adjusting the determination of the behavior state according to the determination adjustment information.
4. A voice filtering device independent of the functions of the in-vehicle computer, characterized in that, It includes a cloud connection module for connecting to a cloud server according to a wake-up instruction input by a user; A voice relevance determination module, which is used to monitor the behavior state of the user when speaking during the process of uploading the collected user voice to the cloud server in real time, and determine whether the user voice is irrelevant to the in-vehicle device functions according to the behavior state; A voice truncation module, which is used to stop uploading the user voice to the cloud server when it is determined that the user voice is irrelevant to the in-vehicle device functions; Among them, the voice relevance determination module further includes A head image acquisition unit, which is used to obtain an image of the driver's head position through a camera device when the user is speaking; A deviation angle determination unit, which is used to determine the deviation angle of the driver's head relative to looking straight ahead according to the driver's head position image; A voice relevance determination unit, which is used to determine whether the user voice is irrelevant to the in-vehicle device functions according to the deviation angle and a preset first angle threshold. Among them, if the deviation angle is greater than the first angle threshold, it is determined that the user voice is irrelevant to the in-vehicle device functions; The voice filtering device irrelevant to the in-vehicle device functions is further used for: Obtaining the identification information of the driver, where the identification information is used to identify the driver's identity; obtaining the individual difference adjustment information of the driver according to the identification information, where the individual difference adjustment information includes individual difference characteristic information associated with determining the deviation angle and adjustment parameters. The individual difference characteristic information records different values of different identities of the driver on the same individual characteristics, and the adjustment parameters are used to adjust the first angle threshold. The individual difference characteristic information and the adjustment parameters have a mapping relationship to set the first angle threshold for different identities of the driver according to the mapping relationship; Setting the first angle threshold according to the individual difference adjustment information of the driver.
5. The device according to claim 4, wherein The voice relevance determination module includes An image acquisition unit, which is used to obtain an image to be recognized through a camera device when the user is speaking; an image input unit, which is used to input the image to be recognized into a behavior recognition model, where the behavior recognition model is trained by using a training set including user behavior actions; A feature extraction unit, which is used to extract behavior features from the image to be recognized through the behavior recognition model, and output and determine the behavior state of the user when speaking according to the extracted behavior features.
6. An in-vehicle device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein When the processor executes the computer-readable instructions, it executes the steps of the voice filtering method irrelevant to the in-vehicle device functions according to any one of claims 1-3.
7. A computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and wherein When the computer-readable instructions are executed by a processor, the steps of the voice filtering method irrelevant to the in-vehicle device functions according to any one of claims 1-3 are implemented.
Citation Information
Patent Citations
Continuous finite impulse response filter coefficient vector update method and device
CN107749304A
Microphone neck ring earphone
CN108235164A