Voice control method and related equipment

By combining the user's voice signal and gaze content for dual recognition, the problem that the on-board voice control system is difficult to accurately judge user intentions in complex scenarios, achieving higher recognition accuracy and operation efficiency, and improving user experience and driving safety.

CN120048255APending Publication Date: 2025-05-27VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510001039.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing vehicle voice control system is difficult to accurately judge the user's true intentions in complex scenarios, and it is prone to problems of error touch or inaccurate response, which affects driving safety and user experience.

Method used

By combining the user's voice signal and gaze content for double recognition, the semantic information of the voice signal is parsed and the intent correlation with the gaze content is determined, and the corresponding vehicle control operation is only responded to when the correlation degree reaches a preset threshold.

Benefits of technology

It improves the accuracy of identifying users' real needs, reduces false touch and conflicts, simplifies interaction steps, optimizes operation efficiency, and makes voice control more natural, smooth, safe and convenient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048255A_ABST
    Figure CN120048255A_ABST
Patent Text Reader

Abstract

The invention discloses a voice control method and related equipment, and relates to the technical field of intelligent driving, and the method comprises the steps: obtaining the watching content of a target user for a visible interface in a vehicle in response to a voice signal inputted by the target user; semantic information of the voice signal is analyzed, and the intention association degree of the semantic information and the gazing content is determined; and responding to the vehicle control operation corresponding to the semantic information according to the intention association degree. According to the method, double recognition is carried out in combination with the user voice signal and the watching content, the real requirements of the user can be accurately and efficiently judged, mistaken touch and conflicts are reduced, the interaction steps are simplified, the operation efficiency is optimized, and voice control is more natural, smoother, safer and more convenient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and more specifically, to a voice control method and related equipment. Background Art

[0002] With the rapid development of intelligent driving technology, voice control has been widely used in the automotive industry as a key interactive method to improve driving experience and safety. By interacting with the vehicle system through voice, the driver can perform a variety of operations, reduce manual intervention, and improve driving safety and ease of operation. However, the complex interactive scenarios and diverse user needs in the vehicle environment make it a technical challenge for the voice control system to accurately understand the user's intentions and avoid false triggering. Therefore, it is particularly important to improve the recognition accuracy of the in-vehicle voice control system for user intentions and optimize the command triggering mechanism.

[0003] In the related art, the “see and say” function of traditional in-vehicle voice control systems mostly relies on a single voice input signal to trigger commands. This method is difficult to accurately judge the user’s true intentions, especially when the screen display content is complex or the user’s focus is unclear, which can easily lead to problems such as false touches or inaccurate responses. Due to the lack of understanding of the user’s gaze content and analysis of the relevance of intentions, the related technology has low response accuracy to voice commands in complex scenarios, resulting in an operation experience that is not smooth and intelligent enough, and even affects driving safety. Especially in multi-task interaction scenarios, this limitation may make it difficult to accurately execute user commands, thereby limiting the application effect and user experience of the voice control system. That is, there is a problem in the prior art of inaccurate recognition of user intentions for in-vehicle voice control methods. Summary of the invention

[0004] A series of simplified concepts are introduced in the summary of the invention, which will be further described in detail in the detailed description. The summary of the invention of this application does not mean to attempt to define the key features and essential technical features of the technical solution claimed for protection, nor does it mean to attempt to determine the scope of protection of the technical solution claimed for protection.

[0005] The voice control method and related devices provided in the present application can perform dual recognition by combining the user's voice signal and gaze content. The method can accurately and efficiently judge the user's real needs, reduce false touches and conflicts, simplify interaction steps, optimize operational efficiency, and make voice control more natural, smooth, safe and convenient.

[0006] In a first aspect, the present application provides a voice control method, comprising: in response to a voice signal input by a target user, obtaining the target user's gaze content on a visible interface within a vehicle; parsing semantic information of the voice signal and determining an intention correlation between the semantic information and the gaze content; and responding to a vehicle control operation corresponding to the semantic information based on the intention correlation.

[0007] In a feasible implementation, the method of obtaining the target user's gaze content on a visible interface inside the vehicle in response to a voice signal input by the target user includes: obtaining the target user's visual movement in response to the voice signal input by the target user, wherein the visual movement includes facial movement and / or eye movement; and determining the gaze content based on the visual movement and the visible interface inside the vehicle.

[0008] In a feasible implementation, determining the gaze content based on the visual surface movement and the visible interface inside the vehicle includes: judging whether the target user is looking at the visible interface inside the vehicle based on the visual surface movement; and when it is judged that the target user is looking at the visible interface inside the vehicle, determining the gaze content based on the visual surface movement and the visible interface inside the vehicle.

[0009] In a feasible implementation, the method of obtaining the visual movement of the target user in response to a voice signal input by the target user includes: obtaining the visual movement of the target user within a first preset time period in response to the voice signal input by the target user, wherein the first preset time period includes a second preset time period before obtaining the voice signal and / or a third preset time period after obtaining the voice signal.

[0010] In a feasible implementation, determining the intention association between the semantic information and the gaze content includes: performing intent extraction on the semantic information to obtain a first vehicle-mounted control and a first control intention; performing content extraction on the gaze content to obtain a second vehicle-mounted control and a second control intention; and determining the intention association based on the first vehicle-mounted control, the first control intention, the second vehicle-mounted control, and the second control intention.

[0011] In a feasible implementation, the vehicle control operation corresponding to the semantic information is responded to according to the intention association degree, including: when the intention association degree is less than or equal to a first preset association degree, the vehicle control operation corresponding to the semantic information is not performed; when the intention association degree is greater than or equal to a second preset association degree, the vehicle control operation corresponding to the semantic information is performed, wherein the second preset association degree is greater than the first preset association degree.

[0012] In a feasible implementation, the responding to the vehicle control operation corresponding to the semantic information according to the intention association also includes: when the intention association is less than the second preset association and greater than the first preset association, obtaining the historical voice signal of the target user within a fourth preset time period; parsing the historical semantic information of the historical voice signal, and determining the semantic association between the historical semantic information and the semantic information; when the semantic association is greater than or equal to the third preset association, executing the vehicle control operation corresponding to the semantic information; when the semantic association is less than the third preset association, not executing the vehicle control operation corresponding to the semantic information.

[0013] In a second aspect, the present application also provides a voice control device, comprising: a content acquisition unit, used to respond to a voice signal input by a target user to obtain the target user's gaze content on a visible interface within the vehicle; a correlation determination unit, used to parse the semantic information of the voice signal and determine the intention correlation between the semantic information and the gaze content; and a control response unit, used to respond to a vehicle control operation corresponding to the semantic information according to the intention correlation.

[0014] In a third aspect, the present application further provides an electronic device, comprising: a memory and a processor, wherein the processor is configured to implement the steps of the voice control method described in the first aspect when executing a computer program stored in the memory.

[0015] In a fourth aspect, the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the voice control method described in the first aspect.

[0016] In a fifth aspect, the present application also provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the voice control method provided in an embodiment of the present application.

[0017] In summary, the present application can accurately judge the real needs of the user by combining the user's voice signal and gaze content, avoid false touches caused by the user's voice input in a non-attention state, effectively filter irrelevant operation requests, and improve the recognition accuracy of the user's real needs; and this dual recognition method based on gaze and voice reduces complex interaction steps, enhances the fluency of user experience, and makes the operation more natural and convenient; by evaluating the intention correlation between the voice signal and the gaze content, the corresponding vehicle control operation is responded only when the correlation reaches the preset threshold, thereby reducing the problem of false touch conflicts caused by fuzzy instructions, and improving the stability and accuracy of the voice control method in multi-tasks and complex scenarios; and can efficiently identify the user's current needs, and immediately perform the response operation after the match is successful, reducing unnecessary waiting and verification steps, optimizing the operation efficiency, making the voice control response faster, and helping the driver to complete the operation faster. In summary, the voice control method provided by the present application combines the user's voice signal and gaze content for dual recognition. The method can accurately and efficiently judge the user's real needs, reduce false touches and conflicts, simplify the interaction steps, optimize the operation efficiency, and make the voice control more natural, smooth, safe and convenient. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present specification. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0019] Figure 1 A flow chart of a voice control method provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of the structure of a voice control device provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "is" and "has" and any variations thereof involved in the present application are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] The term "module" or "unit" in this application refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. In the following description, it is related to "some embodiments", which describes a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0025] See also Figure 1 , Figure 1 1 is a flow chart of a voice control method provided in an embodiment of the present application. The method may specifically include the following steps 101 to 103:

[0026] Step 101, in response to a voice signal input by a target user, obtaining the target user's gaze content on a visible interface in the vehicle;

[0027] Specifically, the target user is the user who issues voice commands and attempts to control the in-vehicle system, usually the driver or a designated passenger in the car; the voice signal is the command or request information issued by the target user through voice input; the visible interface in the car is various display interfaces or control areas in the car, usually including the central control screen, instrument panel, navigation display screen, etc., which displays functions that users can directly control by voice, such as air conditioning, navigation, music control, etc.; the gaze content is the in-vehicle interface area or specific display content that the target user is gazed at, and the gaze content can be a specific interface element that the user is viewing, such as a button, icon, displayed information or a specific area on the screen.

[0028] For example, when the driver activates voice control and says "adjust the temperature", the on-board microphone captures the command, and the eye tracking system detects that the driver is looking at the temperature adjustment button on the center console. It then parses the voice and determines the gaze content to determine whether to perform the operation of adjusting the temperature in the car.

[0029] By implementing step 101, the user's gaze content is captured while the voice signal is input, ensuring that not only the user's voice command is recognized, but also the user's intended focus is captured in combination with the interface he is looking at.

[0030] Step 102, parsing the semantic information of the speech signal and determining the intentional relevance between the semantic information and the gaze content;

[0031] Specifically, semantic information is the specific meaning or operational intention extracted by parsing the user's voice signal; for example, when the user issues a voice command "turn up the volume", the voice signal is parsed as an operational intention such as "adjust the volume" or "increase the volume". Intention relevance is the degree of association between the operational intention expressed in the voice signal and the content of the user's gaze, reflecting the degree of match between the user's voice command and the content of their gaze. A correlation threshold is usually used to determine whether to perform the operation. For example, when the user gazes at the "volume adjustment" area and issues a "turn up the volume" command, a high correlation between the semantic information and the gaze content is detected, that is, the intention correlation is high, and the operation will be performed; if the user gazes at the navigation interface but issues a volume command, the intention correlation is low and the command is not executed.

[0032] By implementing step 102, the semantic information of the voice signal is parsed and the intent is matched with the gaze content, so that the degree of correlation between the voice content and the gaze content can be determined, and then it can be determined whether the user intends to issue an operation instruction to the current gaze interface, which helps to reduce the risk of false triggering and is particularly effective in multi-tasking scenarios, ensuring the accuracy of operation responses and optimizing user experience.

[0033] Step 103, responding to the vehicle control operation corresponding to the semantic information according to the intention association degree;

[0034] Specifically, first, the intention relevance determined in step 102 may be evaluated. If the intention relevance is higher than a preset threshold, it may be considered that the user's voice command is highly consistent with the content of his gaze, and the corresponding command is executed. If the intention relevance is lower than a preset threshold, the user may be asked to confirm his intention again to ensure the accuracy of the operation; or several possible operation options may be provided for the user to choose from to clarify the user's intention; or the instruction may be ignored and no operation is performed.

[0035] Through the implementation of step 103, after determining that the correlation between the voice and the gaze content reaches the set threshold, the corresponding vehicle control operation is triggered to ensure that the operation is performed only when the correlation is high enough, avoiding false touches and unnecessary interference, and saving system computing resources.

[0036] In summary, the embodiment of the present application can accurately judge the real needs of the user by combining the user's voice signal and gaze content, avoid false touches caused by the user's voice input in a non-attention state, effectively filter irrelevant operation requests, and improve the recognition accuracy of the user's real needs; and this dual recognition method based on gaze and voice reduces complex interaction steps, enhances the fluency of user experience, and makes the operation more natural and convenient; by evaluating the intention correlation between the voice signal and the gaze content, only responding to the corresponding vehicle control operation when the correlation reaches the preset threshold, thereby reducing the problem of false touch conflicts caused by fuzzy instructions, and improving the stability and accuracy of the voice control method in multi-tasks and complex scenarios; and can efficiently identify the user's current needs, and immediately execute the response operation after the match is successful, reducing unnecessary waiting and verification steps, optimizing the operation efficiency, making the voice control response faster, and helping the driver to complete the operation faster. In summary, the voice control method provided by the embodiment of the present application performs dual recognition by combining the user's voice signal and gaze content. The method can accurately and efficiently judge the real needs of the user, reduce false touches and conflicts, simplify the interaction steps, optimize the operation efficiency, and make the voice control more natural, smooth, safe and convenient.

[0037] In some embodiments, the aforementioned step 101 may include: in response to a voice signal input by the target user, obtaining the target user's visual movement, wherein the visual movement may include facial movement and / or eye movement; and determining the gaze content based on the visual movement and the visible interface in the vehicle.

[0038] Specifically, visual direction refers to the target user's facial direction or eye movement, which can reflect the user's current gaze direction. Facial movement refers to the change in the user's head direction, which helps the system roughly determine whether the user is looking at a specific area in the car; eye movement refers to the subtle rotation of the user's eyeballs, which can further accurately locate the specific location of the user's gaze. By combining visual direction with in-car interface information, it is possible to identify whether the user is paying attention to specific interface content and lock the specific area of ​​their gaze.

[0039] For example, when the driver activates voice control and says "open navigation", the on-board microphone captures the command. At the same time, the monitoring system in the car detects that the driver's face is facing the central control screen and his eyes are aimed at the area where the navigation button is located. The facial direction and eye movement data can be combined to confirm that the driver is looking at the navigation area, and the gaze content is determined as the navigation interface. The voice intent is then analyzed through subsequent steps to determine whether to execute the operation of turning on the navigation.

[0040] Through the implementation of the above embodiments, combined with the user's facial and eye movement and other visual movement information, it is possible to accurately determine whether the user is paying attention to the in-vehicle interface; by obtaining the visual movement, the confirmation of the gaze content is improved, avoiding misoperation caused by line of sight deviation, and ensuring that the voice control response is more in line with the user's actual needs.

[0041] In some embodiments, the aforementioned determining the gaze content based on the visual surface movement and the visible interface inside the car may include: judging whether the target user is looking at the visible interface inside the car based on the visual surface movement; when it is judged that the target user is looking at the visible interface inside the car, determining the gaze content based on the visual surface movement and the visible interface inside the car.

[0042] Specifically, first, it is possible to determine whether the user is looking at the in-car display interface by analyzing the user's visual movements (including facial and eye direction); if it is detected that the user is not looking at the interface, the subsequent gaze content recognition steps will be skipped to save computing resources; if it is detected that the user is indeed looking at the interface, the visual movements will be further combined to accurately identify the specific area or content that the user is looking at (such as the navigation icon or temperature adjustment button on the central control screen).

[0043] Exemplarily, when the driver activates voice control and issues a "turn on the air conditioner" command, the driver monitoring system (DMS) determines whether the driver is looking at the central control screen; if the driver's face and eyes are not aimed at the screen, the gaze content recognition process is skipped and no further processing is performed; if it is confirmed that the driver is looking at the central control screen, the air-conditioning control area in the direction of the line of sight is determined as the gaze content, and whether to perform the air-conditioning turning-on operation is determined based on the voice command.

[0044] Through the implementation of the above embodiments, it is first determined whether the user is looking at the display interface, and only when it is confirmed that the user is looking at it, the specific gaze content is further determined; that is, when the user is not looking at the interface, there is no need to process the complex gaze content recognition process, which reduces unnecessary image processing, analysis and matching steps, and reduces high-frequency gaze content recognition, thereby greatly reducing the computational burden of the voice control method, improving resource utilization efficiency, and significantly reducing the power consumption of the voice control method, especially in energy-sensitive application scenarios such as vehicle-mounted systems, which can better save energy and achieve greener technology applications.

[0045] In some embodiments, the aforementioned obtaining the visual movement of the target user in response to the voice signal input by the target user may include: obtaining the visual movement of the aforementioned target user within a first preset time period in response to the voice signal input by the target user, wherein the first preset time period may include a second preset time period before obtaining the voice signal and / or a third preset time period after obtaining the voice signal.

[0046] Specifically, when the user issues a voice command, not only the current visual movement is recorded, but also the facial and eye movements of the user before and after the voice command is issued are obtained to ensure that the user's complete attention trajectory is captured; the second preset time period refers to a period of time before the voice input, and the user's attention transfer is captured through the visual movements within this period; the third preset time period continues to track the user's line of sight for a short period of time after the voice input to confirm whether the content of their gaze is consistent, so as to facilitate a more accurate understanding of the user's intention; the first preset time period may include the second preset time period and / or the third preset time period.

[0047] For example, when the driver says the command "turn on the music", the driver's visual movement data will be obtained within one second before and after the voice signal. If the system detects that the driver's eyes are fixed on the music control area before or after the voice input, it can be determined that the driver's intention is to turn on the music.

[0048] By implementing the above-mentioned embodiments, the visual movement of the user within a certain period of time before and after the voice signal of the user is obtained, so as to ensure that the user's attention can be captured before and after the user issues a command, thereby enhancing the accuracy of the recognition of the gaze content, especially the moment when the user's line of sight shifts, so as to capture the user's focus target, thereby increasing the recognition success rate of the user's command.

[0049] In some embodiments, the aforementioned determination of the intention association between semantic information and gaze content may include: performing intent extraction on the semantic information to obtain a first vehicle-mounted control and a first control intention; performing content extraction on the gaze content to obtain a second vehicle-mounted control and a second control intention; and determining the intention association based on the first vehicle-mounted control, the first control intention, the second vehicle-mounted control, and the second control intention.

[0050] Specifically, the semantic information in the user's voice is first parsed to extract the controls and intentions that the user wants to operate (i.e., the first vehicle-mounted control and the first control intention). For example, in the "volume up" command, the control extracted is "car audio" and the intention is "increase the volume"; at the same time, the user's gaze content is parsed to identify the controls on the gaze interface and the user's possible operation intentions (i.e., the second vehicle-mounted control and the second control intention); then, based on the degree of matching between the controls and the intentions of the voice and gaze content, the intention correlation between the two is calculated to determine the degree of fit between the voice command and the gaze content.

[0051] For example, when the driver gazes at the "temperature control" area on the central control screen and issues a command to "increase the temperature", the voice is extracted for intent, and "car air conditioning" is identified as the first car control, and "increase temperature" is the first control intention; the gaze content is extracted to identify "car air conditioning" as the second car control, and "temperature adjustment" as the second control intention; after comparing the two sets of controls and intentions, it is found that the correlation is high, and it is determined that the voice command and the gaze content are consistent, so the temperature increase operation is executed; if the user is looking at the "volume control" area, the correlation is low, and the command will not be responded to.

[0052] By implementing the above-mentioned embodiments, the intention in the speech semantic information and the control information in the gaze content are extracted, and the intention correlation between the two is calculated, which can effectively screen out the core needs of users and reduce responses to irrelevant content; this intent matching-based approach improves the accuracy of the response of the voice control method and ensures that the operation response is more in line with user expectations.

[0053] In some embodiments, the aforementioned response to the vehicle control operation corresponding to the semantic information based on the intention association degree may include: when the intention association degree is less than or equal to a first preset association degree, not performing the vehicle control operation corresponding to the semantic information; when the intention association degree is greater than or equal to a second preset association degree, performing the vehicle control operation corresponding to the semantic information, wherein the second preset association degree is greater than the first preset association degree.

[0054] Specifically, two thresholds can be set for the intention correlation: the first preset correlation is a lower threshold, indicating that the voice command has a low correlation with the gaze content, and the command can be ignored and no vehicle control operation is performed; the second preset correlation is a higher threshold, indicating that the voice command is highly correlated with the gaze content, and the corresponding vehicle operation can be directly responded to and performed. By setting these two thresholds, low-correlation commands can be filtered out to ensure that the operation is triggered only when it is highly correlated, thereby reducing the possibility of false touches.

[0055] Through the implementation of the above embodiments, different thresholds of intention correlation are set to distinguish between triggering or not triggering operations. The operation response can be flexibly adjusted according to different correlation thresholds, reducing the erroneous triggering of low-correlation instructions, effectively improving the reliability of instruction execution, and further ensuring the user's interactive experience and safety.

[0056] In some embodiments, the aforementioned response to the vehicle control operation corresponding to the semantic information based on the intention association degree may also include: when the intention association degree is less than the second preset association degree and greater than the first preset association degree, obtaining the historical voice signal of the target user within a fourth preset time period; parsing the historical semantic information of the historical voice signal, and determining the semantic association degree between the historical semantic information and the semantic information; when the semantic association degree is greater than or equal to the third preset association degree, executing the vehicle control operation corresponding to the semantic information; when the semantic association degree is less than the third preset association degree, not executing the vehicle control operation corresponding to the semantic information.

[0057] Specifically, when it is detected that the intentional correlation between the voice command and the gaze content is at a medium level (i.e., less than the second preset correlation but greater than the first preset correlation), the user's historical voice information within the fourth preset time period will be further retrieved; the historical voice signal within this period is parsed to extract historical semantic information, and then semantically matched with the current semantic information to calculate the semantic correlation between the two. If the correlation exceeds the third preset correlation, it is determined that the current command and the historical command are related, and the corresponding operation is performed; if it is lower than the threshold, the operation is not performed to avoid potential mis-touch.

[0058] Exemplarily, when the driver gazes at the volume adjustment area and speaks the "turn up" command, it is detected that the current intention association is between the first and second preset thresholds. At this time, the driver's recent historical voice information is retrieved and analyzed to see whether there is a similar "volume adjustment" command; if the recent voice history includes semantic information related to "volume", the calculated association exceeds the third preset threshold, the user's intention can be confirmed, and the volume increase operation is executed; if no relevant historical information is found, the association is low, and the command will not be responded to, so as to reduce misoperation.

[0059] Through the implementation of the above-mentioned embodiments, under medium intention correlation, the correlation analysis of historical voice information is used to further determine the necessity of executing the operation, and the context correlation of historical information is introduced, which helps the voice control method to further confirm user needs in uncertain scenarios, improve judgment accuracy, reduce the risk of false touches, and ensure the intelligence and continuity of the interaction process.

[0060] Furthermore, as an implementation of the aforementioned method embodiment, the present application also provides a voice control device for implementing the aforementioned method embodiment. The device embodiment corresponds to the aforementioned method embodiment. For ease of reading, the present voice control device embodiment will no longer describe the details of the aforementioned method embodiment one by one, but it should be clear that the device in the present application embodiment can correspond to and implement all the contents of the aforementioned method embodiment. Figure 2 As shown, the voice control device 20 includes: a content acquisition unit 201, a correlation determination unit 202 and a control response unit 203, wherein the content acquisition unit 201 is used to respond to the voice signal input by the target user to obtain the target user's gaze content on the visible interface in the vehicle; the correlation determination unit 202 is used to parse the semantic information of the voice signal and determine the intention correlation between the semantic information and the gaze content; the control response unit 203 is used to respond to the vehicle control operation corresponding to the semantic information according to the intention correlation.

[0061] In some embodiments, the content acquisition unit 201 is also used to obtain the target user's visual movement in response to a voice signal input by the target user, where the visual movement includes facial movement and / or eye movement; and determine the gaze content based on the visual movement and the visible interface in the car.

[0062] In some embodiments, the content acquisition unit 201 is further used to determine whether the target user is looking at a visible interface in the car according to the visual surface movement; when it is determined that the target user is looking at a visible interface in the car, the gaze content is determined according to the visual surface movement and the visible interface in the car.

[0063] In some embodiments, the content acquisition unit 201 is also used to acquire the visual movement of the target user within a first preset time period in response to a voice signal input by the target user, wherein the first preset time period includes a second preset time period before the voice signal is acquired and / or a third preset time period after the voice signal is acquired.

[0064] In some embodiments, the association determination unit 202 is also used to perform intent extraction on semantic information to obtain a first vehicle-mounted control and a first control intention; perform content extraction on gaze content to obtain a second vehicle-mounted control and a second control intention; and determine the intention association based on the first vehicle-mounted control, the first control intention, the second vehicle-mounted control, and the second control intention.

[0065] In some embodiments, the control response unit 203 is also used to not perform the vehicle control operation corresponding to the semantic information when the intention association degree is less than or equal to the first preset association degree; and to perform the vehicle control operation corresponding to the semantic information when the intention association degree is greater than or equal to the second preset association degree, wherein the second preset association degree is greater than the first preset association degree.

[0066] In some embodiments, the control response unit 203 is also used to obtain the historical voice signal of the target user within a fourth preset time period when the intention association is less than the second preset association and greater than the first preset association; parse the historical semantic information of the historical voice signal, and determine the semantic association between the historical semantic information and the semantic information; when the semantic association is greater than or equal to the third preset association, execute the vehicle control operation corresponding to the semantic information; when the semantic association is less than the third preset association, do not execute the vehicle control operation corresponding to the semantic information.

[0067] The present application also provides a computer-readable storage medium, which stores computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute any step of the voice control method provided in the present application.

[0068] In some embodiments, the computer-readable storage medium may be a memory such as RAM, read-only memory (ROM), flash memory, magnetic surface memory, optical disk, or compact disc read-only memory (CD-ROM); or it may be various devices including one or any combination of the above memories.

[0069] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0070] In some embodiments, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).

[0071] In some embodiments, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.

[0072] like Figure 3As shown, the present application also provides an electronic device 30, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor, and when the processor 320 executes the computer program 311, any step of the above-mentioned voice control method is implemented.

[0073] The present application also provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or the computer executable instruction from the computer-readable storage medium, and the processor executes the computer program or the computer executable instruction, so that the electronic device performs any step of the voice control method described above in the present application.

[0074] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A voice control method, characterized in that: include: In response to a voice signal input by a target user, obtaining a gaze content of the target user on a visible interface in the vehicle; Analyzing semantic information of the speech signal and determining the intentional relevance of the semantic information to the gaze content; A vehicle control operation corresponding to the semantic information is responded to according to the intention association degree.

2. The voice control method according to claim 1, characterized in that: The step of obtaining the target user's gaze content on the visible interface in the vehicle in response to the voice signal input by the target user includes: In response to a voice signal input by a target user, obtaining a visual movement of the target user, wherein the visual movement includes facial movement and / or eye movement; The gaze content is determined according to the viewing surface movement and the visible interface in the vehicle.

3. The voice control method according to claim 2, characterized in that: The determining the gaze content according to the visual surface movement and the visible interface in the vehicle includes: Determining whether the target user is looking at the visible interface in the vehicle according to the visual surface movement; When it is determined that the target user is gazing at the visible interface within the vehicle, the gazing content is determined according to the viewing surface movement and the visible interface within the vehicle.

4. The voice control method according to claim 2, characterized in that: The step of obtaining the target user's visual orientation in response to a voice signal input by the target user includes: In response to a voice signal input by a target user, the visual movement of the target user within a first preset time period is obtained, wherein the first preset time period includes a second preset time period before the voice signal is obtained and / or a third preset time period after the voice signal is obtained.

5. The voice control method according to claim 1, characterized in that: The determining the intentional association between the semantic information and the gaze content includes: Performing intention extraction on the semantic information to obtain a first vehicle-mounted control and a first control intention; Extracting the gaze content to obtain a second vehicle-mounted control and a second control intention; The intention association degree is determined based on the first in-vehicle control, the first control intention, the second in-vehicle control and the second control intention.

6. The voice control method according to any one of claims 1 to 5, characterized in that: The step of responding to the vehicle control operation corresponding to the semantic information according to the intention association degree includes: When the intention association degree is less than or equal to a first preset association degree, not performing a vehicle control operation corresponding to the semantic information; When the intention association degree is greater than or equal to a second preset association degree, a vehicle control operation corresponding to the semantic information is performed, wherein the second preset association degree is greater than the first preset association degree.

7. The voice control method according to claim 6, characterized in that: The step of responding to the vehicle control operation corresponding to the semantic information according to the intention association degree further includes: When the intention correlation is less than the second preset correlation and greater than the first preset correlation, acquiring a historical voice signal of the target user within a fourth preset time period; parsing the historical semantic information of the historical speech signal, and determining the semantic relevance between the historical semantic information and the semantic information; When the semantic association degree is greater than or equal to a third preset association degree, executing a vehicle control operation corresponding to the semantic information; When the semantic association degree is less than the third preset association degree, the vehicle control operation corresponding to the semantic information is not performed.

8. A voice control device, characterized in that: include: A content acquisition unit, configured to acquire the gaze content of the target user on the visible interface in the vehicle in response to a voice signal input by the target user; a relevance determination unit, configured to analyze the semantic information of the speech signal and determine the intention relevance between the semantic information and the gaze content; A control response unit is used to respond to the vehicle control operation corresponding to the semantic information according to the intention association degree.

9. An electronic device, comprising: A memory and a processor, wherein the processor is used to implement the steps of the voice control method according to any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the voice control method according to any one of claims 1 to 7 are implemented.