A voice interaction method, device, equipment, medium and vehicle
By working collaboratively on the user end and in the cloud, the system filters out target voice commands and responds to their actions, solving the problem of repetitive operations in voice interaction and improving the user experience.
Patent Information
- Application Number
- CN202310487275.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-04-28
AI Technical Summary
During voice interaction, if a user repeatedly speaks the same button name, it will result in repeated clicks, causing the operation to not meet the user's expectations and affecting the user experience.
The user sends the collected voice information and display page information to the cloud. The cloud generates N voice commands and collection times. The user selects the target voice command according to the preset time protection interval and collection time, and responds to the target command to control the display page action.
This effectively avoids repetitive operations in voice interaction and improves the user experience.
Smart Images

Figure CN118398012B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent voice technology, and in particular to a voice interaction method, device, equipment, medium, and vehicle. Background Technology
[0002] The "see-and-say" method in intelligent voice technology allows users to click buttons displayed on their mobile devices by speaking their names, replacing the need to physically tap the buttons. This voice interaction method significantly improves convenience and security, particularly in applications such as in-vehicle systems, aerial photography, and mobile robots in shopping malls.
[0003] However, the "see-and-say" method in intelligent voice recognition has the following problems: When a user repeatedly speaks the voice command corresponding to the same button name, it is expected that a click effect will occur once. However, after the first click effect, repeated click effects will occur. This means that after the button name is hit the first time, subsequent click effects will continue after the button changes, resulting in a discrepancy with the user's expectations and severely impacting the user experience. Therefore, how to avoid repeated voice operations during voice interaction has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a voice interaction method, apparatus, device, medium, and vehicle to solve the problem of repetitive voice operations during voice interaction.
[0005] In a first aspect, embodiments of the present invention provide a voice interaction method, the voice interaction method being applied to a user terminal, the voice interaction method comprising:
[0006] The user terminal sends the collected user voice information and the page information displayed on the user terminal's display page when the voice information is collected to the cloud.
[0007] The user terminal obtains N first voice commands generated by the cloud based on the voice information and the page information, and the collection time of the corresponding first voice commands;
[0008] The user terminal selects the target voice command from the N first voice commands according to the preset time protection interval and the acquisition time of each first voice command;
[0009] The user terminal responds to the target voice command to control the operation of the display page.
[0010] Secondly, embodiments of the present invention provide a voice interaction device, which is applied to a user terminal, and the voice interaction device includes:
[0011] The information collection module is used to send the collected user's voice information and the page information displayed on the user's terminal when the voice information is collected to the cloud.
[0012] The information acquisition module is used to acquire N first voice commands generated by the cloud based on the voice information and the page information, and the acquisition time of the corresponding first voice commands;
[0013] The information filtering module is used to filter out the target voice command from the N first voice commands according to the preset time protection interval and the acquisition time of each first voice command;
[0014] The information response module is used to respond to the target voice command in order to control the operation of the display page.
[0015] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the voice interaction method as described in the first aspect.
[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the voice interaction method as described in the first aspect.
[0017] Fifthly, embodiments of the present invention provide a vehicle including a vehicle-mounted infotainment system, the vehicle-mounted infotainment system being used to implement the voice interaction method as described in the first aspect.
[0018] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:
[0019] This invention provides a voice interaction method applied to a user terminal. The user terminal sends the collected user voice information and the page information displayed on the user terminal's screen at the time of voice information collection to the cloud. The user terminal obtains N first voice commands generated by the cloud based on the voice information and page information, along with the corresponding collection time of the first voice commands. The user terminal filters out a target voice command from the N first voice commands according to a preset time protection interval and the collection time of each first voice command. The user terminal responds to the target voice command to control the display page actions. By filtering out the target voice command from multiple first voice commands based on the collection time of the first voice command and the preset time protection interval, the method effectively avoids repetitive voice operations during voice interaction and improves the user's operating experience. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of an application environment for a voice interaction method provided in Embodiment 1 of the present invention;
[0022] Figure 2 This is a flowchart illustrating a voice interaction method provided in Embodiment 1 of the present invention;
[0023] Figure 3 This is an example diagram of a time function provided in Embodiment 2 of the present invention;
[0024] Figure 4 This is a schematic diagram of the structure of a voice interaction device provided in Embodiment 3 of the present invention;
[0025] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 5 of the present invention. Detailed Implementation
[0026] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0027] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0028] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0029] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0030] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0031] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0032] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0033] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0034] The voice interaction method provided in Embodiment 1 of this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server. The user terminal includes, but is not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. For example, in a scenario where voice interaction is used in a vehicle, the user terminal can be the vehicle's central control unit. The central control unit connects to the server via a wireless network to send relevant data to the server, and the driver and passengers control the interface displayed on the central control unit via voice commands, such as clicking and swiping.
[0035] See Figure 2 This is a flowchart illustrating a voice interaction method provided in Embodiment 1 of the present invention. The above-described voice interaction method can be applied to... Figure 1 The user end in the middle. For example... Figure 2 As shown, the voice interaction method may include the following steps:
[0036] In step S201, the user terminal sends the collected user voice information and the page information displayed on the user terminal's display page when the voice information is collected to the cloud.
[0037] The user terminal includes a sound acquisition device capable of capturing sound and a display device for displaying the corresponding page. Voice information includes the captured voice and the corresponding acquisition time. Page information includes the scene identifier of the page displayed on the user terminal and the control identifiers of corresponding controls or buttons on the page. The cloud is a cloud platform that performs data statistical analysis with the user terminal. When a user operates the user terminal using the "visible and speakable" method of intelligent voice, the user continuously speaks N voice control commands on the user terminal, where N is an integer greater than 1, such as "play a song," "please turn off," "fast forward," etc. The sound acquisition device will collect the voice information in real time and send each voice message to the cloud. Simultaneously, it will obtain the page information displayed on the user terminal's display device at the time the voice information is collected and send the page information to the cloud. The "visible and speakable" aspect of intelligent voice refers to the ability of the user to control the page by reading the displayed content aloud.
[0038] Step S202: The user terminal obtains N first voice commands generated by the cloud based on voice information and page information, and the acquisition time of the corresponding first voice commands.
[0039] The cloud stores a speech recognition database and a speech command database. The user sends the collected user speech information and the page information displayed on the user's screen when the speech information is collected to the cloud. The cloud recognizes and analyzes the speech information and page information based on the speech recognition database, and retrieves the relevant first speech command and the corresponding collection time of the first speech command from the speech command database. Then, it generates N first speech commands and the corresponding collection time of the first speech command, and then sends the generated N first speech commands and the corresponding collection time of the first speech command to the user.
[0040] Optionally, when the cloud generates N first voice commands and collects the corresponding first voice commands based on voice information and page information, it includes:
[0041] The voice information is parsed in the cloud to obtain at least one word command, as well as the command type and acquisition time of the corresponding word command;
[0042] The page information is parsed to determine the scene code and control code that match each word command, forming the first voice command for the corresponding word command;
[0043] Based on the acquisition time of the word command, determine the acquisition time of the corresponding first voice command;
[0044] By iterating through N words, we can obtain N first speech commands and the acquisition time of the corresponding first speech commands.
[0045] The instruction type refers to the type of action performed on the display screen as instructed by the corresponding instruction. For example, click and swipe actions can refer to the two common types of actions users use when operating a touchscreen: click-type actions and swipe-type actions. In this invention, the instruction type corresponding to the voice instruction can include click-type and swipe-type. A command word is a command term for the user terminal, a specific instruction that enables the user terminal to understand the user's intent. When the user terminal recognizes the command word spoken by the user, it executes the corresponding instruction operation. Therefore, a voice command word control table is pre-set in the cloud. This voice command word control table includes multiple preset commands and a control code corresponding to each command word. The control code of the command word is used to identify the type of instruction intended by the user. Then, by matching the voice information with each command word in the voice command word control table, the command word corresponding to the voice information is determined based on the matching result. Furthermore, the corresponding instruction type is determined based on the control code corresponding to the command word. The acquisition time of the voice information is used as the acquisition time of the corresponding command word.
[0046] It should be noted that the control codes for click types are the same, while the control codes for swipe types are not unique. Since swipe types include swiping up, swiping down, swiping left, and swiping right, different control codes will correspond to different swipe directions.
[0047] Scene encoding refers to the identification information of the page displayed on the user's end. In other words, the page displayed on the user's end can be uniquely identified based on this identification information. Each page displayed on the user's end corresponds to a unique scene code. Control encoding refers to the identification information of a control or button. Therefore, by parsing the page information, the scene code and control code matching each word are determined, forming the first voice command for that word.
[0048] Since each page displayed on the user's end corresponds to a unique scene code, and different controls or buttons on each page also correspond to a unique control code, after parsing the control code corresponding to each word in the cloud, the scene code of the page is combined with the control code corresponding to any control or button on the page to form the operation instruction code of the first voice command of the corresponding word. This operation instruction code can be mapped to a specific control or button on a specific page. For example, if the scene code of the current page is 001, the control code for confirming the operation is 002, and the control code for canceling the operation is 003, then the operation instruction code for confirming the operation on the current page is 001002, and the operation instruction code for canceling the operation on the current page is 001003.
[0049] For any given word command, the acquisition time of the word command is determined as the acquisition time of the corresponding first speech instruction. This process is repeated for N word commands to obtain N first speech instructions and their corresponding acquisition times.
[0050] In step S203, the user terminal selects the target voice command from N first voice commands based on the preset time protection interval and the acquisition time of each first voice command.
[0051] The time protection interval can be a preset duration with an initial value, and the unit can be milliseconds, seconds, etc. It is used to filter out the target voice command from multiple first voice commands within the preset time protection interval to avoid repeated operation of voice commands.
[0052] Optionally, before the user terminal selects the target voice command from N first voice commands based on a preset time protection interval and the acquisition time of each first voice command, the method further includes:
[0053] The user client obtains the preset time protection interval from the cloud.
[0054] The preset time protection interval is stored in the cloud. Therefore, the user terminal obtains the preset time protection interval from the cloud, and then selects the target voice command from N first voice commands based on the preset time protection interval and the collection time of each first voice command.
[0055] Optionally, the user terminal selects the target voice command from N first voice commands based on a preset time protection interval and the acquisition time of each first voice command, including:
[0056] The user terminal determines the earliest acquisition time among all the acquisition times of the first voice command, constructs at least one target time range based on the earliest acquisition time and a preset time protection interval, and determines all the first voice commands within each target time range.
[0057] For any given target time range, select the target voice command from all first voice commands within that target time range.
[0058] The process involves acquiring the acquisition time of each first voice command, sorting the acquisition times of all first voice commands according to their chronological order, and obtaining a sorting result. The first acquisition time in the sorting result is taken as the earliest acquisition time. At least one target time range is constructed using the earliest acquisition time and a preset time protection interval, and all first voice commands within each target time range are determined. For any target time range, a target voice command is selected from all the first voice commands within that target time range.
[0059] Optionally, after clustering all first voice commands within the target time range to obtain at least one clustering result, the process includes:
[0060] For any clustering result, detect whether there are duplicate first voice commands in the clustering results;
[0061] If a duplicate first voice command is detected in the clustering results, the duration of the preset time protection interval is increased to obtain an increased result;
[0062] The client sends the added results to the cloud so that the cloud can use the added results to update the preset time protection interval.
[0063] The preset time protection interval is designed to detect repeated voice commands given by users within a short period of time. By maintaining a detection time for repeated voice commands, the system aims to prevent discrepancies between the voice operation result and the user's expectations based on the detection results. Therefore, for any clustering result, the system checks whether there are repeated first voice commands within the clustering result. If repeated first voice commands are detected, it indicates that the user habitually gives more repeated first voice commands. In this case, the preset time protection interval is increased, resulting in an increased time protection interval. This increased time protection interval is then sent to the cloud as the new time protection interval, allowing the cloud to update the preset time protection interval to the new time protection interval. This makes the target time range constructed using the time protection interval more rigorous.
[0064] Optionally, after detecting whether there are duplicate first voice commands in the clustering results, the following steps are also included:
[0065] If no duplicate first voice command is detected in the clustering results, the duration of the preset time protection interval is reduced to obtain a reduction result;
[0066] The reduced results are sent to the cloud so that the cloud can update the preset time protection interval using the reduced results.
[0067] Similarly, the preset time protection interval is designed to detect repeated voice commands given by the user within a short period of time. By maintaining a detection time for repeated voice commands, the system aims to prevent the voice operation result from deviating from the user's expectations based on the detection results. Therefore, for any clustering result, if no repeated first voice command is detected in the clustering result, it means that the first voice command given by the user is relatively clear and there will be no repeated voice operation. In this case, the duration of the preset time protection interval is reduced, resulting in a reduced time protection interval. This reduced time protection interval is then sent to the cloud as the new time protection interval, so that the cloud can update the preset time protection interval to the new time protection interval, thereby improving the response efficiency of the user end.
[0068] Optionally, the user terminal constructs at least one target time range based on the earliest acquisition time and a preset time protection interval, including:
[0069] The user end constructs a time series based on the acquisition time of each first voice command;
[0070] The time series is segmented using a preset time protection interval to obtain at least one target time range.
[0071] In this process, a sliding window with a preset time protection interval is used to slide through the sorting results in steps with the preset time protection interval to obtain at least one target time range. The acquisition time within the sliding window when the sliding stops constitutes a target time range.
[0072] Optionally, for any target time range, the target voice command is selected from all first voice commands within that target time range, including...
[0073] For any target time range, cluster all first voice commands within the target time range to obtain at least one clustering result;
[0074] Determine the target speech command corresponding to each clustering result;
[0075] Iterate through all clustering results to obtain the target speech command corresponding to each clustering result.
[0076] In particular, considering that during voice interaction, when a user repeatedly speaks the voice command corresponding to the same button name, it is expected that a single click will occur. However, after the first click, repeated clicks will occur. This means that after the button name is hit for the first time, subsequent clicks will continue even after the button changes, resulting in a discrepancy with the user's expectations and severely impacting the user experience. For example, during voice interaction, there is a "play / pause" button below a music player. Currently, the button is paused. Clicking this button will play the music. If the user wants to click the "play" button via voice command, after saying "play play," that is, two consecutive voice commands are both "play." After the first voice command clicks the "play" button, the music status changes to play. If clicked again, the music will pause. However, the second voice command takes effect, still clicking in the same location, but it changes to the "pause" button. The music status changes from play to pause, which does not meet the user's expectations. Therefore, in this embodiment of the invention, for any target time range, all first voice commands within the target time range are clustered to obtain at least one clustering result. Then, the target voice command corresponding to each clustering result is determined from each clustering result. All clustering results are traversed to obtain the target voice command corresponding to each clustering result.
[0077] For example, for any target time range, the repeated first voice commands within the target time range are clustered into one category, and the individual first voice commands are clustered into a separate category. For example, assuming a target time range contains 5 first voice commands such as ABAAC, 3 A's are clustered into one cluster, 1 B's into one cluster, and 1 C's into one cluster, then clustering all the first voice commands within the target time range yields 3 clustering results.
[0078] Optionally, the target speech command corresponding to each clustering result is determined from each clustering result, including:
[0079] When the instruction type corresponding to the clustering result is the target type, a first voice instruction is determined from the clustering result as the target voice instruction of the corresponding clustering result;
[0080] When the instruction type corresponding to the clustering result is not the target type, all first voice instructions in the clustering result are determined as the target voice instructions.
[0081] The target type refers to the click type, and the command type includes click and swipe types. Users often have a habit of continuously swiping when manually operating touchscreen devices. However, swipe operations differ from click operations. Click operations change the state of the page, space, or button after the first click, and repeated clicks may produce unexpected effects, leading to accidental operations. Swipe operations, on the other hand, involve swiping up, down, left, or right on the screen, generally only affecting the page content. Because users frequently perform continuous swipe operations when manually operating the screen, they also frequently use swipe operations when using the "see and speak" method of intelligent voice commands. Voice commands corresponding to swipe types will all take effect and produce a swipe effect. Therefore, for any clustering result, when the instruction type corresponding to the clustering result is a click type, a first voice instruction is determined from the clustering result as the target voice instruction for the corresponding clustering result, and the remaining first voice instructions have no effect; conversely, when the instruction type corresponding to the clustering result is not a click type, all first voice instructions in the clustering result are determined as target voice instructions, that is, when the instruction type corresponding to the clustering result is a swipe type, and the first voice instructions corresponding to the swipe type can all take effect and produce a swipe effect.
[0082] In an optional embodiment, the preset time protection interval of cloud storage can be fine-tuned according to the user's voice habits. For example, if the user is accustomed to giving a large number of repetitive voice commands in a short period of time, since the voice commands given by the user in a short period of time may essentially be an operation, after clustering all the first voice commands within the target time range and obtaining at least one clustering result, the preset time protection interval is updated according to the number of repetitions of the first voice commands in each clustering result. This allows the preset time protection interval of cloud storage to be adjusted according to the user's personal habits, which can better filter target voice commands and improve the user experience.
[0083] Optionally, the preset time protection interval is updated based on the number of repetitions of the first voice command in each clustering result, including:
[0084] For any clustering result, detect whether there are duplicate first voice commands in the clustering results;
[0085] If a duplicate first voice command is detected in the clustering results, the duration of the preset time protection interval is increased to obtain an increased result;
[0086] The client sends the added results to the cloud so that the cloud can use the added results to update the preset time protection interval.
[0087] Specifically, for any clustering result, it checks whether there are duplicate first voice commands in the clustering result. If duplicate first voice commands are detected in the clustering result, it means that the user habitually gives more duplicate first voice commands. In this case, the duration of the preset time protection interval needs to be increased to obtain the increased time protection interval. The increased time protection interval is sent to the cloud as the new time protection interval so that the cloud can update the preset time protection interval to the new time protection interval. This makes the time protection interval more in line with the user's personal habits, and the target time range constructed using the time protection interval is more rigorous.
[0088] Optionally, updating the preset time protection interval based on the number of repetitions of the first voice command in each clustering result also includes:
[0089] If no duplicate first voice command is detected in the clustering results, the duration of the preset time protection interval is reduced to obtain a reduction result;
[0090] The reduced results are sent to the cloud so that the cloud can update the preset time protection interval using the reduced results.
[0091] Similarly, if no duplicate first voice command is detected in the clustering results, it means that the first voice command given by the user is relatively clear and there will be no repeated voice operation. Then, the duration of the preset time protection interval is reduced to obtain the reduced time protection interval. The reduced time protection interval is sent to the cloud as the new time protection interval so that the cloud can update the preset time protection interval to the new time protection interval, thereby improving the response efficiency of the user terminal and enhancing the user experience.
[0092] In step S204, the user terminal responds to the target voice command to control the display page action.
[0093] Among them, the target voice command refers to the user terminal determining whether there are repeated first voice commands within the preset time protection interval when the user speaks multiple first voice commands in succession. If there are, the target voice command is obtained from the first voice commands. The user terminal responds to the target voice command to implement the control action effect. The other first voice commands have no effect. Therefore, the target voice command determined from each cluster result is sent to the user terminal. The user terminal controls the display page action according to the received target voice command.
[0094] For example, when there is a repeated first voice command within the preset time protection interval, the first first voice command among the repeated first voice commands will take effect. For instance, if the five first voice commands ABAAC are received within the 2500ms time protection interval, assuming they correspond to the "Play", "OK", "Play", "Play" and "Back" buttons on the same display page, the user terminal will only respond to the three first voice commands ABC, that is, click the "Play", "OK" and "Back" buttons, and will not click the second and third "Play" buttons.
[0095] This invention provides a voice interaction method applied to a user terminal. The user terminal sends the collected user voice information and the page information displayed on the user terminal's screen when the voice information is collected to the cloud. The user terminal obtains N first voice commands generated by the cloud based on the voice information and the page information, along with the corresponding collection time of the first voice commands. The user terminal filters out a target voice command from the N first voice commands according to a preset time protection interval and the collection time of each first voice command. The user terminal responds to the target voice command to control the display page action. By filtering out the target voice command from multiple first voice commands based on the collection time of the first voice command and the preset time protection interval, the method effectively avoids repetitive voice operations during voice interaction and improves the user's operating experience.
[0096] Optionally, based on the above Figure 1Based on the voice interaction method of Embodiment 1, Embodiment 2 of the present invention provides a method for updating a preset time protection interval, which specifically includes:
[0097] Set the time function time(x) corresponding to the time protection interval. The value of the time function is the duration of the time protection interval. Therefore, the time function time(x) is:
[0098]
[0099] Where a is a pre-defined first constant, and a>0; b is a pre-defined second constant, and b>0; c is a pre-defined third constant, and c≥0; x is the independent variable of the time function, and the value of x is initialized to 0 by default, i.e. time(0)=a+c.
[0100] Within a time protection interval, it is determined whether a repeated first voice command occurs within that time protection interval. If a repeated first voice command occurs, the independent variable of the time function is set to x+1 to increase the duration of the time protection interval; if no repeated first voice command occurs, the independent variable of the time function is set to x-1 to decrease the duration of the time protection interval.
[0101] It should be noted that, depending on the device and version, a first constant 'a', a second constant 'b', and a third constant 'c' are pre-set on the server side (these can be adjusted in real time according to different situations). This yields the time function `time(x)` for different devices, which takes effect immediately. After a period of use, the time protection interval learns to suit the user's individual usage habits, optimizing the user experience. For example, assuming the first constant 'a' is 2500, the second constant 'b' is 0.01, and the third constant 'c' is 0, the time function would look like this: Figure 3 As shown, the characteristic of this time function is that since a is a positive constant greater than 0, the time function is monotonically increasing. When the value of x approaches positive infinity, the value of the time function time(x) is 2a+c, and when the value of x approaches negative infinity, the value of the time function time(x) is c. Therefore, by setting the constant of this time function, the value of the time function time(x) can be kept within a range that conforms to the user's habits.
[0102] The preset time protection interval is updated by detecting whether a repeated first voice command occurs within the preset time protection interval. If a repeated first voice command occurs, the duration of the time protection interval is increased; conversely, if no repeated first voice command occurs, the duration of the time protection interval is decreased. Specifically, the required increase or decrease in duration can be calculated after each complete voice interaction to update the cloud-based time protection interval, which is then used in the next voice interaction. Furthermore, the update of the time protection interval can be instantaneous. That is, after constructing at least one target time range in step S203, cluster analysis is first performed on all first voice commands within the existing target time range to update the preset time protection interval. Subsequently, the updated time protection interval is used to re-divide all first voice commands within the remaining target time range according to the acquisition time, resulting in a new target time range. This clustering and updating process is repeated until all the sorting results of the acquisition time have been processed.
[0103] Corresponding to the voice interaction method in Embodiment 1 above, Figure 4 A structural block diagram of the voice interaction device provided in Embodiment 3 of the present invention is shown. For ease of explanation, only the parts related to the embodiments of the present invention are shown.
[0104] See Figure 4 The voice interaction device includes:
[0105] The information collection module 41 is used to send the collected user's voice information and the page information displayed on the user's terminal when the voice information is collected to the cloud.
[0106] The information acquisition module 42 is used to acquire N first voice commands generated by the cloud based on the voice information and the page information, and the acquisition time of the corresponding first voice commands.
[0107] The information filtering module 43 is used to filter out the target voice command from the N first voice commands according to the preset time protection interval and the acquisition time of each first voice command.
[0108] The information response module 44 is used to respond to the target voice command in order to control the operation of the display page.
[0109] Optionally, the information filtering module 43 includes:
[0110] The instruction segmentation submodule is used by the user terminal to determine the earliest acquisition time among all the acquisition times of the first voice instructions, construct at least one target time range with the earliest acquisition time and the preset time protection interval, and determine all the first voice instructions within each target time range.
[0111] The target confirmation submodule is used to filter out the target voice command from all the first voice commands within any target time range.
[0112] Optionally, the target confirmation submodule includes:
[0113] A clustering unit is used to cluster all first voice commands within any target time range to obtain at least one clustering result.
[0114] The confirmation unit is used to determine the target speech command corresponding to each clustering result from each clustering result;
[0115] The traversal unit is used to traverse all clustering results and obtain the target speech command corresponding to each clustering result.
[0116] Optionally, the confirmation unit includes:
[0117] The first confirmation subunit is used to determine a first voice instruction from the clustering result as the target voice instruction of the corresponding clustering result when the instruction type corresponding to the clustering result is the target type;
[0118] The second confirmation subunit is used to determine all first voice commands in the clustering result as target voice commands when the command type corresponding to the clustering result is not the target type.
[0119] Optionally, the clustering unit includes:
[0120] The first detection subunit is configured to, after clustering all first voice commands within the target time range and obtaining at least one clustering result, detect whether there are duplicate first voice commands in any clustering result.
[0121] The first calculation subunit is used to increase the duration of the preset time protection interval if a duplicate first voice command is detected in the clustering result, thereby obtaining an increased result;
[0122] The first sending subunit is used by the user terminal to send the added result to the cloud, so that the cloud uses the added result to update the preset time protection interval.
[0123] Optionally, the clustering unit includes:
[0124] The second calculation unit is used to reduce the duration of the preset time protection interval after detecting whether there is a duplicate first voice command in the clustering result. If no duplicate first voice command is detected in the clustering result, the reduction result is obtained.
[0125] The second sending subunit is used to send the reduction result to the cloud so that the cloud uses the reduction result to update the preset time protection interval.
[0126] Optionally, the instruction division submodules include:
[0127] A time construction unit is used by the user terminal to construct a time sequence based on the acquisition time of each first voice command;
[0128] A time segmentation unit is used to segment the time series using the preset time protection interval to obtain at least one target time range.
[0129] The voice interaction device includes:
[0130] The time update module is used to cluster all first voice commands within the target time range and obtain at least one clustering result, and then update the preset time protection interval based on the number of repetitions of the first voice commands in each clustering result.
[0131] Optionally, the time update module includes:
[0132] The detection submodule is used to detect whether there are duplicate first voice commands in any clustering result;
[0133] An additional submodule is added to increase the duration of the preset time protection interval if a duplicate first voice command is detected in the clustering result, thereby obtaining an increased result.
[0134] The first update submodule is used by the user terminal to send the addition result to the cloud, so that the cloud uses the addition result to update the preset time protection interval.
[0135] Optionally, the time update module includes:
[0136] The reduction submodule is used to reduce the duration of the preset time protection interval if no duplicate first voice command is detected in the clustering result, thereby obtaining the reduction result;
[0137] The second update submodule is used to send the reduction result to the cloud so that the cloud uses the reduction result to update the preset time protection interval.
[0138] Optionally, the information acquisition module 42 includes:
[0139] The first parsing submodule is used to parse the voice information from the cloud to obtain at least one word command and the corresponding command type and acquisition time.
[0140] The second parsing submodule is used to parse the page information, determine the scene code and control code that match each word command, and form a first voice command corresponding to the word command;
[0141] The time determination submodule is used to determine the acquisition time of the corresponding first voice instruction based on the acquisition time of the word instruction;
[0142] The instruction generation submodule is used to traverse the N words to obtain N first voice instructions and the acquisition time of the corresponding first voice instructions.
[0143] Optionally, the voice interaction device includes:
[0144] The time acquisition module is used to obtain the preset time protection interval from the cloud before the user terminal selects the target voice instruction from the N first voice instructions based on the preset time protection interval and the acquisition time of each first voice instruction.
[0145] It should be noted that the information interaction and execution process between the above modules, sub-modules, units, and sub-units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and the resulting technical effects, please refer to the method embodiments section, which will not be repeated here.
[0146] Embodiment 4 of the present invention provides a vehicle including a vehicle-mounted unit, which is used to implement the steps in any of the above-described voice interaction method embodiments.
[0147] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 5 of the present invention. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, which, when executed by the processor, implements the steps in any of the above-described embodiments of the voice interaction methods.
[0148] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 5 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.
[0149] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0150] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of a computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0151] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0152] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.
[0153] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0154] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0155] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0156] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A voice interaction method, characterized in that, The voice interaction method is applied to the user terminal, and the voice interaction method includes: The user terminal sends the collected user voice information and the page information displayed on the user terminal's display page when the voice information is collected to the cloud. The user terminal obtains N first voice commands generated by the cloud based on the voice information and the page information, and the collection time of the corresponding first voice commands; The user terminal selects the target voice command from the N first voice commands according to the preset time protection interval and the acquisition time of each first voice command; The user terminal responds to the target voice command to control the actions of the display page; The user terminal, based on a preset time protection interval and the acquisition time of each first voice command, filters out the target voice command from the N first voice commands, including: The user terminal determines the earliest acquisition time among all the acquisition times of the first voice commands, constructs at least one target time range based on the earliest acquisition time and a preset time protection interval, and determines all the first voice commands within each target time range. For any given target time range, select the target voice command from all first voice commands within that target time range.
2. The voice interaction method according to claim 1, characterized in that, For any given target time range, select the target voice command from all first voice commands within that target time range, including: For any target time range, all first voice commands within the target time range are clustered to obtain at least one clustering result; Determine the target speech command corresponding to each clustering result; Iterate through all clustering results to obtain the target speech command corresponding to each clustering result.
3. The voice interaction method according to claim 2, characterized in that, From each clustering result, determine the target speech command corresponding to that clustering result, including: When the instruction type corresponding to the clustering result is the target type, a first voice instruction is determined from the clustering result as the target voice instruction of the corresponding clustering result; When the instruction type corresponding to the clustering result is not the target type, all first voice instructions in the clustering result are determined as the target voice instructions.
4. The voice interaction method according to claim 3, characterized in that, After clustering all first voice commands within the target time range to obtain at least one clustering result, the process includes: For any clustering result, detect whether there are duplicate first voice commands in the clustering result; If a duplicate first voice command is detected in the clustering result, the duration of the preset time protection interval is increased to obtain an increased result; The user terminal sends the added result to the cloud so that the cloud uses the added result to update the preset time protection interval.
5. The voice interaction method according to claim 3, characterized in that, After detecting whether there are duplicate first voice commands in the clustering results, the method further includes: If no duplicate first voice command is detected in the clustering results, the duration of the preset time protection interval is reduced to obtain a reduction result; The reduction result is sent to the cloud so that the cloud can use the reduction result to update the preset time protection interval.
6. The voice interaction method according to any one of claims 1-5, characterized in that, The user terminal constructs at least one target time range based on the earliest acquisition time and the preset time protection interval, including: The user terminal constructs a time series based on the acquisition time of each first voice command; The time series is segmented using the preset time protection interval to obtain at least one target time range.
7. The voice interaction method according to claim 6, characterized in that, After clustering all first voice commands within the target time range to obtain at least one clustering result, the method further includes: The preset time protection interval is updated based on the number of repetitions of the first voice command in each clustering result.
8. The voice interaction method according to claim 7, characterized in that, The preset time protection interval is updated based on the number of repetitions of the first voice command in each clustering result, including: For any clustering result, detect whether there are duplicate first voice commands in the clustering result; If a duplicate first voice command is detected in the clustering result, the duration of the preset time protection interval is increased to obtain an increased result; The user terminal sends the added result to the cloud so that the cloud uses the added result to update the preset time protection interval.
9. The voice interaction method according to claim 7, characterized in that, The preset time protection interval is updated based on the number of repetitions of the first voice command in each clustering result, and further includes: If no duplicate first voice command is detected in the clustering results, the duration of the preset time protection interval is reduced to obtain a reduction result; The reduction result is sent to the cloud so that the cloud can use the reduction result to update the preset time protection interval.
10. The voice interaction method according to claim 1, characterized in that, The cloud generates N first voice commands and corresponding acquisition times for the first voice commands based on the voice information and the page information, including: The cloud analyzes the voice information to obtain at least one word command, as well as the corresponding command type and acquisition time. The page information is parsed to determine the scene code and control code that match each word command, and a first voice command corresponding to the word command is formed. Based on the acquisition time of the given words, determine the acquisition time of the corresponding first voice command; Iterate through all the words to obtain N first speech commands and the acquisition time of the corresponding first speech commands.
11. The voice interaction method according to claim 1, characterized in that, Before the user terminal filters out the target voice command from the N first voice commands according to the preset time protection interval and the acquisition time of each first voice command, the method further includes: The user terminal obtains the preset time protection interval from the cloud.
12. A voice interaction device, characterized in that, The voice interaction device is used on the user terminal, and the voice interaction device includes: The information collection module is used to send the collected user's voice information and the page information displayed on the user's terminal when the voice information is collected to the cloud. The information acquisition module is used to acquire N first voice commands generated by the cloud based on the voice information and the page information, and the acquisition time of the corresponding first voice commands; The information filtering module is used to filter out the target voice command from the N first voice commands according to the preset time protection interval and the acquisition time of each first voice command; An information response module is used to respond to the target voice command in order to control the operation of the display page; The information filtering module includes: The instruction segmentation submodule is used by the user terminal to determine the earliest acquisition time among all the acquisition times of the first voice instructions, construct at least one target time range with the earliest acquisition time and a preset time protection interval, and determine all the first voice instructions within each target time range. The target confirmation submodule is used to filter out the target voice command from all the first voice commands within any target time range.
13. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the voice interaction method as described in any one of claims 1 to 10.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice interaction method as described in any one of claims 1 to 10.
15. A vehicle, including a vehicle-mounted infotainment system, the infotainment system being configured to implement the voice interaction method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Method for enabling voice dialogue interface, computer storage medium and device
CN110248019A
Application control method, electronic equipment, device and medium
CN115048161A