Voice personalized customization method and device, electronic equipment and storage medium
By collecting and analyzing user facial images and further input operations, the execution results of the intelligent cockpit voice system are judged and corrected, solving the problem of the voice execution results not matching the user's intentions and improving user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA FAW CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-04-24
AI Technical Summary
Existing intelligent cockpit voice systems cannot make judgments based on the specific semantic habits of users, resulting in voice execution results that do not match the user's intentions and reducing the user experience.
By collecting user facial images and analyzing expressions to determine whether the execution result of voice commands is satisfactory, and if not, collecting further input from the user to determine the true intent, the system will redirect the input to the user's true intent.
It effectively solves the problem of monotonous and rigid voice execution, and can meet the various voice intentions of different users, thereby improving user satisfaction.
Smart Images

Figure CN121922129A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle-to-everything (V2X) big data technology, and in particular to methods, devices, electronic devices and storage media for personalized voice customization. Background Technology
[0002] Currently, the voice commands in smart cockpits can only execute actions based on limited semantics and cannot make judgments based on the specific semantic habits of users.
[0003] For example, with the same voice command: "Turn on sound," user A's intention is to unmute the current device, while user B's intention is to turn on the sound settings. Currently, only user A's intention to unmute is supported, so user B's intention cannot be executed, which reduces the user experience. Summary of the Invention
[0004] The purpose of this invention is to provide a method, device, electronic device, and storage medium for personalized voice customization, which can solve the problem of monotonous and rigid voice execution. Without infinitely expanding the semantic coverage of the cloud, it can still meet the various voice intentions of different users and improve user satisfaction.
[0005] This invention provides the following solution:
[0006] According to one aspect of the present invention, a voice personalization customization method is provided, the voice personalization customization method comprising:
[0007] When executing a user's voice command, capture the user's facial image;
[0008] By analyzing facial expressions, it can be determined whether the user is satisfied with the execution result of the voice command;
[0009] If the user is not satisfied with the result, collect the user's next input.
[0010] Determine the user's true intention based on their subsequent input.
[0011] Direct voice commands toward the user's true intent.
[0012] Optionally, based on the user's subsequent input, determine the user's true intent, including:
[0013] The system searches for all possible objects that can be used to execute voice commands.
[0014] The system uses the user's subsequent input to determine the target of the voice command.
[0015] Optionally, retrieve all possible objects for the voice command through a system search, including:
[0016] Within the system settings, retrieve all possible objects.
[0017] Optionally, retrieve all possible objects for the voice command through a system search, including:
[0018] In the applications installed on the system, obtain all possible objects.
[0019] Optionally, based on the user's subsequent input, determine the user's true intent, including:
[0020] By analyzing the target object, obtain all actions related to the target object;
[0021] The system uses the user's subsequent input to determine the actual action that the voice command needs to perform.
[0022] Optionally, based on the user's subsequent input, determine the user's true intent, including:
[0023] By analyzing the historical voice command execution process, all possible command parameters associated with the voice command can be obtained;
[0024] The system uses the user's subsequent input to determine the actual command parameters associated with the current voice command.
[0025] Optionally, if the user is not satisfied with the execution result, collect the user's subsequent input operations, including:
[0026] Collect user input within a specified time period.
[0027] According to a second aspect of the present invention, a voice personalization customization device is provided, the voice personalization customization device comprising:
[0028] The first acquisition module is used to acquire the user's facial image when executing the user's voice command;
[0029] The first judgment module is used to determine whether the user is satisfied with the execution result of the voice command by analyzing the facial expression of the facial image;
[0030] The second data acquisition module is used to collect the user's subsequent input operations if the user is not satisfied with the execution result.
[0031] The second judgment module is used to determine the user's true intention based on the user's subsequent input operations;
[0032] The pointing module is used to direct voice commands toward the user's true intent.
[0033] According to three aspects of the present invention, an electronic device is provided, the electronic device comprising:
[0034] Processor, communication interface, memory, and communication bus.
[0035] The processor, communication interface, and memory communicate with each other through a communication bus.
[0036] The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the voice personalization method described above.
[0037] According to four aspects of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the voice personalization method described above.
[0038] The above solution achieves the following beneficial technical effects:
[0039] By making full use of multimodal information, this invention effectively discovers the user's true intent during each voice command, thereby solving the problem of monotonous and rigid voice execution. Without needing to infinitely expand the semantic coverage of the cloud, it can still meet the various voice intents of different users and improve user satisfaction. Attached Figure Description
[0040] Figure 1 This is a flowchart of a voice personalization customization method provided in one or more embodiments of the present invention;
[0041] Figure 2 This is a flowchart of the second judgment operation in the voice personalization customization method provided by one or more embodiments of the present invention;
[0042] Figure 3 This is a flowchart of the second judgment operation in the voice personalization customization method provided by one or more embodiments of the present invention;
[0043] Figure 4 This is a flowchart of the second judgment operation in the voice personalization customization method provided by one or more embodiments of the present invention;
[0044] Figure 5 This is a flowchart of a voice personalization customization method provided in one or more embodiments of the present invention;
[0045] Figure 6 This is a structural diagram of a voice personalization device provided in one or more embodiments of the present invention;
[0046] Figure 7 This is a structural diagram of an electronic device provided in one or more embodiments of the present invention. Detailed Implementation
[0047] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Figure 1 This is a flowchart of a voice personalization method provided in one or more embodiments of the present invention. See also... Figure 1 The personalized voice customization method includes the following steps:
[0049] S11: When executing a voice command input by the user, capture the user's facial image.
[0050] S12, by analyzing facial expressions in the facial image, determines whether the user is satisfied with the execution result of the voice command.
[0051] S13, If the user is not satisfied with the execution result, collect the user's next input operation.
[0052] S14, determine the user's true intention based on the user's subsequent input.
[0053] S15 directs voice commands toward the user's true intent.
[0054] Currently, the wave of digital and intelligent technologies is profoundly impacting people's vehicle usage habits.
[0055] Users in the cockpit frequently use voice commands to adjust the equipment. However, existing in-vehicle infotainment systems often fail to accurately recognize the user's actual intent when inputting voice commands.
[0056] For a simple example, when a user says "play music," they often already have some expectation in mind about the specific piece of music to play. For instance, classical music fans might prefer to listen to Bach, while pop music fans might prefer to listen to Zhang Jie's songs, and so on. However, the car's infotainment system usually cannot capture these user expectations when they issue the command.
[0057] Furthermore, there are some instances where the voice commands are not properly understood during execution. For example, the voice command "turn on the air conditioner" should, in summer, switch the air conditioner to cooling mode. The same command, in winter, should switch it to heating mode. Simply setting the air conditioner to blow cold air regardless of the season misinterprets the user's intended meaning.
[0058] Considering the factors mentioned above, voice commands in the cockpit often result in outcomes that do not match the user's original expectations.
[0059] The purpose of this embodiment is to discover these intentional distortions in the actual cockpit, and with reference to some reference data, to discover the user's original intention as much as possible, and then to complete the execution of voice commands according to the original intention.
[0060] To determine whether the user has approved the execution of the voice command, a camera is installed in the cockpit in this embodiment. Furthermore, the camera captures the facial image of the user issuing the voice command in real time.
[0061] After capturing facial images of users inside the cockpit, the system identifies the users' emotions through facial recognition. Typically, there are two possible outcomes for the emotion assessment: a positive emotion assessment or a negative emotion assessment.
[0062] A negative emotion assessment result indicates that the user is dissatisfied with the current voice command execution. In other words, the voice command execution process needs improvement.
[0063] It should be understood that when there is a discrepancy between the execution result of a voice command and the user's actual intention—that is, when the user is dissatisfied with the result—the user will typically perform further operations on the vehicle's central control screen to achieve their intended goal. For example, if a user inputs the voice command "turn on the air conditioning," their actual intention is to turn on the air conditioning in cooling mode, but the actual result is to turn on the air conditioning in heating mode, the user will usually manually operate the central control screen to turn on the air conditioning in heating mode.
[0064] In other words, when a discrepancy is found between the execution result of a voice command and the actual user intent, the user's actual intent can be discovered by collecting further operational actions of the user on the central control screen, thereby correcting the voice command recognition and execution process.
[0065] Specifically, if a user is dissatisfied with the result of a voice command, the system collects data on the user's subsequent actions on the central control screen. By analyzing these actions, the system determines the user's true intent behind the voice command and then redirects the user's voice commands back to that intent.
[0066] Through the above series of operations, it is possible to effectively detect situations where the execution result of semantic commands does not match the user's actual intent. When this happens, timely correction can be made to redirect the voice commands back to the user's true intent, thus solving the problem of monotonous and rigid voice execution.
[0067] Figure 2 This is a flowchart of the second judgment operation in the voice personalization customization method provided in one or more embodiments of the present invention. See also Figure 2 Based on the user's subsequent input, determine the user's true intent, including the following steps:
[0068] S21 retrieves all possible objects for the voice command through a search within the system.
[0069] S22, using the user's subsequent input, determines the object to which the voice command is actually executed.
[0070] There could be many reasons why a user's voice commands might be misinterpreted.
[0071] The solution provided in this embodiment addresses a situation where the action recognition in a voice command is correct, but the target of the action is incorrect.
[0072] For example, a user says "play music," and the central control screen opens its own music player app and starts playing music. However, what the user actually needs is to listen to the music playing on the FM radio at that moment. Therefore, the user will inevitably be dissatisfied with the vehicle control system's execution of this voice command.
[0073] The ultimate dissatisfaction stemmed from a misinterpretation of the target application for the "play" command in the voice commands. In other words, the command should have been used to listen to a web radio app, not to play a music player.
[0074] The solution to this error is to find the object to which the action was executed.
[0075] Typically, the object to which the action should be directed can be found by searching system configuration items or installed software.
[0076] Taking Ubuntu as an example, you can obtain information about the installed software packages by executing the following command, thereby determining the target of the action:
[0077] dpkg --list
[0078] apt list --installed
[0079] snap list
[0080] If more precise results are needed, the above commands can be used in conjunction with the grep command to identify potential targets.
[0081] A similar process can be used in the system's configuration settings to discover the object that the action should target.
[0082] It should be understood that the number of objects that the action should point to by searching installed software or system configuration items may not be multiple. In other words, the objects obtained by searching installed software or system configuration items are usually a group of objects.
[0083] To accurately determine the user's intent object within this set of objects, it is necessary to refer to the user's further input operations.
[0084] For example, after determining that the user is not satisfied with the result of the current voice command, the user performs a series of input operations. If one of these input operations targets an object from a previously identified set of objects, then it can be further determined that the object pointed to by the user's input operation is the actual user intent object.
[0085] Figure 3 This is a flowchart of the second judgment operation in the voice personalization customization method provided in one or more embodiments of the present invention. See also Figure 3 Based on the user's subsequent input, determine the user's true intent, including the following steps:
[0086] S31: By analyzing the target object, obtain all actions related to the target object.
[0087] S32 uses the user's subsequent input to determine the actual action that the voice command needs to perform.
[0088] The foregoing embodiments of this application provide a solution for correcting the execution result of voice commands when the action object is incorrectly identified.
[0089] However, deviations from the original user intent during the actual execution of voice commands are not entirely due to errors in action object recognition. In some cases, the recognition of the action itself is problematic.
[0090] For another example, consider this voice command: "Raise the seat." After execution, it's discovered that the user is not satisfied with the result.
[0091] The voice command was directed at the seat the user was using. Therefore, the target of the action was not misidentified. The problem lies in the understanding of the action itself.
[0092] This problem can be solved by collecting the possible actions of the object. After collecting all the object's historical actions, its possible actions were found to be: backrest moving forward, backrest moving backward, armrest rising, armrest lowering, seat cushion rising, and seat cushion lowering.
[0093] In the previous voice command execution, the action performed was to move the backrest forward. Analysis of the user's facial expressions confirmed that this action did not indicate the user's true intention. Therefore, the next action that actually indicates the user's intention might be to raise the seat cushion or the armrests.
[0094] By capturing the user's manual actions, it was discovered that the user actually needed to raise the seat cushion, thus pinpointing the user's true intention. Therefore, in the future, when receiving the same voice command from the user, the action of raising the seat cushion should be executed directly.
[0095] Figure 4 This is a flowchart of the second judgment operation in the voice personalization customization method provided in one or more embodiments of the present invention. See also Figure 4 Based on the user's subsequent input, determine the user's true intent, including the following steps:
[0096] S41, by analyzing the historical voice command execution process, obtains all possible command parameters associated with the voice command.
[0097] S42 uses the user's subsequent input to determine the actual command parameters associated with this voice command.
[0098] If a user is dissatisfied with the execution result of a voice command, in addition to the object recognition error and action recognition error described in the foregoing embodiments of this application, another possibility is that the recognition of the action and the object is accurate, but there is a problem with the execution parameters of the voice command.
[0099] For example, a user inputs the voice command "Call Mom." The user is dissatisfied with the result.
[0100] Analysis revealed that the voice command's corresponding action was to make a call, and the target was a telephone. Both of these were correct. Therefore, the problem lay in the action's execution parameter, "mother."
[0101] It should be understood that when making a phone call, there must be someone being called. A search of the system's address book revealed three people related to the mother: one whose name was simply "Mom," another was Liao Jianing's mother, and the third was her aunt.
[0102] By tracking further user input, it was confirmed that the call was made to Liao Jianing's mother. Therefore, in future calls receiving the same voice command, the call parameters should be directed to Liao Jianing's mother.
[0103] Figure 5 This is a flowchart of a voice personalization method provided in one or more embodiments of the present invention. See also... Figure 5 The personalized voice customization method includes the following steps:
[0104] S51 captures the user's facial image when executing the user's voice command.
[0105] S52 determines whether the user is satisfied with the execution result of the voice command by analyzing the facial expression of the facial image.
[0106] S53: If the user is not satisfied with the execution result, collect the input operations within the time period set by the user.
[0107] S54 determines the user's true intention based on the user's subsequent input.
[0108] S55 directs voice commands toward the user's true intent.
[0109] It should be understood that when a user is dissatisfied with the execution result of a voice command, the time limit for obtaining further input from the user should not be extended indefinitely.
[0110] This limitation is primarily based on the consideration that when users find the execution result of a voice command unsatisfactory, they will generally immediately perform corresponding input operations based on their original intent. After completing the input corresponding to their original intent, the subsequent operations are unrelated to the execution of the voice command.
[0111] Therefore, providing a time limit for further input operations can better ensure that the collected input operations are precisely those related to the user's original intent, avoiding the collection of invalid operations and thus invalidating the collected data.
[0112] Figure 6 This is a structural diagram of a voice personalization device provided in one or more embodiments of the present invention. See also... Figure 6 Personalized voice control devices include:
[0113] The first acquisition module 61 is used to acquire the user's facial image when executing the user's voice input command.
[0114] The first judgment module 62 is used to determine whether the user is satisfied with the execution result of the voice command by analyzing the facial expression of the facial image.
[0115] The second acquisition module 63 is used to acquire the user's subsequent input operations if the user is not satisfied with the execution result.
[0116] The second judgment module 64 is used to judge the user's true intention based on the user's subsequent input operations.
[0117] The pointing module 65 is used to direct voice commands toward the user's true intent.
[0118] It is worth noting that although only some basic functional modules are disclosed in the embodiments of this invention, it does not mean that the composition of this system is limited to the above-mentioned basic functional modules. On the contrary, what this embodiment intends to express is that, based on the above-mentioned basic functional modules, those skilled in the art can arbitrarily add one or more functional modules in combination with existing technology to form an infinite number of embodiments or technical solutions. That is to say, this system is open rather than closed. The fact that this embodiment only discloses a few basic functional modules should not be considered as the scope of protection of the claims of this invention being limited to the disclosed basic functional modules. At the same time, for the convenience of description, the above device is described separately according to its functions as various units and modules. Of course, in implementing this invention, the functions of each unit and module can be implemented in one or more software and / or hardware.
[0119] like Figure 7 As shown, the present invention also provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of a voice personalization customization method.
[0120] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. For example... Figure 7 The structure shown in this embodiment of the invention includes an electronic device comprising one or more processors 710 and a memory 720; the processors 710 in this electronic device may be one or more. Figure 7 Taking a processor 710 as an example; a memory 720 is used to store one or more programs; the one or more programs are executed by the one or more processors 710, so that the one or more processors 710 implement the voice personalization customization method as described in any one of the embodiments of the present invention.
[0121] The electronic device may also include an input device 730 and an output device 740.
[0122] The processor 710, memory 720, input device 730, and output device 740 in this electronic device can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0123] The memory 720 in this electronic device serves as a computer-readable storage medium, capable of storing one or more programs. These programs can be software programs, computer-executable programs, or modules, such as the program instructions / modules corresponding to the voice personalization customization method provided in this embodiment. The processor 710 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 720, thereby implementing the voice personalization customization method described in the above embodiment.
[0124] The memory 720 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 720 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 720 may further include memory remotely located relative to the processor 710, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0125] Input device 730 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 740 may include display devices such as a display screen.
[0126] The present invention also provides a computer-readable storage medium storing a computer program executable by an electronic device, which, when run on the vehicle infotainment system, causes the electronic device to perform the steps of a voice personalization method.
[0127] Specifically, the computer storage medium in this embodiment of the invention can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be—but is not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for personalized voice customization, characterized in that, The voice personalization method includes: When executing a user's voice command, capture the user's facial image; By analyzing facial expressions, it can be determined whether the user is satisfied with the execution result of the voice command; If the user is not satisfied with the result, collect the user's next input. Determine the user's true intention based on their subsequent input. Direct voice commands toward the user's true intent.
2. The method according to claim 1, characterized in that, Based on the user's subsequent input, determine the user's true intent, including: The system searches for all possible objects that can be used to execute voice commands. The system uses the user's subsequent input to determine the target of the voice command.
3. The method according to claim 2, characterized in that, The system searches for all possible objects for the voice command, including: Within the system settings, retrieve all possible objects.
4. The method according to claim 2, characterized in that, The system searches for all possible objects for the voice command, including: In the applications installed on the system, obtain all possible objects.
5. The method according to claim 1, characterized in that, Based on the user's subsequent input, determine the user's true intent, including: By analyzing the target object, obtain all actions related to the target object; The system uses the user's subsequent input to determine the actual action that the voice command needs to perform.
6. The method according to claim 1, characterized in that, Based on the user's subsequent input, determine the user's true intent, including: By analyzing the historical voice command execution process, all possible command parameters associated with the voice command can be obtained; The system uses the user's subsequent input to determine the actual command parameters associated with the current voice command.
7. The method according to claim 1, characterized in that, If the user is not satisfied with the result, collect the user's subsequent input, including: Collect user input within a specified time period.
8. A voice personalization customization device, characterized in that, The personalized voice customization device includes: The first acquisition module is used to acquire the user's facial image when executing the user's voice command; The first judgment module is used to determine whether the user is satisfied with the execution result of the voice command by analyzing the facial expression of the facial image; The second data acquisition module is used to collect the user's subsequent input operations if the user is not satisfied with the execution result. The second judgment module is used to determine the user's true intention based on the user's subsequent input operations; The pointing module is used to direct voice commands toward the user's true intent.
9. An electronic device, characterized in that, The electronic device includes: Processor, communication interface, memory, and communication bus. The processor, communication interface, and memory communicate with each other through a communication bus. The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the voice personalization method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the voice personalization method according to any one of claims 1 to 7.