Method, apparatus, electronic device and storage medium for determining voice interaction information
By obtaining and analyzing the voiceprint information and personnel types of drivers and passengers in the vehicle, and determining exclusive voice interaction information, the problem of single voice information feedback from the on-board voice system is solved, and the user's emotional experience and diversity of voice interaction is improved.
Patent Information
- Application Number
- CN202310854800.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-07-12
AI Technical Summary
The existing vehicle voice system has single content and lacks temperature in voice interaction, resulting in poor user experience.
By obtaining the target voiceprint information and target personnel types of drivers and passengers in the vehicle, the target voice information used to interact with the drivers and passengers is determined based on the corresponding relationship between the preset voiceprint information, personnel type and voice information, including the emotional type of voice, the pronunciator of the voice and the voice content.
It realizes the setting of exclusive voice interaction information for different personnel and personnel types, improves the emotional experience of users in the car, makes the voice interaction information rich and diverse, and solves the problem of single form of voice information in the process of voice interaction.
Smart Images

Figure CN117174066B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice interaction, and particularly to a method, device, electronic device and storage medium for determining voice interaction information. Background Art
[0002] With the development of in-vehicle voice technology, more and more members in the vehicle can enjoy the convenience and speed brought by the voice interaction method like the driver, realizing an equal interaction experience for the whole vehicle. However, during the voice interaction process between the driver and passengers and the in-vehicle voice product, even after the same user completes multiple voice interactions at a high frequency, the voice dialogue content feedback by the in-vehicle voice system to the user is still a single and emotionless system response, resulting in a poor user experience.
[0003] Therefore, the prior art has the problem of a single form of voice interaction information. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a method, device, electronic device and storage medium for determining voice interaction information to solve the problem of single voice interaction information in the prior art.
[0005] In the first aspect of the embodiments of the present application, a method for determining voice interaction information is provided, including:
[0006] Obtaining target voiceprint information and target personnel type of the driver and passengers in the vehicle;
[0007] Determining target voice information for interacting with the driver and passengers according to the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type and voice information, where the voice information includes at least one of the emotional type of the voice, the speaker of the voice, and the voice content.
[0008] In the second aspect of the embodiments of the present application, a device for determining voice interaction information is provided, including:
[0009] An obtaining module, configured to obtain target voiceprint information and target personnel of the driver and passengers in the vehicle;
[0010] A determining module, configured to determine target voice information for interacting with the driver and passengers according to the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type and voice information, where the voice information includes at least one of the emotional type of the voice, the speaker of the voice, and the voice content.
[0011] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0012] In a fourth aspect of the embodiments of the present application, a storage medium is provided. The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0013] The beneficial effects of the embodiments of the present application are as follows:
[0014] Obtain the target voiceprint information and target personnel type of the passengers in the vehicle, and determine the target voice information for interacting with the passengers according to the target voiceprint information, target personnel type, and the pre-set correspondence relationship between voiceprint information, personnel type, and voice information, where the voice information includes at least one of the emotional type of the voice, the speaker of the voice, and the voice content; realizing the setting of exclusive voice interaction information for different personnel and personnel types, thereby enhancing the emotional experience of in-vehicle users, making the voice interaction information rich and diverse, and solving the problem of the single form of voice information feedback by the in-vehicle voice system during the voice interaction process in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a schematic flowchart of a method for determining voice interaction information provided by an embodiment of the present application;
[0017] Figure 2 is a schematic flowchart of another method for determining voice interaction information provided by an embodiment of the present application;
[0018] Figure 3 is a schematic structural diagram of a device for determining voice interaction information provided by an embodiment of the present application;
[0019] Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In the following description, specific details such as specific system architectures and technologies are presented for purposes of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.
[0021] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0022] In addition, it should be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the elements.
[0023] A method and device for determining voice interaction information according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0024] Figure 1 is a schematic flowchart of a method for determining voice interaction information provided by an embodiment of the present application.
[0025] As Figure 1 shown, the method for determining voice interaction information includes:
[0026] Step 101, obtain the target voiceprint information and target personnel type of the occupants in the vehicle.
[0027] Specifically, when obtaining the target voiceprint information of the occupants in the vehicle, the voice information that appears in the vehicle can be obtained and the voice information can be recognized to obtain the target voiceprint information. The voice information is the voice that can be recognized by the in-vehicle voice product, including but not limited to human voices and mechanical sounds. After obtaining the target voiceprint information, the target voiceprint information can be cached in the cloud.
[0028] The personnel type includes the gender, age, etc. of the personnel. For example, the personnel type can include children, male adults, female adults, elderly women, elderly men, etc. The age of the personnel does not require specific analysis to obtain an accurate age value. It only needs to determine whether it is an elderly person or a child based on facial features or voice features. The target personnel type is the personnel type corresponding to the vehicle occupants.
[0029] It should be noted that, for the accuracy of voice interaction information determination, after the in-vehicle voice system is awakened, voiceprint information and personnel type can be continuously obtained.
[0030] By obtaining the target voiceprint information and target personnel type of the vehicle occupants inside the vehicle, the identity recognition of the vehicle occupants inside the vehicle is achieved.
[0031] Step 102, according to the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type, and voice information, determine the target voice information for interacting with the vehicle occupants.
[0032] Among them, the voice information includes at least one of the emotional type of the voice, the speaker of the voice, and the voice content.
[0033] In this embodiment, the corresponding relationship between voiceprint information, personnel type, and voice information can be pre-set, so that through this corresponding relationship, the voice information matching the specific voiceprint information and specific personnel type can be queried, which provides convenience for the query of the target voice information corresponding to the target voiceprint information and target personnel type.
[0034] Specifically, the voiceprint information and personnel type can be manually input in advance, or can be obtained and cached by the in-vehicle voice product and the cockpit monitoring system (Drive Monitoring System / Occupancy Monitoring System, DMS / OMS). The voice information can be the content such as the text, tone mode, and speaker set by the user according to their preferences, or can be the voice information corresponding to the user generated by the autonomous learning algorithm based on the user's conversation habits.
[0035] It should be noted that each user can correspond to different voice information, or users with similar conversation habits can correspond to the same voice information, and no specific limitation is made here. In addition, the voice information corresponding to different personnel can be set by the user or automatically matched by the in-vehicle voice product.
[0036] It should also be noted that the emotional type of the voice includes but is not limited to relaxed, soothing, gentle, naughty, nervous, plain, and respectful. The emotional type can be custom-added by the user or set by the vehicle system.
[0037] The speaker of the voice can be a robot or a natural person. Natural persons include, but are not limited to, children, adult females, and adult males.
[0038] The voice content can be user-defined content that matches the person, system default content, or content generated by the in-vehicle voice system through autonomous learning. The voice content matches different types of people. For example, the voice content matching the elderly is more respectful, and the words used can be "you" and "please". Another example is that the voice content matching the vehicle owner can be more relaxed or playful, and the words used can be "yes sir", "start, start" or combined with dialects such as "ok". No limitation is imposed here.
[0039] The voice content can be stored in a preset copywriting library, and when the driver or passenger wakes up the in-vehicle voice system, the copywriting matching the person is selected from the copywriting library.
[0040] Through the above corresponding relationship, the target voice information for interacting with the driver or passenger is determined. The target voice information is the voice information corresponding to the target voiceprint information and the target person type of the driver or passenger in the above corresponding relationship, realizing the personalization of voice interaction information and meeting the needs of different users for different voice interaction information.
[0041] According to the technical solution provided by the embodiment of the present application, by obtaining the target voiceprint information and the target person of the driver or passenger in the vehicle, and according to the target voiceprint information, the target person type, and the pre-set corresponding relationship between the voiceprint information, the person type, and the voice information, the target voice information for interacting with the driver or passenger is determined, where the voice information includes at least one of the emotional type of the voice, the speaker of the voice, and the voice content. It realizes setting matching voice interaction information for different people, improves the emotional experience of users when using voice interaction, enriches the form of voice interaction in the voice interaction process, and solves the problem that the form of voice information feedback by the in-vehicle voice system is single in the related art during the voice interaction process of users.
[0042] In some embodiments, before determining the target voice information for interacting with the driver or passenger according to the target voiceprint information, the target person type, and the pre-set corresponding relationship between the voiceprint information, the person type, and the voice information, it further includes:
[0043] Silently registering the voiceprint information of the driver or passenger according to the historical voice interaction data between the driver or passenger and the in-vehicle voice system in the vehicle;
[0044] Generating voice information to be recommended to the driver or passenger according to the historical voice interaction data;
[0045] Establish the correspondence relationship among the voiceprint information, personnel type, and voice information of the driver and passengers.
[0046] Specifically, the historical voice interaction data is the data stored in the cloud after the driver and passengers interact with the in-vehicle voice system. When the same driver and passengers complete voice interactions multiple times, audio features can be collected, the voiceprint information of the driver and passengers can be obtained and registered, that is, silent voiceprint registration is performed on the driver and passengers.
[0047] In addition, in this embodiment, according to the historical voice interaction data between the user and the in-vehicle voice system, the usage habits and preferences of the driver and passengers can be analyzed to generate voice information to be recommended to the driver and passengers. Of course, after generating the voice information, the correspondence relationship among the voiceprint information, personnel type, and voice information of the driver and passengers can be established.
[0048] In the subsequent voice interaction process, the voice information can be intelligently recommended to the driver and passengers as the target voice information for voice interaction. It should be noted that intelligent recommendation is to recommend more suitable voice information, functions, or preferences, etc. for the driver and passengers according to the driver and passengers' language habits, common emotions, common functions, etc.
[0049] According to the technical solution provided by the embodiment of the present application, silent registration of the voiceprint information of the driver and passengers is realized, and there is no need for the user to specifically record voice for voiceprint registration, thus realizing the convenience of voiceprint registration; in addition, voice information to be recommended to the driver and passengers is generated through historical voice interaction data, so that the voice information can meet the voice interaction needs of the user and improve the voice interaction experience of the driver and passengers.
[0050] In some embodiments, after obtaining the target voiceprint information of the driver and passengers in the vehicle, the following at least one item is further included:
[0051] First: If it is determined that there are personnel among the driver and passengers whose voiceprint information has not been registered according to the pre-registered voiceprint information and the target voiceprint information, the default voice information is determined as the target voice information for interacting with all the driver and passengers.
[0052] Specifically, in this item, the pre-registered voiceprint information and the obtained target voiceprint information of the driver and passengers can be compared. If the pre-registered voiceprint information does not include the obtained target voiceprint information, it is determined that there are personnel among the driver and passengers whose voiceprint information has not been registered. For example, as an example, assume that the pre-registered voiceprint information includes voiceprint A, voiceprint B, and voiceprint C, and the obtained target voiceprint information is voiceprint D, which indicates that the driver and passengers corresponding to the voiceprint D have not registered their voiceprints.
[0053] The default voice information may be a voice information preset in the vehicle voice system, or a voice information customized by the user, which is not specifically limited here.
[0054] If there are passengers whose voiceprint information is not registered, the default voice information will be determined as the target voice information for interacting with all passengers, thereby avoiding mistakes in serious situations due to voice information not matching the current scenario.
[0055] Second: If it is determined based on the pre-registered voiceprint information and the target voiceprint information that the driver and passengers have all registered their voiceprint information, and the number of target person types is less than the number of drivers and passengers, the default voice information is determined as the target voice information for interacting with all drivers and passengers.
[0056] Specifically, if the pre-registered voiceprint information includes all the acquired target voiceprint information, it is determined that all the drivers and passengers have registered their voiceprint information.
[0057] In addition, the present embodiment can determine the number of passengers in the vehicle through in-vehicle cameras (such as DMS and OMS cameras).
[0058] Since there may be multiple passengers of the same type in the vehicle, the number of target passenger types obtained may be less than the number of passengers. If the number of target passenger types is less than the number of passengers, the default voice information is determined as the target voice information for interacting with all passengers. Since the default voice information will not have a bad impact on any person, it is ensured that the voice interaction can meet the interaction needs of most passengers.
[0059] Third: if the target person type includes a specific type, the default voice information is determined as the target voice information for interacting with drivers and passengers other than the specific type, and based on the corresponding relationship, the target voice information for interacting with the specific type of person is determined, and the specific type includes at least one of the elderly and children.
[0060] Specifically, the specific type includes at least one of the elderly and children. Whether the driver or passenger is a specific type of person can be determined by combining information such as sound information and facial features.
[0061] When the occupants of the vehicle include at least one of the elderly and children, in addition to the elderly and children, the default voice information is determined as the target voice information for interaction, thereby avoiding overly intimate greetings and overly lively speakers in front of the elderly or children, thereby avoiding the use of other voice information to create a bad experience for the elderly or children.
[0062] In addition, according to the correspondence relationship between voiceprint information, personnel types, and voice information, this method can determine the target voice information for interacting with the elderly or children, thereby increasing the fun of voice interaction for the elderly or children and improving the voice interaction experience.
[0063] According to the technical solution provided by the embodiments of the present application, different target voice information determination methods are set for different scenarios, taking into account the voice interaction methods in various personnel scenarios, so that the determined target voice information can meet the interaction needs of the occupants while avoiding any adverse effects on any occupant.
[0064] In some embodiments, after obtaining the target voiceprint information and target personnel types of the occupants in the vehicle, it further includes:
[0065] When the target personnel types of all occupants include at least two personnel types, according to the pre-set priority of the personnel types, the target voiceprint information and target personnel types corresponding to the high priority are screened out;
[0066] According to the target voiceprint information, target personnel types corresponding to the high priority, and the correspondence relationship, determine the target voice information for interacting with all occupants.
[0067] Specifically, in this embodiment, the priority of the personnel types can be pre-set. For example, the priority of the elderly is higher than that of adults, and the priority of children is higher than that of adults, etc.
[0068] The priority can be set by the occupants operating the vehicle device, or can be the system default, or can be set only by a certain type of personnel. For example, when setting, a password or a specific voiceprint personnel needs to be input to have the setting permission, which is not specifically limited here.
[0069] If the target personnel types corresponding to all occupants include at least two personnel types, at this time, the target personnel types with high priority can be screened out, and the corresponding target voiceprint information can be obtained; then, according to the target voiceprint information and target personnel types corresponding to the high priority, the corresponding target voice information is determined, and this target voice information is determined as the target voice information for interacting with all occupants, thereby realizing the priority to meet the voice interaction needs of the personnel types with high priority.
[0070] In some embodiments, the determining the target voice information for interacting with the occupants according to the target voiceprint information, target personnel types, and the pre-set correspondence relationship between voiceprint information, personnel types, and voice information includes:
[0071] If it is determined that all the vehicle occupants have registered voiceprint information based on the pre-registered voiceprint information and the target voiceprint information, and it is determined according to the target personnel type that the elderly and children are not included among all the vehicle occupants, then for each vehicle occupant, the target voice information for interacting with the vehicle occupant is determined according to the target voiceprint information, the target personnel type of the vehicle occupant, and the corresponding relationship.
[0072] Specifically, if all the vehicle occupants have registered voiceprint information and the elderly and children are not included among all the vehicle occupants, that is, there are no persons in the vehicle who need special care (such as the elderly and children), then for each vehicle occupant, the target voice information for interacting with each vehicle occupant can be determined.
[0073] In this way, by determining the target voice information for interacting with each vehicle occupant, each vehicle occupant can interact with the in-vehicle voice system through the matching voice information, thus meeting the voice interaction needs of each person and realizing the diversity and interestingness of voice interaction.
[0074] In some embodiments, after determining the target voice information for interacting with the vehicle occupant according to the target voiceprint information, the target personnel type, and the pre-set corresponding relationship between the voiceprint information, the personnel type, and the voice information, the following is further included:
[0075] When no change operation of the vehicle occupant to the corresponding target voice information is detected, the confidence level of the target voice information is increased according to a preset ratio;
[0076] When a change operation of the vehicle occupant to the corresponding target voice information is detected, the confidence level of the target voice information is decreased according to the preset ratio;
[0077] In the case where the confidence level is less than a preset threshold, the voice information corresponding to the vehicle occupant in the corresponding relationship is updated.
[0078] Specifically, the change operation can be changed by means of voice interaction or by means of manual operation, and no specific limitation is made here.
[0079] The preset ratio can be preset by the system. The same confidence gradient can be set for different voice information, or different confidence gradients can be set. For example, the confidence gradients for the voice information corresponding to adults and the elderly are different, but no limitation is made here. In addition, as an example, the preset ratio can be set to 1%, 2%, etc., that is, the confidence level of the preset ratio can be increased or decreased on the basis of the original confidence level.
[0080] When no change operation of the driver or passenger on the corresponding target voice information is detected, it indicates that the target voice information can meet the needs of the driver or passenger. At this time, the confidence level can be increased according to a preset ratio. When a change operation of the driver or passenger on the corresponding target voice information is detected, it indicates that the target voice information cannot meet the interaction needs of the driver or passenger. At this time, the confidence level can be decreased according to a preset ratio, and when the confidence level is less than a preset threshold, the voice information corresponding to the driver or passenger is changed.
[0081] In addition, in this embodiment, an initial value can be set for the confidence level of the voice information, and it can be increased or decreased based on the change operation of the driver or passenger on the corresponding voice information on the basis of the initial value. Of course, the initial value can also not be set, and the in-vehicle voice system can judge the confidence level of the voice information used during the voice interaction process, which is not limited here.
[0082] The preset threshold of the confidence level can be set according to the tolerance of the driver or passenger. For example, if the driver or passenger has a high tolerance for the voice information, the preset threshold can be set lower, otherwise it can be set higher.
[0083] According to the technical solution provided by the embodiment of the present application, it is realized to correct and learn the voice information corresponding to the driver or passenger according to the feedback of the driver or passenger, so that the voice information more suitable for the driver or passenger can be adjusted in time, and the voice interaction experience of the driver or passenger is improved.
[0084] In some embodiments, after determining the target voice information for interacting with the driver or passenger according to the target voiceprint information, the target personnel type, and the corresponding relationship between the pre-set voiceprint information, personnel type, and voice information, it further includes:
[0085] When receiving a control instruction from any driver or passenger, update the target voice information to the preset voice information; wherein, the control instruction is used to indicate interacting with the driver or passenger through the preset voice information.
[0086] Specifically, any driver or passenger refers to any person in the vehicle, including those with registered voiceprints and those without registered voiceprints.
[0087] The preset voice information is the voice information stored in the in-vehicle voice system.
[0088] When receiving a control instruction from any driver or passenger, update the target voice information to the preset voice information, that is, perform voice interaction through the preset voice information, realizing one-key unified update of the voice interaction information.
[0089] It should be noted that the control instruction can also be used to indicate stopping the current voice interaction information, thus realizing one-key closing of the voice interaction.
[0090] According to the technical solution provided by the embodiment of the present application, the content and method of voice interaction information are enriched. Exclusive voices are set for different personnel, and the exclusive voice information can be changed or turned off, enhancing the experience of the driver and passengers.
[0091] Next, Figure 2 the determination process of the target voice information in a specific embodiment of the present application will be described.
[0092] Specifically, after the vehicle is started, the number of passengers N1 in the vehicle can be identified through the DMS and OMS cameras;
[0093] Then, receive the voiceprint data sent from the vehicle cloud, and perform local voiceprint recognition to complete the voiceprint verification and personnel type recognition of the passengers in the vehicle. Assume that the number of personnel types obtained is N2;
[0094] Then, if it is identified that there are passengers with unregistered voiceprints, the default voice information is determined as the target voice information for interacting with all passengers;
[0095] If it is identified that all passengers have registered voiceprints, it is detected whether the number of personnel types is less than the number of passengers N1 in the vehicle;
[0096] If it is less, the default voice information is determined as the target voice information for interacting with all passengers; if it is not less, it is detected whether the personnel types include the elderly or children;
[0097] If so, only the target voice information is matched for the elderly and children, and other personnel use the default voice information for interaction; if the elderly and children are not included, the target voice information is determined according to the target voiceprint information, target personnel type of the passengers, and the pre-set corresponding relationship between the voiceprint information, personnel type and voice information.
[0098] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present application, which will not be elaborated here one by one.
[0099] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0100] Figure 3 It is a schematic diagram of a device for determining voice interaction information provided by an embodiment of the present application. As Figure 3 shown, the device for determining voice interaction information includes:
[0101] An acquisition module 301, configured to acquire the target voiceprint information and target personnel of the passengers in the vehicle;
[0102] A determination module 302, configured to determine target voice information for interacting with the driver or passenger according to the target voiceprint information, the target personnel type, and a pre-set correspondence relationship among voiceprint information, personnel type, and voice information, where the voice information includes at least one of an emotional type of the voice, a speaker of the voice, and voice content.
[0103] The device provided by the embodiment of the present application can implement all the method steps of the above method embodiment and achieve the same technical effect, which will not be elaborated here.
[0104] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiment of the present application.
[0105] Figure 4 is a schematic diagram of an electronic device 4 provided by an embodiment of the present application. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above various method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.
[0106] The electronic device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the electronic device 4, and do not constitute a limitation to the electronic device 4, and may include more or fewer components than those shown in the figure, or different components.
[0107] The processor 401 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0108] The memory 402 can be an internal storage unit of the electronic device 4, for example, the hard disk or memory of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4, for example, a plug-in hard disk equipped on the electronic device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 402 can also include both the internal storage unit of the electronic device 4 and the external storage device. The memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0109] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0110] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0111] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for determining voice interaction information, characterized in that, it includes: Obtain the target voiceprint information and target personnel type of the passengers in the vehicle; According to the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type, and voice information, determine the target voice information for interacting with the passengers, where the voice information includes at least one of the emotional type of the voice, the speaker of the voice, and the voice content; After obtaining the target voiceprint information and target personnel type of the passengers in the vehicle, it further includes at least one of the following: If it is determined that there are personnel with unregistered voiceprint information among the passengers according to the pre-registered voiceprint information and the target voiceprint information, then determine the default voice information as the target voice information for interacting with all passengers; If it is determined that all the passengers have registered voiceprint information according to the pre-registered voiceprint information and the target voiceprint information, and the number of the target personnel types is less than the number of passengers, then determine the default voice information as the target voice information for interacting with all passengers; If the target personnel type includes a specific type, then determine the default voice information as the target voice information for interacting with the passengers except the specific type, and according to the corresponding relationship, determine the target voice information for interacting with the personnel of the specific type, where the specific type includes at least one of the elderly and children.
2. The method according to claim 1, characterized in that, before determining the target voice information for interacting with the passengers according to the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type, and voice information, it further includes: Silently register the voiceprint information of the passengers according to the historical voice interaction data between the passengers and the in-vehicle voice system in the vehicle; Generate the voice information to be recommended to the passengers according to the historical voice interaction data; Establish the corresponding relationship between the voiceprint information, personnel type, and voice information of the passengers.
3. The method according to claim 1, characterized in that, after obtaining the target voiceprint information and target personnel type of the passengers in the vehicle, it further includes: When the target personnel types of all passengers include at least two personnel types, screen out the target voiceprint information and target personnel type corresponding to the high priority according to the pre-set priority of the personnel types; According to the target voiceprint information, target personnel type corresponding to the high priority, and the corresponding relationship, determine the target voice information for interacting with all passengers.
4. The method according to claim 1, characterized in that, determining the target voice information for interacting with the passengers according to the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type, and voice information includes: If it is determined based on the pre-registered voiceprint information and the target voiceprint information that all the vehicle occupants have registered voiceprint information, and it is determined according to the target personnel type that there are no elderly or children among all the vehicle occupants, then for each vehicle occupant, based on the target voiceprint information, target personnel type of the vehicle occupant, and the corresponding relationship, determine the target voice information for interacting with the vehicle occupant.
5. The method according to claim 1, wherein, after determining the target voice information for interacting with the vehicle occupant based on the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type, and voice information, further includes: when no change operation of the vehicle occupant to the corresponding target voice information is detected, increase the confidence level of the target voice information according to a preset ratio; when a change operation of the vehicle occupant to the corresponding target voice information is detected, decrease the confidence level of the target voice information according to the preset ratio; in the case where the confidence level is less than a preset threshold, update the voice information corresponding to the vehicle occupant in the corresponding relationship.
6. The method according to claim 1, wherein, after determining the target voice information for interacting with the vehicle occupant based on the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type, and voice information, further includes: when a control instruction of any vehicle occupant is received, update the target voice information to preset voice information; wherein, the control instruction is used to indicate interacting with the vehicle occupant through the preset voice information.
7. A device for determining voice interaction information, wherein, includes: an acquisition module, configured to acquire the target voiceprint information and target personnel of vehicle occupants; a determination module, configured to determine the target voice information for interacting with the vehicle occupant based on the target voiceprint information, target personnel type, and the pre-set corresponding relationship between voiceprint information, personnel type, and voice information, where the voice information includes at least one of the emotional type of the voice, the speaker of the voice, and the voice content; after acquiring the target voiceprint information and target personnel type of vehicle occupants, further includes at least one of the following: if it is determined based on the pre-registered voiceprint information and the target voiceprint information that there are vehicle occupants who have not registered voiceprint information, then determine the default voice information as the target voice information for interacting with all vehicle occupants; if it is determined based on the pre-registered voiceprint information and the target voiceprint information that all vehicle occupants have registered voiceprint information, and the number of the target personnel type is less than the number of vehicle occupants, then determine the default voice information as the target voice information for interacting with all vehicle occupants; If the target personnel type includes a specific type, determine the default voice information as the target voice information for interacting with passengers other than the specific type, and determine the target voice information for interacting with the personnel of the specific type according to the corresponding relationship, where the specific type includes at least one of the elderly and children.
8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A storage medium storing a computer program, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Voice processing method and system, client, equipment and storage medium
CN110310642A
Voice interaction method and device
CN111292733A