Voiceprint information processing method and device, equipment, storage medium and program product

By monitoring the user's voice data during the vehicle door closing event and performing voiceprint registration without any sense of touch, the problem of complexity in voiceprint registration in the smart cockpit is solved, and the user experience and matching success rate are improved.

CN120673766APending Publication Date: 2025-09-19ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510823000.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The voiceprint registration process of existing smart cockpits requires users to input voice in a quiet environment, which increases operational complexity and reduces user experience.

Method used

By monitoring the vehicle door closing event to obtain user voice data, voiceprint and portrait information are registered without the user's active input, and voice data in the vehicle's internal environment is used for non-sensing voiceprint registration.

Benefits of technology

It reduces the complexity of user operations, improves the success rate of voiceprint matching and the user experience of the smart cockpit, and ensures targeted service provision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673766A_ABST
    Figure CN120673766A_ABST
Patent Text Reader

Abstract

The invention provides a voiceprint information processing method and device, equipment, a storage medium and a program product, and relates to the technical field of vehicles, and the method comprises the steps: obtaining user voice data in the internal environment of a vehicle under the condition that a vehicle door closing event is monitored, marking the user voice data as voice data of a first user associated with the vehicle door closing event; performing voiceprint information extraction based on the user voice data to obtain first voiceprint information of the first user, and performing portrait information extraction based on the user voice data to obtain first portrait information of the first user; and registering the first voiceprint information, and registering the first portrait information. By adopting the method and the device, the voiceprint information of the user can be subjected to non-inductive registration, so that the complexity of user operation is reduced, and the overall use experience of the user on the intelligent cabin is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle technology, and in particular to a method, device, equipment, storage medium, and program product for processing voiceprint information. Background Art

[0002] With the popularization of voice technology in smart cockpits, the extraction of voiceprint information from voice data for identification and authentication of voice interaction objects is becoming more and more widely used in smart cockpits.

[0003] In the related art, the application of voiceprint technology in smart cockpits generally requires users to register their voiceprints in advance to establish a user portrait. Only after the user has registered their voiceprints can the smart cockpit accurately identify and authenticate the user through voiceprint matching in the subsequent voice interaction process to provide users with targeted services. However, the process of voiceprint registration for users in smart cockpits is generally performed conscious, and the conscious voiceprint registration process significantly increases the complexity of user operations. For example, the process of guiding users to register their voiceprints is based on a special interface, and during this process, the user is required to speak several voices in a relatively quiet environment. This is to say that the way smart cockpits in the related art register users' voiceprints to establish user portraits may reduce the user's overall experience of the smart cockpit. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a method, device, equipment, storage medium and program product for processing voiceprint information, aiming to improve the user experience by seamlessly registering the user's voiceprint information.

[0005] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for processing voiceprint information, the method comprising:

[0006] When a vehicle door closing event is detected, acquiring user voice data in the vehicle interior environment, and marking the user voice data as voice data of a first user associated with the vehicle door closing event;

[0007] Extracting voiceprint information based on the user voice data to obtain first voiceprint information of the first user, and extracting portrait information based on the user voice data to obtain first portrait information of the first user;

[0008] Register the first voiceprint information, and register the first portrait information.

[0009] In some embodiments, before registering the first voiceprint information and registering the first portrait information, the method further includes:

[0010] Comparing the first voiceprint information with the registered voiceprint information to obtain a first comparison result;

[0011] The registering of the first voiceprint information and the registering of the first portrait information include:

[0012] When the first comparison result indicates that the first voiceprint information is unregistered voiceprint information, the first voiceprint information is registered as the voiceprint information of the first user, and the first portrait information is registered as the portrait information of the first user.

[0013] In some embodiments, after comparing the first voiceprint information with the registered voiceprint information to obtain a first comparison result, the method further includes:

[0014] When the first comparison result indicates that the first voiceprint information is similar to the voiceprint information of the second user in the registered voiceprint information, the voiceprint information of the second user is updated based on the first voiceprint information, and the portrait information of the second user is updated based on the first portrait information.

[0015] In some embodiments, updating the voiceprint information of the second user based on the first voiceprint information includes:

[0016] Calculating a first product of the first voiceprint information and a first forgetting factor parameter, and calculating a second product of the second user's voiceprint information and a second forgetting factor parameter; the sum of the first forgetting factor parameter and the second forgetting factor parameter is 1, and the first forgetting factor parameter is less than the second forgetting factor parameter;

[0017] The first product is superimposed on the second product to obtain the updated voiceprint information of the second user.

[0018] In some embodiments, the method further comprises:

[0019] Obtain the number of calls for the first voiceprint information; the number of calls is used to represent the amount of user voice data;

[0020] The numerical values ​​of the first forgetting factor parameter and the second forgetting factor parameter are dynamically adjusted based on the number of calls; the numerical value of the first forgetting factor parameter is in inverse proportion to the number of calls.

[0021] In some embodiments, after registering the first voiceprint information and the first portrait information, the method further includes:

[0022] Obtaining a second comparison result between voiceprint information of a third user and voiceprint information of a fourth user in the registered voiceprint information; the third user includes the first user;

[0023] When the second comparison result indicates that the voiceprint information of the third user is similar to the voiceprint information of the fourth user, the voiceprint information of the third user is merged with the voiceprint information of the fourth user, and the portrait information of the third user is merged with the portrait information of the fourth user.

[0024] In some embodiments, merging the voiceprint information of the third user with the voiceprint information of the fourth user includes:

[0025] Calculating a first forgetting factor parameter and a second forgetting factor parameter based on the number of calls of the respective identifiers of the voiceprint information of the third user and the voiceprint information of the fourth user; the number of calls of the identifier of the voiceprint information of the third user is less than the number of calls of the identifier of the voiceprint information of the fourth user, and the first forgetting factor parameter is less than the second forgetting factor parameter;

[0026] Calculating a third product of the voiceprint information of the third user and the first forgetting factor parameter, and calculating a fourth product of the voiceprint information of the fourth user and the second forgetting factor parameter;

[0027] The third product is superimposed on the fourth product to obtain combined voiceprint information.

[0028] In some embodiments, after acquiring user voice data in the vehicle interior environment and marking the user voice data as voice data of the first user associated with the vehicle door closing event, the method further includes:

[0029] The vehicle door closing event and the user voice data are uploaded to a cloud device, so that the cloud device executes the step of extracting voiceprint information based on the user voice data to obtain the first voiceprint information of the first user and subsequent steps.

[0030] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a device for processing voiceprint information, the device comprising:

[0031] A data acquisition module is configured to acquire user voice data in the vehicle interior environment when a vehicle door closing event is detected, and mark the user voice data as voice data of a first user associated with the vehicle door closing event;

[0032] an extraction module, configured to extract voiceprint information based on the user voice data to obtain first voiceprint information of the first user, and to extract portrait information based on the user voice data to obtain first portrait information of the first user;

[0033] A registration module is used to register the first voiceprint information and the first portrait information.

[0034] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes a voiceprint information processing device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the voiceprint information processing method described in the first aspect above.

[0035] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a vehicle, which is equipped with a voiceprint information processing device, and the voiceprint information processing device includes a memory and a processor, and the memory stores a computer program. When the processor executes the computer program, it implements the voiceprint information processing method described in the first aspect above.

[0036] To achieve the above-mentioned purpose, the fifth aspect of an embodiment of the present application proposes a cloud device, which is equipped with a voiceprint information processing device. The voiceprint information processing device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the voiceprint information processing method described in the first aspect above.

[0037] To achieve the above-mentioned purpose, the sixth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the voiceprint information processing method described in the first aspect above.

[0038] To achieve the above-mentioned purpose, the seventh aspect of the embodiments of the present application proposes a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the voiceprint information processing method provided in the first aspect above.

[0039] The voiceprint information processing method, apparatus, device, vehicle, cloud device, computer-readable storage medium and computer program product proposed in the present application obtain user voice data in the vehicle's internal environment when a vehicle door closing event is detected, and mark the user voice data as the voice data of the first user associated with the vehicle door closing event; extract voiceprint information based on the user voice data to obtain the first voiceprint information of the first user, and extract portrait information based on the user voice data to obtain the first portrait information of the first user; register the first voiceprint information, and register the first portrait information.

[0040] Compared to the method of displaying a special interface to guide the user to speak several voices in a quiet environment for voiceprint registration, the embodiment of the present application associates the vehicle door closing event with the collection of user voice data. When the vehicle door closing event is detected, the user voice data of the first user associated with the vehicle door closing event is obtained in the vehicle interior environment. Thereafter, the first voiceprint information of the first user is extracted from the user voice data, and at the same time, the first portrait information of the first user is extracted from the user voice data, and the first voiceprint information and the first portrait information are registered.

[0041] In this way, the embodiment of the present application can utilize the vehicle door closing event to accumulate the user voice data collected in the vehicle's internal environment after the vehicle door closing event as the first user's voice data, without displaying a special interface to guide the user to speak the voice. In this way, the user's voiceprint can be registered imperceptibly during the user's daily voice interaction with the vehicle, which can effectively reduce the complexity of user operations and thereby enhance the user's overall experience of the smart cockpit.

[0042] In addition, the embodiment of the present application always accumulates the user voice data to the first user after detecting the vehicle door closing event to extract the first voiceprint information of the first user. In this way, as the voice data of the first user accumulates more, the quality of the first voiceprint information will also be higher, and the success rate of subsequent recognition and authentication of users through voiceprint matching will also be greater, thereby accurately and stably providing users with targeted smart cockpit services, further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a flowchart of voiceprint extraction in voiceprint technology;

[0044] Figure 2 A flowchart illustrating the steps of the voiceprint information processing method provided in some embodiments of the present application;

[0045] Figure 3A schematic diagram of the door closing detection process involved in some embodiments of the method for processing voiceprint information provided in the embodiments of the present application;

[0046] Figure 4 A schematic diagram of the voiceprint extraction process involved in some embodiments of the voiceprint information processing method provided in the embodiments of the present application;

[0047] Figure 5 A flowchart of steps involved in other embodiments of the method for processing voiceprint information provided in the embodiment of the present application;

[0048] Figure 6 A schematic diagram of the voice-to-cloud migration process involved in some embodiments of the voiceprint information processing method provided in the embodiments of this application;

[0049] Figure 7 A schematic flow chart of steps in some further embodiments of the method for processing voiceprint information provided in the embodiments of the present application;

[0050] Figure 8 A schematic flow chart of steps in some further embodiments of the method for processing voiceprint information provided in the embodiments of the present application;

[0051] Figure 9 for Figure 8 Schematic diagram of the detailed steps of step S801;

[0052] Figure 10 A schematic diagram of the voiceprint storage process involved in some embodiments of the voiceprint information processing method provided in the embodiments of the present application;

[0053] Figure 11 A schematic flow chart of the steps for dynamically adjusting forgetting factor parameters involved in some embodiments of the method for processing voiceprint information provided in the embodiments of the present application;

[0054] Figure 12 A schematic flow chart of steps in some further embodiments of the method for processing voiceprint information provided in the embodiments of the present application;

[0055] Figure 13 for Figure 12 Schematic diagram of the detailed steps of step S1202;

[0056] Figure 14 A schematic diagram of the overall framework involved in a complete embodiment of the method for processing voiceprint information provided in an embodiment of the present application;

[0057] Figure 15 A schematic diagram of the structure of a device for processing voiceprint information provided in an embodiment of the present application;

[0058] Figure 16This is a hardware structure diagram of the voiceprint information processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0060] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0062] First, a brief description of the professional technical terms involved in the method for processing voiceprint information provided in the embodiments of the present application is given.

[0063] Voiceprint technology.

[0064] Similar to a person's fingerprint, each person's voice has unique characteristics. People can identify someone they know by listening to their voice, so the voice contains identifying information about that person. In analogy with a fingerprint, academia and industry often refer to the data extracted from a person's voice that represents the speaker's identity as a "voiceprint."

[0065] The current mainstream voiceprint extraction method is to extract voiceprint information from speech signals through deep neural networks. Figure 1 As shown in the figure, the voiceprint model is a deep neural network model. Its input is the speech signal s, which can also be the audio feature FBank, audio feature MFCC and other features of the speech signal. Its output is an M-dimensional (for example, 512-dimensional) vector e, which represents the voiceprint information of this speech signal s.

[0066] The use of voiceprint technology mainly includes two parts: voiceprint registration and voiceprint authentication (also known as voiceprint retrieval). Among them, voiceprint registration means: assuming there are N users to be registered, voiceprint registration is based on Figure 1 In the process, the voices s1, s2, ..., s of different users to be registered are used.N , extract speech s1,s2,...,s N The corresponding voiceprints are e1, e2, ..., e N , and create voiceprints e1, e2, ..., e N The one-to-one correspondence between user IDs 1, 2, ..., N is stored in the voiceprint database.

[0067] The voiceprint authentication and retrieval process is equivalent to inputting a voice of an unknown user s, extracting its voiceprint e, and comparing the voiceprint e with the registered data in the voiceprint library. For example, the comparison method that can be used is to calculate the cosine similarity, as shown in formula (1). Formula (1) is used to calculate the difference between the voiceprint e and the registered voiceprints e1, e2, ..., e N The cosine similarity between θ1, θ2, ..., θ n , when the maximum similarity θ n ≥λ, the authentication is considered successful (λ is a predetermined threshold. The value of λ is related to the training of the voiceprint model and is not elaborated here), and the user ID corresponding to the registered voiceprint is returned, that is, n in the above 1,2,...,N representation.

[0068]

[0069] Among them, represents the two-norm.

[0070] Next, the overall concept of the voiceprint information processing method provided in the embodiment of the present application is described.

[0071] With the prevalence of voice technology in smart cockpits, the application of voiceprint technology is also expanding. For example, voiceprint information can be used to determine the identity of the user entering the vehicle and adjust the seat parameters to the user's preferred sitting position based on previous user habits. Another example is using voiceprint information to verify whether the speaker is the legitimate vehicle owner and remotely enable the vehicle computer to start or shut down for the legitimate owner. In other words, one of the main applications of voiceprint technology in smart cockpits is user profiling. Since a vehicle is typically used by multiple users, the cockpit system first uses voiceprint technology to determine the user's ID from the human-computer conversation. It also extracts each user's characteristics from the conversation content, such as nickname, interests, conversation style, and vehicle computer operating habits. This creates a comprehensive description of the user's behavior and habits, which is referred to in the industry as a "user profile." For example, the seat parameters mentioned above are also an attribute of the user profile.

[0072] In related technologies, voiceprint user profiling technology for smart cockpits requires a visible voiceprint registration process. This involves an interface guiding the user to speak several voice commands in a relatively quiet environment to complete the voiceprint registration process. Only users who have registered their voiceprint can achieve voiceprint matching during subsequent voice interactions. This user-perceived voiceprint registration process significantly increases the complexity of user operations, thereby reducing the user experience of the smart cockpit. Furthermore, the accuracy of voiceprint matching is significantly related to the length and quality of the speech. The longer the input speech, the more stable its statistical properties and the higher the quality of the extracted voiceprint, thus ensuring a high success rate for voiceprint matching. However, certain short interaction commands, such as "open window," "turn on air conditioning," or the wake-up word for the vehicle computer, can significantly reduce the success rate of voiceprint matching due to the lack of effective information.

[0073] In summary, the process of voiceprint registration for users in smart cockpits is generally carried out in a conscious manner, and the conscious voiceprint registration process will significantly increase the complexity of user operations. In addition, since the accurate matching of voiceprints is directly related to the length and quality of the user's voice, the smart cockpit may not be able to complete voiceprint matching for some shorter voice interaction commands output by the user due to the lack of effective information in the voice interaction commands, making it difficult to provide users with targeted command response services. In other words, the way in which smart cockpits in related technologies register users' voiceprints to establish user portraits may reduce the user's overall experience of the smart cockpit.

[0074] To this end, the embodiments of the present application provide a method, apparatus, device, vehicle, cloud device, computer-readable storage medium, and computer program product for processing voiceprint information, aiming to overcome the shortcomings of the above-mentioned related technologies and improve the user experience by seamlessly registering the user's voiceprint information.

[0075] For the user portrait task in the smart cockpit, the embodiment of the present application monitors the vehicle door closing event, and when the vehicle door closing event is detected, obtains the user voice data in the vehicle's internal environment, and marks the user voice data as the voice data of the first user associated with the vehicle door closing event. Thereafter, voiceprint information is extracted based on the user voice data to obtain the first voiceprint information of the first user, and portrait information is extracted based on the user voice data to obtain the first portrait information of the first user. Finally, the first voiceprint information is registered, and the first portrait information is registered.

[0076] Thus, compared to a method that displays a dedicated interface to guide the user to speak several voices in a quiet environment for voiceprint registration, the embodiment of the present application associates the vehicle door closing event with the collection of user voice data. When a vehicle door closing event is detected, user voice data of a first user associated with the vehicle door closing event is obtained from the vehicle interior environment, and the first voiceprint information of the first user is extracted from the user voice data. Simultaneously, the first portrait information of the first user is extracted from the user voice data. The first voiceprint information and the first portrait information are then registered. In this way, by utilizing the vehicle door closing event, all user voice data collected from the vehicle interior environment after the vehicle door closing event is accumulated as the first user's voice data. This eliminates the need to display a dedicated interface to guide the user to speak a few voices. Voiceprint registration can be seamlessly performed during daily voice interactions between the user and the vehicle, effectively reducing the complexity of user operations and, in turn, improving the user's overall experience with the smart cockpit.

[0077] In addition, the embodiment of the present application always accumulates the user voice data to the first user to extract the first voiceprint information of the first user after detecting the vehicle door closing event. In this way, as the voice data of the first user accumulates more, the quality of the first voiceprint information will also be higher, and the success rate of subsequent recognition and authentication of users through voiceprint matching will also be greater, thereby accurately and stably providing users with targeted smart cockpit services, further improving the user experience.

[0078] Next, the voiceprint information processing method, device, equipment, vehicle, computer-readable storage medium and computer program product provided in the embodiments of the present application are specifically described through the following embodiments, and the voiceprint information processing method provided in the embodiments of the present application is first described in detail.

[0079] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0080] It should be noted that the method for processing voiceprint information provided in the embodiments of the present application can be applied to the terminal, can also be applied to the server side, and can also be software running in the terminal or the server side. In some embodiments, the terminal can be an on-board terminal on the vehicle (such as an on-board computing platform), or can be a computer device such as a smart phone, tablet computer, laptop computer, desktop computer, etc. associated with the vehicle. The association between the terminal and the vehicle means that the terminal can communicate and interact with the vehicle based on the network. The server side can be the background server terminal device of the vehicle, which can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers. It can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The software can be an application that implements the method for processing voiceprint information, a computer program, and a storage medium that carries the computer program. It should be understood that based on different design requirements of actual applications, in different feasible embodiments, the terminal, server side, and software that apply the voiceprint information processing method provided in the embodiments of the present application may of course also be in other forms not listed here, and the voiceprint information processing method provided in the embodiments of the present application does not specifically limit this.

[0081] In addition, the present application can also be used in many general or special computer system environments or configurations. For example: vehicles, personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, personal computers (PCs), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0082] For ease of understanding and explanation, the following text will use the terminal device (directly configured on or associated with the vehicle) applying the voiceprint information processing method provided in the embodiments of the present application as an example to describe the various specific embodiments of the present application in detail. The implementation of the voiceprint information processing method provided in the embodiments of the present application in any of the above-mentioned forms of subject matter can refer to the process of applying the voiceprint information processing method provided in the terminal device described below.

[0083] Please refer to Figure 2 , Figure 2 The following is a flowchart of the steps in some embodiments of the method for processing voiceprint information provided in the embodiments of this application. It should be understood that although Figure 2 The following flowcharts show the execution order of some method steps, but based on the different design requirements of actual applications, the voiceprint information processing method provided in the embodiment of the present application can certainly adopt an execution order different from the method steps shown in the figure. Figure 2 The order of the steps in the method shown does not constitute a limitation on the execution logic order of the voiceprint information processing method provided in the embodiment of the present application. Figure 2 Reasonable changes in the order of the steps of the method shown should be included in the scope of protection of the voiceprint information processing method provided in the embodiments of the present application.

[0084] like Figure 2 As shown, in some embodiments, the method for processing voiceprint information provided in the embodiments of the present application may include steps S201 to S203 as shown below.

[0085] Step S201: when a vehicle door closing event is detected, user voice data in the vehicle interior environment is acquired, and the user voice data is marked as voice data of a first user associated with the vehicle door closing event.

[0086] It should be noted that the first user is associated with the vehicle door closing event, which means that the first user is the user who triggered the vehicle door closing event. Generally speaking, after the vehicle's smart cockpit is closed, there is only one user entering the smart cockpit. For example, when a door closing event occurs on the driver's side of the vehicle, there is only one user entering the driver's smart cockpit, and the user does not change until the next door closing event occurs on the driver's side of the vehicle. For another example, when a door closing event occurs on the side of the vehicle's rear business seat (a separate seat on the left or right side of the rear row when the rear seats are not connected), there is only one user entering the business seat smart cockpit, and the user does not change until the next door closing event occurs on that side of the vehicle.

[0087] While the vehicle is parked or running, the terminal device continuously monitors vehicle door closing events based on the vehicle's sensors. Subsequently, when a vehicle door closing event is detected, voice data is further collected from the vehicle's interior environment, thereby obtaining user voice data from the vehicle's interior environment. At this point, since the trigger of the vehicle door closing event and the output of user voice data from the vehicle's interior environment are generally performed by the same person, the terminal device can mark all user voice data collected from the vehicle's interior environment since the current vehicle door closing event is detected as the voice data of the first user who triggered the vehicle door-related event. In other words, the terminal device utilizes the fact that the user remains unchanged after the vehicle door is closed. After detecting a vehicle door closing event, the terminal device will accumulate user voice data collected from the same person until the next vehicle door closing event is detected, i.e., mark it as the voice data of the first user who triggered the current vehicle door closing event.

[0088] In some embodiments, after monitoring a vehicle door closing event, the terminal device can use the first user voice data obtained from the vehicle's internal environment as the wake-up word for the current voiceprint processing (including registration or update, etc.) of the first user who triggered the vehicle door closing event without any perception, thereby accumulating subsequent user voice data obtained from the vehicle's internal environment to the first user.

[0089] For example, Figure 3 In the door closing detection process shown in FIG, when the terminal device detects a vehicle door closing event (the door state δ changes from open to closed when δ=1), the terminal device obtains the user voice data s through the microphone in the vehicle interior environment, and accumulates the user voice data s obtained from the vehicle interior environment from the current vehicle door closing event until the next vehicle door closing event is detected as the first user Voice data s. Among them, A temporary ID is assigned by the terminal device when it detects δ = 1. After assigning the temporary ID, the terminal device sets the door state δ = 0 to indicate that the current vehicle door closing event has been processed. Thereafter, a new temporary ID is assigned until the next time the door state changes from open to closed δ = 1 (the next vehicle door closing event is detected).

[0090] Step S202: extracting voiceprint information based on the user voice data to obtain first voiceprint information of the first user, and extracting portrait information based on the user voice data to obtain first portrait information of the first user.

[0091] After acquiring the user's voice data, the terminal device may input the user's voice data into the voiceprint model to extract the voiceprint information therein, and mark the voiceprint information as the first voiceprint information of the first user. Simultaneously, the terminal device may also perform voice recognition, semantic understanding, and other processing on the user's voice data to extract user profile information such as the user's nickname, interests, hobbies, and usage habits from the user's voice data, and mark the profile information as the first profile information of the first user.

[0092] For example, Figure 4 The voiceprint extraction process shown in the figure, the terminal device combines the current door state δ and uses the voiceprint model to extract the temporary voiceprint from the acquired user voice data s. When δ=1 is detected, the temporary voiceprint Assign a temporary ID As the first user First voiceprint information If δ=0, then the temporary voiceprint Extract temporary voiceprint from user voice data s before allocating The temporary voiceprint Assigned voiceprint ID n. In this way, the terminal device can use the voiceprint information in the user voice data s in the whole process from monitoring the current vehicle door closing event to monitoring the next vehicle door closing event. Accumulated to the first user superior.

[0093] Step S203: registering the first voiceprint information and registering the first portrait information.

[0094] After the terminal device extracts the first voiceprint information and the first portrait information of the first user, it can complete the registration of the first voiceprint information and the first portrait information without the first user being aware of it by storing the first voiceprint information and the first portrait information in a database (such as a voiceprint library, etc.).

[0095] In some embodiments, when storing the first voiceprint information and the first portrait information in the database, the terminal device can establish a one-to-one correspondence between the first voiceprint information and the first portrait information and the first user, so that after further receiving the user voice data of the first user, the terminal device can continue to extract the voiceprint information and the portrait information from the user voice data. If the newly extracted voiceprint information matches the first voiceprint information of the first user in the database, the terminal device can use the newly extracted voiceprint information to update the first voiceprint information in the database, and use the newly extracted portrait information to update the first portrait information in the database. In this way, by continuously updating and iterating the registered voiceprint information (such as the first voiceprint information) and the registered portrait information (such as the first portrait information) in the database, it is possible to achieve strong robustness in voiceprint registration, matching, and authentication using shorter or poorer quality user voice data. That is, even if the user voice data is of low quality, it is not easy to affect the voiceprint extraction and matching process of the terminal device.

[0096] In other embodiments, after the terminal device further receives the user voice data of the first user, if the newly extracted voiceprint information matches the first voiceprint information of the first user in the database, the terminal device can also directly use the first portrait information of the first user in the database to provide targeted smart cockpit services for the first user.

[0097] In an embodiment of the present application, a terminal device continuously monitors vehicle door closing events using vehicle sensors. Upon detecting a vehicle door closing event, voice data is collected from the vehicle's interior environment to obtain user voice data within the vehicle's interior environment. Furthermore, all user voice data collected from the vehicle's interior environment since the current vehicle door closing event is marked as the voice data of the first user who triggered the door-related event. Subsequently, the terminal device inputs the user voice data into a voiceprint model to extract voiceprint information, which is then marked as the first voiceprint information of the first user. Simultaneously, the terminal device performs voice recognition and semantic understanding on the user voice data to extract user profile information, such as the user's nickname, interests, and usage habits, from the voice data, and also marks this profile information as the first profile information of the first user. Finally, the terminal device stores the first voiceprint information and the first profile information in a database (e.g., a voiceprint library), thereby completing the registration of the first voiceprint information and the first profile information without any noticeable changes to the first user.

[0098] In this way, the embodiment of the present application can utilize the vehicle door closing event to accumulate the user voice data collected in the vehicle's internal environment after the vehicle door closing event as the first user's voice data, without displaying a special interface to guide the user to speak the voice. In this way, the user's voiceprint can be registered imperceptibly during the user's daily voice interaction with the vehicle, which can effectively reduce the complexity of user operations and thereby enhance the user's overall experience of the smart cockpit.

[0099] In addition, the embodiment of the present application always accumulates the user voice data to the first user after detecting the vehicle door closing event to extract the first voiceprint information of the first user. In this way, as the voice data of the first user accumulates more, the quality of the first voiceprint information will also be higher, and the success rate of subsequent recognition and authentication of users through voiceprint matching will also be greater, thereby accurately and stably providing users with targeted smart cockpit services, further improving the user experience.

[0100] In some embodiments, the terminal device can locally execute the process shown in steps S201 to S203 to perform the non-sensing registration and update of the user's voiceprint. In other embodiments, the terminal device can also perform the non-sensing registration and update of the user's voiceprint based on the cloud device of vehicle management.

[0101] Please refer to Figure 5 , Figure 5 A flowchart of the steps involved in other embodiments of the method for processing voiceprint information provided in the embodiment of the present application.

[0102] like Figure 5 As shown, in some embodiments, after the above-mentioned step S201: obtaining user voice data in the vehicle interior environment and marking the user voice data as the voice data of the first user associated with the vehicle door closing event, the embodiment of the present application may also include but is not limited to the step S501 shown below.

[0103] Step S501: Upload the vehicle door closing event and the user voice data to a cloud device, so that the cloud device executes the step of extracting voiceprint information based on the user voice data to obtain the first voiceprint information of the first user and subsequent steps.

[0104] After acquiring the user's voice data, the terminal device can use the vehicle's pre-set vehicle-side voice-to-cloud link (or a data-to-cloud link modified from the vehicle-side voice-to-cloud link) to upload the detected vehicle door closing event and the user's voice data to the cloud device. In this way, the cloud device can execute the above-mentioned process shown in steps S201 to S203 on the cloud based on the user's voice data uploaded by the terminal device to perform processes such as non-sensing registration and updating of the user's voiceprint.

[0105] For example, Figure 6 In the voice-to-cloud process shown, after the terminal device detects a vehicle door closing event and obtains the user voice data s in the vehicle interior environment, it may deem that a wake-up word has been detected for the current voiceprint registration or update of the first user in the vehicle interior environment without any perception. The user voice data s and subsequent user voice data s obtained thereafter until the next vehicle door closing event is detected, together with the current vehicle door closing state δ, are uploaded to the cloud device, causing the cloud device to execute the operation process shown in steps S201 to S203 above. After uploading the user voice data s and the door closing state δ to the cloud device, the terminal device also sets the door closing state δ to 0 to indicate that the current door closing state has been processed.

[0106] It should be noted that the operation process of voiceprint extraction, matching, registration and updating by cloud devices in the cloud is basically the same as the operation process of voiceprint extraction, matching, registration and updating by terminal devices described in this article. The same content will not be repeated here. The processing of voiceprint information by cloud devices can directly refer to the processing of voiceprint information by terminal devices described in this article.

[0107] In some embodiments, the terminal device can compare the first voiceprint information currently extracted with the registered voiceprint information in a database (e.g., a voiceprint library), and based on the comparison result, decide whether to directly register the first voiceprint information or update the registered voiceprint information in the database based on the first voiceprint information.

[0108] Please refer to Figure 7 , Figure 7 Schematic diagram of the steps of the voiceprint information processing method provided in some embodiments of the present application.

[0109] like Figure 7 As shown, in some embodiments, before the above-mentioned step S203: registering the first voiceprint information and registering the first portrait information, the method for processing voiceprint information provided in the embodiment of the present application may further include:

[0110] Step S701: Compare the first voiceprint information with the registered voiceprint information to obtain a first comparison result.

[0111] The terminal device compares the first voiceprint information extracted this time with the registered voiceprint information (voiceprint information of one or more users) stored in the database by calculating the similarity (for example, the cosine similarity mentioned above) between the first voiceprint information extracted this time and the registered voiceprint information (voiceprint information of one or more users). At this time, the similarity result calculated by the terminal device can be used as the first comparison result obtained by comparing the first voiceprint information with the registered voiceprint information.

[0112] like Figure 7 As shown, in some embodiments, when the terminal device compares the first voiceprint information with the registered voiceprint information, the above-mentioned step S203: registering the first voiceprint information, and registering the first portrait information, may include step S702 as shown below.

[0113] Step S702: When the first comparison result indicates that the first voiceprint information is unregistered voiceprint information, register the first voiceprint information as the voiceprint information of the first user, and register the first portrait information as the portrait information of the first user.

[0114] After the terminal device calculates the similarity between the first voiceprint information and the registered voiceprint information, and uses the calculated similarity result as the first comparison result of the first voiceprint information and the registered voiceprint information, if the first comparison result is less than the preset similarity threshold, that is, the first voiceprint information does not match the registered voiceprint information (when there are multiple registered voiceprint information, the first voiceprint information does not match any of the registered voiceprint information), and the first comparison result indicates that the first voiceprint information is unregistered voiceprint information not stored in the database, then in this case, the terminal device determines that the current first user is a suspected new user, and directly stores the first voiceprint information in the database to register it as the voiceprint information of the first user, and at the same time stores the first portrait information in the database to register it as the portrait information of the first user.

[0115] In some embodiments, because the user voice data initially acquired by the terminal device may be small or of poor quality, the extracted first voiceprint information, even if it is registered voiceprint information, cannot be successfully matched with the registered voiceprint information (i.e., the first comparison result is less than a preset similarity threshold). Therefore, the terminal device first registers the first voiceprint information and the first portrait information of the first user as a suspected new user. Later, as the first user accumulates more user voice data or the quality improves, the terminal device updates the registered voiceprint information based on the first voiceprint information if the first voiceprint information successfully matches the registered voiceprint information, and also updates the registered portrait information based on the first portrait information.

[0116] Please refer to Figure 8 , Figure 8 A flowchart of the steps of the voiceprint information processing method provided in the embodiments of the present application in some further embodiments.

[0117] like Figure 8 As shown, after the above-mentioned step S701: comparing the first voiceprint information with the registered voiceprint information to obtain a first comparison result, the voiceprint information processing method provided by the embodiment of the present application may further include the following step S801.

[0118] Step S801: When the first comparison result indicates that the first voiceprint information is similar to the voiceprint information of the second user in the registered voiceprint information, the voiceprint information of the second user is updated based on the first voiceprint information, and the portrait information of the second user is updated based on the first portrait information.

[0119] After the terminal device calculates the similarity between the first voiceprint information and the registered voiceprint information, and uses the calculated similarity result as the first comparison result of the first voiceprint information and the registered voiceprint information, the first comparison result can also indicate another situation, that is, if the first comparison result is greater than or equal to the preset similarity threshold, so that the first voiceprint information matches the registered voiceprint information (when there are multiple registered voiceprint information, the first voiceprint information matches one of the registered voiceprint information), at this time, the first comparison result indicates that the first voiceprint information is similar to the voiceprint information of the second user in the registered voiceprint information, thereby indicating that the current first user is the second user (old user) whose voiceprint has been registered by the terminal device without any perception. In this way, the terminal device updates the voiceprint information of the second user in the database based on the first voiceprint information on the basis of the second user, and at the same time updates the portrait information of the second user in the database based on the first portrait information.

[0120] In this embodiment, the terminal device calculates the similarity between the first voiceprint information currently extracted and the registered voiceprint information stored in the database, obtaining a first comparison result for comparing the first voiceprint information with the registered voiceprint information. If the first comparison result is less than a preset similarity threshold, indicating that the first voiceprint information is unregistered voiceprint information not stored in the database, the terminal device determines that the current first user is a suspected new user and directly stores the first voiceprint information in the database to register it as the first user's voiceprint information. Simultaneously, the first portrait information is also stored in the database to register it as the first user's portrait information. If the first comparison result is greater than or equal to the preset similarity threshold, indicating that the first voiceprint information is similar to the voiceprint information of a second user in the registered voiceprint information, this indicates that the current first user is a second user (an old user) whose voiceprint has been previously registered by the terminal device without any prior knowledge. The terminal device then updates the second user's voiceprint information in the database based on the first voiceprint information and simultaneously updates the second user's portrait information in the database based on the first portrait information.

[0121] In this way, the registered voiceprint information and registered portrait information are continuously updated and iterated through the terminal device based on the voiceprint extraction and matching process during the user's daily voice interaction in the smart cockpit, so that the voiceprint registration, matching and authentication using shorter or lower-quality user voice data can be more robust. That is, even if the user voice data obtained by the terminal device is low-quality voice data, it is not easy to affect the voiceprint extraction and matching process based on the voice data.

[0122] In some embodiments, the terminal device can update the registered voiceprint information using an adjustable forgetting factor parameter. In addition, when the terminal device merges the registered portrait information, unique slots such as user nicknames and places of origin can use the new data in the first portrait information to overwrite the old data in the registered portrait information. For slots with list-like properties such as hobbies, the new data can be added to the previous list to enrich the user portrait.

[0123] Please refer to Figure 9 , Figure 9 for Figure 8 Schematic diagram of the detailed steps of step S801.

[0124] like Figure 9 As shown, in some embodiments, in the above-mentioned step S801, the step of "updating the voiceprint information of the second user based on the first voiceprint information" may include the following steps S901 and S902.

[0125] Step S901: Calculate a first product of the first voiceprint information and a first forgetting factor parameter, and calculate a second product of the second user's voiceprint information and a second forgetting factor parameter; the sum of the first forgetting factor parameter and the second forgetting factor parameter is 1, and the first forgetting factor parameter is less than the second forgetting factor parameter.

[0126] It should be noted that the sum of the first forgetting factor parameter and the second forgetting factor parameter is 1, and the first forgetting factor parameter is smaller than the second forgetting factor parameter. For example, when the first forgetting factor parameter is 0.1, the second forgetting factor parameter may be 0.9. For another example, when the first forgetting factor parameter is 0.01, the second forgetting factor parameter may be 0.99.

[0127] When the terminal device updates the voiceprint information of the second user that is similar to the registered voiceprint information based on the currently extracted first voiceprint information, it can use a smaller first forgetting factor parameter to multiply the first voiceprint information to calculate the first product, and use a larger second forgetting factor parameter to multiply the second user's voiceprint information to calculate the second product.

[0128] Step S902: superimpose the first product and the second product to obtain the updated voiceprint information of the second user.

[0129] After calculating the first product and the second product, the terminal device further superimposes the first product and the second product, and the superimposed result can be used as the updated voiceprint information of the second user.

[0130] For example, Figure 10 The voiceprint storage process shown in the figure, the terminal device calculates the first voiceprint information currently extracted (The temporary identification ID is ) are respectively compared with the registered voiceprint information e1, e2, ..., e N The similarity between the first voiceprint information With the registered voiceprint information e1, e2, ..., e N Compare to retrieve the temporary identifier The ID n of the matched registered voiceprint information.

[0131] If the first voiceprint information With the registered voiceprint information e1, e2, ..., e N The similarity of any voiceprint information is less than the threshold, so the first voiceprint information is determined Temporary identification It does not match any ID1, 2, ..., N in the voiceprint database, indicating that the current first user Is a suspected new user n, so the first voiceprint information The first portrait information is newly stored in the database as the voiceprint information e of the suspected new user n n and portrait information.

[0132] In addition, if the first voiceprint information With the registered voiceprint information e1, e2, ..., e N The voiceprint information e of the second user n n The similarity is greater than or equal to the threshold, thereby determining the first voiceprint information Temporary identification Matches with ID n in the voiceprint database, that is, the first user If the user is an old user who has registered, that is, the second user n, then the voiceprint and portrait information are updated based on the second user n. At this time, the terminal device uses the following formula (2) to update the voiceprint information based on the first voiceprint information. Perform the voiceprint information e of the second user n n Updates.

[0133]

[0134] Among them, α and β are adjustable forgetting factor parameters. Specifically, β is the first forgetting factor parameter, and α is the second forgetting factor parameter.

[0135] In some embodiments, the terminal device may dynamically adjust the values ​​of the first forgetting factor parameter and the second forgetting factor parameter based on the number of times the first voiceprint information is called.

[0136] It should be noted that the number of times the temporary ID is called is used to represent the amount of user voice data. Temporary identification ID If the number of calls is 10, it means that the terminal device has obtained the first user's The number of user voice data s is 10. This is because the terminal device needs to store the first voiceprint information in the user voice data s. Accumulated to the first user In the process from the time when the vehicle door closing event is detected until the next time when the vehicle door closing event is detected, each time the first user The user voice data s, the first voiceprint information extracted from the user voice data s Allocate a temporary ID for the previous session Thus the temporary identification ID It will be recorded as having been called once. The number of calls can be used to represent the amount of user voice data s.

[0137] Please refer to Figure 11 , Figure 11 A flowchart of the steps for dynamically adjusting forgetting factor parameters involved in some embodiments of the voiceprint information processing method provided in the embodiments of the present application.

[0138] like Figure 11 As shown, in some embodiments, the terminal device can dynamically adjust the values ​​of the first forgetting factor parameter and the second forgetting factor parameter through steps S1101 and S1102 shown below.

[0139] Step S1101: Obtain the calling times of the temporary identifier of the first voiceprint information.

[0140] Before updating the voiceprint information of the registered second user based on the first voiceprint information, the terminal device may obtain the number of calls of the temporary identifier of the first voiceprint information. In this way, after subsequently adjusting the values ​​of the first forgetting factor parameter and the second forgetting factor parameter based on the number of calls, the adjusted first forgetting factor parameter and the second forgetting factor parameter are used to update the voiceprint information of the second user based on the first voiceprint information.

[0141] In some embodiments, the terminal device may also continuously obtain the number of calls to the temporary identifier of the first voiceprint information while updating the second user's voiceprint information based on the first voiceprint information. In this way, the terminal device can dynamically adjust the values ​​of the first forgetting factor parameter and the second forgetting factor parameter in real time based on the number of calls, and use the adjusted first forgetting factor parameter and the adjusted second forgetting factor parameter to update the second user's voiceprint information based on the first voiceprint information.

[0142] Step S1102: dynamically adjusting the numerical values ​​of the first forgetting factor parameter and the second forgetting factor parameter based on the number of calls; the numerical value of the first forgetting factor parameter is in inverse proportion to the number of calls.

[0143] After obtaining the call count of the temporary identifier of the first voiceprint information each time, the terminal device determines the respective numerical values ​​of the first forgetting factor parameter and the second forgetting factor parameter to be used based on the amount of user voice data represented by the call count, thereby dynamically adjusting the respective numerical values ​​of the first forgetting factor parameter and the second forgetting factor parameter based on the continuous change in the amount of user voice data represented by the call count. The numerical value of the first forgetting factor parameter decreases as the amount of user voice data represented by the call count increases, i.e., the numerical value of the first forgetting factor parameter is inversely proportional to the number of calls. Conversely, the numerical value of the second forgetting factor parameter increases accordingly as the amount of user voice data represented by the call count increases, i.e., the numerical value of the second forgetting factor parameter is directly proportional to the number of calls.

[0144] For example, in the first voiceprint information Temporary identification ID If the number of calls is small (for example, less than 10), it means that the first voiceprint information is currently Accumulated to the first user At the beginning of the accumulation, in order to speed up the voiceprint update, the first forgetting factor parameter β and the second forgetting factor parameter α can be set to the data size of α = 0.9, β = 0.1. Accumulated to a certain scale (such as the temporary identification ID If the number of calls has occurred more than 10 times), the data sizes of the first forgetting factor parameter β and the second forgetting factor parameter α can be set to: α = 0.99, β = 0.01 to obtain a more stable voiceprint update result.

[0145] In this embodiment, the terminal device combines the number of calls of the temporary identifier of the first voiceprint information to dynamically adjust the respective numerical values ​​of the first forgetting factor parameter and the second forgetting factor parameter required to update the registered voiceprint information based on the first voiceprint information. This can speed up the voiceprint update of the registered voiceprint information in the initial stage of accumulating the first voiceprint information to the first user (the amount of user voice data is small). When the first voiceprint information accumulates to a certain scale (the amount of user voice data is large), the registered voiceprint information can be updated more stably based on the first voiceprint information.

[0146] In some embodiments, the terminal device can merge two similar voiceprints in the registered voiceprint information and merge the corresponding portrait information at the same time. It should be noted that as data accumulates, the registered voiceprint information in the database may have two originally dissimilar voiceprints whose similarity exceeds a threshold. For example, after the terminal device registers the first voiceprint information as the voiceprint information of the first user, as the user voice data of the first user continues to accumulate, the quantity and quality of the newly extracted first voiceprint information are continuously updated, so that the first voiceprint information may not only have a similarity with the first user's voiceprint information exceeding a threshold, but also have a similarity with the voiceprint information of other users in the registered voiceprint information exceeding a threshold. At this time, it means that the voiceprint information of the first user successfully matches the voiceprint information of the other user, and the voiceprint and portrait information of the two users should be merged.

[0147] Please refer to Figure 12 , Figure 12 Schematic diagram of the steps of the voiceprint information processing method provided in some embodiments of the present application.

[0148] like Figure 12 As shown, after the above-mentioned step S203: registering the first voiceprint information, and registering the first portrait information, the voiceprint information processing method provided in the embodiment of the present application may also include the following steps S1201 and S1202.

[0149] Step S1201: Obtain a second comparison result between the voiceprint information of a third user and the voiceprint information of a fourth user in the registered voiceprint information; the third user includes the first user.

[0150] It should be noted that the third user including the first user means that the terminal device can compare the voiceprint information of a user whose registered voiceprint information has already completed the non-sensing registration with the voiceprint information of another user (the fourth user) who has also completed the registration. In addition, the terminal device can also compare the voiceprint information of the first user who has currently completed the non-sensing registration or is in the process of registering (the first voiceprint information in the case of registration) with the voiceprint information of the fourth user in the registered voiceprint information excluding the first user.

[0151] After the terminal device registers the first voiceprint information as the first user's voiceprint information each time without any sense, it further calculates the similarity between each two registered voiceprint information in the database, thereby comparing the two registered voiceprint information. In this way, the terminal device can obtain a second comparison result between the voiceprint information of the third user and the voiceprint information of the fourth user in the registered voiceprint information. The second comparison result can be the similarity result between the voiceprint information of the third user and the voiceprint information of the fourth user calculated by the terminal device.

[0152] Step S1202: When the second comparison result indicates that the voiceprint information of the third user is similar to the voiceprint information of the fourth user, the voiceprint information of the third user is merged with the voiceprint information of the fourth user, and the portrait information of the third user is merged with the portrait information of the fourth user.

[0153] The terminal device calculates the similarity between the third user's voiceprint information and the fourth user's voiceprint information, and uses the calculated similarity result as the second comparison result between the two. If the second comparison result is greater than or equal to a preset similarity threshold, the third user's voiceprint information successfully matches the fourth user's voiceprint information. In this case, the second comparison result indicates that the third user's voiceprint information is similar to the fourth user's voiceprint information, thereby indicating that the third user and the fourth user are actually the same user. In this way, the terminal device merges the third user's voiceprint information with the fourth user's voiceprint information, and simultaneously merges the third user's portrait information with the fourth user's portrait information.

[0154] In some embodiments, the terminal device can also merge the voiceprint information of the third user and the voiceprint information of the fourth user through an adjustable forgetting factor parameter. In addition, when the terminal device merges the portrait information of the third user and the portrait information of the fourth user, for unique slots such as user nicknames and places of origin, the new data can still be used to overwrite the old data. For example, the new data in the portrait information of the third user whose registration event is later can be used to overwrite the old data in the portrait information of the fourth user whose registration time is relatively earlier. For slots with a list nature such as interests and hobbies, the data in the two portrait information can be superimposed to enrich the user portrait.

[0155] Please refer to Figure 13 , Figure 13 for Figure 12 Schematic diagram of the detailed step flow of step S1202.

[0156] like Figure 13 As shown, in some embodiments, in the above-mentioned step S1202, the step of "merging the voiceprint information of the third user with the voiceprint information of the fourth user" may include steps S1301 to S1303 as shown below.

[0157] Step S1301: Calculate a first forgetting factor parameter and a second forgetting factor parameter based on the number of calls of the respective identifiers of the voiceprint information of the third user and the voiceprint information of the fourth user; if the number of calls of the identifier of the voiceprint information of the third user is less than the number of calls of the identifier of the voiceprint information of the fourth user, the first forgetting factor parameter is less than the second forgetting factor parameter.

[0158] When the terminal device combines the voiceprint information of the third user with the voiceprint information of the fourth user, it can obtain the number of calls for the identifier of the third user's voiceprint information and the number of calls for the identifier of the fourth user's voiceprint information. If the number of calls for the identifier of the third user's voiceprint information is less than the number of calls for the identifier of the fourth user's voiceprint information, this indicates that the amount of user voice data accumulated by the terminal device for extracting voiceprint information from the third user is less than the amount of user voice data accumulated for extracting voiceprint information from the fourth user.

[0159] Afterwards, the terminal device calculates a first forgetting factor parameter and a second forgetting factor parameter based on the number of calls of the third user's voiceprint information and the fourth user's voiceprint information identifiers. For example, assuming the number of calls of identifier n of the third user's voiceprint information is |n|, and the number of calls of identifier x of the fourth user's voiceprint information is |x|, then the first forgetting factor parameter β = |n| / (|n| + |x|), and the second forgetting factor parameter α = |x| / (|n| + |x|). Since |n| is less than |x|, the first forgetting factor parameter β is less than the second forgetting factor parameter α.

[0160] Step S1302: Calculate a third product of the voiceprint information of the third user and the first forgetting factor parameter, and calculate a fourth product of the voiceprint information of the fourth user and the second forgetting factor parameter.

[0161] After calculating the first forgetting factor parameter and the second forgetting factor parameter, the terminal device can use the smaller first forgetting factor parameter to multiply the voiceprint information of the third user to calculate the third product, and use the larger second forgetting factor parameter to multiply the voiceprint information of the fourth user to calculate the fourth product.

[0162] Step S1303: superimpose the third product and the fourth product to obtain combined voiceprint information.

[0163] After calculating the third product and the fourth product, the terminal device further superimposes the third product and the fourth product. The superimposed result can be used as the voiceprint information after the voiceprint information of the third user and the voiceprint information of the fourth user are merged.

[0164] For example, the terminal device may use the following formula (3) to calculate the voiceprint information e of the third user: n and the fourth user's voiceprint information e x Merge and get the merged voiceprint information e x .

[0165] e x =αe x +βen (3)

[0166] Among them, β is the first forgetting factor parameter, and α is the second forgetting factor parameter.

[0167] In this embodiment, after each time the terminal device unconsciously registers the first voiceprint information as the first user's voiceprint information, it compares the two registered voiceprint information to obtain a second comparison result between the voiceprint information of the third user and the voiceprint information of the fourth user in the registered voiceprint information. If the second comparison result is greater than or equal to a preset similarity threshold, thus successfully matching the voiceprint information of the third user and the fourth user, the second comparison result indicates that the voiceprint information of the third user and the fourth user are similar, thereby indicating that the third user and the fourth user are actually the same user. In this manner, the terminal device merges the voiceprint information of the third user with the voiceprint information of the fourth user, and simultaneously merges the portrait information of the third user with the portrait information of the fourth user.

[0168] In this way, during the user's daily voice interaction in the smart cockpit, the terminal device can continuously update and iterate the registered voiceprint information and registered portrait information based on voiceprint extraction and matching, so that the subsequent use of shorter or lower-quality user voice data for voiceprint registration, matching, authentication and other processing will also be more robust.

[0169] Next, a complete embodiment of the method for processing voiceprint information provided in the embodiments of the present application is proposed.

[0170] Please refer to Figure 14 , Figure 14 A schematic diagram of the overall framework involved in a complete embodiment of the voiceprint information processing method provided in an embodiment of the present application.

[0171] like Figure 14As shown, when the terminal device applies the voiceprint information processing method provided in the embodiment of the present application to process the voiceprint information, it first performs a door closing detection operation to detect the user's action of closing the vehicle door, wherein δ=1 when the vehicle door state changes from open to closed. Afterwards, the terminal device executes the voice cloud process. After monitoring the vehicle door closing event and obtaining the user voice data s in the vehicle interior environment, it is considered that the wake-up word for the voiceprint registration or update of the first user in the vehicle interior environment is detected without any perception, thereby uploading the user voice data s and the subsequent user voice data s obtained thereafter until the next vehicle door closing event is detected, together with the current vehicle door closing state δ, to the cloud device. Afterwards, the terminal device sets the door closing state δ=0 to indicate that the current door closing state has been processed. Afterwards, the cloud device performs a voiceprint / portrait extraction operation, that is, uses a certain voiceprint algorithm (such as the voiceprint model mentioned above) to extract a temporary voiceprint from the user voice data s. When δ=1 is detected, the temporary voiceprint Assign a temporary ID To use the temporary voiceprint As the first user First voiceprint information In addition, when the cloud device detects δ≠1, it is a temporary voiceprint. Assign the voiceprint ID n of the last session, so that the temporary voiceprint will be used in the whole process from the terminal device detecting the current vehicle door closing event to the next vehicle door closing event. All are accumulated to the same user n.

[0172] The cloud device extracts the first voiceprint information After that, the extracted voiceprint (first voiceprint information ) / portrait (first portrait information, which is extracted by the cloud device from the user's voice data s at the same time as the voiceprint) is stored in the database, realizing the user's non-sensical voiceprint / portrait registration or update. Specifically, the cloud device will temporarily identify the ID currently assigned to the first voiceprint information Compare with the ID assigned to the registered voiceprint information in the database. There are two situations: Case 1: Temporary identification ID It does not match any ID in the database, which means the current first user Is a suspected new user, so the cloud device will send the first voiceprint information and the first portrait information are newly stored in the database, thereby converting the first voiceprint information Register as the first user without any hesitation Voiceprint information, and register the first portrait information as the first user Portrait information; Case 2: Temporary identification ID Matches an ID n in the database Description of the first user If it is an old user n who has already registered, the voiceprint and portrait information will be updated based on the old user n.

[0173] Finally, the cloud device also performs the voiceprint / image merging operation. This is because as the data accumulates, the current voiceprint may n (the voiceprint information of the third user, including the first voiceprint information, the registered voiceprint information of the first user or the registered voiceprint information of other users), and the most similar voiceprint e in the database except itself x The similarity of (the fourth user x's voiceprint information) exceeds the specified threshold (cos(e n ,e x )≥λ), indicating that the current voiceprint e n The voiceprint of the fourth user, x, is successfully matched. At this point, the voiceprint and profile information for both IDs n and x should be merged. During the user profile merging process, unique slots like nicknames and places of origin need to have their old data overwritten with the new data; while slots that are lists, such as hobbies, can have the new data added to the existing list.

[0174] In this embodiment, by taking advantage of the fact that the user remains unchanged after the vehicle door is closed, the user's voice data is always accumulated to the voiceprint of the same person after the door is detected. In this way, as more data is accumulated, the voiceprint quality will also be higher, and the matching success rate will be greater. In addition, a temporary identification ID is generated each time the user closes the door. Accumulate voiceprint information at the same time When the temporary voiceprint is equal to the user's voiceprint n When matching, the temporary ID The user ID n is replaced with the actual user ID, and the voiceprint is updated and the user profile is merged. This allows for seamless voiceprint registration and matching for users, eliminating the need for a dedicated interface to guide the user through voiceprint registration. Instead, voiceprint registration and updates can be performed during daily voice interaction with the vehicle's computer. Furthermore, by matching the current voiceprint with registered voiceprints, the registered voiceprint information and registered profile information are continuously updated and iterated, making it robust against voiceprint registration, matching, and authentication using shorter or lower-quality user voice data.

[0175] See also Figure 15 , an embodiment of the present application also provides a voiceprint information processing device that can implement the above-mentioned voiceprint information processing method.

[0176] like Figure 15 As shown, the voiceprint information processing device provided in the embodiment of the present application includes a data acquisition module 1501, an extraction module 1502 and a registration module 1503.

[0177] The data acquisition module 1501 is configured to acquire user voice data in the vehicle interior environment when a vehicle door closing event is detected, and mark the user voice data as voice data of the first user associated with the vehicle door closing event;

[0178] An extraction module 1502 is configured to extract voiceprint information based on the user voice data to obtain first voiceprint information of the first user, and to extract portrait information based on the user voice data to obtain first portrait information of the first user;

[0179] The registration module 1503 is used to register the first voiceprint information and the first portrait information.

[0180] In some embodiments, the voiceprint information processing device provided in the embodiments of the present application further includes:

[0181] a comparison module, configured to compare the first voiceprint information with registered voiceprint information to obtain a first comparison result;

[0182] The registration module 1503 is also used to register the first voiceprint information as the voiceprint information of the first user and to register the first portrait information as the portrait information of the first user when the first comparison result indicates that the first voiceprint information is unregistered voiceprint information.

[0183] In some embodiments, the voiceprint information processing device provided in the embodiments of the present application further includes:

[0184] An updating module is used to update the voiceprint information of the second user based on the first voiceprint information when the first comparison result indicates that the first voiceprint information is similar to the voiceprint information of the second user in the registered voiceprint information, and to update the portrait information of the second user based on the first portrait information.

[0185] In some embodiments, the update module is further used to calculate a first product of the first voiceprint information and a first forgetting factor parameter, and to calculate a second product of the voiceprint information of the second user and a second forgetting factor parameter; the sum of the first forgetting factor parameter and the second forgetting factor parameter is 1, and the first forgetting factor parameter is less than the second forgetting factor parameter; and, the first product and the second product are superimposed to obtain the updated voiceprint information of the second user.

[0186] In some embodiments, the update module is further used to obtain the number of calls of the first voiceprint information; the number of calls is used to represent the amount of user voice data; and based on the number of calls, the numerical values ​​of the first forgetting factor parameter and the second forgetting factor parameter are dynamically adjusted; the numerical value of the first forgetting factor parameter is inversely proportional to the number of calls.

[0187] In some embodiments, the comparison module is further configured to obtain a second comparison result between the voiceprint information of a third user and the voiceprint information of a fourth user in the registered voiceprint information; the third user includes the first user;

[0188] The voiceprint information processing device provided in the embodiment of the present application further includes:

[0189] A merging module is used to merge the voiceprint information of the third user with the voiceprint information of the fourth user when the second comparison result indicates that the voiceprint information of the third user is similar to the voiceprint information of the fourth user, and to merge the portrait information of the third user with the portrait information of the fourth user.

[0190] In some embodiments, the merging module is further used to calculate a first forgetting factor parameter and a second forgetting factor parameter based on the number of calls of the respective identifiers of the voiceprint information of the third user and the voiceprint information of the fourth user; the number of calls of the identifier of the voiceprint information of the third user is less than the number of calls of the identifier of the voiceprint information of the fourth user, and the first forgetting factor parameter is less than the second forgetting factor parameter; calculate a third product of the voiceprint information of the third user and the first forgetting factor parameter, and calculate a fourth product of the voiceprint information of the fourth user and the second forgetting factor parameter; and superimpose the third product and the fourth product to obtain the merged voiceprint information.

[0191] In some embodiments, the voiceprint information processing device provided in the embodiments of the present application further includes:

[0192] The voice cloud module is used to upload the vehicle door closing event and the user voice data to the cloud device, so that the cloud device executes the step of extracting voiceprint information based on the user voice data to obtain the first voiceprint information of the first user and the subsequent steps.

[0193] It should be noted that the specific implementation of the voiceprint information processing device provided in the embodiment of the present application is basically the same as the specific implementation of the voiceprint information processing method described above, and will not be repeated here.

[0194] See also Figure 16An embodiment of the present application also provides a voiceprint information processing device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned visual language model evaluation method when executing the computer program.

[0195] In some embodiments, the voiceprint information processing device can be any smart terminal such as a vehicle-mounted hardware platform (such as a vehicle-mounted computer, etc.), a tablet computer, a smart phone, a wearable device, etc.

[0196] like Figure 16 As shown, the voiceprint information processing device provided in the embodiment of the present application may include:

[0197] The processor 1601 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0198] The memory 1602 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1602 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 1602, and the processor 1601 calls and executes the voiceprint information processing method of the embodiments of this application;

[0199] Input / output interface 1603, used to implement information input and output;

[0200] Communication interface 1604, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0201] Bus 1605 , which transmits information between various components of the device (e.g., processor 1601 , memory 1602 , input / output interface 1603 , and communication interface 1604 );

[0202] The processor 1601 , the memory 1602 , the input / output interface 1603 and the communication interface 1604 are connected to each other in communication within the device via the bus 1605 .

[0203] An embodiment of the present application also provides a vehicle, which is equipped with a voiceprint information processing device. The voiceprint information processing device includes a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned voiceprint information processing method when executing the computer program.

[0204] An embodiment of the present application also provides a cloud device that can be communicatively connected to a vehicle. The cloud device is configured with a voiceprint information processing device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned voiceprint information processing method when executing the computer program.

[0205] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned method for processing voiceprint information is implemented.

[0206] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0207] The embodiment of the present application further provides a computer program product, including a computer program. The steps implemented when the computer program is executed by a processor are basically the same as the specific embodiments of the above-mentioned voiceprint information processing method, and will not be repeated here.

[0208] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0209] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0210] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0211] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0212] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0213] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0214] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0215] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0216] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0217] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0218] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A method for processing voiceprint information, characterized in that: The method comprises: When a vehicle door closing event is detected, acquiring user voice data in the vehicle interior environment, and marking the user voice data as voice data of a first user associated with the vehicle door closing event; Extracting voiceprint information based on the user voice data to obtain first voiceprint information of the first user, and extracting portrait information based on the user voice data to obtain first portrait information of the first user; Register the first voiceprint information, and register the first portrait information.

2. The method according to claim 1, characterized in that Before registering the first voiceprint information and the first portrait information, the method further includes: Comparing the first voiceprint information with the registered voiceprint information to obtain a first comparison result; The registering of the first voiceprint information and the registering of the first portrait information include: When the first comparison result indicates that the first voiceprint information is unregistered voiceprint information, the first voiceprint information is registered as the voiceprint information of the first user, and the first portrait information is registered as the portrait information of the first user.

3. The method according to claim 2, characterized in that After comparing the first voiceprint information with the registered voiceprint information to obtain a first comparison result, the method further includes: When the first comparison result indicates that the first voiceprint information is similar to the voiceprint information of the second user in the registered voiceprint information, the voiceprint information of the second user is updated based on the first voiceprint information, and the portrait information of the second user is updated based on the first portrait information.

4. The method according to claim 3, characterized in that The updating of the voiceprint information of the second user based on the first voiceprint information includes: Calculating a first product of the first voiceprint information and a first forgetting factor parameter, and calculating a second product of the second user's voiceprint information and a second forgetting factor parameter; the sum of the first forgetting factor parameter and the second forgetting factor parameter is 1, and the first forgetting factor parameter is less than the second forgetting factor parameter; The first product is superimposed on the second product to obtain the updated voiceprint information of the second user.

5. The method according to claim 4, characterized in that The method further comprises: Obtain the number of calls for the first voiceprint information; the number of calls is used to represent the amount of user voice data; The numerical values ​​of the first forgetting factor parameter and the second forgetting factor parameter are dynamically adjusted based on the number of calls; the numerical value of the first forgetting factor parameter is in inverse proportion to the number of calls.

6. The method according to claim 1, characterized in that After registering the first voiceprint information and the first portrait information, the method further includes: Obtaining a second comparison result between voiceprint information of a third user and voiceprint information of a fourth user in the registered voiceprint information; the third user includes the first user; When the second comparison result indicates that the voiceprint information of the third user is similar to the voiceprint information of the fourth user, the voiceprint information of the third user is merged with the voiceprint information of the fourth user, and the portrait information of the third user is merged with the portrait information of the fourth user.

7. The method according to claim 6, characterized in that The merging the voiceprint information of the third user with the voiceprint information of the fourth user includes: Calculating a first forgetting factor parameter and a second forgetting factor parameter based on the number of calls of the respective identifiers of the voiceprint information of the third user and the voiceprint information of the fourth user; the number of calls of the identifier of the voiceprint information of the third user is less than the number of calls of the identifier of the voiceprint information of the fourth user, and the first forgetting factor parameter is less than the second forgetting factor parameter; Calculating a third product of the voiceprint information of the third user and the first forgetting factor parameter, and calculating a fourth product of the voiceprint information of the fourth user and the second forgetting factor parameter; The third product is superimposed on the fourth product to obtain combined voiceprint information.

8. The method according to any one of claims 1 to 7, characterized in that After acquiring user voice data in the vehicle interior environment and marking the user voice data as voice data of the first user associated with the vehicle door closing event, the method further includes: The vehicle door closing event and the user voice data are uploaded to a cloud device, so that the cloud device executes the step of extracting voiceprint information based on the user voice data to obtain the first voiceprint information of the first user and subsequent steps.

9. A device for processing voiceprint information, characterized in that: The device comprises: A data acquisition module is configured to acquire user voice data in the vehicle interior environment when a vehicle door closing event is detected, and mark the user voice data as voice data of a first user associated with the vehicle door closing event; an extraction module, configured to extract voiceprint information based on the user voice data to obtain first voiceprint information of the first user, and to extract portrait information based on the user voice data to obtain first portrait information of the first user; A registration module is used to register the first voiceprint information and the first portrait information.

10. A device for processing voiceprint information, characterized in that: The voiceprint information processing device includes a memory and a processor, the memory stores a computer program, and the processor implements the voiceprint information processing method according to any one of claims 1 to 8 when executing the computer program.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for processing voiceprint information according to any one of claims 1 to 8 is implemented.

12. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the method for processing voiceprint information according to any one of claims 1 to 8 is implemented.