Voiceprint result correction method, controller, vehicle, and computer-readable storage medium

By receiving the first speech and correcting the object type of the passenger in the vehicle, and combining natural language processing and voiceprint recognition algorithms, the problem of misidentification of voiceprint recognition algorithms in in-vehicle voice functions is solved, thereby improving vehicle safety and the accuracy of user interaction.

WO2026031820A9PCT designated stage Publication Date: 2026-03-12BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing voiceprint recognition algorithms are affected by environmental factors when determining a user's age, leading to child passengers being misidentified as adults or adult women being misidentified as children. This makes it impossible to accurately determine the passenger's identity and affects the security of in-vehicle voice functions.

Method used

After receiving the first voice message, the system identifies the sender's object type and issues a prompt. Within a preset time, it receives the second voice message for correction. By combining natural language processing and voiceprint recognition algorithms, the system confirms the sender's accurate object type and stores the voiceprint features and object type in the database, thereby improving recognition accuracy.

Benefits of technology

It improves the security of in-vehicle voice functions, ensures that the operation permissions are accurately matched with the identity of the sender, reduces security risks caused by misidentification, and enhances the accuracy of intelligent voice interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025103865_12032026_PF_FP_ABST
    Figure CN2025103865_12032026_PF_FP_ABST
Patent Text Reader

Abstract

A voiceprint result correction method, a controller, a vehicle, and a computer-readable storage medium. The method comprises: upon reception of a first voice, recognizing a first subject type of a speaker of the first voice and issuing prompt information, wherein a subject type represents a first-type subject or a second-type subject, and the prompt information comprises the subject type of the speaker; and within a preset duration of issuing the prompt information, if a second voice is received, re-determining a second subject type of the speaker on the basis of the second voice.
Need to check novelty before this filing date? Find Prior Art

Description

Voiceprint result correction method, controller, vehicle and computer readable storage medium

[0001] Priority information

[0002] This application claims priority to and the benefit of the filing date of the patent application with the China National Intellectual Property Office, with the patent application number of “202411062671.0” submitted on August 5, 2024, and incorporates it herein in its entirety by reference. TECHNICAL FIELD

[0003] The present application relates to the field of voiceprint recognition algorithms, and more particularly, to a voiceprint result correction method, a controller, a vehicle and a non-volatile computer readable storage medium. BACKGROUND

[0004] At present, more and more vehicles support in-vehicle language functions, so that passengers can use in-vehicle voice functions to voice control the vehicle. However, children passengers lack consideration for safety awareness, such as randomly opening the window or using other vehicle control functions through voice instructions, which not only causes interference to the driver, but also easily causes danger. The existing voiceprint recognition algorithm scheme for judging the age of the user is affected by environmental factors, such as the voiceprint of some adult women being easily misidentified as children, and children in the voice change period being misidentified as adults, resulting in the inability to accurately determine the identity of the passenger issuing the voice, thereby causing the vehicle to have a certain safety problem when using the in-vehicle language function. SUMMARY

[0005] The present application provides a voiceprint result correction method, a controller, a vehicle and a non-volatile computer readable storage medium.

[0006] The voiceprint result correction method of the present application includes, in the case of receiving a first voice, identifying the first object type of the issuer of the first voice and issuing a prompt information, the object type being a first class object or a second class object, and the prompt information including the object type of the issuer; within a preset time length of issuing the prompt information, if a second voice is received, the second object type of the issuer is re-determined based on the second voice.

[0007] The controller of the embodiment of the application comprises a processor, a memory and a computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program comprises instructions for executing a voiceprint result correction method. The voiceprint result correction method comprises, in a case where a first voice is received, identifying a first object type of an issuer of the first voice and issuing prompt information, the object type being a first type of object or a second type of object, the prompt information comprising the object type of the issuer; and within a preset time length after the prompt information is issued, if a second voice is received, re-determining a second object type of the issuer based on the second voice.

[0008] The vehicle of the embodiment of the application comprises a controller. The controller comprises a processor, a memory and a computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program comprises instructions for executing a voiceprint result correction method. The voiceprint result correction method comprises, in a case where a first voice is received, identifying a first object type of an issuer of the first voice and issuing prompt information, the object type being a first type of object or a second type of object, the prompt information comprising the object type of the issuer; and within a preset time length after the prompt information is issued, if a second voice is received, re-determining a second object type of the issuer based on the second voice.

[0009] The non-volatile computer readable storage medium of the embodiment of the application comprises a computer program, which, when executed by a processor, causes the processor to execute a voiceprint result correction method. The voiceprint result correction method comprises, in a case where a first voice is received, identifying a first object type of an issuer of the first voice and issuing prompt information, the object type being a first type of object or a second type of object, the prompt information comprising the object type of the issuer; and within a preset time length after the prompt information is issued, if a second voice is received, re-determining a second object type of the issuer based on the second voice.

[0010] The voiceprint result correction method, the controller, the vehicle and the nonvolatile computer readable storage medium of the embodiments of the present application can identify the first object type of the issuer of the first voice and issue prompt information in the case of receiving the first voice, so as to facilitate the user to determine whether the first object type is identified incorrectly according to the prompt information. Within the preset time length of issuing the prompt information, if the second voice is identified, it can be confirmed that the first object type is identified incorrectly, at this time, the object type of the issuer needs to be corrected, and the second object type of the issuer is re-determined based on the second voice, so as to accurately determine the object type of the issuer. In this way, it is beneficial to accurately determine the use permission of the issuer to the in-vehicle voice function according to the second object type of the issuer, and the vehicle only performs the operation corresponding to the first voice in the case that the use permission of the issuer includes the operation corresponding to the first voice, thereby facilitating to improve the safety of the vehicle when the user uses the in-vehicle voice recognition function.

[0011] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter in the description of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0012] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the description of the embodiments of the present application, taken in conjunction with the following drawings:

[0013] FIG. 1 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0014] FIG. 2 is a schematic diagram of a scenario of a voiceprint result correction method according to some embodiments of the present application;

[0015] FIG. 3 is a schematic diagram of a scenario of a voiceprint result correction method according to some embodiments of the present application;

[0016] FIG. 4 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0017] FIG. 5 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0018] FIG. 6 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0019] FIG. 7 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0020] FIG. 8 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0021] FIG. 9 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0022] Fig. 10 is a flowchart of a voiceprint result correction method according to some embodiments of the present application;

[0023] Fig. 11 is a block diagram of a voiceprint result correction device according to some embodiments of the present application;

[0024] Fig. 12 is a structural diagram of a controller according to some embodiments of the present application;

[0025] Fig. 13 is a structural diagram of a vehicle according to some embodiments of the present application;

[0026] Fig. 14 is a connection state diagram of a non-volatile computer readable storage medium and a processor according to some embodiments of the present application. DETAILED DESCRIPTION

[0027] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which the same or similar components have the same or similar reference numerals throughout. The embodiments described below are examples for explaining the present application, and should not be construed as limiting the present application.

[0028] At present, more and more vehicles support in-vehicle language function, so that passengers can use in-vehicle voice function to voice control the vehicle. However, children passengers lack of safety awareness, such as randomly opening the window or using other vehicle control functions through voice instructions, which not only causes disturbance to the driver, but also is easy to cause danger. The existing voiceprint recognition algorithm scheme for judging the age of the user is affected by environmental factors, such as the voiceprint of some adult women is easily misidentified as a child, and the voice of a child in the voice changing period is misidentified as an adult, resulting in the inability to accurately determine the identity of the passenger who issues the voice, thereby failing to guarantee the safety of the vehicle when using the in-vehicle language function.

[0029] To solve the above technical problems, the present application provides a voiceprint result correction method, which will be described in detail below:

[0030] Referring to Fig. 1, the present application provides a voiceprint result correction method, which comprises:

[0031] Step 011: In the case of receiving the first voice, identifying the first object type of the issuer of the first voice and issuing a prompt information, the object type is a first type of object or a second type of object, and the prompt information includes the object type of the issuer;

[0032] Specifically, the in-vehicle voice function is a convenient technology that allows a user, such as a driver or a passenger, to control various functions of the vehicle through voice commands during driving, thereby improving the safety and convenience of driving. The vehicle can determine the meaning corresponding to the voice uttered by the user according to a natural language processing technology (NLP), so that intelligent voice interaction can be achieved between the vehicle and the user. The first voice is a voice with a control instruction in the content, and the control instruction can be extracted from the first voice using a natural language processing technology, and a function corresponding to the control instruction is executed.

[0033] The object type is a first type of object or a second type of object, and different objects have different use permissions for the in-vehicle voice function.

[0034] For example, the first type of object is an adult, and the second type of object is a child. The use permission of the adult is larger, and the adult can basically control all operations. The use permission of the child is smaller, and the child can control fewer operations. Generally, the child can only control operations such as adjusting the volume that do not affect driving safety, or the child can be directly prohibited from using the in-vehicle voice function. For another example, the first type of object is a stranger, and the second type of object is a passenger or a driver who usually rides in the vehicle. The use permission of the passenger or the driver who usually rides in the vehicle is larger, and the passenger or the driver can basically control all operations. The use permission of the stranger or the passenger who less frequently rides in the vehicle is smaller, and the stranger or the passenger who less frequently rides in the vehicle can control fewer operations. Generally, the stranger or the passenger who less frequently rides in the vehicle can only control operations such as adjusting the volume that do not affect driving safety, or the stranger or the passenger who less frequently rides in the vehicle can be directly prohibited from using the in-vehicle voice function.

[0035] When the user needs to use the in-vehicle voice function, the user will utter the first voice to control the vehicle to perform corresponding operations according to the voice. When the first voice is received, the first object type of the utterer of the first voice can be identified. For example, the first type of object is an adult, and the second type of object is a child. The age of the utterer is determined by using a voiceprint recognition algorithm to determine whether the utterer is a first type of object or a second type of object. Alternatively, an image of the utterer can be obtained by a camera component of the vehicle, and whether the utterer is a first type of object or a second type of object can be determined according to the image. For another example, the first type of object is a stranger or a passenger who less frequently rides in the vehicle, and the second type of object is a passenger or a driver who usually rides in the vehicle. The vehicle can be provided with a preset voiceprint database, and the voiceprint database can store voiceprint features. When the first voice is received, the voiceprint features of the first voice can be extracted and compared with the voiceprint features in the voiceprint database. When there is a matching voiceprint feature in the voiceprint database, it is determined that the first object type is a first type of object.

[0036] In a case where the first object type is identified, a prompt information can be sent, and the prompt information includes the object type of the sender. For example, please refer to FIG. 2. In a case where the user is identified as a child, a child mode can be entered, and the screen of the vehicle can display an interface corresponding to the child mode, which is the prompt information. Alternatively, please refer to FIG. 3. The prompt information “adult detected, please perform XX function” can be displayed on the screen. In this way, the sender of the first voice and other passengers in the vehicle can obtain the identification result of the first voice according to the prompt information, so as to facilitate the sender of the first voice and other passengers in the vehicle to confirm whether the identification result is incorrect.

[0037] Step 012: Within the preset time length of sending the prompt information, if a second voice is received, the second object type of the sender is re-determined based on the second voice.

[0038] Specifically, the preset time length can be the maximum time length within which the sender of the first voice and other passengers in the vehicle judge that the identification result is incorrect after receiving the prompt information and send the second voice. The second voice is a voice including information of the object type of the sender of the first voice or information indicating that the identification result is incorrect. For example, the second voice can be “I am not a child”, or the second voice can be “identification error”.

[0039] Within the preset time length of sending the prompt information, if a second voice is received, it can be considered that the first object type of the sender at this time is incorrect, and the object type of the sender needs to be corrected, that is, it can be confirmed that the user initiates a request for correction of the voiceprint identification result at this time. Therefore, the second object type of the sender can be re-determined based on the second voice.

[0040] For example, the sender is an adult, and the first object type is a child. After seeing the prompt information, the sender can send the second voice “I am not a child”, and then the identity of the sender can be determined as an adult based on the second voice, that is, the second object type of the sender is determined as an adult. Alternatively, after seeing the prompt information, the sender can send the second voice “identification error”, and then the first object type can be determined as incorrect based on the second voice.

[0041] The second voice can also be received after the sending of the confirmation information, including a voice of the first voice sender's object type or a voice representing the information that the recognition result is incorrect. At this time, the preset time length is the maximum time length of the first voice sender and other passengers in the vehicle to send the second voice after receiving the prompt information and the confirmation information. The sender can send a reply voice "I am not a child" after seeing the prompt information, and at this time, the first object type error can be determined based on the reply voice. At this time, the confirmation information can be further sent to inquire the identity of the sender, for example, the confirmation information is "Please ask the passenger who sent the control instruction just now whether he is an adult?". At this time, the sender can clearly identify his own identity according to the confirmation information, and send a second voice, for example, the second voice is "I am an adult". Then the second object type of the sender can be determined according to the second voice.

[0042] Of course, the sender of the second voice can also not be the sender of the first voice. For example, it can be stipulated that only the second voice of some users meeting certain conditions is used to determine the second object type, for example, the driver is very important to the vehicle, so the driver can be set with the highest authority, and the object identity of the sender of the first voice is confirmed by the driver. At this time, it can be stipulated that only the second voice sent by the main driver is used to determine the second object type, so as to improve the accuracy of the second object type.

[0043] In this way, the user can determine whether the first object type is recognized incorrectly according to the prompt information, and send the second voice in the case that the first object type is recognized incorrectly. If the vehicle of the present application receives the second voice within the preset time length of sending the prompt information, it can be confirmed that the first object type is recognized incorrectly at this time and needs to be corrected. At this time, the vehicle can determine the second object type of the sender based on the second voice, so as to accurately obtain the object type of the sender, so as to facilitate subsequent determination of the use authority of the sender according to the accurate object type, so as to facilitate confirmation of whether to perform the operation corresponding to the first voice, and further to ensure the safety of the vehicle when the user uses the in-vehicle voice recognition function. For example, in the case of confirming that the sender is a child, it can be judged whether the operation corresponding to the first voice exceeds the use authority of the child, and the operation corresponding to the first voice is performed only in the case of not exceeding, and the operation corresponding to the first voice is not performed in the case of exceeding.

[0044] The voiceprint result correction method of the embodiment of the present application can identify the first object type of the first voice issuer when the first voice is received, and issue a prompt information so that the user can determine whether the first object type is identified incorrectly according to the prompt information. Within a preset time length of issuing the prompt information, if the second voice is identified, it can be confirmed that the first object type is identified incorrectly, at which time the object type of the issuer needs to be corrected, and the second object type of the issuer is re-determined based on the second voice to accurately determine the object type of the issuer. In this way, it is beneficial to accurately determine the use permission of the issuer for the in-vehicle voice function according to the second object type of the issuer. Only when the use permission of the issuer includes the operation corresponding to the first voice, the vehicle performs the operation corresponding to the first voice, thereby improving the safety of the vehicle when the user uses the in-vehicle voice recognition function.

[0045] Referring to FIG. 4, in some embodiments, the control method further comprises:

[0046] Step 013: store the voiceprint feature of the first voice and the second object type in the preset voiceprint database.

[0047] Specifically, the preset voiceprint database can store the voiceprint features of the voices received by the vehicle, and the object types corresponding to each voiceprint feature.

[0048] In the case of confirming the second object type of the issuer, the voiceprint feature of the first voice can be extracted, for example, the voiceprint feature of the first voice can be extracted based on the voiceprint recognition model trained based on the public data set. Alternatively, in the case of confirming the first object type corresponding to the first voice, the voiceprint feature of the first voice can also be extracted, and the first object type can be determined according to the voiceprint feature. At this time, the voiceprint feature of the first voice extracted when the first object type is determined can be directly obtained. Then, the voiceprint feature of the first voice and the second object type are stored in the voiceprint database at the same time, so that the second object type of the first voice is accurately stored in the voiceprint database.

[0049] In this way, in the case that the issuer of the first voice issues the first voice again, the vehicle can obtain the second object type corresponding to the issuer of the first voice in the voiceprint database. It can be understood that the accuracy of the second object type is high at this time, so that the vehicle can accurately determine the use permission of the issuer according to the second object type, and perform the corresponding operation, thereby enabling the vehicle to more accurately and intelligently implement the in-vehicle language function.

[0050] Referring to FIG. 5, in some embodiments, the control method further comprises:

[0051] Step 014: performing similarity matching on the voiceprint feature of the first voice and the voiceprint feature in the preset voiceprint database to obtain a first intermediate voiceprint feature with a similarity greater than a first preset threshold;

[0052] Step 015: deleting the voiceprint feature with the highest similarity in the first intermediate voiceprint feature and the object type corresponding to the voiceprint feature with the highest similarity, and storing the voiceprint feature of the first voice and the second object type into the voiceprint database.

[0053] Specifically, the first preset threshold is the minimum similarity when it can be confirmed that the issuer of the voiceprint feature of the first voice and the issuer of the voiceprint feature in the voiceprint database are probably the same person. It can be understood that in the case of similarity greater than the first preset threshold, the issuers corresponding to the two voiceprint features are considered to be the same person.

[0054] After determining the voiceprint feature of the first voice, the voiceprint feature of the first voice and all voiceprint features in the voiceprint database can be matched in similarity. For example, the cosine similarity of two voiceprint features can be calculated, and the greater the cosine similarity, the greater the similarity between the two voiceprint features. Alternatively, the Euclidean distance of two voiceprint features can also be calculated, and the smaller the distance, the greater the similarity between the two voiceprint features.

[0055] Then the first intermediate voiceprint feature with similarity greater than the first preset threshold is obtained, and it can be understood that the probability that the issuer of the first intermediate voiceprint feature and the issuer of the first voice are the same person is greater. Then the voiceprint feature with the highest similarity in the first intermediate voiceprint feature and the corresponding object type are deleted, and it can be understood that the voiceprint feature with the highest similarity and the issuer of the first voice are the same person, so at this time the voiceprint feature with the highest similarity and the object type corresponding to the voiceprint feature with the highest similarity are deleted from the voiceprint database, and the voiceprint feature of the first voice and the second object type are stored into the voiceprint database. In this way, on the one hand, it can be ensured that the voiceprint database stores the more accurate object type of the issuer of the first voice, and on the other hand, it can be ensured that the voiceprint database does not store more voiceprint features and object types of the same issuer, so that the voiceprint database can store more voiceprint features and object types of users.

[0056] Please refer to FIG. 6, in some embodiments, the control method further comprises:

[0057] Step 016: matching the voiceprint feature of the first voice and the voiceprint features in the voiceprint database to obtain a target voiceprint feature in each voiceprint feature of the voiceprint database;

[0058] Step 017: determining that the object type corresponding to the target voiceprint feature is the object type of the issuer of the first voice.

[0059] Specifically, after the voiceprint feature of the first voice is acquired, the voiceprint feature of the first voice and the voiceprint features in the voiceprint database can be matched to determine a target voiceprint feature in the voiceprint database that has a high similarity with the voiceprint feature of the first voice. For example, a similarity threshold can be set for the similarity, and when a voiceprint feature in the voiceprint database has a similarity with the voiceprint feature of the first voice greater than the similarity threshold, it can be determined that the speaker of the voiceprint feature is more likely to be the same person as the speaker of the first voice. Therefore, the target voiceprint feature can be selected from the voiceprint features with a similarity greater than the similarity threshold, for example, the voiceprint feature with the greatest similarity or the voiceprint feature stored at the latest time can be selected as the target voiceprint feature. It can be understood that the speaker of the target voiceprint feature is more likely to be the same person as the speaker of the first voice.

[0060] Therefore, after the target voiceprint feature is determined, the object type corresponding to the target voiceprint feature can be determined as the object type of the speaker of the first voice, so that the object type of the speaker of the first voice can be accurately determined, and it is convenient for subsequent judgment of whether the operation corresponding to the first voice needs to be performed according to the object type of the speaker of the first voice.

[0061] After the second object type of the speaker of the first voice is confirmed, the second object type is stored in the voiceprint database, so that the voiceprint database stores the more accurate object type of the user corresponding to each voiceprint feature. In this way, when a user corresponding to the voiceprint feature stored in the voiceprint database again speaks the first voice, the object type of the user can be quickly and accurately obtained in the voiceprint database, so that the in-vehicle language function is more accurately and intelligently implemented, and the corresponding voice command is executed.

[0062] When there is no voiceprint feature in the voiceprint database that matches the voiceprint feature of the first voice, it can be confirmed that the voiceprint database does not have the identity type of the speaker of the first voice, and at this time, the existing voiceprint recognition algorithm can be used to determine the first object type of the speaker of the first voice. In the case of a first object type identification error, the user will speak a second voice, and at this time, the second object type of the speaker of the first voice can be re-determined based on the second voice. In this way, the object type of the speaker of the first voice can also be accurately obtained, so that subsequent judgment of whether the operation corresponding to the first voice needs to be performed according to the second object type is facilitated, and the in-vehicle language function can be more accurately and intelligently implemented.

[0063] In this way, the object type of the speaker of the first voice can be accurately determined by combining the NLP algorithm, the voiceprint recognition algorithm and the similarity matching algorithm, so that the accuracy of voiceprint recognition can be enhanced, intelligent voice interaction can be increased, the advantages of intelligent interaction can be embodied, and the intelligent voice interaction experience of the user can be improved.

[0064] Referring to FIG. 7, in some embodiments, the step 016 of matching the voiceprint feature of the first voice with the voiceprint features in the voiceprint database to obtain a target voiceprint feature in each voiceprint feature of the voiceprint database comprises:

[0065] The step 0161 of similarity matching the voiceprint feature of the first voice with the voiceprint features in the voiceprint database to obtain a second intermediate voiceprint feature with a similarity greater than a second preset threshold value;

[0066] The step 0162 of selecting the second intermediate voiceprint feature with the latest storage time as the target voiceprint feature.

[0067] Specifically, one or more voiceprint features of a user can be stored in the voiceprint database. The second preset threshold value is the minimum similarity when it can be confirmed that the issuer of the voiceprint feature of the first voice and the issuer of the voiceprint feature in the voiceprint database are probably the same person. It can be understood that in the case of a similarity greater than the second preset threshold value, the issuers of the two voiceprint features are considered to be the same person.

[0068] The voiceprint feature of the first voice and the voiceprint features in the voiceprint database can be similarity matched to obtain a second intermediate voiceprint feature with a similarity greater than a second preset threshold value. It can be understood that the issuer of the second intermediate voiceprint feature and the issuer of the first voice are probably the same person.

[0069] Then the second intermediate voiceprint feature with the latest storage time can be selected as the target voiceprint feature, that is, the voiceprint feature stored in the voiceprint database most recently is selected as the target voiceprint feature. The voiceprint database stores the voiceprint feature of the first voice after being corrected by the user and the second object type. Therefore, the user can update the second object type of the issuer of the first voice during the correction according to the needs, so that the user can control the object type of the issuer of the first voice in the voiceprint database. In this way, in subsequent use, the vehicle determines the issuer of the first voice according to the object type updated by the user in the voiceprint database most recently, so that the user can adjust the object type of the issuer of the first voice according to the needs to adjust the use permission of the issuer of the first voice, thereby improving the use interest of the user.

[0070] Referring to FIG. 8, in some embodiments, the step 016 of matching the voiceprint feature of the first voice with the voiceprint features in the voiceprint database to obtain a target voiceprint feature in each voiceprint feature of the voiceprint database comprises:

[0071] The step 0163 of similarity matching the voiceprint feature of the first voice with the voiceprint features in the voiceprint database to obtain a third intermediate voiceprint feature with a similarity greater than a third preset threshold value;

[0072] Step 0164: determining the voiceprint feature with the largest similarity in the third intermediate voiceprint features as the target voiceprint feature.

[0073] Specifically, the third preset threshold is the minimum similarity when it can be confirmed that the issuer of the voiceprint feature of the first voice and the issuer of the voiceprint feature in the voiceprint database are probably the same person. It can be understood that in the case where the similarity is greater than the third preset threshold, it can be considered that the issuers corresponding to the two voiceprint features are the same person.

[0074] The voiceprint feature of the first voice and the voiceprint feature in the voiceprint database can be matched for similarity to obtain a third intermediate voiceprint feature with a similarity greater than the third preset threshold. It can be understood that the issuer of the third intermediate voiceprint feature and the issuer of the first voice are probably the same person. At this time, the voiceprint feature with the largest similarity in the third voiceprint feature can be selected as the target voiceprint feature. It can be understood that the voiceprint feature with the largest similarity in the third voiceprint feature has the greatest probability of being the issuer of the first voice, and therefore the object type corresponding to the target voiceprint feature can more accurately represent the object type of the issuer of the first voice.

[0075] In this way, the target voiceprint feature can be determined according to the voiceprint feature with the largest similarity to the voiceprint feature of the first voice among the voiceprint features with a similarity greater than the third preset threshold. It can be understood that the issuer corresponding to the target voiceprint feature at this time and the issuer of the first voice are probably the same person, so that the object type corresponding to the target voiceprint feature can more accurately represent the object type of the issuer of the first voice, thereby facilitating accurate determination of the use permission of the issuer of the first voice.

[0076] Referring to FIG. 9, in some embodiments, the control method further comprises:

[0077] Step 018: If no second voice is received within the preset time length of issuing the prompt information, storing the voiceprint feature of the first voice and the first object type in the preset voiceprint database.

[0078] Specifically, if no second voice is received within the preset time length of issuing the prompt information, it can be considered that the first object type identification is successful, and at this time the accuracy of the first object type is higher. At this time, the voiceprint feature of the first voice and the first object type can be stored in the voiceprint database, so that in the case where the issuer of the first voice issues the first voice again, the object type of the issuer of the first voice can be accurately obtained in the voiceprint database without the need to identify the object type of the issuer of the first voice again using the voiceprint identification algorithm, thereby improving the accuracy and identification speed of voiceprint identification.

[0079] Similar to the case of storing the voiceprint feature of the first voice and the first object type, in the case of storing the voiceprint feature of the first voice and the first object type, the voiceprint feature of the first voice and the voiceprint features in the voiceprint database can also be matched in similarity, and the voiceprint feature with the largest similarity among the voiceprint features with similarity greater than the first preset threshold can be deleted. It can be understood that the voiceprint feature with the highest similarity has the greatest probability of being the same person as the issuer of the first voice. In this way, on the one hand, it can be ensured that the voiceprint database stores a more accurate object type of the issuer of the first voice, and on the other hand, it can be ensured that the voiceprint database does not store more voiceprint features and object types of the same issuer, thereby facilitating the voiceprint database to store more voiceprint features and object types of users.

[0080] Referring to FIG. 10, in some embodiments, step 012: if the second voice is received within the preset time length of issuing the prompt information, the second object type of the issuer is re-determined based on the second voice, including:

[0081] Step 0121: if the second voice of the preset personnel is received within the preset time length of issuing the prompt information, the second object type of the issuer is re-determined based on the second voice, the preset personnel including personnel in a preset position and / or registered personnel.

[0082] Specifically, it can be specified that the second object type of the issuer is determined only according to the second voice of the preset personnel, so as to improve the accuracy of the second object type. For example, the preset personnel includes personnel in a preset position, such as personnel in the main driver position, i.e., the driver, or personnel in the co-driver position. At this time, the position of the issuer of the voice can be determined according to the positioning technology, so that the second voice of the personnel in the preset position can be accurately obtained. Alternatively, the preset personnel can also include registered personnel, i.e., personnel who have successfully registered to be able to drive the vehicle. Each registered personnel has a corresponding identification code on the system of the vehicle, and the voiceprint feature corresponding to the registered personnel is stored in the voiceprint database. Whether the issuer of the voice is a registered personnel can be determined according to the voiceprint feature of the language, so that the second voice of the registered personnel can be accurately obtained within the preset time length of issuing the prompt information.

[0083] If the second voice of the preset personnel is received within the preset time length of issuing the prompt information, the second object type of the issuer of the first voice can be determined according to the second voice. Further, before the second voice is received, a confirmation information can also be issued, for example, the confirmation information is “Please ask the passenger who issued the instruction just now whether it is a child or an adult?”, or the confirmation information is “Please ask the passenger in the back row whether it is a child or an adult?”. The preset personnel can issue the second voice based on the confirmation information, so as to facilitate the extraction of the object type information in the second voice, thereby facilitating the improvement of the accuracy of the second object type.

[0084] Thus, in the case of a first object type identification error, the object type of the first voice issuer can be determined by a preset person, and the second object type of the first voice issuer can be determined according to a second voice issued by the preset person, so as to ensure the accuracy of the second object type.

[0085] Referring to FIG. 11, to better implement the voiceprint result correction method of the embodiments of the present application, the embodiments of the present application further provide a voiceprint result correction device 10. The voiceprint result correction device 10 can include an identification module 11 and a correction module 12. The identification module 11 is configured to, in the case of receiving a first voice, identify a first object type of an issuer of the first voice and issue prompt information, the object type being a first type of object or a second type of object, and the prompt information including the object type of the issuer; and the correction module 12 is configured to, within a preset time length of issuing the prompt information, if a second voice is received, re-determine a second object type of the issuer based on the second voice.

[0086] The voiceprint result correction device 10 further includes a first storage module 13. The first storage module 13 is configured to store the voiceprint feature of the first voice and the second object type into a voiceprint database.

[0087] The voiceprint result correction device 10 further includes a second storage module 14. The second storage module 14 is configured to perform similarity matching on the voiceprint feature of the first voice and the voiceprint features in the preset voiceprint database to obtain first intermediate voiceprint features with a similarity greater than a first preset threshold; delete a voiceprint feature with the highest similarity in the first intermediate voiceprint features and an object type corresponding to the voiceprint feature with the highest similarity, and store the voiceprint feature of the first voice and the second object type into the voiceprint database.

[0088] The voiceprint result correction device 10 further includes a matching module 15. The matching module 15 is configured to perform matching on the voiceprint feature of the first voice and the voiceprint features in the preset voiceprint database to obtain a target voiceprint feature in each voiceprint feature of the voiceprint database; and determine that an object type corresponding to the target voiceprint feature is the object type of the issuer of the first voice.

[0089] The matching module 15 is specifically configured to perform similarity matching on the voiceprint feature of the first voice and the voiceprint features in the voiceprint database to obtain second intermediate voiceprint features with a similarity greater than a second preset threshold; and select a second intermediate voiceprint feature with the latest storage time as the target voiceprint feature.

[0090] The voiceprint result correction device 10 further includes a third storage module 16. The third storage module 16 is configured to, within the preset time length of issuing the prompt information, if no second voice is received, store the voiceprint feature of the first voice and the first object type into the preset voiceprint database.

[0091] The correction module 12 is configured to, if a second voice of a preset person is received within a preset time length of sending the prompt information, determine a second object type of the sender based on the second voice, the preset person including a person at a preset location and / or a registered person.

[0092] The voiceprint result correction apparatus 10 is described above from the perspective of functional modules, which can be implemented in the form of hardware, implemented by instructions in the form of software, or implemented by a combination of hardware and software modules. Specifically, each step of the method embodiment in the embodiments of the present application can be completed by integrated logic circuits of hardware in a processor and / or instructions in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware coding processors for execution, or executed by a combination of hardware and software modules in the coding processor. Alternatively, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps in the above method embodiment.

[0093] Referring to FIG. 12, the controller 100 of the embodiment of the present application includes a processor 20, a memory 30, and a computer program, wherein the computer program is stored in the memory 30 and executed by the processor 20, and the computer program includes instructions for executing the voiceprint result correction method of any of the above-mentioned embodiments.

[0094] Referring to FIG. 13, the vehicle 200 of the embodiment of the present application includes the controller 100 of any of the above-mentioned embodiments.

[0095] Referring to FIG. 14, the embodiment of the present application further provides a computer readable storage medium 300, which stores a computer program 310, and the computer program 310 is executed by a processor 320 to implement the steps of the voiceprint result correction method of any of the above-mentioned embodiments. For brevity, details are not described herein.

[0096] In the description of the present specification, the description of the terms "certain embodiments", "in an example", "exemplarily", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above-mentioned terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0097] Any procedural or methodological descriptions in flow charts or otherwise described herein can be understood to represent modules, segments or portions of code that include executable instructions for implementing the specified logical function or process(es), and the scope of preferred embodiments of the present application includes additional implementations that can not perform the functions in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, as will be understood by those skilled in the art to which embodiments of the present application pertain.

[0098] Although the embodiments of the present application have been shown and described above, it is to be understood that the above-described embodiments are exemplary only, and are not to be construed as limiting the present application, and that variations, modifications, substitutions and changes can be made to the above-described embodiments by one skilled in the art without departing from the scope of the present application.

Claims

1. A voiceprint result correction method, wherein, The method comprises: In the case of receiving a first voice, identifying a first object type of an issuer of the first voice and issuing prompt information, the object type being a first type of object or a second type of object, the prompt information including the object type of the issuer; Within a preset time length after issuing the prompt information, if a second voice is received, re-determining a second object type of the issuer based on the second voice.

2. The voiceprint result correction method of claim 1, wherein, The method further comprises: Storing a voiceprint feature of the first voice and the second object type into a preset voiceprint database.

3. The voiceprint result correction method of claim 1, wherein, The method further comprises: Performing similarity matching on the voiceprint feature of the first voice and voiceprint features in the preset voiceprint database to obtain a first intermediate voiceprint feature with a similarity greater than a first preset threshold; Deleting a voiceprint feature with the highest similarity in the first intermediate voiceprint feature and an object type corresponding to the voiceprint feature with the highest similarity, and storing the voiceprint feature of the first voice and the second object type into the voiceprint database.

4. The voiceprint result correction method according to any one of claims 1-3, wherein, The method further comprises: Performing matching on the voiceprint feature of the first voice and voiceprint features in the preset voiceprint database to obtain a target voiceprint feature in each voiceprint feature of the voiceprint database; Determining an object type corresponding to the target voiceprint feature as an object type of an issuer of the first voice.

5. The voiceprint result correction method of claim 4, wherein, The matching on the voiceprint feature of the first voice and voiceprint features in the preset voiceprint database to obtain a target voiceprint feature in each voiceprint feature of the voiceprint database comprises: Performing similarity matching on the voiceprint feature of the first voice and the voiceprint features in the voiceprint database to obtain a second intermediate voiceprint feature with a similarity greater than a second preset threshold; Selecting a second intermediate voiceprint feature with the latest storage time as the target voiceprint feature.

6. The voiceprint result correction method according to any one of claims 1-3, wherein, The method further comprises: Within the preset time length after issuing the prompt information, if no second voice is received, storing a voiceprint feature of the first voice and the first object type into a preset voiceprint database.

7. The voiceprint result correction method of claim 1, wherein, The re-determining of the second object type of the issuer based on the second voice within the preset time length after issuing the prompt information comprises: Within the preset time length after issuing the prompt information, if a second voice issued by a preset person is received, re-determining a second object type of the issuer based on the second voice, the preset person including a person at a preset location and / or a registered person.

8. A controller wherein, Comprise: a processor, a memory; and a computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program comprises instructions for executing the voiceprint result correction method of any one of claims 1 to 7.

9. A vehicle, wherein, Comprise: the controller of claim 8.

10. A non-transitory computer readable storage medium containing a computer program, wherein, The computer program is executed by the processor, so that the processor executes the voiceprint result correction method of any one of claims 1 to 7.