A speech system, method and device based on a cognitive large model
By analyzing voice features and evaluating confidence, combined with vehicle history records and control instructions, the problems of insufficient voice recognition accuracy and user intent understanding in existing voice systems are solved, achieving higher recognition accuracy and improved user experience.
Patent Information
- Application Number
- CN202410590550.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-05-13
AI Technical Summary
Existing speech systems based on cognitive large models have problems with speech recognition accuracy and insufficient understanding of user intent in the field of in-vehicle intelligence.
By analyzing the person's voice characteristics such as volume, speaking speed and tone, the confidence of the voice command is evaluated, and combined with the vehicle history record and control command keyword library, the specified control command is recognized and executed, while providing feedback prompts to optimize the system's cognitive status.
It improves the accuracy of voice recognition and understanding of user intent, enhances the correctness and execution accuracy of the system's response to user commands, improves driving convenience and safety, and improves user experience through feedback prompts.
Smart Images

Figure CN118280365B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech system, method and device based on a cognitive large model. Background Art
[0002] Voice systems based on cognitive big models are playing a revolutionary role in the field of in-vehicle intelligence. Their technical background integrates multiple key technologies, including speech recognition, natural language processing, intelligent control, and network connectivity, providing vehicles with more intelligent and convenient interactive features. The system efficiently and accurately recognizes voice commands from drivers and passengers, converting them into actionable instructions, such as those for querying vehicle information and controlling the vehicle's comfort zone. This intuitive interaction significantly enhances the convenience and safety of the driving experience.
[0003] For example, the invention patent with publication number CN101281745A discloses an in-vehicle voice interaction system, including a voice acquisition module, a voice recognition core module and a voice feedback module. The voice recognition core module includes an acoustic model and pronunciation dictionary module, a context-independent grammar module, and a path search module. The acoustic model and pronunciation dictionary module are used to establish a set of mapping tables corresponding to changing characteristics such as accents based on a statistical algorithm; the context-independent grammar module is used to construct the grammar and rule structure of the natural continuous speech to be recognized; and the path search module is used to approximate and simplify the observation probability calculation part with the largest computational complexity.
[0004] For example, the invention patent with announcement number: CN103730119B provides a vehicle-mounted human-computer voice interaction system, including: a voice acquisition module; a voice purification module connected to the output end of the voice acquisition module; a voice processing method selection module connected to the output end of the voice purification module; a network module connected between the output end of the voice processing method selection module and the input end of the cloud server; a system terminal processor connected to the output end of the voice processing method selection module; a cloud server; and a vehicle-mounted control unit module.
[0005] Based on the above solutions, it can be seen that there are still some shortcomings in the current analysis of speech systems based on cognitive large models. Summary of the Invention
[0006] In response to the deficiencies in the prior art, the present invention provides a speech system, method, and device based on a cognitive large model, which can effectively solve the problems involved in the above-mentioned background technology.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a speech system based on a cognitive large model, including a speech recognition and analysis module, which is used to identify the voice input instructions of personnel for analysis and processing, and evaluate the confidence of the personnel's voice instructions.
[0008] The voice function control module is used to obtain trusted personnel, analyze the correlation between the trusted personnel's voice commands and various control commands, locate the vehicle control components corresponding to the specified control commands to enable function execution.
[0009] The feedback prompt module is used to comprehensively analyze the cognitive status of the speech system and provide feedback prompts.
[0010] The speech recognition database is used to store the vehicle's historical voice input record information and a keyword library of various control commands.
[0011] Furthermore, the voice input instructions of the identified person are analyzed and processed. The specific analysis process is: the voice input instructions of the identified person are extracted from the person's voice characteristic information, including the person's volume at each time point and the person's average speaking speed, and the vehicle's historical voice input record information is obtained from the voice recognition database, wherein the vehicle's historical voice input record information includes the historical voice reference volume, historical voice reference speed, and historical voice reference tone waveform of the vehicle's personnel.
[0012] The acoustic matching of the preliminary calculation personnel meets the first eigenvalue, which is calculated as follows:
[0013] In the formula, α indicates that the acoustic matching of the person meets the first eigenvalue, H j H0 represents the volume of the person at the j-th time point, H0 represents the historical voice reference volume of the vehicle person, ν3 represents the average speaking speed of the person, Δν represents the historical voice reference speed, a1 and a2 represent the sound matching compensation factors corresponding to the set voice volume and speaking speed respectively, j represents the number of each time point, j = 1, 2, 3, ..., m, and m represents the total number of time points.
[0014] According to the personnel's voice input instructions, the personnel's tone waveform diagram is constructed, and the personnel's tone waveform length is extracted. At the same time, based on the historical voice reference tone waveform of the vehicle's personnel, the personnel's tone waveform diagram is overlapped and compared with the historical voice reference tone waveform of the vehicle's personnel to obtain the maximum length of a single deviation of the personnel's tone waveform and the tone waveform overlap length.
[0015] The acoustic matching of the operator is calculated according to the second eigenvalue, which is calculated as follows:
[0016] Where, the acoustic matching of β personnel meets the second eigenvalue, L1 represents the maximum length of a single deviation of the personnel's pitch waveform, L0 represents the length of the personnel's pitch waveform, and L 2→cIt represents the pitch waveform overlap length between the pitch waveform of the person and the historical voice reference pitch waveform of the vehicle person, b1 represents the interference factor corresponding to the set maximum deviation length of the pitch waveform unit, b2 represents the correction factor corresponding to the set pitch waveform overlap length, and e represents a natural constant.
[0017] The confidence level of the operator's voice command is calculated using the following formula:
[0018]
[0019] Where χ represents the confidence level of the person's voice command, ε1 and ε2 represent the weight factors corresponding to the set acoustic matching meeting the first eigenvalue and the acoustic matching meeting the second eigenvalue, respectively.
[0020] Furthermore, the specific process of obtaining the trusted personnel is: extracting the person's voice command confidence, and comparing the person's voice command confidence with the voice command confidence threshold according to the set voice command confidence threshold. If the person's voice command confidence is higher than the voice command confidence threshold, the person is determined to be a trusted person.
[0021] Furthermore, the analysis process of evaluating the correlation between the voice instructions of the trusted person and various control instructions is as follows: extracting the voice instructions of the trusted person and pre-processing and converting them into text form, and processing the recognized text through natural language processing to obtain the control instruction keywords in the voice instructions of the trusted person.
[0022] According to the keyword library of various control instructions stored in the voice database, the number of overlaps between the voice instructions of the trusted personnel and the keywords of various control instructions is compared to evaluate the correlation between the voice instructions of the trusted personnel and various control instructions. The calculation formula is:
[0023] Where, δ i Indicates the correlation between the voice commands of the trusted personnel and various control commands, S i represents the number of overlaps between the trusted personnel's voice command and the keywords of the i-th type of control command, S0 represents the set number of overlaps in the control command, c1 represents the correction factor for the set number of overlaps, i represents the number of each type of control command, i = 1, 2, 3, ..., n, and n represents the total number of control command types.
[0024] Furthermore, the vehicle control components corresponding to the designated control instructions are located and enabled to execute functions. The specific process is: the correlation between the voice instructions of the trusted personnel and various control instructions is sorted in descending order, the control instruction ranked first is extracted as the designated control instruction, and the vehicle control components corresponding to the designated control instruction are enabled to execute functions.
[0025] Furthermore, the comprehensive analysis of the cognitive status of the speech system to provide feedback prompts is specifically carried out as follows: calculating the recognition function evaluation value of the speech system, and the calculation formula is:
[0026] Where ψ is the recognition function evaluation value of the speech system, K is the first cognitive function evaluation value of the speech system, Υ is the second cognitive function evaluation value of the speech system, τ1 and τ2 represent the weight factors corresponding to the set first cognitive function evaluation value and second cognitive function evaluation value, respectively.
[0027] Based on the recognition function evaluation value of the voice system, it is compared with the set recognition function evaluation threshold. If the recognition function evaluation value of the voice system is lower than the recognition function evaluation threshold, feedback prompts are given.
[0028] Furthermore, the first cognitive function evaluation value of the speech system is specifically analyzed in the following process: setting a monitoring period, counting the recognition accuracy of the speech system and the interval time for receiving feedback responses of commands during the monitoring period, and extracting the influencing factors corresponding to the recognition accuracy and the interval time for receiving feedback responses of units stored in the speech database.
[0029] Calculate the first cognitive function evaluation value of the speech system, and the calculation formula is:
[0030] Where K represents the first cognitive function evaluation value of the speech system, Z represents the recognition accuracy of the speech system, ΔZ represents the recognition accuracy, Ts represents the interval time for receiving feedback response of speech commands, η1 represents the correction coefficient of recognition accuracy, and η2 represents the impact factor corresponding to the interval time for receiving feedback response per unit.
[0031] Furthermore, the second cognitive function evaluation value of the speech system is specifically analyzed by counting the number of times the speech system continues in multiple rounds of dialogue under a specified control instruction and the number of times it corrects errors in the user's speech input during the monitoring period, and calculating the second cognitive function evaluation value of the speech system. The calculation formula is:
[0032] Where, Υ represents the second cognitive function evaluation value of the speech system, λ and They represent the number of coherences in the multi-round dialogue under the specified control command and the number of errors in correcting the user's voice input, Δλ and They represent the defined number of consecutive times and the defined number of errors respectively, and ω1 and ω2 represent the correction factors corresponding to the set number of consecutive times and the number of errors respectively.
[0033] The second aspect of the present invention also provides a speech method based on a cognitive large model, comprising: identifying a person's speech input instructions for analysis and processing, and evaluating the confidence of the person's speech instructions.
[0034] Acquire trusted personnel, analyze the correlation between the trusted personnel's voice commands and various control commands, locate the vehicle control components corresponding to the specified control commands and enable function execution.
[0035] Comprehensively analyze the cognitive status of the speech system and provide feedback prompts.
[0036] The third aspect of the present invention also provides a speech device based on a cognitive big model, comprising: a processor, and a memory and a network interface connected to the processor; the network interface is connected to a non-volatile memory in a server; the processor retrieves a computer program from the non-volatile memory through the network interface during operation, and runs the computer program through the memory to execute the above-mentioned speech method based on a cognitive big model.
[0037] The present invention has the following beneficial effects:
[0038] (1) The present invention analyzes the volume and duration of a person's speech, the person's speaking speed and pitch, and the reasonable analysis and application of these acoustic features can improve the performance of the speech system based on the cognitive large model, help the system understand and optimize these sound features, better adapt to different speech inputs, and improve the accuracy of speech recognition. It also helps to improve the recognition rate of the speech system based on the cognitive large model.
[0039] (2) The present invention helps to evaluate the accuracy of the speech system based on the cognitive big model in understanding the user's intention by analyzing the number and length of overlaps between the person's voice instructions and various control instruction keywords. A higher number of overlaps and overlap lengths usually mean a better match, thereby improving the correctness and execution accuracy of the system to the user's instructions. It can also help improve the response of the speech system based on the cognitive big model to the user's instructions, thereby enhancing the user experience.
[0040] (3) By analyzing the recognition accuracy of control commands and the interval time between command reception and feedback response, the present invention enables the system to more accurately understand and recognize the user's voice commands, which can further help evaluate the performance and user experience of the voice system based on the cognitive large model, thereby improving the reliability, practicality and user satisfaction of the operation.
[0041] (4) The present invention integrates a speech recognition and analysis module, a speech function control module and a feedback prompt module. The speech recognition and analysis module can efficiently capture and accurately identify people's voice commands, thereby improving driving convenience and safety. At the same time, by analyzing the confidence of voice commands, the system's understanding of user intentions and recognition accuracy are improved, thereby effectively optimizing the execution reliability of commands. The feedback prompt module comprehensively analyzes the cognitive status of the speech system and provides real-time feedback and prompts, helping users to better understand the system operation status and improving user experience and trust in the system.
[0042] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a schematic diagram of system module connections of the present invention.
[0044] Figure 2 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "thickness", "top", "middle", "length", "inside", "around" and the like indicating orientation or positional relationship are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0047] See also Figure 1 As shown, an embodiment of the present invention provides a technical solution: a speech system based on a cognitive large model, including a speech recognition and analysis module, which is used to identify the speech input instructions of a person for analysis and processing, and evaluate the confidence of the person's speech instructions.
[0048] The voice function control module is used to obtain trusted personnel, analyze the correlation between the trusted personnel's voice commands and various control commands, locate the vehicle control components corresponding to the specified control commands to enable function execution.
[0049] The feedback prompt module is used to comprehensively analyze the cognitive status of the speech system and provide feedback prompts.
[0050] The speech recognition database is used to store the vehicle's historical voice input record information and a keyword library of various control commands.
[0051] Specifically, the voice input instructions of the person are identified and analyzed and processed. The specific analysis process is: identify the voice input instructions of the person and extract the person's voice characteristic information therefrom, including the person's volume at each time point and the person's average speaking speed, and obtain the vehicle's historical voice input record information from the voice recognition database, where the vehicle's historical voice input record information includes the historical voice reference volume, historical voice reference speed, and historical voice reference tone waveform of the vehicle's personnel.
[0052] The acoustic matching of the preliminary calculation personnel meets the first eigenvalue, which is calculated as follows:
[0053] In the formula, α indicates that the acoustic matching of the person meets the first eigenvalue, H j H0 represents the volume of the person at the j-th time point, H0 represents the historical voice reference volume of the vehicle person, ν3 represents the average speaking speed of the person, Δν represents the historical voice reference speed, a1 and a2 represent the sound matching compensation factors corresponding to the set voice volume and speaking speed respectively, j represents the number of each time point, j = 1, 2, 3, ..., m, and m represents the total number of time points.
[0054] According to the personnel's voice input instructions, the personnel's tone waveform diagram is constructed, and the personnel's tone waveform length is extracted. At the same time, based on the historical voice reference tone waveform of the vehicle's personnel, the personnel's tone waveform diagram is overlapped and compared with the historical voice reference tone waveform of the vehicle's personnel to obtain the maximum length of a single deviation of the personnel's tone waveform and the tone waveform overlap length.
[0055] The acoustic matching of the operator is calculated according to the second eigenvalue, which is calculated as follows:
[0056] Where, the acoustic matching of β personnel meets the second eigenvalue, L1 represents the maximum length of a single deviation of the personnel's pitch waveform, L0 represents the length of the personnel's pitch waveform, and L 2→c It represents the pitch waveform overlap length between the pitch waveform of the person and the historical voice reference pitch waveform of the vehicle person, b1 represents the interference factor corresponding to the set maximum deviation length of the pitch waveform unit, b2 represents the correction factor corresponding to the set pitch waveform overlap length, and e represents a natural constant.
[0057] The confidence level of the operator's voice command is calculated using the following formula:
[0058] Where χ represents the confidence level of the person's voice command, ε1 and ε2 represent the weight factors corresponding to the set acoustic matching meeting the first eigenvalue and the acoustic matching meeting the second eigenvalue, respectively.
[0059] In this embodiment, volume analysis can help improve the performance of the speech system based on the cognitive large model. A higher volume can provide a clear and easily recognizable sound signal and reduce the impact of noise on recognition accuracy. Therefore, volume analysis helps to improve the recognition rate of the system.
[0060] In this embodiment, long-term utterances can provide more information, making it easier for the speech system based on the cognitive large model to capture key information, especially for complex instructions or sentences.
[0061] In this embodiment, analysis of speech speed also contributes to the accuracy of the speech system based on the cognitive macro model. Too fast or too slow speech speed may affect the recognition effect of the system, so a moderate speech speed can improve recognition accuracy and system reliability.
[0062] In this embodiment, for a speech system based on a cognitive macro model, correct tone expression can provide more contextual information, which helps to understand the user's intention more accurately.
[0063] In this embodiment, by analyzing the volume and duration of a person's voice, the person's speaking speed and pitch, the reasonable analysis and application of these acoustic features can improve the performance of the speech system based on the cognitive big model, help the system understand and optimize these sound features, better adapt to different voice inputs, and improve the accuracy of speech recognition, and also help to improve the recognition rate of the speech system based on the cognitive big model.
[0064] Specifically, to obtain a trusted person, the specific process is: extracting the person's voice command confidence, and comparing the person's voice command confidence with the voice command confidence threshold according to the set voice command confidence threshold. If the person's voice command confidence is higher than the voice command confidence threshold, the person is determined to be a trusted person.
[0065] In this implementation, the confidence level of the voice commands of the computing personnel is analyzed, which provides data support for the subsequent evaluation of the correlation between the voice commands of the confident personnel and various control commands, and for locating the specified control commands, thereby helping the speech system based on the cognitive big model to recognize the speech more accurately and improving the accuracy of the speech recognition system.
[0066] Specifically, the correlation between the voice instructions of the trusted personnel and various control instructions is evaluated. The analysis process is as follows: the voice instructions of the trusted personnel are extracted and pre-processed into text form, and the recognized text is processed through natural language processing to obtain the control instruction keywords in the voice instructions of the trusted personnel.
[0067] According to the keyword library of various control instructions stored in the voice database, the number of overlaps between the voice instructions of the trusted personnel and the keywords of various control instructions is compared to evaluate the correlation between the voice instructions of the trusted personnel and various control instructions. The calculation formula is:
[0068]
[0069] Where, δ i Indicates the correlation between the voice commands of the trusted personnel and various control commands, S i represents the number of overlaps between the trusted personnel's voice command and the keywords of the i-th type of control command, S0 represents the set number of overlaps in the control command, c1 represents the correction factor for the set number of overlaps, i represents the number of each type of control command, i = 1, 2, 3, ..., n, and n represents the total number of control command types.
[0070] In this implementation, the number of overlaps in command keywords reflects the degree of match between the user's voice commands and the system's preset commands. A higher number of overlaps usually indicates that the user's voice commands are closer to the control commands defined by the voice system based on the cognitive big model, thereby improving the accuracy and reliability of the system in identifying user intentions.
[0071] In this implementation, the overlap length refers to the continuous length of the voice command that overlaps with the control command keyword. A longer overlap length means that more voice content matches the preset command, which can enhance the system's understanding of the user's intention and further improve the accuracy and reliability of the voice system recognition based on the cognitive large model.
[0072] In this implementation, by analyzing the number and length of overlaps between a person's voice commands and various control command keywords, it helps to evaluate the accuracy of the voice system based on the cognitive big model in understanding the user's intentions. A higher number of overlaps and overlap lengths usually mean a better match, thereby improving the system's correctness and execution accuracy of user commands. It can also help improve the response of the voice system based on the cognitive big model to user commands, thereby improving the user experience.
[0073] Specifically, the vehicle control components corresponding to the designated control instructions are located and the function is enabled. The specific process is: the correlation between the voice instructions of the trusted personnel and various control instructions is sorted in descending order, the control instruction ranked first is extracted as the designated control instruction, and the vehicle control components corresponding to the designated control instruction are enabled and executed.
[0074] In this implementation, the function of enabling the vehicle control components corresponding to the specified control instructions also includes the ability to efficiently capture and accurately identify people's voice commands, so that the driver can easily perform operations such as vehicle body information query and vehicle comfort domain module control.
[0075] Specifically, the cognitive status of the speech system is comprehensively analyzed to provide feedback prompts. The specific process is: calculating the recognition function evaluation value of the speech system, and the calculation formula is:
[0076]
[0077] Where ψ is the recognition function evaluation value of the speech system, K is the first cognitive function evaluation value of the speech system, Υ is the second cognitive function evaluation value of the speech system, τ1 and τ2 represent the weight factors corresponding to the set first cognitive function evaluation value and second cognitive function evaluation value, respectively.
[0078] Based on the recognition function evaluation value of the voice system, it is compared with the set recognition function evaluation threshold. If the recognition function evaluation value of the voice system is lower than the recognition function evaluation threshold, feedback prompts are given.
[0079] Specifically, the first cognitive function evaluation value of the speech system is analyzed as follows: a monitoring period is set, during which the recognition accuracy rate and the interval between command reception and feedback response of the speech system are counted, and the influencing factors corresponding to the recognition accuracy rate and the interval between unit reception and feedback response stored in the speech database are extracted; and the first cognitive function evaluation value of the speech system is calculated using the following formula:
[0080] Where K represents the first cognitive function evaluation value of the speech system, Z represents the recognition accuracy of the speech system, ΔZ represents the recognition accuracy, Ts represents the interval time for receiving feedback response of speech commands, η1 represents the correction coefficient of recognition accuracy, and η2 represents the impact factor corresponding to the interval time for receiving feedback response per unit.
[0081] In this implementation, the recognition accuracy of control commands is one of the important indicators for measuring the performance of the voice system. A higher recognition accuracy means that the system can more accurately understand and recognize the user's voice commands, thereby improving the accuracy of operations and user satisfaction. The improvement in accuracy can be achieved by optimizing the voice recognition algorithm, increasing training data, and other means.
[0082] In this implementation, the time it takes for a command to receive a system reply, that is, the speed of the system response, directly affects the smoothness and practicality of the user experience. A shorter response time can enhance the user's trust and satisfaction with the system and improve the overall user experience. This can be achieved by optimizing the response speed of the voice system based on the cognitive big model, such as improving algorithm efficiency, reducing processing delays, or optimizing network connections.
[0083] In this implementation, by analyzing the recognition accuracy of control commands and the interval time between command reception and feedback responses, the system can more accurately understand and recognize the user's voice commands, which can further help evaluate the performance and user experience of the voice system based on the cognitive big model, thereby improving the reliability, practicality and user satisfaction of the operation.
[0084] Specifically, the second cognitive function evaluation value of the voice system is analyzed as follows: during the monitoring period, the number of times the voice system maintains coherence in multiple rounds of dialogue under specified control instructions and the number of times it corrects errors in user voice input are counted, and the second cognitive function evaluation value of the voice system is calculated using the following formula:
[0085]
[0086] Where, Υ represents the second cognitive function evaluation value of the speech system, λ and They represent the number of coherences in the multi-round dialogue under the specified control command and the number of errors in correcting the user's voice input, Δλ and They represent the defined number of consecutive times and the defined number of errors respectively, and ω1 and ω2 represent the correction factors corresponding to the set number of consecutive times and the number of errors respectively.
[0087] In this implementation, by integrating the voice recognition analysis module, the voice function control module and the feedback prompt module, the voice recognition analysis module can efficiently capture and accurately identify people's voice commands, thereby improving driving convenience and safety. At the same time, by analyzing the confidence of voice commands, the system's understanding and recognition accuracy of user intentions are improved, thereby effectively optimizing the execution reliability of commands. The feedback prompt module conducts a comprehensive analysis of the cognitive status of the voice system and provides real-time feedback and prompts to help users better understand the system operation status, thereby enhancing user experience and trust in the system.
[0088] The second aspect of the present invention also provides a speech method based on a cognitive large model, comprising: identifying a person's speech input instructions for analysis and processing, and evaluating the confidence of the person's speech instructions.
[0089] Acquire trusted personnel, analyze the correlation between the trusted personnel's voice commands and various control commands, locate the vehicle control components corresponding to the specified control commands and enable function execution.
[0090] Comprehensively analyze the cognitive status of the speech system and provide feedback prompts.
[0091] The third aspect of the present invention also provides a speech device based on a cognitive big model, comprising: a processor, and a memory and a network interface connected to the processor; the network interface is connected to a non-volatile memory in a server; the processor retrieves a computer program from the non-volatile memory through the network interface during operation, and runs the computer program through the memory to execute the above-mentioned speech system based on a cognitive big model.
[0092] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0093] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A speech system based on a cognitive large model, characterized in that: include: A speech recognition and analysis module is used to identify and analyze a person's voice input commands, evaluate the confidence level of the person's voice commands, identify the person's voice input commands, and extract the person's voice characteristic information from them, including the person's volume and average speaking speed at each time point. It also obtains historical vehicle voice input record information from a speech recognition database, where the vehicle's historical voice input record information includes the historical voice reference volume, historical voice reference speed, and historical voice reference pitch waveform of the vehicle's personnel. It also preliminarily calculates the person's acoustic match compliance with the first eigenvalue based on the person's voice characteristic information and the vehicle's historical voice input record information. Based on the personnel's voice input command, a pitch waveform diagram of the personnel is constructed, and the length of the personnel's pitch waveform is extracted. At the same time, based on the historical voice reference pitch waveform of the vehicle personnel, the personnel's pitch waveform diagram is overlapped and compared with the historical voice reference pitch waveform of the vehicle personnel to obtain the maximum length of a single deviation of the personnel's pitch waveform and the length of the pitch waveform overlap, and further obtain the acoustic matching compliance with the second eigenvalue; the voice command confidence is obtained by combining the acoustic matching compliance with the first eigenvalue and the acoustic matching compliance with the second eigenvalue; The voice function control module is used to obtain trusted personnel through voice command confidence screening, analyze the correlation between the trusted personnel's voice commands and various control commands, and locate the vehicle control components corresponding to the specified control commands to enable function execution; A feedback prompt module is used to comprehensively analyze the cognitive status of the speech system and provide feedback prompts. The comprehensive analysis of the cognitive status of the speech system and providing feedback prompts refers to counting the recognition accuracy rate of the speech system and the interval time of command reception feedback response during the monitoring period, extracting the recognition accuracy rate, the correction coefficient of the recognition accuracy rate, and the influence factor corresponding to the unit feedback response interval stored in the speech database, and calculating the first cognitive function evaluation value of the speech system; During the monitoring period, the system counts the number of coherences in multiple rounds of conversation under specified control commands and the number of errors in the user's voice input that were corrected. The system also extracts the defined number of coherences and the defined number of errors stored in the voice database, as well as the correction factors corresponding to the number of coherences and the number of errors. The system calculates the second cognitive function evaluation value of the voice system and provides feedback based on the recognition function evaluation value of the voice system. The speech recognition database is used to store the vehicle's historical voice input record information and a keyword library of various control commands.
2. The speech system based on the cognitive macromodel according to claim 1, characterized in that: The speech input command of the recognition personnel is analyzed and processed, and the specific analysis process is as follows: The acoustic matching of the preliminary calculation personnel meets the first eigenvalue, which is calculated as follows: , Where, Indicates that the acoustic matching of the person meets the first eigenvalue, Indicates the volume of people at the jth time point Indicates the historical voice reference volume of the vehicle's personnel. Indicates the average speaking speed of the person. Indicates the historical speech reference speed, 、 They represent the sound matching compensation factors corresponding to the set voice volume and voice speed, Indicates the number of each time point, , Indicates the total number of time points; The acoustic matching of the operator is calculated according to the second eigenvalue, which is calculated as follows: , Where, The acoustic matching of the person is in accordance with the second eigenvalue, Indicates the maximum length of a single deviation of a person's voice waveform. Indicates the length of the person's pitch waveform, The length of the pitch waveform overlap between the pitch waveform of the person and the historical voice reference pitch waveform of the vehicle's personnel, Indicates the interference factor corresponding to the maximum deviation length of the pitch waveform unit. It represents the correction factor corresponding to the set pitch waveform overlap length, and e represents a natural constant; The confidence level of the operator's voice command is calculated using the following formula: , Where, Indicates the confidence level of the person's voice command, and The weight factors respectively represent that the set acoustic matching conforms to the first eigenvalue and the acoustic matching conforms to the second eigenvalue.
3. The speech system based on cognitive large model according to claim 1, characterized in that: The specific process of obtaining the trusted personnel is as follows: The person's voice command confidence is extracted, and the person's voice command confidence is compared with the voice command confidence threshold according to the set voice command confidence threshold. If the person's voice command confidence is higher than the voice command confidence threshold, the person is determined to be a trusted person.
4. The speech system based on cognitive large model according to claim 1, characterized in that: The analysis of the correlation between the voice command of the trustee and various control commands is carried out as follows: Extracting the voice instructions of the trusted personnel and pre-processing them into text form, and processing the recognized text through natural language processing to obtain the control instruction keywords in the voice instructions of the trusted personnel; According to the keyword library of various control instructions stored in the voice database, the number of overlaps between the voice instructions of the trusted personnel and the keywords of various control instructions is compared, and the correlation between the voice instructions of the trusted personnel and various control instructions is analyzed. The calculation formula is: , Where, Indicates the correlation between the voice commands of the trusted personnel and various control commands, represents the number of keywords that overlap between the voice command of the trusted person and the control command of the i-th category, Indicates the number of overlapping limits of the set control instructions. Indicates the correction factor for the set number of overlaps, Indicates the number of each type of control instruction, , Indicates the total number of control instruction types.
5. The speech system based on cognitive large model according to claim 3, characterized in that: The vehicle control components corresponding to the positioning designated control instructions are enabled to execute functions, and the specific process is as follows: The correlation between the voice commands of the trusted personnel and various control commands is sorted in descending order, the control command ranked first is extracted as the designated control command, and the vehicle control components corresponding to the designated control command are enabled for execution.
6. The speech system based on cognitive large model according to claim 1, characterized in that: The comprehensive analysis of the cognitive status of the speech system to provide feedback prompts is as follows: Calculate the recognition function evaluation value of the speech system. The calculation formula is: , Where, is the recognition function evaluation value of the speech system, is the first cognitive function evaluation value of the speech system, is the second cognitive function evaluation value of the speech system, and Respectively represent the weight factors corresponding to the set first cognitive function evaluation value and the second cognitive function evaluation value; Based on the recognition function evaluation value of the voice system, it is compared with the set recognition function evaluation threshold. If the recognition function evaluation value of the voice system is lower than the recognition function evaluation threshold, feedback prompts are given.
7. The speech system based on cognitive large model according to claim 6, characterized in that: The first cognitive function evaluation value of the speech system is specifically analyzed as follows: Setting a monitoring period, counting the recognition accuracy of the voice system and the interval time for receiving feedback responses of commands during the monitoring period, and extracting the influencing factors corresponding to the recognition accuracy and the interval time for receiving feedback responses of units stored in the voice database; Calculate the first cognitive function evaluation value of the speech system, and the calculation formula is: , Where, represents the first cognitive function evaluation value of the speech system, Indicates the recognition accuracy of the speech system, Represents the recognition accuracy, Indicates the interval time between receiving feedback responses for voice commands. Represents the correction coefficient of recognition accuracy, Indicates the impact factor corresponding to the unit feedback response interval.
8. The speech system based on cognitive macromodel according to claim 6, characterized in that: The specific analysis process of the second cognitive function evaluation value of the speech system is as follows: During the monitoring period, the number of times the voice system maintains coherence in multiple rounds of dialogue under specified control instructions and the number of times it corrects errors in user voice input are counted, and the second cognitive function evaluation value of the voice system is calculated. The calculation formula is: , Where, represents the second cognitive function evaluation value of the speech system, and They represent the number of times the voice system is coherent in multiple rounds of dialogue under specified control instructions and the number of times it corrects errors in user voice input. and Respectively represent the number of consecutive definitions and the number of incorrect definitions, and They represent the correction factors corresponding to the set number of consecutive times and the number of errors respectively.
9. A speech method based on a cognitive large model, applied to the speech system based on a cognitive large model according to any one of claims 1 to 8, characterized in that: include: Identify the person's voice input commands, analyze and process them, and evaluate the confidence level of the person's voice commands; Acquire trusted personnel, analyze the correlation between the trusted personnel's voice commands and various control commands, locate the vehicle control components corresponding to the specified control commands and enable function execution; Comprehensively analyze the cognitive status of the speech system and provide feedback prompts.
10. A speech device based on a cognitive large model, characterized in that: include: A processor, and memory and network interfaces connected to the processor; The network interface is connected to the non-volatile memory in the server; When running, the processor retrieves a computer program from the non-volatile memory through the network interface, and runs the computer program through the memory to execute the system described in any one of claims 1 to 8.
Citation Information
Patent Citations
Interactive system for vehicle-mounted voice
CN101281745A
Vehicle Human-Machine Voice Interaction System
CN103730119B
Voice instruction response method and device and storage medium
CN114399992A
Coal mine robot voice interaction system
CN116994595A
Interaction prompt text determination method and device, electronic equipment and storage medium
CN117765942A