Voice model adjustment method, device and equipment

By collecting voice interaction data between the vehicle and the supplier's cloud in the transit system, the host factory can directly obtain complete interactive data, solving the problem that the vehicle and machine factory finds difficulty in optimizing the supplier's cloud voice model and achieving more efficient performance adjustment and optimization.

CN120148489APending Publication Date: 2025-06-13GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510298837.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the online service of smart cockpit voice interaction, it is difficult for the vehicle manufacturer to directly access the detailed data of the SDK on the supplier side, resulting in the inability to optimize the performance of the supplier's voice model in the cloud.

Method used

Through the transit system, the voice commands sent by designated vehicle machines to the supplier's cloud and reply commands sent by the cloud are collected. The host factory can directly obtain complete voice interaction data, thereby determining the performance adjustment parameters and sending them to the supplier's cloud for adjustment.

Benefits of technology

This method breaks the barriers to traditional data acquisition, improves data accessibility, accurately recognizes the performance bottlenecks of voice models, improves performance adjustment efficiency, and achieves targeted optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148489A_ABST
    Figure CN120148489A_ABST
Patent Text Reader

Abstract

The invention provides a supplier cloud voice model adjustment method, device and equipment, and belongs to the technical field of vehicle control. Comprising the following steps: collecting each voice instruction sent to a supplier cloud by a specified vehicle machine; collecting a reply instruction corresponding to each voice instruction sent by the supplier cloud to the specified vehicle machine; determining a performance adjustment parameter of the supplier cloud based on each voice instruction and a reply instruction corresponding to each voice instruction; and sending the performance adjustment parameter to the supplier cloud, so that the supplier cloud performs adjustment based on the performance adjustment parameter. By collecting the voice instruction sent by the specified vehicle machine to the supplier cloud and the reply instruction sent by the supplier cloud to the vehicle machine, the main engine plant can obtain complete voice interaction data, so that the data accessibility is enhanced. Meanwhile, by analyzing the collected voice instruction and the corresponding reply instruction, the performance bottleneck and the problem point can be identified, so that the efficiency of performance optimization is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of vehicle control, and particularly to a method, device, and equipment for adjusting a voice model on a supplier cloud. Background Art

[0002] In the intelligent cockpit voice interaction service, core capabilities such as automatic speech recognition (ASR) and natural language understanding (NLU) are the keys to providing a high-quality user experience. These capabilities are usually provided in two ways: offline and online. The offline capabilities are mainly provided by the in-vehicle unit and are applicable when the in-vehicle network condition is poor or there is no network connection; while the online capabilities rely on cloud services and are applicable when the in-vehicle network is good.

[0003] Currently, the online service of intelligent cockpit voice interaction is usually provided by a supplier directly accessing cloud services through a software development kit (SDK) on the terminal side. However, in this mode, it is difficult for the vehicle manufacturer to directly access the detailed data of the supplier's terminal-side SDK and obtain voice commands and reply commands for the voice commands. Therefore, it is difficult for the host factory to optimize the performance of the voice model on the supplier cloud. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a method, device, and equipment for adjusting a voice model on a supplier cloud, which are used to overcome the problem that it is difficult to optimize the voice model deployed on the supplier cloud. The technical solutions are as follows:

[0005] A method for adjusting a voice model on a supplier cloud, which is applied to a transit system. The method includes:

[0006] Collecting each voice command sent by a specified in-vehicle unit to the supplier cloud;

[0007] Collecting each reply command corresponding to each voice command sent by the supplier cloud to the specified in-vehicle unit;

[0008] Determining a performance adjustment parameter of the supplier cloud based on each voice command and each reply command corresponding to each voice command;

[0009] Sending the performance adjustment parameter to the supplier cloud so that the supplier cloud can make adjustments based on the performance adjustment parameter.

[0010] It should be noted that by collecting the voice commands sent by the specified in-vehicle unit to the supplier cloud and the reply commands sent by the cloud through the relay system, the vehicle manufacturer can directly obtain the complete voice interaction data. This direct access method breaks the barriers of traditional data acquisition, thereby improving data accessibility. The acquisition of complete data can better provide performance adjustment parameters for the supplier cloud. At the same time, since the vehicle manufacturer can directly access and analyze the complete interaction data, they can more accurately identify the performance bottlenecks and problem points of the voice model. This accurate identification reduces unnecessary optimization attempts, thereby improving the performance adjustment efficiency of the supplier cloud. In addition, by analyzing the voice commands and the corresponding reply commands, the vehicle manufacturer can determine which specific commands or scenarios lead to performance problems. Based on these analysis results, the supplier can formulate targeted performance adjustment parameters and conduct targeted optimizations instead of extensive and non-targeted adjustments.

[0011] Optionally, the supplier cloud includes multiple types of voice models, and the performance adjustment parameters include the resource allocation ratios of multiple types of voice models. Determining the performance adjustment parameters of the supplier cloud based on each of the voice commands and the reply commands corresponding to each of the voice commands includes:

[0012] Determining the types of voice models corresponding to each of the voice commands and the reply commands corresponding to each of the voice commands;

[0013] Based on the types of the corresponding voice models, determining the resource allocation ratios of multiple types of voice models in the supplier cloud.

[0014] It should be noted that by setting multiple types of voice models in the cloud service and determining the usage frequency of each model based on the actual interaction data, the supplier can adjust the service according to the needs of different users, thereby enhancing the customization ability of the cloud service. By analyzing the voice commands and their corresponding reply commands to determine the ratios of different types of voice models, the supplier can optimize the resource allocation accordingly to ensure that the more frequently used models receive more resources, thereby improving the overall performance.

[0015] Optionally, determining the resource allocation ratios of multiple types of voice models in the supplier cloud based on the types of the corresponding voice models includes:

[0016] Based on the types of the corresponding voice models, determining the usage proportions of multiple types of voice models;

[0017] Based on the usage proportions, determining the resource allocation ratios of multiple types of voice models in the supplier cloud.

[0018] It should be noted that by analyzing the types of speech models corresponding to each speech instruction and reply instruction, it is possible to identify which model types are more frequently used in actual applications, that is, the usage ratio is higher. Based on the usage ratio, the resource allocation is dynamically adjusted, and more resources are allocated to the model types with higher usage frequencies. This can reduce resource waste because the model types with low usage frequencies do not occupy excessive resources, thereby improving the overall resource utilization rate. Moreover, when resources are more effectively allocated to the model types with higher usage frequencies, these model types can process speech instructions faster. Faster processing speed means shorter response time, thus enhancing the user experience. At the same time, due to more precise resource allocation, the model can operate at a higher efficiency, reducing performance bottlenecks caused by insufficient resources.

[0019] Optionally, the performance adjustment parameter includes the resource allocation ratio for each time period. Determining the performance adjustment parameter of the supplier cloud based on each of the speech instructions and the reply instructions corresponding to each of the speech instructions includes:

[0020] Determine the first timestamp of each of the speech instructions;

[0021] Determine the second timestamp of the reply instruction;

[0022] Based on the first timestamp and the second timestamp, determine the processing volume of speech instructions in each time period;

[0023] Based on the processing volume of speech instructions in each of the time periods, determine the resource allocation ratio for each of the time periods in the supplier cloud.

[0024] It should be noted that by recording the first timestamp of the speech instruction and the second timestamp of the reply instruction, the processing time of each speech instruction from reception to reply can be accurately measured. This real-time monitoring ability enables the system to track the entire process of speech instruction processing and promptly detect any delays or abnormalities during the processing. Moreover, by analyzing the processing volume of speech instructions in each time period, a trend chart of speech instruction processing can be drawn to identify the peak periods of speech instruction processing. Identifying the peak periods of speech instruction processing helps predict the peak of system resource requirements, thereby providing a basis for system resource configuration. Based on the identified processing peak periods, the resource allocation ratio can be dynamically adjusted to ensure that more resources are available for processing speech instructions during peak hours. By optimizing resource configuration, the risk of system overload can be reduced, and the stability and response speed of the system can be improved.

[0025] Optionally, the supplier cloud includes multiple types of speech models. Determining the resource allocation ratio for each of the time periods based on the processing volume of speech instructions in each of the time periods includes:

[0026] Determine the types of speech models corresponding to each of the speech commands;

[0027] Based on the processing volume of speech commands in each of the time periods, and the types of speech models corresponding to each of the speech commands, determine the resource allocation ratios of various types of speech models in each of the time periods in the supplier cloud.

[0028] It should be noted that by determining the types of speech models corresponding to each speech command, the supplier cloud can understand the performance differences of different types of models when processing speech commands. Combining the processing volume of speech commands in each time period, the cloud can identify which types of speech models are used more frequently in a specific time period. Based on this information, the cloud can dynamically adjust the resource allocation, giving priority to allocating to the model types with higher performance requirements, thereby optimizing the performance of the entire system. During the peak processing period, dynamically adjust the resource allocation ratio and allocate more resources to the speech model types with faster processing speeds. Since the resource allocation is more reasonable, these model types can process speech commands faster, thereby reducing the processing delay.

[0029] Optionally, the performance adjustment parameter includes a resource allocation parameter. The determining of the performance adjustment parameter of the supplier cloud based on each of the speech commands and the reply commands corresponding to each of the speech commands includes:

[0030] Determine the third timestamps of each of the speech commands;

[0031] Determine the fourth timestamps of the reply commands corresponding to each of the speech commands;

[0032] Based on the third timestamp and the fourth timestamp, determine the response time corresponding to each of the speech commands;

[0033] Based on the response time corresponding to each of the speech commands, determine the resource allocation parameter of the supplier cloud.

[0034] It should be noted that by recording the third timestamp of the speech command and the fourth timestamp of the reply command, the performance of the cloud can be evaluated in real time to ensure the instant response ability of the system. And determining the response time of each speech command helps to identify and optimize the response speed of the system, improving the instantaneity of user interaction. The response time data can be used as an important basis for resource allocation to ensure that more resources are invested during periods with higher processing delays. By analyzing the response time, performance bottlenecks in the system, such as insufficient computing resources and network latency, can be identified and optimized accordingly.

[0035] Optionally, the supplier cloud includes multiple types of speech models. The determining of the resource allocation parameter of the supplier cloud based on the response time corresponding to each of the speech commands includes:

[0036] Determine the types of speech models corresponding to each of the speech commands;

[0037] Based on the response time corresponding to each of the speech commands and the types of speech models corresponding to each of the speech commands, determine the resource allocation parameters of each type of speech model in the supplier cloud.

[0038] It should be noted that by determining the types of speech models corresponding to each speech command and their response times, the supplier cloud can identify which models are more efficient in processing speech commands, so as to preferentially allocate resources to models with shorter response times, which can reduce the processing time and thus improve the response speed of the entire system.

[0039] Optionally, the supplier cloud includes speech models for multiple application scenarios, and the performance adjustment parameters of the supplier cloud include the speech models for the corresponding application scenarios in the supplier cloud. The determining of the performance adjustment parameters of the supplier cloud based on each of the speech commands and the reply commands corresponding to each of the speech commands includes:

[0040] Determine the semantic information of each of the speech commands and the reply commands corresponding to each of the speech commands;

[0041] Based on the semantic information, determine the application scenario of the specified in-vehicle unit;

[0042] Based on the application scenario of the specified in-vehicle unit, determine the speech models for the corresponding application scenarios in the supplier cloud.

[0043] It should be noted that identifying the application scenario of the specified in-vehicle unit based on semantic information helps the cloud system to provide customized services for users in specific environments. Application scenario recognition can improve the user experience because users can get a more personalized speech interaction experience when using the in-vehicle unit. At the same time, selecting appropriate speech models according to different application scenarios can make the models more suitable for the scenario requirements and improve the performance of the models in specific scenarios.

[0044] A speech model adjustment device for a supplier cloud, the device includes:

[0045] A speech command collection unit, which collects each speech command sent by a specified in-vehicle unit to the supplier cloud;

[0046] A reply command collection unit, which collects the reply commands corresponding to each of the speech commands sent by the supplier cloud to the specified in-vehicle unit;

[0047] An adjustment parameter determination unit, which determines the performance adjustment parameters of the supplier cloud based on each of the speech commands and the reply commands corresponding to each of the speech commands;

[0048] An adjustment parameter uploading unit that sends the performance adjustment parameter to the supplier cloud so that the supplier cloud can make adjustments based on the performance adjustment parameter.

[0049] A voice model adjustment device for a supplier cloud, comprising:

[0050] At least one processor; and,

[0051] A memory communicatively connected to the at least one processor; wherein,

[0052] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:

[0053] Collect each voice command sent by a specified vehicle-mounted device to the supplier cloud;

[0054] Collect each reply command corresponding to each voice command sent by the supplier cloud to the specified vehicle-mounted device;

[0055] Based on each voice command and each reply command corresponding to each voice command, determine the performance adjustment parameter of the supplier cloud;

[0056] Send the performance adjustment parameter to the supplier cloud so that the supplier cloud can make adjustments based on the performance adjustment parameter.

[0057] A non-volatile computer storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a computer, they can implement:

[0058] Collect each voice command sent by a specified vehicle-mounted device to the supplier cloud;

[0059] Collect each reply command corresponding to each voice command sent by the supplier cloud to the specified vehicle-mounted device;

[0060] Based on each voice command and each reply command corresponding to each voice command, determine the performance adjustment parameter of the supplier cloud;

[0061] Send the performance adjustment parameter to the supplier cloud so that the supplier cloud can make adjustments based on the performance adjustment parameter.

[0062] By means of the above technical solutions, a voice model adjustment method, device, equipment and medium for a supplier cloud provided by the present disclosure are applied to a transit system.

[0063] The above description is only an overview of the technical solution of the present disclosure. In order to be able to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the specific implementation manners of the present disclosure are specifically exemplified below. Description of the Drawings

[0064] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present disclosure. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0065] Figure 1 A flowchart showing a method for adjusting a voice model in a supplier cloud provided by an embodiment of the present disclosure is shown;

[0066] Figure 2 A structural diagram showing a voice skill data collection and analysis system provided by an embodiment of the present disclosure is shown;

[0067] Figure 3 A structural diagram showing a device for adjusting a voice model in a supplier cloud provided by an embodiment of the present disclosure is shown. Detailed Description of the Preferred Embodiments

[0068] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0069] In the field of voice interaction services in intelligent cockpits, speech recognition technology and natural language understanding technology play a core role, and they are crucial for ensuring that users can enjoy a high-quality user experience. These two technologies are the key to realizing the voice interaction function in intelligent cockpits, and their technical implementation methods are mainly divided into two modes: offline and online.

[0070] The offline capability is mainly provided by the intelligent cockpit system of the vehicle itself. This mode is suitable for environments where the vehicle has poor network signals or no network connection at all. In the offline mode, the vehicle can independently process the user's voice commands without relying on external network resources, thus ensuring the real-time performance and stability of the voice interaction.

[0071] In contrast, the online capability relies on cloud services and is applicable when the vehicle has a good network connection. In the online mode, the intelligent cockpit system of the vehicle can send the user's voice commands to the cloud in real time. The cloud server performs speech recognition and natural language understanding and then feeds back the processing results to the vehicle.

[0072] Currently, most of the online services for voice interaction in intelligent cockpits are realized by directly accessing cloud services through the software development kits (SDKs) provided by suppliers on the device side. In this mode, vehicle manufacturers often have difficulty directly obtaining detailed data of the supplier's device-side SDK, including voice commands and response information for these commands. This information isolation results in the inability of the vehicle manufacturers to effectively optimize the performance of the voice models on the supplier's cloud side.

[0073] This problem of information asymmetry undoubtedly restricts the further improvement of the user experience and functionality of the intelligent cockpit voice interaction system. Since vehicle manufacturers cannot deeply understand the working principle of the cloud voice model, it is difficult for them to carry out targeted optimization and adjustment. This not only affects the overall performance of the intelligent cockpit voice interaction system but also limits the technological innovation and development of vehicle manufacturers in the field of intelligent cockpits. Therefore, how to break this information isolation and achieve effective communication and collaboration between vehicle manufacturers and suppliers has become an important issue in promoting the development of intelligent cockpit voice interaction technology. For this purpose, this application provides a schematic flowchart of a method for adjusting the voice model on the supplier's cloud side, as Figure 1 shown. This process can be executed by a transfer system. Some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.

[0074] Among them, the above-mentioned transfer system includes two subsystems: a transfer service and a data platform.

[0075] First, the in-vehicle unit sends voice commands to the transfer service, and the transfer service forwards them to the cloud service of the voice supplier. In this process, the transfer service appears to be communicating with the cloud service of the voice supplier to the in-vehicle unit side, and the transfer service is a transparent role. To the cloud service of the voice supplier, the request it receives is from the in-vehicle unit side, and the transfer service is also a transparent role.

[0076] Then, the cloud service of the voice supplier returns various result data (reply commands) to the transfer service, including ASR results, NLU results, and voice skill dialogue results. The transfer service first forwards these results to the in-vehicle unit side as they are and then reports a copy to the data platform.

[0077] After the data platform receives various voice skill result data from the transit service, it will perform parsing, processing, and storage in the database. Then, data development engineers can perform statistical analysis, Badcase analysis, etc. based on the data stored in the database.

[0078] For the voice model adjustment on the supplier cloud, the method flow steps of the embodiments of the present application are as follows:

[0079] S101, collect each voice command sent by a specified vehicle head unit to the supplier cloud.

[0080] In the embodiments of the present application, a dedicated voice command collection module can be deployed in the transit system, and this module is responsible for real-time monitoring of the voice commands sent from the specified vehicle head unit.

[0081] S102, collect each reply command corresponding to each of the voice commands sent by the supplier cloud to the specified vehicle head unit.

[0082] In the embodiments of the present application, a dedicated reply command collection module can be deployed in the transit system, and this module is responsible for real-time monitoring of the reply commands corresponding to each voice command sent from the supplier cloud.

[0083] S103, determine the performance adjustment parameters of the supplier cloud based on each of the voice commands and each of the reply commands corresponding to the voice commands.

[0084] In the embodiments of the present application, collect all voice commands and their corresponding reply commands of a specified vehicle for a period of time, including the execution frequency, success rate, user feedback, etc. of the commands. Conduct a detailed analysis of the collected data to identify performance bottlenecks and user requirements. Based on the analysis results, key performance indicators, that is, performance adjustment parameters, can be set. The performance adjustment parameters can be response time, error rate, system stability, etc. Determine the target value of each performance adjustment parameter for subsequent performance adjustment.

[0085] It should be noted that according to the set key performance indicators, specific performance adjustment strategies can be formulated. For the response time, the algorithm and data processing flow can be optimized to reduce latency. For the error rate, the error handling mechanism can be enhanced to improve the fault tolerance of the system. For system stability, load balancing and redundant design can be implemented to ensure the stability of the system under high load.

[0086] S104, send the performance adjustment parameters to the supplier cloud so that the supplier cloud can make adjustments based on the performance adjustment parameters.

[0087] In the embodiments of the present application, send the performance adjustment parameters to the supplier cloud, and the supplier cloud can make gradual adjustments based on the above performance adjustment strategies and monitor the adjustment effect in real time.

[0088] It should be noted that by collecting the voice commands sent by the specified in-vehicle unit to the supplier cloud and the reply commands sent by the cloud through the transfer system, the vehicle manufacturer can directly obtain the complete voice interaction data. This direct access method breaks the barriers of traditional data acquisition, thereby improving data accessibility. The acquisition of complete data can better provide performance adjustment parameters for the supplier cloud. At the same time, since the vehicle manufacturer can directly access and analyze the complete interaction data, they can more accurately identify the performance bottlenecks and problem points of the voice model. This accurate identification reduces unnecessary optimization attempts, thereby improving the performance adjustment efficiency of the supplier cloud. In addition, by analyzing the voice commands and the corresponding reply commands, the vehicle manufacturer can determine which specific commands or scenarios lead to performance problems. Based on these analysis results, the supplier can formulate targeted performance adjustment parameters for targeted optimization instead of making extensive and non-targeted adjustments.

[0089] Optionally, the supplier cloud includes multiple types of voice models, and the performance adjustment parameters include the resource allocation ratios of multiple types of voice models. When determining the performance adjustment parameters of the supplier cloud based on each of the voice commands and the reply commands corresponding to each of the voice commands, it is possible to first determine the types of voice models corresponding to each of the voice commands and the reply commands corresponding to each of the voice commands; based on the types of voice models corresponding to each of the voice commands and the reply commands corresponding to each of the voice commands, determine the resource allocation ratios of multiple types of voice models in the supplier cloud.

[0090] It should be noted that regarding the above content, the following specific implementation solutions can be adopted:

[0091] Instruction Classification: Classify voice models into different types based on the functions and uses of the voice models. For example, question-and-answer voice models, navigation voice models, entertainment voice models, information query voice models, etc.

[0092] Model Evaluation: Evaluate different types of voice models available in the cloud, including their performance, accuracy rate, response time, etc.

[0093] Instruction-Model Matching: Match each instruction to the most suitable type of voice model according to the content of the instruction.

[0094] Resource Requirement Analysis: Conduct performance tests on each model to determine their resource consumption under different loads. Establish a resource consumption model for each model type, including CPU, memory, network bandwidth, etc.

[0095] User Behavior Analysis: Analyze the usage frequency of users for different skill type instructions. Set priorities for different skill types according to the usage frequency and user feedback.

[0096] Resource allocation strategy: Create a resource pool that includes all available computing and storage resources. Implement a dynamic resource allocation system to determine the resource allocation ratios for multiple types of speech models in the vendor's cloud based on resource demand analysis and user behavior analysis.

[0097] It should be noted that by setting multiple types of speech models in the cloud service and determining the usage frequency of each model based on actual interaction data, the vendor can adjust the service according to the needs of different users, thereby enhancing the customization ability of the cloud service. By analyzing the speech commands and their corresponding response commands to determine the ratios of different types of speech models, the vendor can optimize resource allocation to ensure that more frequently used models receive more resources, thereby improving overall performance.

[0098] Optionally, when determining the resource allocation ratios for multiple types of speech models based on the types of speech models corresponding to each speech command and each response command corresponding to each speech command, the usage ratios of multiple types of speech models can be determined based on the types of speech models corresponding to each speech command and each response command corresponding to each speech command; based on the usage ratios, determine the resource allocation ratios for multiple types of speech models in the vendor's cloud.

[0099] It should be noted that regarding the above content, the following specific implementation solutions can be adopted:

[0100] Usage ratio analysis: Count the number of times or usage frequency of each speech model type being called. Calculate the total number of times all speech model types are called. Calculate the usage ratio of each speech model type.

[0101] Resource demand assessment: Evaluate the resource consumption of each speech model type, including CPU, memory, network bandwidth, etc. Determine the key performance indicators affecting resource allocation, such as response time, accuracy rate, concurrent processing ability, etc.

[0102] Determination of resource allocation ratios: Set priorities for different speech model types according to business requirements and user experience. Based on usage ratio analysis, resource demand assessment, and priorities, determine the resource allocation ratios.

[0103] For the determination of the above resource allocation ratios, the following specific implementation solutions can be adopted:

[0104] Requirement analysis and planning phase: Clarify the roles and importance of different speech model types in the business. For example, some models may be used for core functions, while others may be used for auxiliary functions. Collect user feedback to understand users' usage habits and satisfaction with different speech model types.

[0105] Data collection and analysis phase: Collect the actual usage data of different speech model types through methods such as log analysis and user behavior tracking, and calculate their usage proportions. Evaluate the resource consumption of different speech models when processing speech requests, including CPU, memory, network bandwidth, etc.

[0106] Priority setting phase: According to business requirements and user experience, classify speech model types into multiple categories, such as core, important, auxiliary, etc. Combine the usage proportion and resource requirement evaluation to assign priorities to each model type. Generally, model types with a high usage proportion and low resource consumption will be assigned higher priorities.

[0107] Resource allocation ratio determination phase: Determine the total amount of resources that can be allocated according to the overall resource status of the system. Based on the priorities, calculate the resource ratio that should be allocated to each model type. For example, a model type with a usage proportion of 30% and the highest priority may receive 20% of the resources, while an auxiliary model type with a usage proportion of 10% may only receive 5% of the resources.

[0108] It should be noted that by analyzing the speech model types corresponding to each speech instruction and reply instruction, it can be identified which model types are more frequently used in actual applications, that is, have a higher usage proportion. Based on the usage proportion, dynamically adjust the resource allocation and allocate more resources to the model types with higher usage frequencies. This can reduce resource waste because model types with low usage frequencies will not occupy too many resources, thereby improving the overall resource utilization rate. Moreover, when resources are more effectively allocated to model types with higher usage frequencies, these model types can process speech instructions faster. Faster processing speed means shorter response time, thus enhancing the user experience. At the same time, due to more precise resource allocation, the model can run at a higher efficiency, reducing performance bottlenecks caused by insufficient resources.

[0109] Optionally, the performance adjustment parameters include the resource allocation ratios for each time period. When determining the performance adjustment parameters of the supplier cloud based on each speech instruction and the reply instruction corresponding to each speech instruction, the first timestamp of each speech instruction can be determined; the second timestamp of the reply instruction corresponding to each speech instruction can be determined; based on the first timestamp and the second timestamp, the speech instruction processing volume within each time period can be determined; based on the speech instruction processing volume within each time period, the resource allocation ratio for each time period in the supplier cloud can be determined.

[0110] It should be noted that regarding the above content, the following specific implementation schemes can be adopted:

[0111] Voice command timestamp collection: Ensure that the system records the first timestamp of all voice commands, i.e., the time when the command is received. Also record the second timestamp of the corresponding response command, i.e., the time when the system finishes the response.

[0112] Time period division: Divide the time into multiple consecutive small time periods, such as every minute, every hour, or every day. Archive the timestamps of each voice command and response command into the corresponding time periods.

[0113] Voice command processing volume statistics: For each time period, count the processing volume of all voice commands within that time period. Record the processing volume of voice commands for each time period to provide a data basis for subsequent resource allocation.

[0114] Resource consumption analysis: Monitor the usage of cloud resources, including CPU, memory, network bandwidth, etc. Archive the resource consumption data for each time period.

[0115] Determination of resource allocation ratio for each time period: Analyze the relationship between the processing volume of voice commands and resource consumption for each time period. Calculate the resource allocation ratio based on the processing volume of voice commands for each time period.

[0116] For the determination of the resource allocation ratio for each time period above, the following specific implementation plan can be adopted:

[0117] Data collection phase: Ensure that the system log can record the processing volume of voice commands and the corresponding resource consumption data for each time period, including CPU usage rate, memory usage, network bandwidth, etc. Use performance monitoring tools to collect data on voice command processing in real time, including processing time, error rate, response speed, etc.

[0118] Data analysis phase: Divide the data according to different time periods, such as peak hours, off-peak hours, and low-night-peak hours, etc. Analyze the relationship between the processing volume of voice commands and resource consumption for each time period to determine whether there are obvious trends or patterns.

[0119] Resource consumption model establishment: Based on the collected data, establish a mathematical model to describe the relationship between the processing volume of voice commands and resource consumption. This may be a linear model, a polynomial model, or a more complex non-linear model. Optimize the model parameters by minimizing the error or maximizing the model fitting degree.

[0120] Resource allocation ratio calculation: According to the established model, calculate the required resource allocation ratio for each time period. For example, if the model shows that the resource consumption during peak hours is twice that during off-peak hours, then the resource allocation ratio during peak hours should be twice that during off-peak hours. Considering the dynamic changes in system load, it may be necessary to adopt a dynamic resource allocation strategy to adjust the resource allocation ratio based on real-time monitoring data.

[0121] It should be noted that by recording the first timestamp of the voice command and the second timestamp of the response command, the processing time of each voice command from reception to response can be accurately measured. This real-time monitoring ability enables the system to track the entire process of voice command processing and promptly detect any delays or anomalies during the processing. Moreover, by analyzing the volume of voice command processing in each time period, a trend graph of voice command processing can be plotted to identify the peak periods of voice command processing. Identifying the peak periods of voice command processing helps predict the peak of system resource requirements, thereby providing a basis for system resource allocation. Based on the identified processing peak periods, the resource allocation ratio can be dynamically adjusted to ensure that more resources are available for processing voice commands during peak hours. By optimizing resource allocation, the risk of system overload can be reduced, and the stability and response speed of the system can be improved.

[0122] Optionally, the supplier cloud includes multiple types of voice models. When determining the resource allocation ratio for each time period based on the volume of voice command processing in each time period, the type of voice model corresponding to each voice command can be determined; based on the volume of voice command processing in each time period and the type of voice model corresponding to each voice command, the resource allocation ratio for each type of voice model in each time period in the supplier cloud can be determined.

[0123] It should be noted that regarding the above content, the following specific implementation solutions can be adopted:

[0124] Time period division: Divide the time into multiple consecutive small time periods, such as every minute, every hour, or every day. Ensure that each voice command and its corresponding model processing have timestamp records.

[0125] Voice command processing volume statistics: Collect the processing volume of all voice commands in each time period. For each time period, count how many commands are processed by different types of voice models.

[0126] Resource consumption analysis: Monitor the resource consumption of each voice model type, including CPU, memory, network bandwidth, etc. Record the resource consumption data of each model type in each time period.

[0127] Resource allocation ratio calculation: According to the processing volume and resource consumption data of different types of voice models in each time period, calculate the resource requirements of each model type. Based on the processing volume and resource requirements, determine the resource allocation ratio for each type of voice model in each time period.

[0128] It should be noted that by determining the types of speech models corresponding to each speech instruction, the supplier cloud can understand the performance differences of different types of models when processing speech instructions. Combining the speech instruction processing volumes in each time period, the cloud can identify which types of speech models are used more frequently in a specific time period. Based on this information, the cloud can dynamically adjust the resource allocation, giving priority to allocating resources to the model types with higher performance requirements, thereby optimizing the performance of the entire system. During the peak processing period, dynamically adjust the resource allocation ratio and allocate more resources to the speech model types with faster processing speeds. Due to the more reasonable resource allocation, these model types can process speech instructions faster, thereby reducing the processing latency.

[0129] Optionally, the performance adjustment parameter includes a resource allocation parameter. When determining the performance adjustment parameter of the supplier cloud based on each speech instruction and the reply instruction corresponding to each speech instruction, the third timestamp of each speech instruction can be determined; the fourth timestamp of the reply instruction corresponding to each speech instruction can be determined; based on the third timestamp and the fourth timestamp, the response time corresponding to each speech instruction can be determined; based on the response time corresponding to each speech instruction, the resource allocation parameter of the supplier cloud can be determined.

[0130] It should be noted that regarding the above content, the following specific implementation solutions can be adopted:

[0131] Timestamp collection: Determine a third timestamp for each speech instruction, which is usually the time point when the instruction starts to be processed. Determine a fourth timestamp for the reply instruction corresponding to each speech instruction, which is usually the time point when the instruction processing is completed and the reply is sent.

[0132] Response time calculation: The response time is the time interval from when the instruction is received to when the reply is completed. For each speech instruction, calculate the response time as follows:

[0133] Response time = Fourth timestamp - Third timestamp.

[0134] Response time statistics: Statistically analyze the response times of all speech instructions in each time period (such as every minute, every hour). Calculate the average response time, minimum response time, maximum response time, and response time distribution for each time period.

[0135] Resource allocation parameter determination: Analyze the relationship between the response time and the cloud resource consumption (such as CPU, memory, network bandwidth). Define resource allocation parameters, such as CPU utilization rate, memory occupancy rate, response time target, etc. Based on the relationship between the response time and the cloud resource consumption, as well as the resource allocation parameters, formulate a resource allocation strategy.

[0136] It should be noted that the determination of resource allocation parameters can be achieved through the following specific implementation plans:

[0137] Data analysis: Analyze the time series data of response time and resource consumption, and find the correlation between the two. Use statistical analysis methods (such as regression analysis, correlation coefficient calculation, etc.) to determine the quantitative relationship between response time and resource consumption.

[0138] Definition of resource allocation parameters: Set a target CPU utilization rate, such as not exceeding 80%. Set a target memory occupancy rate, such as not exceeding 70%. Set an average response time target, such as not exceeding 500 milliseconds.

[0139] Analysis of the relationship between resource consumption and response time: Based on the collected data, establish a model of response time and resource consumption. Determine the thresholds of resource consumption. When the resource consumption exceeds these thresholds, the response time may increase.

[0140] Formulation of resource allocation strategies: Formulate dynamic resource adjustment strategies. When it is detected that the resource consumption is close to the threshold, automatically increase the resource allocation. Assign different priorities to different voice command processes to ensure that critical tasks (such as emergency commands) are processed first. Manage the cloud resource pool and adjust the resource allocation according to real-time requirements.

[0141] It should be noted that by recording the third timestamp of the voice command and the fourth timestamp of the reply command, the performance of the cloud can be evaluated in real time to ensure the instant response ability of the system. Determining the response time of each voice command helps to identify and optimize the response speed of the system and improve the instantaneity of user interaction. The response time data can be used as an important basis for resource allocation to ensure that more resources are invested during periods with higher processing delays. By analyzing the response time, performance bottlenecks in the system, such as insufficient computing resources and network latency, can be identified and optimized accordingly.

[0142] Optionally, the supplier cloud includes multiple types of voice models. When determining the resource allocation parameters of the supplier cloud based on the response time corresponding to each voice command, the type of voice model corresponding to each voice command can be determined first; based on the response time corresponding to each voice command and the type of voice model corresponding to each voice command, determine the resource allocation parameters of each type of voice model in the supplier cloud.

[0143] It should be noted that the determination of resource allocation parameters can be achieved through the following specific implementation plans:

[0144] Data analysis: Aggregate the response time data of different voice model types. Analyze performance indicators such as the average response time, maximum response time, and minimum response time of different voice model types.

[0145] Resource allocation parameter definition: Define the threshold of CPU occupancy rate according to the model type. Define the threshold of memory occupancy rate according to the model type. Set the average response time target for each model type.

[0146] Resource allocation strategy formulation: Establish a resource demand model to predict the resource consumption of different model types under different loads. Allocate resource priorities according to the model type and response time target. Develop an adaptive adjustment strategy to dynamically adjust resource allocation based on real-time response time and resource consumption.

[0147] It should be noted that by determining the voice model type corresponding to each voice command and its response time, the supplier cloud can identify which models are more efficient in processing voice commands, so as to preferentially allocate resources to models with shorter response times, which can reduce the processing time and thus improve the response speed of the entire system.

[0148] Optionally, the supplier cloud includes voice models for multiple application scenarios, and the performance adjustment parameters of the supplier cloud include the voice models for the corresponding application scenarios in the supplier cloud. When determining the performance adjustment parameters of the supplier cloud based on each voice command and the reply command corresponding to each voice command, the semantic information of each voice command and the reply command corresponding to each voice command can be determined first; based on the semantic information, determine the application scenario of the specified in-vehicle unit; based on the application scenario of the specified in-vehicle unit, determine the voice model for the corresponding application scenario in the supplier cloud.

[0149] It should be noted that the determination of resource allocation parameters can be achieved through the following specific implementation methods:

[0150] Semantic information extraction of voice commands and reply commands: Use speech recognition technology to convert voice commands into text. Perform semantic analysis on the converted text to extract intent, entity, and context information. Use a machine learning model to identify the user's intent, such as asking about the weather, navigation, playing music, etc. Extract key entities from the semantic analysis, such as locations, times, contacts, etc.

[0151] Application scenario recognition: Define the characteristics and keywords of each application scenario, and the application scenarios can include business scenarios, home scenarios, military scenarios, etc.

[0152] Scenario matching: Match the voice commands with predefined scenarios according to the extracted intent and entity information.

[0153] Voice model selection: Build a library containing voice models for different application scenarios in the cloud. Select the most suitable voice model from the model library according to the matched application scenario.

[0154] Further, send the performance adjustment parameters to the supplier cloud. When the supplier cloud makes adjustments based on the performance adjustment parameters, the voice model corresponding to the application scenario in the supplier cloud can be sent to the supplier cloud, so that the supplier cloud can make adjustments based on the voice model corresponding to the application scenario in the supplier cloud.

[0155] It should be noted that identifying the application scenario of the specified in-vehicle unit according to semantic information helps the cloud system provide customized services for users in a specific environment. Application scenario recognition can improve the user experience because users can obtain a voice interaction experience that better meets their needs when using the in-vehicle unit. At the same time, selecting a suitable voice model according to different application scenarios can make the model more suitable for the scenario requirements and improve the performance of the model in a specific scenario.

[0156] It should be noted that in the intelligent cockpit voice interaction service, core capabilities such as voice ASR and NLU are generally built by the vehicle manufacturer relying on the capabilities of suppliers. These capabilities are provided in two ways: offline (provided by the in-vehicle unit side) and online (provided by the cloud side), which work respectively when the in-vehicle network is poor and good. When the voice interaction works in an online manner and the end side calls the cloud capabilities, in the past, the supplier side SDK directly accessed the supplier cloud service, and it was difficult for the host factory to statistically analyze and optimize the usage of voice skills, such as daily usage volume, statement hit rate, usage proportion of each voice skill, unsupported statements with high usage, etc.

[0157] The problem solved by this technical solution is that it is difficult for the vehicle manufacturer to collect, statistically analyze, and optimize the voice skill data in the intelligent cockpit voice interaction function.

[0158] This system adopts a distributed microservices architecture, supports multi-node deployment, supports load balancing and fault tolerance, provides good dynamic scaling ability, can dynamically increase or decrease nodes according to the size of the business volume, and balances the contradiction between operating costs and meeting the dynamic changes in the user request volume.

[0159] To address the aforementioned problems, this technical solution aims to provide a system to achieve the collection, statistical analysis, and optimization of voice skill data. This system does not have any impact on the interaction logic between the in-vehicle voice system and the voice cloud system, and no additional code changes are required. Only the service address used for connecting the end side to the cloud needs to be changed to the address of the transit service. Therefore, it is very convenient, has low invasiveness, and is easy to implement.

[0160] Figure 2 This provides a schematic diagram of the structure of a voice skill data collection and analysis system for this application. Combining Figure 2 , the design of this technical solution is described as follows:

[0161] 1. This system consists of two subsystems: a transfer service and a data platform.

[0162] 2. First, the in-vehicle unit sends voice commands to the transfer service, and the transfer service forwards them to the cloud service of the voice provider. In this process, the transfer service appears to the in-vehicle unit as communicating with the cloud service of the voice provider, and the transfer service plays a transparent role. To the cloud service of the voice provider, the requests it receives are from the in-vehicle unit, and the transfer service also plays a transparent role.

[0163] 3. Then, the cloud service of the voice provider returns various result data to the transfer service, including ASR results, NLU results, and voice skill dialogue results. The transfer service first forwards these results as they are to the in-vehicle unit and then reports a copy to the data platform.

[0164] 4. After receiving various voice skill result data from the transfer service, the data platform will parse, process, and store them in the database. Then, data development engineers can perform statistical analysis, Badcase analysis, etc. based on the stored data. Product managers can discover points that need to be optimized through these data analysis dashboards and iteratively optimize the voice function.

[0165] This technical solution can achieve data statistical analysis of intelligent cockpit voice skill calls, analysis of unsupported statements, and optimization suggestions. The key technical points that need to be protected include:

[0166] Key Point 1: The communication mechanism between this system and the end side;

[0167] Key Point 2: The communication mechanism between this system and the cloud service of the voice provider;

[0168] Other effects: This system adopts a distributed microservices architecture, supports multi-node deployment, supports load balancing and fault tolerance, provides good dynamic scaling ability, and can dynamically increase or decrease nodes according to the volume of business, balancing the contradiction between operating costs and meeting the dynamic changes in the volume of user requests.

[0169] Possible alternative solutions:

[0170] Solution 1: The cloud of the voice provider realizes data statistical analysis work. This method requires purchasing its services and has poor customization capabilities.

[0171] Solution 2: The cloud of the voice provider forwards the cloud voice skill call data to the vehicle manufacturer, and the vehicle manufacturer then conducts statistical analysis based on this data. This method requires the cooperation of the supplier and an additional development fee.

[0172] Figure 3The present application provides a schematic structural diagram of a voice model adjustment device for a supplier cloud. The device includes: a voice command collection unit 301, a reply command collection unit 302, an adjustment parameter determination unit 303, and an adjustment parameter upload unit 304.

[0173] The voice command collection unit 301 collects each voice command sent by a specified vehicle head unit to the supplier cloud;

[0174] The reply command collection unit 302 collects each reply command corresponding to the voice commands sent by the supplier cloud to the specified vehicle head unit;

[0175] The adjustment parameter determination unit 303 determines performance adjustment parameters of the supplier cloud based on each voice command and the reply command corresponding to each voice command;

[0176] The adjustment parameter upload unit 304 sends the performance adjustment parameters to the supplier cloud so that the supplier cloud can perform adjustments based on the performance adjustment parameters.

[0177] The present application provides a voice model adjustment device for a supplier cloud, including:

[0178] At least one processor; and,

[0179] A memory communicatively connected to the at least one processor; wherein,

[0180] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:

[0181] Collect each voice command sent by a specified vehicle head unit to the supplier cloud;

[0182] Collect each reply command corresponding to the voice commands sent by the supplier cloud to the specified vehicle head unit;

[0183] Determine performance adjustment parameters of the supplier cloud based on each voice command and the reply command corresponding to each voice command;

[0184] Send the performance adjustment parameters to the supplier cloud so that the supplier cloud can perform adjustments based on the performance adjustment parameters.

[0185] The present application provides a non-volatile computer storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a computer, the computer is enabled to:

[0186] Collect each voice command sent by a specified vehicle head unit to the supplier cloud;

[0187] Collect the response instructions corresponding to each of the voice instructions sent by the supplier cloud to the specified in-vehicle unit;

[0188] Based on each of the voice instructions and the response instructions corresponding to each of the voice instructions, determine the performance adjustment parameters of the supplier cloud;

[0189] Send the performance adjustment parameters to the supplier cloud so that the supplier cloud can make adjustments based on the performance adjustment parameters.

[0190] In this embodiment, the vehicle can be divided into functional modules according to the above method examples. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0191] In the case of dividing each functional module according to each function, the vehicle can include: XX module, XX module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.

[0192] The vehicle provided in this embodiment is used to execute the above method for adjusting the voice model of a supplier cloud, so the same effects as the above implementation method can be achieved.

[0193] In the case of adopting an integrated unit, the vehicle can include a processing module and a storage module. Among them, the processing module can be used to control and manage the actions of the vehicle. The storage module can be used to support the vehicle to execute mutual program codes and data, etc.

[0194] Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor can also be a combination that realizes computing functions, such as including a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory.

[0195] This embodiment also provides a computer-readable storage medium. Computer program code is stored in the computer-readable storage medium (including but not limited to disk memory, CD-ROM, optical memory, etc.). When the computer program code runs on a computer, it causes the computer to execute the above relevant method steps to implement the method for adjusting the voice model of a supplier cloud provided in the above embodiment.

[0196] This embodiment also provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the above-related steps to implement the voice model adjustment method for the supplier cloud provided in the above embodiment.

[0197] Among them, for the beneficial effects of the above embodiment, reference can be made to the beneficial effects in the corresponding method provided above, which will not be elaborated here.

[0198] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0199] In the embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0200] In the description of the present disclosure, it should be understood that if terms such as "upper", "lower", "front", "rear", "left" and "right" are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated position or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present disclosure.

[0201] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0202] The above are only embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, various modifications and variations can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims of the present disclosure.

Claims

1. A method for adjusting a speech model, characterized in that: Applied to a transfer system, the method comprises: Collect all voice commands sent by the designated vehicle computer to the supplier’s cloud; Collecting reply instructions corresponding to each of the voice instructions sent by the supplier cloud to the designated vehicle computer; Determining a performance adjustment parameter of the provider cloud based on each of the voice commands and a reply command corresponding to each of the voice commands; The performance adjustment parameters are sent to the provider cloud so that the provider cloud makes adjustments based on the performance adjustment parameters.

2. The method according to claim 1, characterized in that The supplier cloud includes multiple types of voice models, the performance adjustment parameters include resource allocation of the multiple types of voice models, and the performance adjustment parameters of the supplier cloud are determined based on each of the voice commands and the reply commands corresponding to each of the voice commands, including: Determine the types of voice models corresponding to the voice commands and the reply commands corresponding to the voice commands; Based on the corresponding types of voice models, determine the resource allocation ratio of multiple types of voice models in the provider cloud.

3. The method according to claim 2, characterized in that The determining, based on the type of the corresponding voice model, resource allocation ratios of multiple types of voice models in the supplier cloud comprises: Based on the type of the corresponding speech model, determining the usage ratio of multiple speech models of the type; Based on the usage ratio, determine the resource allocation ratio of multiple speech models of the type in the provider cloud.

4. The method according to claim 1, characterized in that: The performance adjustment parameters include resource allocation ratios for each time period, and the performance adjustment parameters of the provider cloud are determined based on each of the voice commands and the reply commands corresponding to each of the voice commands, including: Determining a first timestamp for each of the voice commands; determining a second timestamp of the reply instruction; Determining the amount of voice command processing in each time period based on the first timestamp and the second timestamp; Based on the voice command processing volume in each of the time periods, a resource allocation ratio for each of the time periods in the provider cloud is determined.

5. The method according to claim 4, characterized in that The supplier cloud includes multiple types of voice models, and the resource allocation ratio of each time period is determined based on the voice command processing volume in each time period, including: Determining the type of voice model corresponding to each of the voice instructions; Based on the amount of voice command processing in each of the time periods and the type of voice model corresponding to each of the voice commands, the resource allocation ratio of each type of voice model in each of the time periods in the supplier cloud is determined.

6. The method according to claim 1, characterized in that The performance adjustment parameters include resource allocation parameters, and the determining of the performance adjustment parameters of the provider cloud based on each of the voice commands and the reply commands corresponding to each of the voice commands includes: Determining a third timestamp for each of the voice commands; Determining a fourth timestamp of a reply instruction corresponding to each of the voice instructions; Determining a reaction time corresponding to each of the voice commands based on the third timestamp and the fourth timestamp; Based on the response time corresponding to each of the voice commands, a resource allocation parameter of the provider cloud is determined.

7. The method according to claim 6, characterized in that The supplier cloud includes multiple types of voice models, and determining the resource allocation parameters of the supplier cloud based on the reaction time corresponding to each of the voice commands includes: Determining the type of voice model corresponding to each of the voice instructions; Based on the reaction time corresponding to each of the voice commands and the type of voice model corresponding to each of the voice commands, resource allocation parameters for each type of voice model in the provider cloud are determined.

8. The method according to claim 1, characterized in that The supplier cloud includes voice models of multiple application scenarios, and the performance adjustment parameters of the supplier cloud include voice models of corresponding application scenarios in the supplier cloud. The performance adjustment parameters of the supplier cloud are determined based on each of the voice commands and a reply command corresponding to each of the voice commands, including: Determining the semantic information of each of the voice commands and the reply command corresponding to each of the voice commands; Based on the semantic information, determining an application scenario of the designated vehicle computer; Based on the application scenario of the designated vehicle computer, a speech model of the corresponding application scenario in the supplier cloud is determined.

9. A speech model adjustment device, characterized in that: Applied to a transfer system, the device comprises: The voice command collection unit collects the voice commands sent by the designated vehicle computer to the supplier's cloud; A reply instruction collection unit, which collects reply instructions corresponding to each of the voice instructions sent by the supplier cloud to the designated vehicle computer; an adjustment parameter determination unit, which determines a performance adjustment parameter of the provider cloud based on each of the voice commands and a reply command corresponding to each of the voice commands; The adjustment parameter uploading unit sends the performance adjustment parameter to the supplier cloud so that the supplier cloud performs adjustment based on the performance adjustment parameter.

10. A speech model adjustment device, characterized in that: Applied to transit systems, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the speech model adjustment method as described in any one of claims 1-8.