Model iteration method and device, equipment and storage medium

By acquiring voice interaction data from multiple cloud platforms, and using an automatic annotation model to identify and filter high-quality samples for model iteration, the problem of unreliable data quality in vehicle-mounted voice recognition models has been solved, thus improving recognition accuracy.

CN121789650APending Publication Date: 2026-04-03SAIC GM WULING AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing in-vehicle voice recognition models suffer from unreliable data quality and are susceptible to noise contamination when processing voice recognition error samples, which affects recognition accuracy.

Method used

By acquiring voice interaction data from different cloud platforms, an automatic annotation model is used to identify error types, generate preliminary annotation results, and perform preset credibility analysis to select high-quality samples for model iteration.

Benefits of technology

It effectively filters out low-quality, high-risk mislabeled data, avoids noise data contamination, and ensures a steady improvement in the recognition accuracy of the speech processing model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789650A_ABST
    Figure CN121789650A_ABST
Patent Text Reader

Abstract

The invention discloses a model iteration method, device and equipment and a storage medium, and relates to the technical field of vehicle voice interaction, and the method comprises the steps: obtaining to-be-processed voice interaction data from at least two different cloud platforms; performing identification error type judgment on the to-be-processed voice interaction data through the automatic labeling model, and generating a preliminary labeling result; performing preset credibility analysis on the preliminary labeling result to obtain a credibility evaluation value which is used for representing the credibility of the preliminary labeling result; and performing model iteration on the voice processing model according to the credibility evaluation value and the preliminary labeling result. According to the method and the device, through automatic credibility grading of voice interaction error labeling samples from multiple platforms, low-quality and high-risk error labeling data can be effectively filtered, the model is prevented from being polluted by noise data, and the recognition precision of a voice processing model is ensured to be stably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle voice interaction technology, and in particular to a model iteration method, apparatus, device and storage medium. Background Technology

[0002] With the widespread adoption of intelligent connected vehicles, in-vehicle voice assistants have become one of the primary interaction methods for users within the car. However, due to the complex in-vehicle environment (such as background noise, dialects, and ambiguity in short commands) and the high difficulty of semantic understanding, voice recognition often encounters problems such as "not matching any skill" or "mismatched matching" (falsely triggering irrelevant functions), resulting in an inability to accurately recognize user intent and severely impacting the user's in-vehicle interaction experience. In such cases, it is necessary to collect voice recognition error samples (bad cases) to optimize the recognition capabilities of the voice recognition model.

[0003] However, existing in-vehicle voice recognition models suffer from unreliable data quality when processing voice recognition error samples, making them susceptible to noise contamination and affecting recognition accuracy.

[0004] Therefore, how to optimize the recognition accuracy of speech recognition models has become an urgent problem to be solved. Summary of the Invention

[0005] The main objective of this application is to provide a model iteration method, apparatus, device, and storage medium, which aims to solve the technical problem of how to optimize the recognition accuracy of a speech recognition model.

[0006] To achieve the above objectives, this application proposes a model iteration method, which includes: Acquire unprocessed voice interaction data from at least two different cloud platforms; The voice interaction data to be processed is identified by an automatic annotation model to determine the error type and generate preliminary annotation results. A preset credibility analysis is performed on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results; The speech processing model is iterated based on the credibility assessment value and the preliminary annotation results.

[0007] Furthermore, to achieve the above objectives, this application also proposes a model iteration device, which includes: The data processing module is used to acquire voice interaction data to be processed from at least two different cloud platforms; The automatic annotation module is used to identify error types in the voice interaction data to be processed using an automatic annotation model and generate preliminary annotation results. The annotation credibility analysis module is used to perform a preset credibility analysis on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results. The model iteration module is used to iterate the speech processing model based on the credibility evaluation value and the preliminary annotation results.

[0008] In addition, to achieve the above objectives, this application also proposes a model iteration device, which includes: a memory, a processor, and a model iteration program stored in the memory and executable on the processor, the model iteration program being configured to implement the steps of the model iteration method as described above.

[0009] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a program implementing the model iteration method is stored, and the program implementing the model iteration method is executed by a processor to implement the steps of the model iteration method as described above.

[0010] This application provides a model iteration method, apparatus, device, and storage medium. The method includes: acquiring speech interaction data to be processed from at least two different cloud platforms; determining the error types of the speech interaction data to be processed using an automatic annotation model to generate preliminary annotation results; performing a preset credibility analysis on the preliminary annotation results to obtain a credibility evaluation value, which characterizes the credibility level of the preliminary annotation results; and iterating the speech processing model based on the credibility evaluation value and the preliminary annotation results. This application rapidly identifies error types in multi-source cloud platform data through an automatic annotation model and achieves accurate model iteration by filtering high-quality samples through credibility analysis. Therefore, this application can effectively filter low-quality, high-risk mislabeled data by automatically classifying the credibility of speech interaction error samples, avoiding noise data contamination of the model, and ensuring a steady improvement in the recognition accuracy of the speech processing model. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a first flowchart illustrating the first embodiment of the model iteration method of this application; Figure 2 This is a second flowchart illustrating the first embodiment of the model iteration method of this application; Figure 3 This is a flowchart illustrating the second embodiment of the model iteration method of this application; Figure 4 This is a flowchart illustrating the third embodiment of the model iteration method of this application; Figure 5 This is a schematic diagram of the process of the third embodiment of the model iteration method of this application; Figure 6 This is a schematic diagram of the module structure of the model iteration device in an embodiment of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the model iteration method in this application embodiment.

[0014] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0015] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0016] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0017] The main solution of this application is: to acquire speech interaction data to be processed from at least two different cloud platforms; to determine the error type of the speech interaction data to be processed through an automatic annotation model and generate preliminary annotation results; to perform a preset credibility analysis on the preliminary annotation results and obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results; and to iterate the speech processing model based on the credibility evaluation value and the preliminary annotation results.

[0018] Currently, in-vehicle voice recognition development and optimization technologies typically involve pulling voice logs from the cloud, using AI (Artificial Intelligence) models to automatically identify and label bad cases (voice recognition error samples), and then importing the labeled results into the voice model to improve recognition accuracy. However, in practical applications, the following key drawbacks exist: the quality of error samples collected by existing in-vehicle voice recognition models during actual use is low, and high-risk mislabeled data may directly enter the training set, easily polluting the voice recognition model and causing a decrease in the model's recognition accuracy.

[0019] To address this issue, this application improves the recognition accuracy of the speech recognition model by efficiently utilizing and accurately optimizing erroneous sample data from in-vehicle speech recognition systems. Specifically, this application uses an automatic annotation model to quickly identify error types in multi-source cloud platform data and employs credibility analysis to filter high-quality samples for precise model iteration. Therefore, this application can effectively filter low-quality, high-risk mislabeled data through automated credibility grading of voice interaction error samples, preventing noisy data from contaminating the model and ensuring a steady improvement in the recognition accuracy of the speech processing model.

[0020] It should be noted that the execution subject in this embodiment can be a model iteration system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a model iteration device capable of performing the above functions. This embodiment does not specifically limit it in this way. The following uses a model iteration device (hereinafter referred to as an iteration device) as the execution subject to describe this embodiment and the following embodiments.

[0021] Based on this, the embodiments of this application provide a model iteration method, referring to... Figure 1 , Figure 1 This is a first flowchart illustrating the first embodiment of the model iteration method of this application.

[0022] In this embodiment, the model iteration method includes steps S10 to S40: Step S10: Obtain voice interaction data to be processed from at least two different cloud platforms; It is important to understand that the aforementioned voice interaction data to be processed can be log data generated during voice interaction between the iterative device and the corresponding connected vehicle system on at least two different cloud platforms (such as public cloud and private cloud). This voice interaction data to be processed may contain key information such as Automatic Speech Recognition (ASR) text, actual skills performed, user intent tags, and conversation context, and serves as the foundational data for subsequent error identification and model iteration.

[0023] Understandably, existing voice interaction data acquisition suffers from inefficient and unreliable issues. This stems from the fact that current voice recognition systems rely on periodically exporting table files or manually downloading log data, which is susceptible to network fluctuations, leading to data loss or delays and failing to effectively guarantee data integrity. Therefore, in a feasible implementation, referring to... Figure 2 , Figure 2 This is a second flowchart illustrating the first embodiment of the model iteration method of this application. In this embodiment, step S10 may include steps A1 to A3: Step A1: Receive voice interaction log messages from at least two different cloud platforms through a message queue subscription mechanism. The voice interaction log messages have different protocol formats. Step A2: Perform protocol parsing processing on the voice interaction log message to extract key field information, including voice recognition text, execution skill identifier, and user intent label; It is important to understand that the aforementioned message queue subscription mechanism refers to a multi-source real-time data reception mechanism based on the Kafka message queue. In this embodiment, the cloud can push voice interaction log data to a specified message topic, and then the iterative device can consume the data in real time by subscribing to this topic to avoid data loss or delay. For example, this embodiment can build a Kafka message queue system and then store the full volume of connected vehicle voice log data in the cloud (public cloud, private cloud). The cloud backend can push new data to the "raw_data_topic" topic of Kafka daily. Simultaneously, the iterative device can deploy a raw data receiving module and subscribe to the "raw_data_topic" topic to consume voice interaction log messages from different cloud platforms in real time, avoiding the data loss problem caused by relying on timed table export or manual download in existing solutions.

[0024] At this point, the aforementioned voice interaction log messages can be raw log data generated during the vehicle-to-everything (V2X) voice interaction process stored on the cloud platform. These logs may include information such as voice signals, ASR text, executed skills, user intent, device ID, and timestamps. Furthermore, log messages from different cloud platforms have different protocol formats. For example, public clouds may use the HTTP (Hypertext Transfer Protocol) protocol, while private clouds may use the MQTT (Message Queuing Telemetry Transport) protocol.

[0025] Therefore, the aforementioned protocol parsing process refers to the process by which the iterative device, after connecting to the communication protocols (such as HTTP, MQTT, etc.) of different cloud platforms, parses the voice interaction log messages obtained from different cloud platforms and extracts key field information. At this time, the iterative device can obtain the same key data fields representing the core content of the voice interaction from voice log data of different protocol formats, namely the aforementioned key field information, and remove redundant data. In this embodiment, the key field information may include speech recognition text (e.g., ASR text), execution skill identifiers (e.g., unique identifiers for skills such as weather query and navigation included in the speech recognition system), and user intent labels (e.g., pre-set intent classification labels such as "weather query" or "air conditioning control").

[0026] Step A3: Standardize the format of the key field information and add data reception timestamp information to generate voice interaction data to be processed.

[0027] Understandably, the aforementioned format standardization process refers to the process by which the iterative device standardizes the key field information parsed from different cloud platforms according to a preset data structure to eliminate format differences and generate unified structure data (such as unified field names, data types, and format standards) that facilitates subsequent automatic voice type labeling. In this case, the data reception timestamp information can be a time stamp (accurate to milliseconds) added by the iterative device when receiving voice interaction log messages, allowing for subsequent tracking of message loss or delays based on this timestamp information, ensuring data traceability. Finally, the iterative device can merge the format-standardized and timestamped data into a unified structured master table, generating the aforementioned voice interaction data to be processed. This data can then be sent via Kafka to another pre-set consumption topic, "merged_data_topic". Subsequently, the front-end page can subscribe to "merged_data_topic" to load database data in real time, generating an interactive data dashboard that supports filtering, chart display, and trend analysis by text, skill, device identifier, time period, and other dimensions.

[0028] At this point, this embodiment can provide an efficient and reliable method for acquiring voice interaction data from multiple cloud platforms, solving the problems of inefficient data acquisition, easy loss, high latency, and inconsistent formats in existing technologies, and effectively ensuring the integrity, consistency, and traceability of the data to be processed.

[0029] Step S20: The voice interaction data to be processed is identified by an automatic annotation model to determine the error type and generate preliminary annotation results; It is important to understand that the aforementioned automatic annotation model can be a pre-trained AI model used to determine the error type of voice interaction data. Its core function can be to identify whether the data belongs to the "not in the domain" (failed to match any skill) or "domain error" (mistakenly triggering irrelevant functions) badcase type, and output preliminary error sample annotation results (such as annotating as "ASR recognition error" or "intent classification error"). The iterative device can simultaneously input the structured master table data (i.e., the aforementioned voice interaction data to be processed) into the automatic annotation model to output preliminary annotation results while generating the master table data.

[0030] Step S30: Perform a preset credibility analysis on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results; Step S40: Iterate the speech processing model based on the credibility assessment value and the preliminary annotation results.

[0031] It is easy to understand that the aforementioned credibility assessment value can be a numerical value used to quantify the credibility of the preliminary annotation results, with a value range of 0 to 1. The higher the value, the more credible the annotation result. The aforementioned speech processing model refers to the core model used for speech recognition and intent understanding in the vehicle-mounted voice interaction system (i.e., the "large speech model" in the disclosure document). Its performance directly affects the accuracy of voice interaction and needs to be iteratively optimized through high-quality samples. Therefore, in this embodiment, the iterative device can perform credibility screening on the preliminary annotation results based on the credibility assessment value, and then perform effective data iterative optimization of the speech processing model through the screened high-quality annotation samples.

[0032] In summary, existing vehicle-mounted speech recognition models rely on low-quality error samples for iteration, leading to decreased model recognition accuracy. Furthermore, issues such as inefficient data acquisition, lack of annotation quality assessment mechanisms, and a single feedback path further exacerbate model performance degradation. This embodiment addresses these problems by ensuring data integrity through a multi-source cloud platform data access mechanism, rapidly identifying error types using an automatic annotation model, filtering high-quality samples through multi-dimensional credibility analysis, and achieving precise model iteration based on a hierarchical feedback mechanism. Therefore, this embodiment introduces an annotation credibility scoring mechanism to predict the quality of AI annotation results, avoiding blind trust in automated output, effectively filtering low-quality, high-risk mislabeled data, preventing noisy data from contaminating the model, and ensuring a steady improvement in the speech processing model's recognition accuracy. Simultaneously, through Kafka + multi-protocol data format unification + timestamp mechanism, end-to-end data traceability is achieved, facilitating the identification of data interruptions and delays, and improving data acquisition success rate. Furthermore, this embodiment improves the efficiency of error sample processing, reduces ineffective manual intervention, enables continuous optimization of the speech recognition model, and effectively improves recognition accuracy.

[0033] This embodiment provides a model iteration method, which includes: receiving voice interaction log messages from at least two different cloud platforms through a message queue subscription mechanism, wherein the voice interaction log messages have different protocol formats; performing protocol parsing processing on the voice interaction log messages to extract key field information, including speech recognition text, execution skill identifiers, and user intent tags; performing format unification processing on the key field information and adding data reception timestamp information to generate voice interaction data to be processed. An automatic annotation model is used to determine the error type of the voice interaction data to be processed, generating preliminary annotation results; a preset credibility analysis is performed on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results; and a preset credibility analysis is performed on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results. This application uses an automatic annotation model to quickly identify error types in multi-source cloud platform data and uses credibility analysis to filter high-quality samples to achieve accurate model iteration. This embodiment can improve the processing efficiency and accuracy of erroneously labeled samples, while enhancing system observability and operational capabilities. Furthermore, this embodiment can improve the efficiency of erroneous sample processing, reduce ineffective manual intervention, achieve continuous optimization of the speech recognition model, and effectively improve recognition accuracy.

[0034] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.

[0035] Based on the first embodiment, please refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the model iteration method of this application. In this embodiment, step S30 includes steps S31 to S33: Step S31: Extract multi-dimensional credibility influence parameters based on the preliminary annotation results; It is easy to understand that, in this embodiment, the credibility influence parameter can be a multi-dimensional key indicator used to evaluate the credibility of the preliminary annotation results. For example, it may include a first-dimensional parameter based on ASR confidence, a second-dimensional parameter based on intent classification entropy, a third-dimensional parameter based on the similarity of historical similar annotations, and a fourth-dimensional parameter based on contextual logical consistency.

[0036] Therefore, in a feasible implementation, in this embodiment, step S31 may include steps B1 to B5: Step B1: Determine the first dimension parameter based on the speech recognition confidence parameter carried in the speech interaction data to be processed; It is important to understand that the aforementioned first-dimensional parameter can be a reliability influence parameter determined based on the confidence level of speech recognition (ASR) in the voice interaction data to be processed, and its value range can be 0~1. In this embodiment, the first-dimensional parameter can be represented by S1. S1 can directly reflect the accuracy of the ASR recognition result and is the basic indicator for labeling reliability. For example, in this embodiment, S1 = ASR native confidence level. If the ASR does not output a confidence level, S1 is defaulted to 0.3, which is considered a low confidence base score. If the ASR recognized text is empty or semantically fragmented (such as containing only meaningless interjections "ah" or "oh"), S1 can be forced to 0.1 to avoid interference from invalid data.

[0037] Step B2: Obtain the intent classification entropy value corresponding to the preliminary annotation result, and map the intent classification entropy value to the second dimension parameter according to the preset conversion rule; The aforementioned second-dimensional parameter can be a credibility influence parameter obtained by inverse mapping based on the intent classification entropy value. Its value ranges from 0 to 1 and is denoted by S2. The intent classification entropy value can be a numerical value representing the uncertainty of intent classification output during semantic recognition. It reflects the uncertainty of the user intent classification recognition result during speech recognition (the higher the entropy value, the more ambiguous the classification, and the lower the credibility). Therefore, a higher intent classification entropy value indicates a more ambiguous intent classification, and the lower the corresponding value of the second-dimensional parameter S2. In this embodiment, S2 can directly reflect the credibility of the annotation result in intent judgment.

[0038] For example, assuming the intent classification entropy value is H (the theoretical range of entropy value is 0~+∞, but in practical applications, due to the limited number of intent categories, H is usually ≤3), the iterative device can correspondingly map the intent classification entropy value to the second dimension parameter (S2) according to the following preset conversion rules: when H≤0.5 (clear classification), set S2=1.0; when 0.5<H≤2.0 (moderately ambiguous classification), set S2=1-(H-0.5) / (2.0-0.5)=1.333-0.666H (linear inverse mapping); when H>2.0 (extremely ambiguous classification), set S2=0.2 (minimum base score to avoid extreme bias caused by excessively low scores).

[0039] Step B3: Calculate the similarity between the preliminary annotation results and historical similar annotation results to generate similarity parameters, and determine the third dimension parameters based on the historical similar annotation results and the similarity parameters; At this point, the aforementioned third-dimensional parameter can be a reliability influence parameter determined based on the similarity between the preliminary annotation result and historical similar annotation results. Its value range can also be 0 to 1, and it is denoted by S3. The higher the similarity value of S3, the more the annotation result conforms to historical patterns, and the higher its reliability. Furthermore, the aforementioned historical similar annotation results can be annotation data from the system's historical annotation library that has the same "annotation type + intent category" as the currently generated preliminary annotation result (e.g., both are annotation results for "ASR recognition error + weather query intent").

[0040] For example, the iterative device can filter similar data from the system's historical annotation library that matches the "annotation type + intent category" of the current preliminary annotation result, and use this as a historical similar annotation result. Then, the iterative device can use a cosine similarity algorithm to calculate the average similarity between the current annotated text, i.e., the preliminary annotation result, and similar historical data, and represent it as Sim (Sim ranges from 0 to 1). Finally, the iterative device can perform a score mapping, even if S3 = Sim (directly reusing the similarity result without additional conversion). In addition, if a special case occurs, i.e., if the number of available historical similar annotation results is less than the preset number of similar results (e.g., 3), and the corresponding data volume is insufficient, S3 can be defaulted to 0.6 (corresponding to a medium-credibility base score, balancing the impact of data scarcity).

[0041] Step B4: Perform logical consistency analysis on the contextual dialogue information of the voice interaction data to be processed and the preliminary annotation results, and determine the fourth dimension parameter based on the consistency analysis results; Step B5: Summarize the first dimension parameter, the second dimension parameter, the third dimension parameter, and the fourth dimension parameter into a credibility influence parameter.

[0042] It is easy to understand that the above-mentioned contextual dialogue information refers to the dialogue content of the first 3 rounds and the last round in the conversation corresponding to the voice interaction data to be processed, which is used to verify the logical rationality of the annotation results. Therefore, the fourth dimension parameter can be a reliability influence parameter determined based on the logical consistency between the contextual dialogue information of the voice interaction data to be processed and the preliminary annotation results. The value range can also be 0~1, and it is represented as S4. S4 can reflect the fit between the annotation results and the conversation scenario.

[0043] At this point, during the logical consistency analysis of the contextual dialogue information and the preliminary annotation results of the voice interaction data to be processed, the iterative device needs to determine the degree of consistency of the fourth dimension parameter with the semantic, intent, and entity information of the contextual dialogue information and the preliminary annotation results, and determine the second dimension parameter based on the determined degree of consistency.

[0044] For example, when the consistency analysis result is completely consistent (corresponding to no conflict between the initial annotation result and the context semantics, intent, and entity information, such as the user saying "check Beijing temperature" and annotating it as "ASR misidentifies 'Beijing' as 'background'"), the fourth dimension parameter S4 can be set to 1.0; when the consistency analysis result is partially consistent (corresponding to no direct conflict between the initial annotation result and the context, but weak correlation, such as the user saying "check weather" and annotating it as "ASR misidentifies 'temperature' as 'humidity'"), S4 can be set to 0.7; when the consistency analysis result is irrelevant (corresponding to no valid context information to verify the initial annotation result, such as a single-turn isolated dialogue), S4 can be set to 0.5; when the consistency analysis result is logically contradictory (corresponding to a conflict between the initial annotation result and the context, such as the user saying "check flights" but annotating it as "intent classification error: check weather"), S4 can be set to 0.2.

[0045] In summary, this embodiment directly obtains the native confidence score of ASR and verifies its validity to obtain the first dimension parameter; calculates the intent classification entropy value and maps it inversely to obtain the second dimension parameter; calculates the similarity with historical similar annotation results using the cosine similarity algorithm to obtain the third dimension parameter; and analyzes the logical consistency between the context dialogue and the annotation results to obtain the fourth dimension parameter. Finally, the iterative device can classify and summarize the first, second, third, and fourth dimension parameters according to their dimensions to form a complete set of confidence-influencing parameters for subsequent high-precision weighted fusion calculations.

[0046] Step S32: Obtain the dynamic dimension weights corresponding to the current business scenario; Step S33: Perform weighted fusion processing on the credibility influence parameters according to the dynamic dimension weights to obtain the credibility evaluation value.

[0047] It is understandable that the aforementioned dynamic dimension weights can be automatically adjusted by the iterative device or manually adjusted by maintenance personnel based on the current business scenario (e.g., dialect recognition scenario, noisy environment scenario, etc.). This embodiment can use the Analytic Hierarchy Process (AHP) to evaluate the importance of each dimension parameter in the current business scenario and determine the weight allocation for each dimension (e.g., increasing the weight of the historical similarity dimension in a dialect scenario). In this embodiment, the default weights can be set as follows: first dimension parameter weight 0.3, second dimension 0.25, third dimension 0.25, fourth dimension 0.2, and the total weight sum is 1.0. That is, the specific values ​​of the aforementioned dynamic dimension weights in this embodiment can be flexibly adjusted according to actual needs.

[0048] At this point, the iterative device can perform a weighted summation of the multi-dimensional credibility influence parameters based on dynamic dimension weights, so as to integrate the scattered single-dimensional evaluation results into a unified credibility evaluation value. The weighted summation formula used for the credibility evaluation value S can be: S = (S1 × W1) + (S2 × W2) + (S3 × W3) + (S4 × W4).

[0049] In summary, existing technologies lack a quality assessment mechanism for annotation results, relying solely on AI models for automatic annotation, which leads to high-risk mislabeled data contaminating the model and reducing its performance. Furthermore, fixed-weight evaluation cannot adapt to different business scenarios and suffers from insufficient accuracy. This embodiment addresses these issues by quantitatively evaluating annotation result quality across four core dimensions: ASR recognition accuracy, intent classification clarity, historical data fit, and contextual consistency. This avoids blindly trusting AI-generated annotation results. Moreover, the analytic hierarchy process (AHP) is used to determine dynamic dimensional weights suitable for different business scenarios, improving the relevance and accuracy of the evaluation results. This provides a reliable basis for subsequent high-quality sample selection and reduces interference from invalid data on the model.

[0050] This embodiment discloses a method for determining a first-dimensional parameter based on the speech recognition confidence parameter carried in the speech interaction data to be processed; obtaining the intent classification entropy value corresponding to the preliminary annotation result and mapping the intent classification entropy value to a second-dimensional parameter according to a preset conversion rule; calculating the similarity between the preliminary annotation result and historical similar annotation results to generate a similarity parameter, and determining a third-dimensional parameter based on the historical similar annotation results and the similarity parameter; performing a logical consistency analysis on the contextual dialogue information of the speech interaction data to be processed and the preliminary annotation result, and determining a fourth-dimensional parameter based on the consistency analysis result; and summarizing the first-dimensional parameter, second-dimensional parameter, third-dimensional parameter, and fourth-dimensional parameter into a credibility influence parameter. The dynamic dimension weights corresponding to the current business scenario are obtained; the credibility influence parameter is weighted and fused according to the dynamic dimension weights to obtain a credibility evaluation value. This embodiment can extract multi-dimensional credibility influence parameters by comprehensively covering the key factors affecting annotation credibility, then using the analytic hierarchy process (AHP) to determine the dynamic dimension weights, and obtaining an accurate credibility evaluation value through weighted fusion calculation, thereby improving the accuracy of high-quality sample screening.

[0051] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the contents that are the same as or similar to those in the first and second embodiments described above can be referred to the above description and will not be repeated hereafter.

[0052] It is important to understand that existing speech recognition model solutions suffer from a single and rigid feedback mechanism. Regardless of whether the annotation is correct or not, a uniform feedback path is used, and there is a lack of differentiated processing strategies for data with different confidence levels, which restricts the efficiency of model iteration.

[0053] Therefore, in this embodiment, based on the first embodiment and / or the second embodiment, please refer to... Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the model iteration method of this application. In this embodiment, step S40 includes steps S41 to S42: Step S41: Based on the credibility assessment value, perform credibility grading processing on the preliminary annotation results to obtain grading annotation data corresponding to different credibility levels; Step S42: Based on the credibility level, the hierarchical annotation data is transmitted to the speech processing model through the corresponding feedback channel for model iteration.

[0054] It is easy to understand that the above-mentioned credibility grading process can refer to the process by which the iterative device compares the magnitude of the credibility assessment value (S) with the preset grading standard and divides the preliminary labeling results into different credibility levels. In one embodiment, the credibility levels include low credibility, medium credibility and high credibility. Therefore, in this embodiment, the iterative device can clearly divide the credibility levels into three levels: high credibility (corresponding to S≥0.8), medium credibility (corresponding to 0.5≤S<0.8), and low credibility (corresponding to S<0.5).

[0055] At this point, the aforementioned graded annotation data can be a set of preliminary annotation results belonging to different confidence levels after confidence grading processing, and in this embodiment, each level of data can correspond to a different feedback channel. Specifically, the dedicated data transmission channels set up for graded annotation data at different confidence levels may include a first feedback channel (high_confirmed_topic) for real-time iteration of high-confidence data, a second feedback channel (medium_confirmed_topic) for periodic iteration of medium-confidence data, and a channel for back-training of low-confidence erroneous data. Finally, the iteration device can import the graded annotation data into the corresponding feedback channel according to the confidence level, transmit it to the speech processing model, update the model parameters, and complete the iteration; if low-confidence data has been reviewed incorrectly, it is back-trained to the automatic annotation model.

[0056] It is important to understand that existing manual review processes are passive and inefficient. They rely on static spreadsheets, requiring reviewers to examine each item individually, lacking a prioritization mechanism and easily leading to high-value bad cases being overlooked. To address this issue, in one feasible implementation, step S42 may include steps C1-C3: Step C1: Based on the review priority corresponding to the credibility level, sort the hierarchical annotation data to generate a review task sequence; Step C2: Based on the credibility level, determine the corresponding manual review mode to process the review task sequence and generate the review confirmation result; the manual review mode includes at least the multi-auditor cross-confirmation mode corresponding to the low credibility level, the single-auditor confirmation mode corresponding to the medium credibility level, and the sampling review mode corresponding to the high credibility level. Step C3: Based on the credibility level and the review confirmation result, the graded annotation data is transmitted to the speech processing model through the corresponding feedback channel for model iteration.

[0057] It is easy to understand that this embodiment can define different review criteria based on the credibility level: for high-credibility data, its performance in all dimensions is excellent, and the annotation results can be directly accepted; for medium-credibility data, its core dimensions meet the standards and require simple review and confirmation; while for low-credibility data, its annotation has obvious contradictions or insufficient data support and requires in-depth review and correction. In this case, the above review priority can be a pre-set order of manual review for data annotated at different credibility levels, with the priority from high to low being: low credibility level > medium credibility level > high credibility level, ensuring that high-risk, low-quality data is reviewed and corrected first.

[0058] Understandably, the manual review task list generated after sorting the graded labeled data according to review priority, i.e., the above-mentioned review task sequence, allows reviewers to process the data sequentially, preventing high-value bad cases from being overlooked. Specifically, the iterative device can sort all graded labeled data according to the priority of "low confidence > medium confidence > high confidence," placing low-confidence data first and high-confidence data last, forming an ordered review task sequence, which is then displayed to reviewers through a front-end interactive dashboard.

[0059] It should be understood that, in this embodiment, the aforementioned manual review mode can be a review rule pre-designed for graded labeled data of different credibility levels. For example, for the aforementioned three credibility levels, in this embodiment, the manual review mode may include a multi-reviewer cross-confirmation mode (for low credibility levels), a single-reviewer confirmation mode (for medium credibility levels), and a sampling review mode (for high credibility levels). For example, in this embodiment, a multi-reviewer cross-confirmation mode can be used for low credibility data, meaning at least two reviewers independently review the data, and a consensus is required for a valid result; a single-reviewer confirmation mode is used for medium credibility data, meaning only one reviewer is needed; and a sampling review mode can be used for high credibility data, meaning a no-review mode or a sampling review mode is employed, i.e., a sample review is conducted according to a preset ratio, with a default sampling ratio of 10%, and no errors are assumed to indicate all data is correct.

[0060] At this point, the aforementioned review and confirmation result can be the output of the reviewer after verifying the graded labeled data. It can include two types: "Confirmed Correct" and "Confirmed Incorrect." This review and confirmation result can serve as the basis for whether the aforementioned preliminary labeled data enters the model's iteration process or is used for back-training. Accordingly, the reviewer can perform the verification operation of the labeled data on the front-end interface: for data that is confirmed correct, click "Pass" to mark it as a "high-quality training sample"; for data that is confirmed incorrect, click "Reject," which can be used later to record the error type and feed it back to the automatic labeling model for incremental training.

[0061] In this implementation, review priorities can be set according to credibility levels to generate an ordered sequence of review tasks. Three review modes—multi-auditor cross-checking, single-auditor review, and random sampling—are designed for different levels, and data routing is achieved based on review results. Therefore, this embodiment allows auditors to prioritize high-risk, low-credibility data, improving the detection rate of critical bad cases. Furthermore, the differentiated review modes optimize human resource allocation and reduce ineffective review workload. Finally, this embodiment achieves precise data routing based on a combination of review results and levels, further ensuring the quality of samples for model iteration.

[0062] It should be noted that in this embodiment, the model can be iterated by combining the credibility level and the review result. For example, data with no errors in high credibility sampling or data that passes medium credibility review can enter the corresponding feedback channel for iteration; data with low credibility that passes cross-review enters the iteration channel, and data with review errors is returned to the automatic labeling model; data with errors found in high credibility sampling is re-evaluated for credibility level and reviewed according to the corresponding mode.

[0063] It should be understood that in this embodiment, different levels of data can follow different data feedback paths. In one feasible implementation, step C3 may include steps C31 to C33: Step C31: When the confidence level is high confidence, the corresponding hierarchical annotation data is transmitted to the first feedback channel for real-time model iteration of the speech processing model. It should be noted that the aforementioned first feedback channel can be a pre-set dedicated data transmission channel for real-time model iteration of high-confidence graded labeled data. This channel can correspond to a pre-configured Kafka message topic "high_confirmed_topic" in the cloud. The speech processing model can then consume channel data from "high_confirmed_topic" in real time and update its parameters. In this embodiment, for graded labeled data classified as high-confidence (S≥0.8), after manual review or random sampling without errors, the iteration device can automatically push it to the Kafka topic "high_confirmed_topic". The speech processing model can subscribe to this topic in real time, consume the data, import it into the generalized phrasing library, and instantly update the model's intent recognition parameters, achieving real-time iteration.

[0064] Step C32: When the credibility level is medium and the review confirmation result is confirmed as correct, the corresponding hierarchical annotation result is transmitted to the second feedback channel for the periodic model iteration of the speech processing model. It's easy to understand that the aforementioned confirmation information could be feedback from auditors confirming that the labeling results match the actual error type and are logically consistent after reviewing the graded labeling data. The second feedback channel could be a dedicated data transmission channel for periodic model iteration of the medium-confidence graded labeling data, corresponding to a pre-configured Kafka message topic "medium_confirmed_topic". The speech processing model can periodically (e.g., daily) consume data from the "medium_confirmed_topic" channel for model iteration. In other words, after graded labeling data classified as medium-confidence level is confirmed correct by a single auditor, the iteration device can push it to the Kafka topic "medium_confirmed_topic". The speech processing model can then consume the data from this topic in batches according to a preset period (e.g., daily at midnight), integrate it, and update the model's generalization rules to achieve periodic iteration.

[0065] Step C33: If the credibility level is low and the review confirmation result is a confirmation error, the preliminary annotation result is returned to the automatic annotation model for incremental training.

[0066] It is easy to understand that the feedback information obtained by the aforementioned error confirmation reviewers after verifying the graded annotation data indicates that the annotation results do not match the actual error type and contain logical contradictions. In this case, the aforementioned incremental training can refer to feeding back the preliminary annotation results, which are confirmed as errors, to the automatic annotation model. This allows the automatic annotation model to correct its annotation logic based on the error samples, improving the accuracy of subsequent annotations. Specifically, in this embodiment, after two reviewers cross-confirm the errors in the graded annotation data classified as low confidence levels, the iterative device can record the annotation error type (such as ASR recognition error annotation deviation, intention classification error misjudgment, etc.) and feed back the graded annotation result and the corresponding annotation error type to the automatic annotation model. This allows the automatic annotation model to adjust its annotation algorithm logic based on the sample, completing incremental training and improving the accuracy of subsequent annotations.

[0067] Therefore, in this embodiment, high-confidence data can be used in real-time iteration to improve the model's response speed, medium-confidence correct data can be used in periodic iteration to ensure the stability of model updates, and low-confidence erroneous data can be fed back to the automatically labeled model for incremental training, gradually improving the ability to recognize complex scenes. At the same time, this embodiment avoids a large influx of erroneous data into the model at once through hierarchical and time-sharing iteration, reducing the risk of model degradation.

[0068] In summary, referring to Figure 5 The model iteration process in this embodiment will be explained and illustrated. Figure 5 For example, a process diagram. Figure 5 As shown, the iterative device equipped with the vehicle-mounted voice badcase closed-loop processing system can first obtain the voice interaction data to be processed from the cloud data system, corresponding to at least two different cloud platforms (including public cloud and private cloud). In this process, the iterative device can use the Kafka message queue subscription mechanism to achieve multi-source data access. The cloud data can store the full amount of voice log data, and the background will push the newly added data to the specified message topic (raw_data_topic) on a daily schedule.

[0069] The iterative device can be adapted to raw data receiving modules of different cloud platform protocols, i.e. Figure 5 The multi-source voice log access and fusion module shown consumes messages from raw_data_topic, then extracts key fields such as ASR text, actual execution skills, and user intent tags from them. After format normalization, it merges them into a unified structure (such as SQL (Structured Query Language) format) master table data, thus obtaining the above-mentioned voice interaction data to be processed.

[0070] Then, the visualization data analysis and credibility screening module in the iterative device can use the automatic annotation model to identify the error type of the voice interaction data to be processed. That is, the normalized voice interaction data to be processed is input into the automatic annotation model. The automatic annotation model can first determine whether the user command hits the vehicle's preset skill, and then further determine whether there are problems such as "not falling into the domain" or "occurrence error". Finally, it outputs the preliminary annotation results containing information such as error type (such as ASR recognition error, intent classification error) and associated intent category.

[0071] Subsequently, the iterative device can introduce the Confidence Scoring Model (CSM) to perform a preset confidence analysis on the initial annotation results. Based on four core dimensions—ASR confidence, intent classification entropy, similarity to historical bad cases, and consistency with contextual logic—and combined with dynamically allocated dimension weights, a confidence evaluation value is calculated using a weighted summation algorithm.

[0072] Finally, the iterative device can use the human-machine collaborative verification and analysis feedback module to iterate the speech processing model based on the credibility assessment value and the preliminary annotation results. Specifically, the iterative device can classify the preliminary annotation results into different credibility levels based on the credibility assessment value, and then use differentiated review mechanisms and feedback channels for different credibility levels to transmit the verified high-quality annotation data to the speech processing model (i.e.,...). Figure 5 The large speech model shown is used to update the generalization rules and intent recognition parameters of the speech processing model, thereby achieving iterative optimization of the speech processing model.

[0073] Therefore, this embodiment can display data in order of credibility on the front-end dashboard, allowing reviewers to prioritize high-risk samples and improve the detection rate of critical bad cases. It also employs a "three-level admission mechanism (high / medium / low credibility)" and a tiered feedback mechanism to ensure that only fully validated data enters the speech model training process, effectively preventing noisy data from contaminating the model and reducing the risk of model degradation. Simultaneously, it avoids a large influx of erroneous data into the model at once, ensuring a steady improvement in speech recognition accuracy and building a secure and controllable data feedback channel. Finally, this embodiment can also correct the current results using erroneously labeled data and reverse-train the automatic labeling model, gradually improving its ability to recognize complex scenarios such as ambiguous semantics and regional accents, achieving the self-evolution capability of the closed-loop system.

[0074] This embodiment discloses a method for transmitting graded labeled data based on confidence level to a speech processing model for model iteration via corresponding feedback channels. High-confidence graded labeled data automatically enters the first feedback channel, where the speech processing model consumes it in real time and imports it into a generalized statement library. Medium-confidence graded labeled data enters the second feedback channel after review and confirmation, for periodic model iteration. Low-confidence graded labeled data enters the corresponding channel after cross-review and confirmation of correctness; if incorrect, it is returned to the automatic labeling model for incremental training, ultimately achieving accurate iteration of the speech processing model. Simultaneously, the preliminary annotation results are graded according to the credibility assessment value to obtain graded annotation data corresponding to different credibility levels. Credibility levels include low, medium, and high credibility. Based on the review priority corresponding to the credibility level, the graded annotation data is sorted to generate a review task sequence. The review task sequence is processed according to the corresponding manual review mode determined by the credibility level to generate review confirmation results. The manual review modes include at least the multi-reviewer cross-confirmation mode for low credibility level, the single-reviewer confirmation mode for medium credibility level, and the sampling review mode for high credibility level. Based on the credibility level and the review confirmation result, the graded annotation data is transmitted to the speech processing model for model iteration through the corresponding feedback channel. When the credibility level is high, the corresponding graded annotation data is transmitted to the first feedback channel for real-time model iteration of the speech processing model. When the credibility level is medium and the review confirmation result is correct, the corresponding graded annotation result is transmitted to the second feedback channel for periodic model iteration of the speech processing model. When the credibility level is low and the review confirmation result is incorrect, the preliminary annotation results are returned to the automatic annotation model for incremental training. This embodiment classifies data by credibility assessment value and sets up dedicated feedback channels for different levels of data to achieve differentiated transmission and iteration, avoid interference from noisy data, and ensure the accuracy of model iteration. At the same time, combined with review and sorting processing, it realizes refined data management and improves the controllability of the overall iteration process.

[0075] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the model iteration method of this application. Any simple transformations based on this technical concept are all within the protection scope of this application.

[0076] This application also provides a model iteration device, please refer to... Figure 6 , Figure 6 This is a schematic diagram of the module structure of the model iteration device according to an embodiment of this application. In this embodiment, the model iteration device includes: The data processing module 601 is used to acquire voice interaction data to be processed from at least two different cloud platforms; Automatic annotation module 602 is used to determine the error type of the voice interaction data to be processed through an automatic annotation model and generate preliminary annotation results; The annotation credibility analysis module 603 is used to perform a preset credibility analysis on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results. The model iteration module 604 is used to perform model iteration on the speech processing model based on the credibility evaluation value and the preliminary annotation results.

[0077] Optionally, in this embodiment, the labeling credibility analysis module 603 is further used for the step of performing a preset credibility analysis on the preliminary labeling results to obtain a credibility evaluation value, including: extracting multi-dimensional credibility influence parameters based on the preliminary labeling results; obtaining dynamic dimension weights corresponding to the current business scenario; and performing weighted fusion processing on the credibility influence parameters according to the dynamic dimension weights to obtain the credibility evaluation value.

[0078] Optionally, in this embodiment, the annotation credibility analysis module 603 is further configured to: determine a first dimension parameter based on the speech recognition confidence parameter carried in the speech interaction data to be processed; obtain the intent classification entropy value corresponding to the preliminary annotation result, and map the intent classification entropy value to a second dimension parameter according to a preset conversion rule; calculate the similarity between the preliminary annotation result and historical similar annotation results to generate a similarity parameter, and determine a third dimension parameter based on the historical similar annotation results and the similarity parameter; perform logical consistency analysis on the contextual dialogue information of the speech interaction data to be processed and the preliminary annotation result, and determine a fourth dimension parameter based on the consistency analysis result; and summarize the first dimension parameter, the second dimension parameter, the third dimension parameter, and the fourth dimension parameter into a credibility influence parameter.

[0079] Optionally, in this embodiment, the model iteration module 604 is further configured to perform credibility grading processing on the preliminary annotation results according to the credibility evaluation value to obtain grading annotation data corresponding to different credibility levels; and transmit the grading annotation data to the speech processing model for model iteration through the corresponding feedback channel based on the credibility level.

[0080] Optionally, in this embodiment, the credibility level includes low credibility, medium credibility, and high credibility; the model iteration module 604 is further configured to sort the hierarchical annotation data based on the review priority corresponding to the credibility level to generate a review task sequence; process the review task sequence according to the corresponding manual review mode determined by the credibility level to generate the review confirmation result; the manual review mode includes at least the multi-reviewer cross-confirmation mode corresponding to the low credibility level, the single-reviewer confirmation mode corresponding to the medium credibility level, and the sampling review mode corresponding to the high credibility level; and transmit the hierarchical annotation data to the speech processing model for model iteration through the corresponding feedback channel according to the credibility level and the review confirmation result.

[0081] Optionally, in this embodiment, the model iteration module 604 is further configured to: transmit the corresponding hierarchical annotation data to a first feedback channel for real-time model iteration of the speech processing model when the confidence level is high confidence level; transmit the corresponding hierarchical annotation result to a second feedback channel for periodic model iteration of the speech processing model when the confidence level is medium confidence level and the review confirmation result is confirmed as correct; and return the preliminary annotation result to the automatic annotation model for incremental training when the confidence level is low confidence level and the review confirmation result is confirmed as incorrect.

[0082] Optionally, in this embodiment, the data processing module 601 is further configured to receive voice interaction log messages from at least two different cloud platforms through a message queue subscription mechanism, wherein the voice interaction log messages have different protocol formats; perform protocol parsing processing on the voice interaction log messages to extract key field information, wherein the key field information includes speech recognition text, execution skill identifier, and user intent tag; perform format unification processing on the key field information and add data reception timestamp information to generate voice interaction data to be processed.

[0083] The model iteration apparatus provided in this application, employing the model iteration method in the above embodiments, can solve the technical problem of model iteration. Compared with the prior art, the beneficial effects of the model iteration apparatus provided in this application are the same as those of the model iteration method provided in the above embodiments, and other technical features in the model iteration apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0084] This application provides a model iteration device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the model iteration method in Embodiment 1 above.

[0085] The following is for reference. Figure 7 The diagram illustrates a structural schematic suitable for implementing the model iteration device in the embodiments of this application. The model iteration device in the embodiments of this application can refer to a physical device with data storage, data processing, and model running capabilities, typically a mobile terminal such as a cloud server, industrial computer, vehicle local processing unit, or vehicle terminal (e.g., vehicle navigation terminal), or a fixed terminal such as a digital TV or desktop computer. Figure 7 The model iteration device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0086] like Figure 7 As shown, the model iteration device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the model iteration device. The processing unit 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the model iteration device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows model iteration devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0087] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a model iteration program product comprising a model iteration program carried on a computer-readable medium, the model iteration program containing program code for performing the methods shown in the flowcharts. In such embodiments, the model iteration program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the model iteration program is executed by processing device 1001, it performs the functions defined above in the methods of the embodiments disclosed in this application.

[0088] The model iteration device provided in this application, employing the model iteration method in the above embodiments, can solve the technical problem of model iteration. Compared with the prior art, the beneficial effects of the model iteration device provided in this application are the same as those of the model iteration method provided in the above embodiments, and other technical features in this model iteration device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0089] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0090] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0091] This application provides a storage medium having computer-readable program instructions (i.e., a model iteration program) stored thereon, the computer-readable program instructions being used to execute the model iteration method in the above embodiments.

[0092] The storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of the storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0093] The aforementioned storage medium may be included in the model iteration device; or it may exist independently and not be assembled into the model iteration device.

[0094] The aforementioned storage medium carries one or more programs, which, when executed by the model iteration device, enable the model iteration device to iterate.

[0095] Model iteration program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of system, method, and model iterative program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0097] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0098] The readable storage medium provided in this application is a storage medium that stores computer-readable program instructions (i.e., a model iteration program) for executing the above-described model iteration method, which can solve the technical problem of model iteration. Compared with the prior art, the beneficial effects of the storage medium provided in this application are the same as the beneficial effects of the model iteration method provided in the above embodiments, and will not be repeated here.

[0099] The above are only some embodiments of this application and do not limit the scope of the solution of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of this application.

Claims

1. A model iteration method, characterized in that, The method includes: Acquire unprocessed voice interaction data from at least two different cloud platforms; The voice interaction data to be processed is identified by an automatic annotation model to determine the error type and generate preliminary annotation results. A preset credibility analysis is performed on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results; The speech processing model is iterated based on the credibility assessment value and the preliminary annotation results.

2. The model iteration method as described in claim 1, characterized in that, The step of performing a preset credibility analysis on the preliminary annotation results to obtain a credibility evaluation value includes: Based on the preliminary annotation results, multi-dimensional credibility influence parameters are extracted; Obtain the dynamic dimension weights corresponding to the current business scenario; The credibility impact parameters are weighted and fused according to the dynamic dimension weights to obtain the credibility evaluation value.

3. The model iteration method as described in claim 2, characterized in that, The step of extracting multi-dimensional credibility influence parameters based on the preliminary annotation results includes: The first dimension parameter is determined based on the speech recognition confidence parameter carried in the speech interaction data to be processed; Obtain the intent classification entropy value corresponding to the preliminary annotation result, and map the intent classification entropy value to the second dimension parameter according to the preset conversion rule; The similarity between the preliminary annotation results and historical similar annotation results is calculated to generate similarity parameters, and a third dimension parameter is determined based on the historical similar annotation results and the similarity parameters. A logical consistency analysis is performed on the contextual dialogue information of the voice interaction data to be processed and the preliminary annotation results, and the fourth dimension parameter is determined based on the consistency analysis results; The first dimension parameter, the second dimension parameter, the third dimension parameter, and the fourth dimension parameter are summarized into a credibility influence parameter.

4. The model iteration method as described in claim 1, characterized in that, The step of iterating the speech processing model based on the credibility assessment value and the preliminary annotation results includes: The initial annotation results are processed according to the credibility assessment value to obtain tiered annotation data corresponding to different credibility levels; Based on the credibility level, the graded labeled data is transmitted to the speech processing model through the corresponding feedback channel for model iteration.

5. The model iteration method as described in claim 4, characterized in that, The credibility levels include low credibility, medium credibility, and high credibility; The step of transmitting the graded labeled data to the speech processing model for model iteration through the corresponding feedback channel based on the credibility level includes: Based on the review priority corresponding to the credibility level, the hierarchical labeled data is sorted to generate a review task sequence; The manual review mode is determined according to the credibility level to process the review task sequence and generate the review confirmation result; the manual review mode includes at least the multi-reviewer cross-confirmation mode corresponding to the low credibility level, the single-reviewer confirmation mode corresponding to the medium credibility level, and the sampling review mode corresponding to the high credibility level. Based on the credibility level and the review confirmation result, the graded annotation data is transmitted to the speech processing model through the corresponding feedback channel for model iteration.

6. The model iteration method as described in claim 5, characterized in that, The step of transmitting the graded annotation data to the speech processing model for model iteration through the corresponding feedback channel based on the credibility level and the review confirmation result includes: When the confidence level is high, the corresponding hierarchical annotation data is transmitted to the first feedback channel for real-time model iteration of the speech processing model. When the credibility level is medium and the review confirmation result is confirmed as correct, the corresponding hierarchical annotation result is transmitted to the second feedback channel for the periodic model iteration of the speech processing model. If the credibility level is low and the review confirmation result is a confirmation error, the preliminary annotation result is returned to the automatic annotation model for incremental training.

7. The model iteration method as described in claim 1, characterized in that, The step of acquiring voice interaction data to be processed from at least two different cloud platforms includes: Voice interaction log messages are received from at least two different cloud platforms via a message queue subscription mechanism, and the voice interaction log messages have different protocol formats. The voice interaction log messages are parsed to extract key field information, which includes speech recognition text, execution skill identifiers, and user intent tags. The key field information is formatted and standardized, and data reception timestamp information is added to generate voice interaction data to be processed.

8. A model iteration device, characterized in that, The model iteration device includes: The data processing module is used to acquire voice interaction data to be processed from at least two different cloud platforms; The automatic annotation module is used to identify error types in the voice interaction data to be processed using an automatic annotation model and generate preliminary annotation results. The annotation credibility analysis module is used to perform a preset credibility analysis on the preliminary annotation results to obtain a credibility evaluation value, which is used to characterize the credibility of the preliminary annotation results. The model iteration module is used to iterate the speech processing model based on the credibility evaluation value and the preliminary annotation results.

9. A model iteration device, characterized in that, The device includes: a memory, a processor, and a model iteration program stored in the memory and executable on the processor, the model iteration program being configured to implement the steps of the model iteration method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a model iteration program, which, when executed by a processor, implements the steps of the model iteration method as described in any one of claims 1 to 7.