Hybrid expert model negative feedback iteration method and device, storage medium and program product

By employing cross-modal scene semantic analysis and negative feedback-driven targeted incremental training, the update lag and stability issues of hybrid expert models in vehicle-road-cloud collaborative environments are resolved, enabling rapid and accurate model iteration and improved stability.

CN122087614APending Publication Date: 2026-05-26TUS CLOUD CONTROL (BEIJING) TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TUS CLOUD CONTROL (BEIJING) TECH LTD
Filing Date
2025-12-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing hybrid expert model architectures struggle to achieve multi-source data fusion in vehicle-road-cloud collaborative environments, resulting in delayed updates of abnormal sub-models and insufficient model stability, making it difficult to meet the rapid iteration requirements in real-time scenarios.

Method used

Scene label information is generated through cross-modal scene semantic analysis, expert sub-models are dynamically scheduled, responsible expert sub-models are identified by combining negative feedback data, and targeted optimization and updates are carried out by parameter freezing and incremental training to build a closed-loop adaptive update system.

Benefits of technology

It enables rapid and accurate iteration of hybrid expert models in dynamic environments, improves scene understanding and overall decision-making performance, ensures model stability and response speed, and reduces the risk of global parameter drift.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087614A_ABST
    Figure CN122087614A_ABST
Patent Text Reader

Abstract

The invention provides a hybrid expert model negative feedback iteration method and device, a storage medium and a program product, and the method comprises the steps: carrying out the cross-modal scene semantic analysis of vehicle end feedback data collected from a vehicle end, roadside sensing data of a roadside device, and macroscopic traffic data of a cloud system, and generating corresponding scene label information; based on the scene label information and the macroscopic traffic data, determining an activation weight of each expert sub-model in the hybrid expert model; negative feedback data in the vehicle running process are collected, and in combination with the influence degree of the target expert sub-model in the negative feedback scene, a responsibility expert sub-model needing to be optimized is determined; and performing directional incremental training on the responsible expert sub-model based on the negative feedback data, performing comprehensive evaluation on the expert sub-model after incremental training, and executing parameter replacement or version updating when an updating condition is met. According to the method, rapid and accurate local optimization is realized while the overall reliability of the model is ensured, and global parameter drift is effectively inhibited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, and in particular to a hybrid expert model negative feedback iterative method, device, storage medium and program product. Background Technology

[0002] With the development of vehicle-road-cloud integration technology, cloud-based driving models are gradually evolving from traditional static deployment to dynamic updates, and are beginning to introduce the Mixture of Experts (MoE) architecture. By setting up multiple expert sub-models to handle different driving tasks or scenarios, the aim is to improve decision-making accuracy.

[0003] However, existing hybrid expert model architectures have shortcomings in multi-scenario adaptation and feedback-driven optimization. Current MoE architectures in practical applications often focus on data processing for a single objective. When a certain expert sub-model continues to perform abnormally in a specific scenario, it usually still relies on retraining the entire architecture, lacking a targeted optimization mechanism for that expert sub-model. This results in high model update costs and slow response times, making it difficult to meet the rapid iteration requirements of real-time scenarios.

[0004] Furthermore, in vehicle-road-cloud collaborative application scenarios, it is necessary to process multi-source information such as roadside perception data, macro traffic data, and vehicle feedback data simultaneously. However, existing model update strategies mostly rely on offline full-scale iterative updates or online small-scale parameter adjustments. The former has a long update cycle and slow response, while the latter, although improving the response speed, lacks an effective stability control mechanism, which can easily cause overall model parameter drift. It is difficult to achieve a reasonable balance between update efficiency and model stability, thus restricting the model's adaptability in complex dynamic scenarios.

[0005] Therefore, there is an urgent need for a cloud-based driving model update solution that can achieve multi-source data fusion, targeted model iteration, and stable and controllable updates in a vehicle-road-cloud collaborative environment. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this application provides a hybrid expert model negative feedback iteration method, device, storage medium, and program product, which at least solves the problem that existing hybrid expert models are unable to accurately locate and directionally iterate abnormal sub-models based on multi-source data under vehicle-road-cloud collaborative conditions, resulting in update lag and insufficient model stability.

[0007] To achieve the above objectives and other advantages, some embodiments of this application provide the following aspects:

[0008] In a first aspect, some embodiments of this application provide a hybrid expert model negative feedback iterative method, including:

[0009] Cross-modal scene semantic analysis is performed on vehicle-side feedback data collected from vehicles, roadside perception data from roadside equipment, and macro-traffic data from cloud systems to generate corresponding scene label information;

[0010] Based on the scene label information and the macro traffic data, the activation weights of each expert sub-model in the hybrid expert model are determined, and the target expert sub-model is scheduled online according to the activation weights.

[0011] Collect negative feedback data during vehicle operation to reflect the model decision-making bias, conduct a comprehensive analysis on the degree of influence of the negative feedback data on each of the target expert sub-models, and determine at least one responsible expert sub-model that needs to be optimized and updated from the target expert sub-models;

[0012] Parameter freezing control is performed on other target expert sub-models besides the responsible expert sub-model, and the current model parameters of the responsible expert sub-model are used as initial parameters. Targeted incremental training is performed on the responsible expert sub-model based on the negative feedback data.

[0013] Based on a preset performance evaluation mechanism, the running performance of the responsible expert sub-model after incremental training is comprehensively evaluated. When the evaluation result meets the preset update conditions, the responsible expert sub-model is subjected to parameter replacement or version update. When the evaluation result does not meet the preset update conditions, data supplementation or rule optimization instructions are output to re-enter the negative feedback-driven adaptive iteration process.

[0014] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising:

[0015] One or more processors; and a memory storing computer program instructions that, when executed, cause the processors to perform the hybrid expert model negative feedback iterative method as described above.

[0016] Thirdly, some embodiments of this application also provide a computer-readable storage medium having a computer program and / or instructions stored thereon, which, when executed by a processor, implement the hybrid expert model negative feedback iterative method as described above.

[0017] Fourthly, some embodiments of this application also provide a computer program product, including a computer program and / or instructions that, when executed by a processor, implement the hybrid expert model negative feedback iterative method as described above.

[0018] Compared with existing technologies, the solution provided in this application integrates vehicle-side feedback data, roadside perception data, and macro-level traffic data in a vehicle-road-cloud collaborative environment to construct a unified semantic analysis and expert scheduling mechanism for driving scenarios. This enables the hybrid expert model to select the most suitable expert sub-model to participate in inference based on real-time scenario characteristics. Using negative feedback data generated during vehicle operation, an impact analysis is performed on the target expert sub-models participating in inference, identifying the expert sub-models that make the main contribution to decision-making bias. Then, a parameter freezing combined with incremental learning approach is used to perform targeted training only on the responsible expert sub-model. This on-demand, real-time iterative model update strategy not only avoids the high computational cost and long-term delays caused by full model retraining but also effectively suppresses global parameter drift, ensuring the performance stability of other expert sub-models. This achieves rapid and accurate local optimization while maintaining the overall reliability of the model. Based on the above mechanism, this application can form a closed-loop adaptive update system oriented towards scenario semantics, traffic conditions, and negative feedback signals. This enables the cloud-based driving model to have higher scenario understanding capabilities, more efficient iteration speed, and stronger model stability in dynamic environments, significantly improving the overall decision-making performance in vehicle-road-cloud integrated scenarios. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other implementation methods can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating the hybrid expert model negative feedback iterative method provided in the embodiments of this application;

[0021] Figure 2 This is a schematic diagram of the MoE model architecture provided in an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] Some embodiments of this application relate to a hybrid expert model negative feedback iterative method, see reference Figure 1 As shown, the method may include the following steps:

[0025] Step S1: Perform cross-modal scene semantic analysis on the vehicle-side feedback data collected from the vehicle, the roadside perception data from the roadside equipment, and the macro traffic data from the cloud system to generate corresponding scene label information.

[0026] Specifically, vehicle-side feedback data can include image information collected by the vehicle's built-in sensors, radar point clouds, vehicle status data, and behavior prediction results generated by the vehicle-side model; roadside perception data from roadside equipment can include environmental images, target detection results, trajectory information, and road structure descriptions acquired by roadside perception units such as road cameras, millimeter-wave radar, and lidar; and macroscopic traffic data from the cloud system can include current road network traffic flow, average driving speed, time period indicators, road congestion levels, and information related to traffic conditions such as weather and events.

[0027] To improve the alignment compatibility between different modal data, in the scene understanding stage of the MoE model, this embodiment uses pre-trained modal encoders to extract features from the above three types of data: vehicle-side feedback data encoders are used to embed and represent the temporal information of vehicle behavior to generate vehicle-side dynamic feature vectors; roadside perception input encoders are used to extract semantics from images, point clouds, or multi-target information to generate roadside environment feature vectors; and macroscopic traffic data encoders are used to embed structured traffic situation fields to generate macroscopic traffic state feature vectors.

[0028] After obtaining the three types of feature vectors mentioned above, a cross-modal semantic fusion model (such as a cross-modal Transformer based on cross-attention mechanism or a visual language model VLM) is used to align and jointly encode the vehicle dynamic features, roadside environment features, and macro-traffic state features. This allows the model to simultaneously focus on vehicle behavior, the surrounding environment, and the overall traffic situation, forming a unified scene semantic vector that reflects the current driving state. Finally, a scene semantic classifier is used to perform scene recognition on the scene semantic vector to obtain corresponding scene label information, which is used to characterize the specific scene category to which the current driving state belongs, such as complex intersection conflict scenarios, abnormal weather scenarios, ambiguous traffic sign scenarios, three-vehicle parallel or queuing congestion scenarios, pedestrian crossing scenarios, etc., thereby providing accurate scene basis for expert sub-model scheduling and incremental iteration.

[0029] In some embodiments, scene label information can be a single scene label or a composite scene label consisting of multiple labels such as weather attributes, road attributes, traffic participant categories, and behavioral semantics, in order to support a detailed depiction of complex driving states.

[0030] Step S2: Based on scene label information and macro traffic data, determine the activation weights of each expert sub-model in the hybrid expert model, and schedule the target expert sub-model online according to the activation weights.

[0031] In a preferred embodiment, step S2 specifically includes:

[0032] Step S201: Based on scene label information, determine the candidate set of expert sub-models corresponding to the current driving scene from the hybrid expert model;

[0033] Step S202: Perform feature encoding on the macroscopic traffic data to generate a macroscopic traffic feature vector that represents the current traffic situation;

[0034] Step S203: Based on the macro traffic feature vector and the feature sensitivity weight vector of each expert sub-model in the expert sub-model candidate set, calculate the matching degree between the current traffic situation and each expert sub-model. The feature sensitivity weight vector is used to characterize the weight combination of the response degree of each expert sub-model to each feature parameter of the macro traffic data.

[0035] Step S204: Normalize the matching degree and generate activation weights for each expert sub-model to determine the target expert sub-model participating in the current driving scenario reasoning.

[0036] Specifically, in the expert routing phase of the MoE model, refer to Figure 2 As shown, the system selects a set of expert candidates from the complete expert sub-model cluster based on scene label information. For example, when the scene label is characterized as "intersection congestion + irregular traffic participants", expert sub-models closely related to the scene, such as "complex conflict identification" and "key target identification", can be selected first, thereby limiting the scope of participation in weight calculation and improving scheduling efficiency.

[0037] Macroscopic traffic data is characterized by feature encoding to obtain a macroscopic traffic feature vector F=[f1,f2,…,f k These characteristics can include traffic flow density, average vehicle speed, time-of-day attributes (such as peak / off-peak), road grade, and the impact of road segment events.

[0038] For each expert submodel E in the expert submodel candidate set i 'i' is used to identify the i-th expert sub-model in the expert sub-model set, and its preset feature sensitivity weight vector W is called. i =[w i1 ,w i2 ,…,w ikThe value ] represents the expert's response weights to different macroscopic traffic features. Based on this, the matching degree between the expert sub-model and the current scenario is calculated as follows:

[0039]

[0040] Among them, W i represents the feature sensitivity weight vector; k represents the total number of dimensions of the macro-traffic feature vector; F represents the macro-traffic feature vector; Represents the expert sub-model E i The degree of fit with the current traffic situation; the higher the value, the more suitable it is for participating in reasoning in the current scenario.

[0041] The matching degree is normalized to obtain the activation weights corresponding to each expert sub-model. In one specific implementation, the system can use the Softmax function combined with a temperature parameter α to normalize the matching degree, so that the weight distribution can show stronger or weaker concentration depending on the magnitude of α, thereby realizing an online selection mechanism that gives higher weights to experts with high matching and lower weights to experts with low matching. Normalized activation weights The expression is as follows:

[0042]

[0043] Where α represents the temperature coefficient of the Softmax function; n represents the number of expert sub-models in the hybrid expert model; This represents the normalized activation weights.

[0044] A normalization algorithm is used to map the matching degree values ​​of different expert sub-models to... A continuous weight space within the interval ensures that the sum of all activation weights equals 1. This normalization process not only highlights expert sub-models that better fit the current traffic situation but also preserves the low-weight participation of other expert sub-models, thus forming a continuous and adjustable expert scheduling mechanism based on real-time traffic characteristics. After generating activation weights, the online routing module of the hybrid expert model can dynamically configure the participation ratio of expert sub-models according to the magnitude of each activation weight. This allows for rapid adaptation to the current driving scenario through weight fine-tuning rather than model retraining, providing expert collaborative input that matches the traffic situation for subsequent inference processes.

[0045] In a preferred embodiment, step S2 further includes:

[0046] The system obtains the performance evaluation information of each expert sub-model under similar historical driving scenarios, and introduces a preset balance coefficient to fuse the performance evaluation information with the matching degree to generate a comprehensive evaluation result for each expert sub-model. The comprehensive evaluation result is used as the updated matching degree for normalization processing.

[0047] Performance evaluation information may include performance indicators such as the expert sub-model's recognition accuracy, decision success rate, and risk avoidance rate in similar past driving scenarios, used to characterize the expert sub-model's historical performance in the corresponding scenario. In this embodiment, this performance indicator is denoted as recognition accuracy A. i .

[0048] To achieve a dynamic balance between real-time scene matching performance and the historical performance of expert sub-models, a preset balance coefficient β (where 0 ≤ β ≤ 1) is introduced. By adjusting the size of the balance coefficient, the weight ratio of real-time feature matching degree and historical performance evaluation in the fusion result can be controlled: when β is larger, more emphasis is placed on the adaptability of current traffic features; when β is smaller, more emphasis is placed on the stable performance of expert sub-models in historical scenarios.

[0049] After obtaining the above parameters, the real-time feature matching degree of the expert sub-model will be... Its historical performance evaluation value A i The data is then fused to obtain the comprehensive score corresponding to the expert sub-model. The overall score can be calculated as follows:

[0050]

[0051] Among them, E i Let F be the i-th expert sub-model; F is the macroscopic traffic feature vector; For expert sub-model E i Real-time matching degree with the current traffic situation; A i For expert sub-model E i Performance evaluation values ​​(such as accuracy) under similar historical scenarios; β is a preset balance coefficient; This is a comprehensive score that combines real-time matching accuracy with historical performance.

[0052] After obtaining the comprehensive scores of each expert sub-model, the comprehensive scores are further normalized using the Softmax function. This process maps the comprehensive scores of all expert sub-models to activation weights, ensuring that the weight values ​​satisfy a probability distribution characteristic where the sum is 1. The activation weights can be calculated using the following formula:

[0053]

[0054] in, α represents the activation weights corresponding to the expert sub-model Ei; n represents the number of candidate expert sub-models; α is the temperature coefficient of the Softmax function. The larger α is, the more concentrated the weights are on the experts with higher scores. The overall score for the expert sub-model.

[0055] Through the processing in steps S201-S204 above, the activation weights of each expert sub-model in the current driving scenario can be calculated based on the macroscopic traffic feature vector and the feature sensitivity weight vectors of each expert sub-model. Unlike inference methods based on fixed experts or static routes, this application does not require retraining the hybrid expert model as a whole for each scenario change. Instead, it dynamically adjusts the activation weights of each expert sub-model, enabling the model to automatically exhibit inference preferences that match the scenario features under different traffic conditions. That is, the model can quickly adapt to real-time scenarios while keeping the parameters unchanged, thus realizing an online optimization logic that can adapt to scenario changes without full retraining and only through weight fine-tuning.

[0056] Step S3: Collect negative feedback data during vehicle operation to reflect the decision-making bias of the model, conduct a comprehensive analysis on the impact of the negative feedback data on each target expert sub-model, and determine at least one responsible expert sub-model that needs to be optimized and updated from the target expert sub-models.

[0057] In a preferred embodiment, step S3 specifically includes:

[0058] Step S301: Perform hierarchical statistics on negative feedback data according to driving scenarios and event types, and generate the negative feedback intensity corresponding to each target expert sub-model based on the distribution of negative feedback events under different negative feedback scenarios.

[0059] During the online negative feedback processing phase of vehicle operation, different types of negative feedback events have significantly different impacts on model performance. For example, manual intervention usually represents a serious decision-making bias, while minor identification errors may only be general noise. Therefore, it is necessary to distinguish between event types to avoid treating errors of different severity levels equally. Furthermore, negative feedback events have a clear scene dependence; the distribution of negative feedback in special driving scenarios such as rainy days, nighttime, and congested roads is often significantly different. Only by performing cluster analysis according to driving scenarios can problem scenarios be identified and it determined whether the performance of a particular expert sub-model degrades under specific scenarios.

[0060] In a preferred embodiment, step S301 specifically includes:

[0061] Step S3011: Cluster the negative feedback data according to driving scenarios to obtain the number of negative feedback scenarios associated with each target expert sub-model;

[0062] Step S3012: In each negative feedback scenario, count the number of negative feedback events associated with each target expert sub-model, and introduce a severity coefficient to characterize the severity of each type of negative feedback event.

[0063] Step S3013: Based on the number of negative feedback events and the corresponding severity coefficients, and combined with the correlation coefficients that characterize the degree of correlation between the current negative feedback scenario and each target expert sub-model, the negative feedback events of each target expert sub-model under different negative feedback scenarios are weighted and summarized to generate the negative feedback intensity corresponding to each target expert sub-model.

[0064] In this embodiment, different driving scenarios exhibit significantly different risk characteristics and perception challenges, such as visual degradation in rainy weather, insufficient lighting at night, multi-agent interactions at complex intersections, and high-speed movement characteristics on highways. Simply accumulating all negative feedback events directly would not only lead to a mixture of logical biases from different sources but also make it difficult for the system to accurately determine in which driving scenario a particular expert sub-model exhibits performance deficiencies. Therefore, scenario-level clustering of the negative feedback data is necessary to classify and organize negative feedback events semantically.

[0065] Specifically, based on the scene labels and environmental semantic features (such as weather, road type, lighting, traffic flow density, etc.) contained in the negative feedback data, a pre-defined scene clustering model is used to divide all negative feedback samples into several driving scene clusters, so that negative feedback events within the same cluster correspond to driving scenes with consistent semantics and similar environmental features. For example, negative feedback events generated under rainy and weak visual conditions are clustered into the "rainy scene cluster", negative feedback events related to intersection conflicts are clustered into the "complex intersection scene cluster", and negative feedback events related to high-speed merging or lane changing are clustered into the "high-speed merging scene cluster".

[0066] After completing scene clustering, and combining the expert sub-model call information recorded by the MoE core routing module during the online inference process, the number of times each target expert sub-model participated in inference and triggered negative feedback in each driving scene cluster is statistically analyzed. By identifying the driving scene type corresponding to each negative feedback event and tracing back the expert sub-models scheduled by the system when the event occurred, the statistical analysis of a specific target expert sub-model can be performed. The number of times negative feedback occurs in each driving scenario cluster; the sum of these occurrences is recorded as the number of negative feedback scenarios. ,in Representing the target expert sub-model The number is used to characterize how many different driving scenarios the expert sub-model has experienced decision bias in. For example, in a certain operating cycle, the vehicle experienced three clusters of driving scenarios: "rainy intersection scenario," "nighttime highway scenario," and "traffic jam following scenario." If the expert sub-model... If an expert submodel participated in inference and triggered prediction errors in the "rainy intersection scenario," participated in inference and triggered emergency braking in the "nighttime highway scenario," but did not generate negative feedback in the "congested following scenario," then this expert submodel is... Number of negative feedback scenarios =2, to reflect how many different driving scenarios the expert sub-model performs poorly.

[0067] Based on the routing records of the hybrid expert model during the online inference phase, the number of events in which the target expert sub-model participated in inference and triggered negative feedback in this negative feedback scenario was counted, and this number was uniformly recorded as the event count. For example, in a negative feedback scenario of "rainy day + intersection", if the expert sub-model... If the event involved in the reasoning process and triggered two predicted trajectory deviation events and one target miss event, then the event count for this scenario is as follows. =3.

[0068] To ensure that the final calculated negative feedback intensity can distinguish between negative feedback events of different risk levels, this embodiment classifies negative feedback events into multiple event types based on their risk attributes, according to the event labels in the negative feedback data. Examples include "emergency braking trigger," "excessive predicted trajectory deviation," "missed target detection," "lane line recognition failure," "slow or fast speed decision," and "unreasonable obstacle avoidance strategy." Subsequently, a corresponding severity coefficient is preset for each event type. ,parameter This is an index for negative feedback event types. Severity coefficient. This is used to characterize the potential impact of such events on decision security or strategy quality, with high-risk events corresponding to larger severity coefficients and low-risk events corresponding to smaller severity coefficients.

[0069] After obtaining the severity coefficients of different negative feedback events, the number of events... Severity coefficient of event type Combined with the correlation coefficient, which represents the degree of correlation between the negative feedback scenario and the target expert sub-model. The negative feedback impact of each objective expert sub-model under different negative feedback scenarios is weighted and summarized. Therefore, the [number]th [model] can be calculated. The overall negative feedback strength of the objective expert sub-model in all negative feedback scenarios The corresponding metric relationship can be expressed as:

[0070]

[0071] in, Indicates the first The number of negative feedback scenarios associated with each expert sub-model; This indicates that the expert sub-model is in the first... The total number of negative feedback events triggered in each negative feedback scenario; Indicates the first The negative feedback scenario for the first The correlation coefficient of each expert sub-model is used to characterize the degree of participation and responsibility of the expert sub-model in this scenario. This represents the severity coefficient corresponding to different event types.

[0072] Furthermore, in order to more accurately assess the degree of responsibility of different expert sub-models in various negative feedback scenarios, for each negative feedback scenario... With each target expert sub-model Calculate its correlation coefficient This coefficient comprehensively describes the degree of dependence, capability fit, and historical performance between the expert sub-model and a specific scenario. Its calculation process is as follows:

[0073] Using the expert activation weights output by the MoE core routing module during the online inference phase, for scenarios belonging to negative feedback... Statistical analysis was performed on all samples to calculate the expert sub-model. The average activation probability in this scenario reflects the user's actual engagement level. This average activation probability can be expressed as:

[0074]

[0075] in, For the first An expert sub-model in the scene The average activation probability in; For the sample Routed to the Activation weights of each expert sub-model; For negative feedback scenarios The number of samples.

[0076] Based on the similarity between scene semantic feature vectors and expert capability feature vectors, the task fit of the expert sub-model in the scene is evaluated. The scene can be obtained by calculating the cosine similarity between the two. With expert sub-model Static matching degree between :

[0077]

[0078] in, For the scene semantic feature vectors; For expert sub-model The capability feature vector.

[0079] To further reflect the historical performance of the expert sub-model in similar scenarios, historical performance evaluation metrics are introduced, such as the expert's precision and recall in similar scenarios, as an experience enhancement term. This term can be expressed as:

[0080]

[0081] in, Expert Sub-model In the scene Or historical performance metrics in similar scenarios; The accuracy rate is based on historical data statistics.

[0082] Based on preset weighting coefficients, the three types of indicators are linearly weighted to generate the final correlation coefficient between the scenario and the expert. :

[0083]

[0084] in, Let be the weight coefficient, and satisfy... .

[0085] Step S302: Based on the negative feedback intensity and the performance of the target expert sub-model in similar historical driving scenarios, determine the update priority of each target expert sub-model to identify at least one responsible expert sub-model.

[0086] Negative feedback intensity can characterize the number and severity of errors made by expert sub-models in the current driving scenario. However, this indicator is also affected by factors such as scenario complexity and data distribution fluctuations. Relying solely on negative feedback intensity for judgment can easily misjudge some "inherently difficult scenarios" as "deteriorating model capabilities." Therefore, it is necessary to further incorporate the operational performance of expert sub-models in similar historical driving scenarios to reflect the normal capability level and long-term performance of the expert sub-model. When an expert sub-model exhibits high negative feedback intensity in the current round of negative feedback data, and its historical performance in similar scenarios shows a significant decline, it indicates that the expert sub-model has experienced genuine capability degradation in this type of scenario, and targeted updates should be prioritized. Conversely, if the negative feedback intensity is high but the historical performance is already low, it may be a normal error caused by the difficulty of the scenario, and direct model updates are not advisable.

[0087] In a preferred embodiment, step S302 specifically includes:

[0088] Step S3021: Obtain the performance evaluation information of each target expert sub-model under similar historical driving scenarios;

[0089] Step S3022: Integrate the negative feedback intensity and operational performance evaluation information corresponding to each target expert sub-model, and introduce a preset balance coefficient to construct an update priority index that comprehensively reflects the update requirements of the target expert sub-model.

[0090] Step S3023: Based on the update priority index, sort the target expert sub-models, and determine at least one expert sub-model with an update priority of a preset number of preceding values ​​from the sorting results as the responsible expert sub-model.

[0091] Specifically, the cloud-based training service obtains performance evaluation information of the expert sub-model on historical scene datasets with similar semantic features to the current negative feedback scenario from the expert sub-model's historical running records and offline validation set statistical results. This performance evaluation information may include metrics such as scene recognition accuracy, recall rate, and policy success rate, which reflect the capability level that the expert sub-model should possess under normal circumstances.

[0092] By introducing a preset balance coefficient, the above-mentioned operational performance evaluation information and its corresponding negative feedback intensity are fused to construct an update priority index that comprehensively reflects the update needs of the expert sub-model. The update priority can be calculated according to the following expression:

[0093]

[0094] in, Indicates the first Update priority metrics for each objective expert sub-model; This indicates the intensity of the negative feedback in the expert sub-model; This represents the performance evaluation metric of the expert sub-model under historically similar driving scenarios; This represents the balance coefficient, used to adjust the weight ratio between the negative feedback strength and historical performance. This represents the maximum value of the negative feedback intensity among all target expert sub-models, used for normalization; This represents the maximum historical performance index among all target expert sub-models, used for normalization.

[0095] In the expression, the first term represents the "proportion of negative feedback intensity," with a larger value indicating a more concentrated recent negative feedback from the expert; the second term represents the "proportion of historical performance degradation," with a larger value indicating worse performance; the two are weighted and combined to form the final update priority. A higher update priority means that the expert sub-model needs more optimization and updates.

[0096] Based on the calculated update priority index, all target expert sub-models are ranked from highest to lowest according to their update demand. The ranking results reflect a comprehensive picture of the severity, frequency, and performance stability of each expert sub-model in the current driving scenario and similar historical scenarios. According to a preset threshold for the number of preceding sub-models (e.g., the top-K), several expert sub-models with update priorities in the preceding interval are selected from the ranking results as the responsible expert sub-models for this iteration. This approach not only ensures that at least one expert sub-model with the most prominent update demand is selected for subsequent targeted optimization, but also supports parallel or phased optimization training of multiple high-priority expert sub-models when multiple expert sub-models perform poorly simultaneously, thereby improving the model's global adaptability and robustness in complex dynamic scenarios.

[0097] Through the processing in steps S301-S302 above, this embodiment can perform hierarchical statistical analysis and multidimensional quantitative analysis on negative feedback data generated during vehicle operation. It integrates the occurrence characteristics, severity, and expert participation of different negative feedback events in different driving scenarios into the evaluation system, thereby constructing a negative feedback intensity index that reflects the true degradation degree of each target expert sub-model. By comprehensively integrating the negative feedback intensity with the operating performance of expert sub-models in similar historical scenarios, and introducing a balance coefficient to coordinate the current degradation status of the model with its historical capability level, it can effectively identify the expert sub-models that most need priority optimization in the current iteration cycle. Through this update priority evaluation mechanism, the model can more accurately locate the expert sub-models causing the current decision-making bias, achieving early perception and rapid response to performance degradation. This ensures that incremental training resources are concentrated on the most critical sub-models, significantly improving the targeting and efficiency of model iteration. Simultaneously, avoiding indiscriminate updates to non-responsible experts helps maintain the stability of the overall model structure, reduces the risk of performance fluctuations in multiple scenarios, and improves the global consistency and reliability after iteration.

[0098] Step S4: Perform parameter freezing control on the target expert sub-models other than the responsible expert sub-model, and use the current model parameters of the responsible expert sub-model as the initial parameters to perform targeted incremental training on the responsible expert sub-model based on the negative feedback data.

[0099] After determining at least one expert sub-model as the responsible expert sub-model in step S3, parameter freezing control is performed on the other target expert sub-models that are not selected. That is, all trainable parameters of these expert sub-models are locked in a read-only state, so that they do not participate in gradient backpropagation and parameter updates in subsequent model update processes, thereby avoiding the spread of parameter offset caused by local scenarios to the entire hybrid expert architecture.

[0100] Subsequently, the model parameters of the responsible expert sub-model in the current iteration are used as the initial parameters for incremental training, and a negative feedback sample set is constructed based on the negative feedback data obtained in step S3. This negative feedback sample set includes image features, temporal behavioral features, vehicle state data, and environmental semantic labels corresponding to negative feedback events such as prediction errors, trajectory deviations, and decision anomalies in the current driving scenario. This sample set is used to perform targeted incremental training on the responsible expert sub-model. By performing several iterations of optimization with a small learning rate, the model focuses on correcting feature interpretation errors or decision strategy biases that lead to negative feedback, without requiring full retraining of the entire expert set.

[0101] During training, stability strategies such as gradient pruning, parameter regularization, and freezing of highly sensitive layers can be introduced to prevent overfitting or semantic drift caused by the update process. Through these methods, the responsible expert sub-model can improve its performance in the current driving scenario while maintaining its original capabilities, thus achieving fast, stable, and efficient local optimization of the hybrid expert model.

[0102] Furthermore, under the stability design of this embodiment, through global collaborative tuning and parameter alignment mechanisms, the parameter deviation of the responsible expert sub-model can be effectively limited to within 10%, while ensuring that the accuracy fluctuation of the non-responsible expert sub-model in its corresponding scenario does not exceed ±3%. With the help of the above control strategy, the risk of full-architecture parameter drift commonly found in traditional full-model retraining can be avoided, thereby ensuring the stability and reliability of the overall performance of the hybrid expert model.

[0103] Step S5: Based on the preset performance evaluation mechanism, comprehensively evaluate the running performance of the responsible expert sub-model after incremental training. When the evaluation result meets the preset update conditions, perform parameter replacement or version update on the responsible expert sub-model; when the evaluation result does not meet the preset update conditions, output data supplementation or rule optimization instructions to re-enter the negative feedback driven adaptive iteration process.

[0104] In a preferred embodiment, step S5 involves comprehensively evaluating the performance of the incrementally trained expert sub-model based on a preset performance evaluation mechanism, specifically including:

[0105] Step S501: Obtain the scene recognition accuracy of the responsible expert sub-model on the validation set and the parameter deviation relative to the baseline model parameters;

[0106] Step S502: Compare the scene recognition accuracy with the preset accuracy threshold to obtain the first evaluation index used to characterize the degree of compliance of the responsible expert sub-model in terms of accuracy.

[0107] Step S503: Compare the parameter deviation with the preset deviation threshold to obtain a second evaluation index used to characterize the degree of compliance of the responsible expert sub-model in the parameter stability dimension;

[0108] Step S504: According to the preset weight ratio, the first evaluation index and the second evaluation index are weighted and fused to generate a comprehensive evaluation score to characterize the effect of this round of incremental training.

[0109] Specifically, the scene recognition accuracy (Acc) of the responsible expert sub-model on the validation set, and the parameter deviation (Δθ) between the sub-model's parameters and the baseline version are obtained. The scene recognition accuracy (Acc) reflects the improvement in model accuracy after incremental training; the parameter deviation (Δθ) is obtained by calculating the L2 distance between the current model parameters and the baseline model parameters, and is used to quantify the magnitude of parameter changes after training to evaluate the stability of model updates.

[0110] To quantify the degree of accuracy achievement, a first evaluation index is constructed to characterize the proportion of model accuracy achieved relative to the threshold. The first evaluation index can be defined as:

[0111]

[0112] in, This indicates the scene recognition accuracy of the responsible expert sub-model on the validation set; This indicates the preset target accuracy rate; This represents the primary evaluation indicator; the closer its value is to 1, the higher the degree of accuracy achieved.

[0113] To facilitate the use of a unified numerical scale in the comprehensive evaluation, a second evaluation index is constructed to characterize the degree to which parameter stability meets the standards. The second evaluation index can be defined as:

[0114]

[0115] in, This indicates the parameter bias of the responsible expert sub-model after this round of updates; This indicates the preset parameter deviation threshold. This represents the second evaluation index. The closer its value is to 1, the smaller the parameter deviation and the higher the stability.

[0116] To balance the importance of accuracy and stability, a preset weighting coefficient ω (0 ≤ ω ≤ 1) is introduced. By weighted fusion of the two evaluation metrics, a comprehensive evaluation score E is generated to characterize the effectiveness of this round of incremental training. Specifically, the calculation expression is as follows:

[0117]

[0118] The system uses preset evaluation criteria to quantitatively determine whether the responsible expert sub-models after this round of incremental training meet the online conditions. This is used to characterize whether the model has achieved the expected improvement in both recognition accuracy and parameter stability. When the comprehensive evaluation score meets the following... If the responsible expert sub-model demonstrates both sufficient improvement in recognition ability and acceptable parameter stability during the current incremental training, then the current iteration is deemed successful, triggering the model version management module to perform parameter replacement or version update operations, so that the cloud-based hybrid expert model can absorb the optimization results of the responsible expert sub-model in real time. Conversely, if the comprehensive evaluation score meets the following criteria: If the current incremental training fails to achieve the expected optimization effect, the current iteration is deemed substandard. The system will not immediately update the online model, but will instead output data supplementation instructions or rule optimization instructions to guide the vehicle to collect more typical negative feedback samples or adjust the scene rule configuration, and re-enter the negative feedback-driven adaptive iteration process. This ensures that the model update process is robust and controllable, and does not lead to erroneous optimization due to insufficient or biased local data.

[0119] Through the aforementioned targeted incremental training mechanism, this embodiment significantly reduces the time overhead required for model iteration. Since only small-scale parameter updates are performed on the responsible expert sub-model, without requiring full retraining of the entire hybrid expert architecture, the training scale is reduced from the previous "full model level" to the "single expert level," thereby drastically reducing the number of parameters involved in the update process and lowering the overall training computation by 60%–70%. In a vehicle-cloud collaborative environment, the iteration cycle triggered by negative feedback events is shortened from days in traditional methods to hours, enabling the model to respond quickly to vehicle-side operational anomalies, making it more suitable for intelligent driving scenarios with high real-time requirements.

[0120] Through steps S501 to S504, after the responsible expert sub-model completes incremental training, its performance can be quantitatively evaluated from both accuracy and parameter stability dimensions. A comprehensive evaluation score is then used to controllably determine the effect of this iteration. This allows the system to strike a balance between performance improvement and model stability, ensuring that parameter replacement or version updates are triggered only when the model maintains overall structural stability while improving its target capabilities. This evaluation mechanism significantly reduces the risk of degradation caused by erroneous updates, improves the reliability and controllability of model iteration, and further enhances the continuous evolution capability of the hybrid expert model in dynamic driving scenarios.

[0121] It should be noted that in the negative feedback closed-loop directional incremental iteration method of this application, although the collection and iteration triggering logic of negative feedback events occur during the vehicle operation phase, the actual training, verification and version update are all executed in the cloud or edge computing environment to ensure the stability and security of the training process.

[0122] The updated expert sub-model was evaluated on a validation set in various scenarios to verify its adaptability in typical challenging driving scenarios. Validation set results show that the recognition and decision-making performance in complex conflict scenarios is improved by approximately 10%–20%; the perception and processing capabilities in abnormal weather scenarios are improved by approximately 15%–25%; and the resolution capability for fuzzy traffic signs is improved by approximately 10%–20%. Because only a small number of parameters are updated for the responsible expert sub-model, without requiring full retraining of the entire hybrid expert architecture, the training scale is reduced from the previous "full model level" to the "single expert level," significantly reducing the number of parameters involved in the update process and lowering the overall training computation by 60%–70%.

[0123] Because this application only performs targeted incremental learning on the responsible expert sub-model and combines lightweight sample construction, parameter freezing, and rapid verification mechanisms, the entire closed-loop update cycle is significantly shortened. Tests show that the overall time from negative feedback triggering to the online deployment of a new expert sub-model version can be reduced from days in traditional methods to hours. This significantly improves the model's response speed to vehicle operational anomalies while ensuring training safety, making it more suitable for intelligent driving scenarios with high real-time requirements.

[0124] In summary, the hybrid expert model negative feedback iteration method provided in this application integrates vehicle-side feedback data, roadside perception data, and macro-level traffic data in a vehicle-road-cloud collaborative environment to construct a unified semantic analysis and expert scheduling mechanism for driving scenarios. This enables the hybrid expert model to select the most suitable expert sub-model to participate in inference based on real-time scenario characteristics. Utilizing negative feedback data generated during vehicle operation, an impact analysis is performed on the target expert sub-models participating in inference, identifying the expert sub-models that make the main contribution to decision bias. Then, using parameter freezing combined with incremental learning, targeted training is performed only on the responsible expert sub-model. This on-demand, real-time iterative model update strategy not only avoids the high computational cost and long-term delays caused by full model retraining but also effectively suppresses global parameter drift, ensuring the performance stability of other expert sub-models. This achieves rapid and accurate local optimization while maintaining the overall reliability of the model. Based on this mechanism, this application can form a closed-loop adaptive update system oriented towards scenario semantics, traffic conditions, and negative feedback signals. This enables the cloud-based driving model to have higher scenario understanding capabilities, more efficient iteration speed, and stronger model stability in dynamic environments, significantly improving the overall decision-making performance in vehicle-road-cloud integrated scenarios.

[0125] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0126] The electronic device includes: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform a hybrid expert model negative feedback iterative method as provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0127] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103, and output device 1104 may be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0128] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0129] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).

[0130] In this embodiment, a computer-readable medium stores a computer program / instruction, which, when executed by a processor, implements a hybrid expert model negative feedback iterative method provided in any one or more of the above embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more computer-readable instructions.

[0131] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.

[0132] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0133] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0134] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0135] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0136] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0137] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0138] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0139] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0140] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A hybrid expert model negative feedback iteration method, characterized in that, The method comprises the following steps: Cross-modal scene semantic analysis is performed on vehicle-end feedback data collected from a vehicle end, roadside perception data of a roadside device, and macro-traffic data of a cloud system to generate corresponding scene label information; Based on the scene label information and the macro-traffic data, the activation weights of each expert sub-model in the hybrid expert model are determined, and the target expert sub-model is scheduled online according to the activation weights; Negative feedback data reflecting model decision bias during vehicle operation is collected, the influence of each target expert sub-model is comprehensively analyzed based on the negative feedback data, and at least one responsible expert sub-model that needs to be optimized and updated is determined from the target expert sub-models; Parameter freezing control is performed on the target expert sub-models other than the responsible expert sub-model, the current model parameters of the responsible expert sub-model are taken as initial parameters, and the responsible expert sub-model is subjected to directional incremental training based on the negative feedback data; Based on a preset performance evaluation mechanism, the running performance of the responsible expert sub-model after incremental training is comprehensively evaluated, and when the evaluation result meets the preset update condition, the parameter replacement or version update of the responsible expert sub-model is performed; when the evaluation result does not meet the preset update condition, data supplement or rule optimization instructions are output to re-enter the adaptive iteration process driven by negative feedback.

2. The hybrid expert model negative feedback iteration method of claim 1, wherein, The step of determining the activation weights of each expert sub-model in the hybrid expert model based on the scene label information and the macro-traffic data, and scheduling the target expert sub-model according to the activation weights, comprises the following steps: Based on the scene label information, a candidate set of expert sub-models corresponding to the current driving scene is determined from the hybrid expert model; The macro-traffic data is feature-encoded to generate a macro-traffic feature vector for representing the current traffic situation; Based on the macro-traffic feature vector and the feature sensitivity weight vector of each expert sub-model in the candidate set of expert sub-models, the matching degree between the current traffic situation and each expert sub-model is calculated, wherein the feature sensitivity weight vector represents the weight combination of the response degree of each expert sub-model to each feature parameter of the macro-traffic data; The matching degree is normalized to generate the activation weight corresponding to each expert sub-model, so as to determine the target expert sub-model participating in the inference of the current driving scene.

3. The hybrid expert model negative feedback iteration method of claim 2, wherein, The step of determining the activation weights of each expert sub-model in the hybrid expert model based on the scene label information and the macro-traffic data, and scheduling the target expert sub-model according to the activation weights, further comprises the following steps: The running performance evaluation information of each expert sub-model in a historical similar driving scene is obtained, and a preset balance coefficient is introduced to fuse the running performance evaluation information and the matching degree to generate a comprehensive evaluation result corresponding to each expert sub-model, which is used as an updated matching degree for normalization.

4. The hybrid expert model negative feedback iteration method of claim 1, wherein, The steps of collecting negative feedback data during vehicle operation to reflect model decision-making bias, comprehensively analyzing the impact of the negative feedback data on each of the target expert sub-models, and determining at least one responsible expert sub-model from the target expert sub-models that needs to be optimized and updated include: The negative feedback data is classified and statistically analyzed according to driving scenarios and event types. Based on the distribution of negative feedback events in different negative feedback scenarios of each target expert sub-model, the negative feedback intensity corresponding to each target expert sub-model is generated. Based on the negative feedback intensity and the performance of the target expert sub-model in similar historical driving scenarios, the update priority of each target expert sub-model is determined to identify at least one responsible expert sub-model.

5. The hybrid expert model negative feedback iteration method of claim 4, wherein, The step of performing hierarchical statistics on the negative feedback data according to driving scenarios and event types, and generating the negative feedback intensity corresponding to each of the target expert sub-models based on the distribution of negative feedback events in different negative feedback scenarios, includes: The negative feedback data is clustered and statistically analyzed according to driving scenarios to obtain the number of negative feedback scenarios associated with each target expert sub-model; In each negative feedback scenario, the number of negative feedback events associated with each target expert sub-model is counted, and a severity coefficient is introduced for each type of negative feedback event to characterize the severity of the negative feedback event. Based on the number of negative feedback events and their corresponding severity coefficients, and combined with the correlation coefficients that characterize the degree of correlation between the current negative feedback scenario and each of the target expert sub-models, the negative feedback events of each of the target expert sub-models under different negative feedback scenarios are weighted and summarized to generate the negative feedback intensity corresponding to each of the target expert sub-models.

6. The hybrid expert model negative feedback iteration method of claim 4, wherein, The step of determining the update priority of each target expert sub-model based on the negative feedback intensity and the performance of the target expert sub-model in similar historical driving scenarios, in order to determine at least one responsible expert sub-model, includes: Obtain the performance evaluation information of each target expert sub-model under historically similar driving scenarios; The negative feedback intensity corresponding to each of the target expert sub-models is fused with the operational performance evaluation information, and a preset balance coefficient is introduced to construct an update priority index that comprehensively reflects the update requirements of the target expert sub-models. Based on the update priority index, the target expert sub-models are sorted, and at least one expert sub-model with an update priority at a preset number of preceding values ​​is determined from the sorting results as the responsible expert sub-model.

7. The hybrid expert model negative feedback iteration method of claim 1, wherein, The step of comprehensively evaluating the performance of the incrementally trained responsible expert sub-model based on a preset performance evaluation mechanism includes: Obtain the scene recognition accuracy and parameter deviation relative to the baseline model parameters of the responsible expert sub-model on the validation set; The scene recognition accuracy is compared with a preset accuracy threshold to obtain a first evaluation index that characterizes the degree to which the responsible expert sub-model meets the accuracy standard. The parameter deviation is compared with a preset deviation threshold to obtain a second evaluation index that characterizes the degree of compliance of the responsible expert sub-model in the parameter stability dimension. According to the preset weight ratio, the first evaluation index and the second evaluation index are weighted and fused to generate a comprehensive evaluation score that characterizes the effect of this round of incremental training.

8. An electronic device, comprising: The electronic device includes: One or more processors; and a memory storing computer program instructions that, when executed, cause the processors to perform the hybrid expert model negative feedback iterative method as described in any one of claims 1-7.

9. A computer readable storage medium having stored thereon a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by the processor, they implement the hybrid expert model negative feedback iterative method as described in any one of claims 1-7.

10. A computer program product comprising computer programs and / or instructions, characterized in that, When the computer program and / or instructions are executed by the processor, they implement the hybrid expert model negative feedback iterative method as described in any one of claims 1-7.