A semantic constraint heterogeneous teacher collaborative decoding method for a vehicle-mounted intelligent cockpit
Patent Information
- Application Number
- CN202611116987.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-07-27
AI Technical Summary
[0003]基于此,有必要提供一种用于车载智能座舱的语义约束异构教师协同解码方法,用以解决车载复杂指令、故障解释和风险提示场景下协同解码不准确的问题,该方法包括:
[0004] Beneficial effects: This method constructs a structured context representation based on the decoding results corresponding to user input commands, vehicle operating status data, cabin environment data, and network status data. It then determines the number of teacher models corresponding to the current decoding state based on the collaborative demand score of the structured context representation. An initial collaborative fragment set is constructed based on the candidate collaborative fragment set generated by the corresponding number of teacher models. A relation matrix is constructed based on the initial collaborative fragment set and the evidence state set. The collaborative fragment coverage difference and evidence coverage difference are calculated based on the relation matrix, and the initial collaborative fragment set is completed based on these differences to obtain the final collaborative fragment set. The final collaborative fragment set is then further decoded, and the resulting final collaborative decoding result is injected into the vehicle-side first language model to generate the final response. This method improves processing consistency, reduces computational overhead, and enhances the accuracy and reliability of the final response.
Smart Images

Figure CN122635568B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle-mounted task decoding technology, and in particular to a semantically constrained heterogeneous teacher collaborative decoding method for vehicle-mounted smart cockpits. Background Technology
[0002] Voice assistants, knowledge-based question answering, fault explanations, and control suggestion generation in in-vehicle intelligent cockpits are crucial components of human-machine interaction and in-vehicle service capabilities in intelligent vehicles. However, existing in-vehicle collaborative reasoning and generation technologies suffer from significant bottlenecks: First, model invocation lacks intelligent scheduling; using a single vehicle-side model is insufficient to handle complex reasoning, fault explanations, and multi-round combined instructions, while statically calling multiple cloud models in parallel leads to redundant computing power and bandwidth, making it impossible to balance response speed and computational efficiency. Second, vehicle-side generation and cloud-assisted reasoning are not effectively decoupled, easily resulting in repeated uploading of context or repeated calls to teacher models in subsequent stages, leading to inconsistent results and increased computational overhead. Third, collaborative fragment generation lacks coverage and balance constraints; no adaptive completion mechanism has been established for evidence state coverage and coverage differences, resulting in insufficient or uneven coverage of vehicle state, environmental state, and safety constraints in collaborative fragments, leading to inaccurate collaborative decoding of complex in-vehicle tasks. Summary of the Invention
[0003] Therefore, it is necessary to provide a semantically constrained heterogeneous teacher collaborative decoding method for in-vehicle intelligent cockpits to solve the problem of inaccurate collaborative decoding in scenarios involving complex in-vehicle commands, fault interpretation, and risk warnings. This method includes: S1: Perform local autoregressive decoding on the acquired user input command, and combine the decoding result, vehicle operating status data, cabin environment data, and network status data to construct a structured context representation; based on the structured context representation, user input command, and current decoding result, extract the evidence state set; and calculate the collaborative requirement score of the structured context representation according to the first preset rule. S2: Determine the number of teacher models corresponding to the current decoding state based on the collaborative demand score; the teacher models are used to generate a set of candidate collaborative fragments based on structured context representation; S3: When the number of teacher models is 1, the candidate collaborative fragment set generated by the corresponding teacher model is used as the initial collaborative fragment set. When the number of teacher models is two or more, a unified representation mapping and semantic deduplication are performed on the candidate collaborative fragment sets generated by all teacher models to obtain the initial collaborative fragment set. S4: Construct a relation matrix based on the initial set of collaborative fragments and the set of evidence states; calculate the collaborative fragment coverage difference and the evidence coverage difference based on the relation matrix; and complete the initial set of collaborative fragments based on the collaborative fragment coverage difference and the evidence coverage difference to obtain the final set of collaborative fragments. S5: Continue decoding the final collaborative fragment set and inject the resulting final collaborative decoding result into the vehicle-side first language model to generate the final response.
[0004] Beneficial effects: This method constructs a structured context representation based on the decoding results corresponding to user input commands, vehicle operating status data, cabin environment data, and network status data. It then determines the number of teacher models corresponding to the current decoding state based on the collaborative demand score of the structured context representation. An initial collaborative fragment set is constructed based on the candidate collaborative fragment set generated by the corresponding number of teacher models. A relation matrix is constructed based on the initial collaborative fragment set and the evidence state set. The collaborative fragment coverage difference and evidence coverage difference are calculated based on the relation matrix, and the initial collaborative fragment set is completed based on these differences to obtain the final collaborative fragment set. The final collaborative fragment set is then further decoded, and the resulting final collaborative decoding result is injected into the vehicle-side first language model to generate the final response. This method improves processing consistency, reduces computational overhead, and enhances the accuracy and reliability of the final response. Attached Figure Description
[0005] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0006] Figure 1 This is a flowchart of a semantic constraint heterogeneous teacher collaborative decoding method for an in-vehicle smart cockpit in this application embodiment. Detailed Implementation
[0007] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0008] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0009] like Figure 1 As shown, this embodiment provides a semantically constrained heterogeneous teacher collaborative decoding method for in-vehicle intelligent cockpits, including: S1: Perform local autoregressive decoding on the acquired user input instructions, and combine the decoding results, vehicle operating status data, cabin environment data, and network status data to construct a structured context representation; based on the structured context representation, user input instructions, and the current decoding results, extract the evidence state set; and calculate the collaborative requirement score of the structured context representation according to the first preset rule.
[0010] In this embodiment, the user input command includes a voice command or a text command, such as: "The tire pressure light is on, can I continue driving now? Also, please navigate to the nearest repair shop." In this embodiment, the onboard first language model is used to perform local autoregressive decoding on the acquired user input command to obtain the current decoding result.
[0011] In this embodiment, the first language model on the vehicle side is a lightweight autoregressive language model deployed on the vehicle's infotainment system.
[0012] In this embodiment, the structured context representation includes a multi-state coupled representation and an evidence state set; the multi-state coupled representation includes a semantic state set, a vehicle state set, an environment state set, a network state set, and a rule state set; the multi-state coupled representation is denoted as... .
[0013] After constructing the structured context representation, evidence states related to the current decoding task are further extracted based on user input commands, the current decoding result, and the state sets in the structured context representation, forming an evidence state set. This evidence state set does not participate in the calculation of the total state scale or multi-state coupling relationships of the structured context representation, but is used for constructing the relationship matrix between the initial collaborative fragment set and the evidence state set, calculating coverage differences, and verifying the consistency of the continued decoding results.
[0014] Specifically, the semantic state set is obtained by performing intent recognition and slot extraction on user input commands and current decoding results; the vehicle state set is obtained by reading vehicle operating status data output by the vehicle bus or diagnostic interface; the environmental state set is obtained by information output by positioning, map, and cockpit sensors; the network state set is obtained by network monitoring data output by the vehicle communication module; the rule state set is obtained by matching with a preset rule base; and the evidence state set is obtained by filtering key states related to the current decoding position and usable for verifying candidate decoding results from the multi-state coupled representation.
[0015] The semantic state set, vehicle state set, environment state set, network state set, and rule state set obtained based on the above user input instructions are shown in Tables 1 to 5. Based on the user input instructions, the current decoding result, and each state set, an evidence state set is further extracted. The evidence state set is used for subsequent relation matrix construction and evidence consistency verification.
[0016] Table 1 Semantic State Set ; Table 2 Vehicle State Set ; Table 3 Set of Environmental States ; Table 4 Network State Set ; Table 5. Set of Rule States ; Specifically, the step of calculating the collaborative requirement score of the structured context representation according to the first preset rule includes: Based on the predicted probabilities of each candidate output in the decoding results, the decoding prediction entropy is calculated, and the prediction probability interval between the two candidate outputs with the highest prediction probabilities is also calculated. The semantic uncertainty is then calculated based on the decoding prediction entropy and the prediction probability interval, using the following formula: ; ; ; in, Indicates semantic uncertainty. This represents the decoding prediction entropy. Indicates the predicted probability interval. This represents a constant (taken as 0.01) to prevent the denominator from being zero. Indicates the first The predicted probabilities of each candidate output are obtained by the onboard first language model based on the current decoding results. Indicates the number of candidate outputs. This represents the predicted probability of the candidate output with the highest predicted probability. This represents the predicted probability of the candidate output with the second-highest predicted probability.
[0017] The total state size is obtained by summing the moduli of each state set in the structured context representation, calculated as follows: ; ; in, Indicates the overall state size. Indicates the number of semantic states. Indicates the number of vehicle statuses. Indicates the number of environmental states. Indicates the number of network states. Indicates the number of rule states. In this embodiment, the module is represented as follows: , .
[0018] The strength of the state relationship is obtained by summing the coupling weights between any two states in the set of all states. The formula is as follows: ; in, Indicates the strength of the state relationship. Representing state With state The coupling weights between them.
[0019] In this embodiment, a structural coupling coefficient is calculated based on the strength of the state relationship and the total state scale, which is used to couple and correct semantic uncertainty.
[0020] For example, in this embodiment, the coupling weight between "fault explanation intent" and "abnormal tire pressure" is set to 0.95, the coupling weight between "risk assessment intent" and "current vehicle speed" is set to 0.88, and the coupling weight between "safety priority rule" and "prohibition of dangerous driving suggestion" is set to 0.92. The coupling weights between other states are set according to semantic relevance, state dependency, and rule constraint. After summing, we get: The results indicate that the current decoding state in this embodiment belongs to a multi-state strongly coupled scenario, which is suitable for entering the subsequent collaborative demand scoring and heterogeneous teacher model scheduling stage.
[0021] In this embodiment, for the critical question of "can we continue driving?", the first language model on the vehicle side provides the following three candidate directions for continuing to be generated and their probabilities: "It is recommended to reduce speed and go to a repair shop as soon as possible," probability 0.46; "It is recommended to stop immediately and wait for rescue," probability 0.38; "Unable to determine, please contact after-sales service", probability 0.16.
[0022] therefore, , , The results indicate that the vehicle-side first language model has high semantic uncertainty at the current risk assessment position.
[0023] Based on the intent classification model, fault dictionary, and risk rule base, control intent sets, fault intent sets, and risk state sets are identified. The task risk level is then calculated by weighted summation of the modulo values of these sets, yielding the following formula: ; ; ; in, Indicates the level of task risk. Indicates the number of control intentions. Indicates the number of fault intentions. Indicates security risk load. Represents the set of control intentions. Represents a set of fault intents. Represents a set of risk states. Indicates the first The risk weight corresponding to each security risk state Indicates the modulus. , , These are the first risk fusion weight, the second risk fusion weight, and the third risk fusion weight, respectively.
[0024] In this embodiment: Control Intent Set Including navigation requests, therefore
[0025] Fault Intent Set Including explanations for abnormal tire pressure, therefore
[0026] Security Risk Status Set This includes two risk states: "abnormal tire pressure" and "current vehicle speed is too high." If each risk state... Corresponding risk weights In this embodiment, the risk weight corresponding to abnormal tire pressure is set to 0.9; the risk weight corresponding to excessive current vehicle speed is set to 0.6.
[0027] therefore, , , .
[0028] Based on the field extraction template and privacy label table, the set of fields to be uploaded and the set of sensitive fields are identified. The privacy-sensitive load is calculated based on the classification of each field in the sensitive field set as a sensitive field and the corresponding sensitivity weight of each field. The total upload load is calculated based on the sensitivity weight of each field in the set of fields to be uploaded. Finally, the privacy sensitivity is calculated based on the privacy-sensitive load and the total upload load, using the following formula: ; ; ; in, Indicates privacy sensitivity, Indicates a privacy-sensitive payload. Indicates the total upload load. This represents a constant used to prevent the denominator from being zero. Indicates the set of fields to be uploaded. Indicates the set of sensitive fields. Indicates the first Sensitive weights of each field, Indicates an indicator function.
[0029] In this embodiment, the set of fields to be uploaded includes: tire pressure value, vehicle speed range, malfunction indicator lamp status, and coarse-grained location. Among these, "coarse-grained location" is a sensitive field. If each field... Corresponding sensitive weights In this embodiment, the weights are set as follows: tire pressure value is 1, vehicle speed range is 1, fault light status is 1, and coarse-grained location is 2.
[0030] therefore, , , .
[0031] Based on the round-trip time, uplink data volume, downlink data volume, uplink bandwidth, and downlink bandwidth in the network state set, the communication delay is calculated; based on the packet loss rate, jitter value, and link switching count in the network state set, the network disturbance cost is calculated; the communication delay and the network disturbance cost are both dimensionless, and the dimensionless communication delay and the network disturbance cost are then weighted and summed to obtain the network cost, calculated as follows: ; ; ; in, Indicates network cost. Indicates communication delay. This represents dimensionless communication delay. This represents the cost of network disturbance. This represents the dimensionless cost of network perturbation. Indicates the first network cost fusion weight. Indicates the second network cost fusion weight. Indicates round-trip time delay. Indicates the amount of data uploaded. Indicates the amount of downlink data. Indicates uplink bandwidth. Indicates downlink bandwidth. Indicates packet loss rate, Indicates the jitter value. Indicates the number of link switching. , , These represent the first network perturbation fusion weight, the second network perturbation fusion weight, and the third network perturbation fusion weight, respectively.
[0032] It should be noted that the above Characterizing transmission delay overhead, the The two represent link disturbance costs, which have different origins; in this application, they are represented by preset weights. and right and By performing a fusion representation under a unified evaluation scale, the network cost can be obtained. .
[0033] In this embodiment: Uplink data volume Uplink bandwidth Downlink data volume KB, downlink bandwidth Packet loss rate jitter value Link switching times Furthermore, take , .
[0034] in, ; ; therefore, , , .
[0035] Calculate the semantic mismatch distance between the vehicle-side first language model and the semantic representation extracted from any candidate teacher model based on user input commands, and sum all semantic mismatch distances to obtain the teacher mismatch degree. The calculation formula is as follows: ; ; in, Indicates the degree of teacher mismatch. This indicates the number of candidate teacher models. This represents the semantic mismatch distance between the first language model on the vehicle and the semantic representation extracted by the m-th candidate teacher model based on user input commands. This indicates that the vehicle-side first language model extracts semantic representations based on user input commands. Let m represent the semantic representation of the m-th candidate teacher model extracted based on user input commands. Indicates transpose. This represents the L2 norm.
[0036] For example, in this embodiment, there are three candidate teacher models: Teacher Model 1: Vehicle Diagnostic Enhancement Model; Teacher Model 2: General Question Answering Model; Teacher Model 3: Navigation Enhancement Model. Let the cosine similarity between these models and the semantic representations extracted by the vehicle-side first language model be: Teacher Model 1: 0.92; Teacher Model 2: 0.71; Teacher Model 3: 0.80, respectively. Then the corresponding mismatch distance is: , , .
[0037] therefore, This indicates that among all the teacher models, the vehicle diagnostic enhancement model is the best match for the current decoding state.
[0038] In this embodiment, the teacher model is an auxiliary decoding model deployed on a cloud server, edge node, or cockpit high-computing-power domain controller. A heterogeneous teacher model refers to one that differs in at least one of the following: model parameter scale, training corpus, task adapter, external toolchain, or knowledge base. For example, the vehicle diagnostic enhancement teacher model focuses on fault explanation and repair suggestion generation, the navigation enhancement teacher model focuses on route planning and repair point retrieval, and the general question-answering teacher model focuses on general language completion and explanatory expression; each teacher model receives a unified structured context representation as input and outputs candidate collaborative fragments.
[0039] The collaborative requirement score is obtained by weighted summation of the normalized semantic uncertainty, task risk, privacy sensitivity, network cost, and teacher mismatch, calculated as follows: ; ; in, Indicates the score for collaborative needs. This represents the normalized semantic uncertainty. This indicates the normalized task risk level. This indicates the normalized privacy sensitivity. This represents the normalized network cost. This represents the normalized teacher fit. , , , , These represent the first fusion weight, the second fusion weight, the third fusion weight, the fourth fusion weight, and the fifth fusion weight, respectively. Indicates the first i Each fusion weight.
[0040] To facilitate demonstration in a single example, this embodiment provides a set of normalized intervals for reproducible experiments: ; Therefore, we can obtain the normalized semantic uncertainty. Normalized task risk Normalized privacy sensitivity Normalized network cost Normalized teacher fit .
[0041] In this embodiment, the following is taken ,but .
[0042] S2: Determine the number of teacher models corresponding to the current decoding state based on the collaborative requirement score; the teacher models are used to generate a set of candidate collaborative fragments based on the structured context representation.
[0043] Specifically, the steps include: The first collaboration threshold and the second collaboration threshold are determined based on the distribution of the collaboration requirement scores, and the first collaboration threshold is less than the second collaboration threshold. In this embodiment, the collaboration requirement scores are arranged from smallest to largest, the first collaboration threshold is the 1 / 3 quantile of the collaboration requirement score set, and the second collaboration threshold is the 2 / 3 quantile of the collaboration requirement score set.
[0044] When the collaborative demand score is less than or equal to the first collaborative threshold, the number of teacher models is a first preset value (e.g., 1). When the collaborative demand score is greater than the first collaborative threshold and less than or equal to the second collaborative threshold, the number of teacher models is a second preset value (e.g., 2). When the collaborative demand score is greater than or equal to the second collaborative threshold, the number of teacher models is a third preset value (e.g., 3).
[0045] In this embodiment, the first collaboration threshold Second collaborative threshold ,and Therefore, the number of teacher models is determined to be 2.
[0046] In this embodiment, examples of candidate collaborative fragment sets generated by the two teacher models are shown in Table 6; Table 6 Candidate Collaborative Fragment Set ; S3: When the number of teacher models is 1, the candidate collaborative fragment set generated by the corresponding teacher model is used as the initial collaborative fragment set. When the number of teacher models is two or more, a unified representation mapping and semantic deduplication are performed on the candidate collaborative fragment sets generated by all teacher models to obtain the initial collaborative fragment set.
[0047] Specifically, the unified representation mapping and semantic deduplication of the candidate collaborative fragment set generated by all teacher models includes: Each candidate co-contractor in the set of all candidate co-contractors is encoded into a corresponding semantic vector, and the cosine similarity between any two semantic vectors is calculated. When the cosine similarity is greater than or equal to a preset similarity threshold, the candidate collaborative fragment corresponding to either of the two semantic vectors is deleted until all candidate collaborative fragments are traversed to obtain an initial collaborative fragment set.
[0048] In this embodiment, when the cosine similarity is greater than or equal to a preset similarity threshold... When determining if two candidate problems are duplicates, deduplication can be performed by comparing their quality scores. Specifically, the candidate problem with the higher quality score among the duplicates is retained. The quality score formula is: ; Where s(q) represents the quality score of candidate question q; h represents the semantic content corresponding to candidate question q; represent the weight coefficients corresponding to the relevance scoring item, the consistency of evidence scoring item, and the safety scoring item, respectively. All are non-negative real numbers.
[0049] Specifically, the The relevance score between the semantic content h corresponding to the candidate question q and the current user input command and the current decoding state is calculated as follows: ; in, This represents a semantic encoding function. Let h represent the semantic vector corresponding to the semantic content h of the candidate question. This represents the semantic vector of the text corresponding to the user input command x or the current decoding state. This represents the cosine similarity. The larger the value of , the stronger the relevance of the candidate problem to the current task context.
[0050] The The evidence consistency score for the semantic content h corresponding to candidate question q is calculated as follows: ; Where E represents the set of evidence states, and |E| represents the number of evidence states in the set of evidence states. Indicates the state of the j-th piece of evidence. This represents the matching relationship value between semantic content h and the j-th evidence state. When semantic content h successfully matches the j-th evidence state... Otherwise, r_{h,j}=0; Indicates an indicator function. The larger the value of , the higher the consistency between the content corresponding to the candidate question and the evidence state set.
[0051] The The security score of the semantic content h corresponding to the candidate question q is calculated as follows: ; Where R represents the set of rule states, and |R| represents the number of rule states in the set of rule states. This represents the state of the k-th rule. The semantic content h represents the rule state. The violation judgment result is as follows: when the semantic content h violates the k-th rule state... ,otherwise The larger the value of , the higher the security and compliance level of the content corresponding to the candidate question.
[0052] Based on the above quality scoring formula, for candidate questions whose semantic similarity meets the duplicate judgment condition, their corresponding s(q) values are compared. Candidate questions with higher quality scores are retained, while candidate questions with lower quality scores are deleted. Thus, candidate question deduplication is achieved under the premise that the candidate questions are relevant to the current task context, consistent with the evidence state, and meet the safety constraints.
[0053] For example, this embodiment sets For the statements "Abnormal right front tire pressure, it is recommended to reduce speed and have it inspected as soon as possible" and "It is not recommended to continue driving at high speed when tire pressure is too low," the semantic similarity was calculated to be 0.83. Therefore, it was determined that the two statements had high repetition, and the former with the higher quality score was retained. For the statements "You can navigate to the nearest repair shop, which is expected to arrive in 8 minutes" and "It is recommended to first drive at a low speed to the nearby repair shop for inspection," the similarity was 0.77, which was below the threshold, and was retained. Therefore, the initial set of collaborative fragments was obtained after deduplication. There are a total of 3 coordinating segments.
[0054] S4: Construct a relation matrix based on the initial set of collaborative fragments and the set of evidence states, calculate the collaborative fragment coverage difference and the evidence coverage difference based on the relation matrix, and complete the initial set of collaborative fragments based on the collaborative fragment coverage difference and the evidence coverage difference to obtain the final set of collaborative fragments.
[0055] Specifically, the construction of the relationship matrix based on the initial set of collaborative fragments and the set of evidence states includes: Each initial collaborative fragment in the initial collaborative fragment set is hard-matched or soft-matched with the name dictionary of any evidence state in the evidence state set. If the match is successful, the relationship value between the corresponding initial collaborative fragment and the evidence state is set to 1; otherwise, it is set to 0. The hard matching is to determine whether the name in the name dictionary of the evidence state is contained in the initial collaborative fragment. The soft matching is to determine whether the maximum cosine similarity between the semantic vector of the initial collaborative fragment and the semantic vector of the name in the name dictionary of the evidence state is greater than or equal to a semantic threshold.
[0056] The matching expression is: ; ; ; ; in, Indicates the first i Initial collaborative fragments With the j Individual Evidence Status Relationship value, This represents the matching function. This represents a hard-match function. This represents a soft-matching function. Indicates an indicator function, Indicates the first j Individual Evidence Status The first in the name dictionary A name, Represents semantic encoding, Represents cosine similarity. This indicates that it exists. Indicates that it is included in, Indicates or, Indicates the semantic threshold.
[0057] The relationship matrix is constructed by sequentially traversing each initial collaborative fragment in the initial collaborative fragment set and each evidence state in the evidence state set, based on the relationship value between any initial collaborative fragment and any evidence state.
[0058] Furthermore, the calculation of collaborative fragment coverage difference and evidence coverage difference based on the relation matrix includes: Summing the relation values corresponding to all evidence states that successfully match any initial collaborative fragment yields the number of covered evidence states for that initial collaborative fragment. The formula is as follows: ; in, Indicates the first i The number of coverage evidence states for each initial collaborative fragment. Indicates the relationship with the first i The number of evidence states that were successfully matched by an initial collaborative fragment. Indicates the relationship with the first i The first initial cooperative fragment successfully matched. j The relation value corresponding to each piece of evidence state.
[0059] Summing the relation values corresponding to all initial collaborative fragments that successfully match any evidence state yields the number of collaborative fragments involved in the corresponding evidence state, calculated as follows: ; in, Indicates the first j The number of collaborative fragments involved in each piece of evidence state, Indicates the relationship with the first j The number of initial collaborative fragments that successfully match the evidence state.
[0060] Calculate the first mean of the number of covered evidence states for all initial collaborative fragments, and calculate the second mean of the number of collaborative fragments involved in all evidence states. The calculation formula is as follows: ; ; in, This represents the first mean. This represents the second mean.
[0061] The first variance between the number of coverage evidence states for each initial co-occurrence fragment and the first mean is calculated. This first variance is the co-occurrence fragment coverage difference, and the calculation formula is as follows: ; in, This indicates differences in the coverage of collaborative fragments.
[0062] The second variance between the number of coordinated segments involved in each evidence state and the second mean is calculated. The second variance is the evidence coverage difference, and the calculation formula is as follows: ; in, This indicates differences in the coverage of evidence.
[0063] Furthermore, the step of completing the initial set of collaborative fragments based on the differences in collaborative fragment coverage and evidence coverage includes: The first coefficient of variation is calculated based on the difference in coverage of the cooperating segments and the first mean, and the calculation formula is as follows: ; in, This represents the first coefficient of variation. This represents a constant used to prevent the denominator from being zero.
[0064] Based on the aforementioned evidence coverage difference and the second mean, the formula for calculating the second coefficient of variation is as follows: ; in, This represents the second coefficient of variation. This represents a constant used to prevent the denominator from being zero.
[0065] The percentage of evidence states involving a number greater than 0 in the total set of evidence states is taken as the evidence state coverage rate, calculated as follows: ; in, This indicates the coverage rate of evidence status.
[0066] Evidence states involving fewer than or equal to a preset evidence state coverage threshold are considered undercovered evidence states. Based on the first coefficient of variation, the second coefficient of variation, and the evidence state coverage, the completion trigger command value is determined using the following formula: ; in, This indicates that the trigger command value has been completed. Indicates an indicator function, This represents the first coefficient of variation. This represents the second coefficient of variation. Indicates the coverage rate of evidence status. Indicates the first threshold. This represents the second threshold. This represents the third threshold. Indicates or; For each of the undercovered evidence states, supplementary collaborative fragments are generated in a targeted manner, and the relation matrix is updated until the completion trigger instruction value is not 1 or the maximum number of iterations is reached, thus obtaining the final set of collaborative fragments.
[0067] In this embodiment, the final set of cooperative fragments can also be selected based on an objective function, which is expressed as: ; Where H represents a subset of candidate collaborative fragments selected from the initial collaborative fragment set; These represent the coverage weight, relevance weight, balance penalty weight, and redundancy penalty weight, respectively, and each weight is a non-negative real number.
[0068] Specifically, Cov(H) represents the coverage of the candidate collaborative fragment subset H to the evidence state set, and is calculated as follows: ; Where E represents the set of evidence states, and |E| represents the number of evidence states in the set of evidence states. This represents the i-th collaborative fragment in the subset H of candidate collaborative fragments. This represents the j-th evidence state in the evidence state set E. This represents the relationship value between the i-th collaborative fragment and the j-th evidence state. When the i-th collaborative fragment and the j-th evidence state successfully match... ,otherwise ; This indicates an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise.
[0069] Rel represents the average relevance score of each co-contract in the candidate co-contract subset H, and is calculated as follows: ; Where |H| represents the number of cooperating fragments in the candidate cooperating fragment subset H. The quality score represents the quality score of the i-th collaborative segment. The quality score is used to characterize the degree of matching between the collaborative segment and the current user input instruction, the structured context representation, and the target decoding task. The higher the quality score, the more suitable the corresponding collaborative segment is to be retained as the final collaborative segment.
[0070] The The coefficient of variation is used to characterize the balance of the number of collaborative fragments involved in the evidentiary status, and its calculation formula is as follows: ; in, This represents the mean number of coordinated segments involved in each piece of evidence state. The standard deviation of the number of coordinated fragments involved in each piece of evidence status. This represents a constant used to prevent the denominator from being zero. The larger the value, the worse the balance of coverage among different evidentiary states.
[0071] Red(H) represents the semantic redundancy within the candidate cooperative fragment subset H, and is calculated as follows: ; in, Representing cooperative fragments The corresponding semantic vector, Representing cooperative fragments The corresponding semantic vector, Semantic vectors With semantic vectors The cosine similarity between them. The larger Red(H) is, the higher the degree of semantic repetition within the candidate co-segment subset H.
[0072] Based on the above objective function, while ensuring the coverage of evidence status and the relevance of collaborative fragments, penalties are imposed on the imbalance of evidence coverage and semantic redundancy of fragments, thereby selecting the final set of collaborative fragments with higher coverage, stronger relevance, more balanced distribution and lower redundancy.
[0073] S5: Continue decoding the final collaborative fragment set and inject the resulting final collaborative decoding result into the vehicle-side first language model to generate the final response.
[0074] Specifically, the further decoding of the final co-located fragment set includes: Any final collaborative fragment in the final collaborative fragment set is processed by a teacher model to generate a further decoding result corresponding to the final collaborative fragment. The further decoding result is the final collaborative decoding result corresponding to the final collaborative fragment. Alternatively, any final collaborative fragment in the final collaborative fragment set can be processed by several teacher models to generate several candidate continued decoding results corresponding to the final collaborative fragment. Based on the consistency of evidence and the consistency between models, the final collaborative decoding result corresponding to the final collaborative fragment can be selected from the several candidate continued decoding results.
[0075] Furthermore, the step of selecting the final collaborative decoding result from several candidate decoding results based on evidence consistency and inter-model consistency includes: Based on the structured context representation, an evidence consistency verification function is constructed, with the expression as follows: ; ; ; in, This represents the evidence consistency verification function, used to determine the candidate for further decoding results. Whether the evidence consistency check has been passed; Indicates the first t The final collaborative fragment, Indicates the first t The first final collaborative fragment corresponding to the first k The candidate continues decoding the result. Indicates the candidate continues decoding result Evidence path, Representing a structured context, This indicates an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise. Indicates and, This represents a path compliance judgment function used to determine the evidence path. Do all the evidence nodes in the document belong to the set of compliant evidence paths? If for any evidence node r, the following condition is met: ,but ,otherwise ; Represents any evidence node , This represents the set of compliant evidence paths. This represents the evidence state consistency discrimination function, used to determine the candidate for continuing decoding. With structured context representation Whether the consistency between the states of relevant evidence in the case meets the preset requirements; This indicates the preset consistency threshold. This represents a security discrimination function used to determine the candidate to continue decoding. Does it satisfy the safety constraints in the rule state set? When not violating preset security rules, ,otherwise Based on the above evidence consistency verification function, the candidate continuing decoding result is determined to pass the evidence consistency verification only when it simultaneously meets the evidence state consistency condition, the evidence path compliance condition, and the security constraint condition.
[0076] All evidence states that successfully match the corresponding final collaborative fragment are selected to form a set of evidence states involved in the collaborative fragment, and candidate decoding results are obtained based on the evidence path. The cited evidence states constitute the cited evidence state set; the evidence state consistency ratio is calculated based on the evidence state set involved in the collaborative fragments and the cited evidence state set, using the following formula: ; in, Indicates the candidate continues decoding result The consistency ratio of evidence status This indicates that the collaborative fragment involves a set of evidence states. The number of pieces of evidence in the state, Indicates the set of evidence states. The number of pieces of evidence in the state, Represents a constant; When the evidence consistency ratio is greater than or equal to a preset ratio threshold, the candidate is determined to continue decoding. Verify the consistency of the evidence status; For candidate decoding results that pass the evidence state consistency check, further evidence consistency check is performed. Based on the candidate decoding results that pass the evidence consistency check, the inter-model consistency score is calculated, and the calculation formula is as follows: ; ; ; in, Indicates the candidate continues decoding result The inter-model consistency score This represents the first set of continued decoding results, consisting of candidate continued decoding results that have passed the evidence consistency check. This indicates the number of candidate results for further decoding in the first set of results. This indicates the total number of candidate results that need to be decoded. Indicates the first t The first final collaborative fragment corresponding to the first l The candidate continues decoding the result. This indicates the set of results from the candidate's continued decoding. This represents the semantic implication discriminant function. Indicates semantic similarity. This indicates a preset semantic implication threshold; A comprehensive score is calculated based on the evidence consistency score, the evidence status consistency ratio, and the inter-model consistency score. The calculation formula is as follows: ; ; in, Indicates the candidate continues decoding result Overall score Indicates the first weight. Indicates the second weight. Indicates the third weight. Indicates the fourth weight. Indicates the confidence score. Indicates the first t The first final collaborative fragment corresponding to the first k The original confidence scores of the candidate decoding results. Indicates the first t The first final collaborative fragment corresponding to the first l The original confidence scores of the candidate decoding results; The candidate with the highest comprehensive score is selected as the final collaborative decoding result corresponding to the final collaborative segment.
[0077] In this embodiment, the final response-final collaborative decoding result log is also stored in the policy knowledge base and can be reused: The final collaborative decoding result, the corresponding final response, the semantic vector of the final collaborative decoding result, and the semantic vector of the final response are stored as a quadruple in the policy knowledge base; When storing, the semantic cosine similarity between the semantic vector of the final collaborative decoding result in the newly stored quadruple and the semantic vector of each of the previously stored final collaborative decoding results is calculated. When the semantic cosine similarity is greater than or equal to the second semantic threshold, the final response corresponding to the previously stored final collaborative decoding result is reused or used as a prompt context to assist in the generation.
[0078] The semantic constraint heterogeneous teacher collaborative decoding method for in-vehicle intelligent cockpits provided in this embodiment has the following beneficial effects: By acquiring user-input voice or text commands, as well as vehicle operating status data, cabin environment data, and network status data, the system then uses the vehicle-side first language model to perform local autoregressive decoding on the user input to obtain the current decoding result. Based on this decoding result, vehicle operating status data, cabin environment data, and network status data, a structured context representation is constructed. This structured context representation includes at least a semantic state set, a vehicle state set, an environment state set, a network state set, and a rule state set. After the structured context representation is constructed, an evidence state set is further extracted. Subsequent steps are based on this structured context representation, eliminating the need to repeatedly upload the entire original context or perform secondary parsing of the user input. The collaborative demand score of the structured context representation is then calculated, and the number of teacher models corresponding to the current decoding state is determined according to the collaborative demand score. At least one heterogeneous teacher model, equal to the number of teacher models, is invoked, and each heterogeneous teacher model is utilized. The teacher model generates a candidate collaborative fragment set based on structured context representation, obtaining candidate collaborative fragments corresponding to each heterogeneous teacher model. Then, based on all candidate collaborative fragments, a unified representation mapping and semantic deduplication are performed to delete duplicate collaborative fragments, resulting in an initial collaborative fragment set. A fragment-evidence relationship matrix is then constructed based on the initial collaborative fragment set and the evidence state set. Based on the relationship matrix, collaborative fragment coverage differences and evidence coverage differences are calculated. A second-stage completion generation is performed based on the collaborative fragment coverage differences and evidence coverage differences to obtain the final collaborative fragment set. Finally, at least one continued decoding result is generated based on the final collaborative fragment set. When the number of continued decoding results is greater than 1, the result that is most consistent with the vehicle state constraints, evidence path, and safety rules is selected from multiple continued decoding results to obtain the final collaborative decoding result. When the number of continued decoding results is 1, the continued decoding result is determined as the final collaborative decoding result. The final collaborative decoding result is injected into the vehicle-side first language model to continue generating the final response, and the final response and the final collaborative decoding result log are stored in the policy knowledge base. Specifically, after obtaining the structured context representation, the entire original context is no longer repeatedly uploaded, and secondary parsing of user input is avoided, thus preventing redundant processing, improving processing consistency, and reducing computational overhead. The number of teacher models is determined based on the collaborative requirement score of the current decoding state, comprehensively considering semantic uncertainty, task risk, privacy sensitivity, network cost, and teacher mismatch. Heterogeneous teacher models are invoked according to collaborative requirements, improving the comprehensiveness of collaborative fragment generation while balancing computational efficiency. The diversity and coverage of the collaborative fragment set are improved through unified representation mapping, semantic deduplication, and coverage completion generation. By selecting the results of continued decoding in accordance with vehicle state constraints, evidence paths, and safety rules, the accuracy and credibility of the final response are improved.
[0079] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0080] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A semantically constrained heterogeneous teacher collaborative decoding method for in-vehicle intelligent cockpits, characterized in that, include: S1: Perform local autoregressive decoding on the acquired user input commands, and combine the decoding results, vehicle operating status data, cabin environment data, and network status data to construct a structured context representation; Based on the structured context representation, user input instructions, and the current decoding result, an evidence state set is extracted; the structured context representation includes a multi-state coupled representation, which includes a network state set, which is obtained through network monitoring data output by the vehicle communication module; The collaborative requirement score for the structured context representation is calculated according to the first preset rule, including: Based on the predicted probabilities of each candidate output in the decoding results, calculate the decoding prediction entropy and the prediction probability interval between the two candidate outputs with the highest prediction probabilities; calculate the semantic uncertainty based on the decoding prediction entropy and the prediction probability interval. Based on the intent classification model, fault dictionary and risk rule base, the control intent set, fault intent set and risk state set are identified, and the modulus of the control intent set, fault intent set and risk state set is weighted and summed to obtain the task risk degree; Based on the field extraction template and privacy label table, the set of fields to be uploaded and the set of sensitive fields are identified. The privacy-sensitive load is calculated based on the status of each field in the sensitive field set as a sensitive field and the sensitivity weight of each field. The total upload load is calculated based on the sensitivity weight of each field in the set of fields to be uploaded. The privacy sensitivity is calculated based on the privacy-sensitive load and the total upload load. Based on the round-trip time, uplink data volume, downlink data volume, uplink bandwidth, and downlink bandwidth in the network state set, the communication latency is calculated; based on the packet loss rate, jitter value, and link switching count in the network state set, the network disturbance cost is calculated; the communication latency and the network disturbance cost are respectively dimensionlessized, and the dimensionless communication latency and the network disturbance cost are weighted and summed to obtain the network cost. Calculate the semantic mismatch distance between the vehicle-side first language model and the semantic representation extracted by any candidate teacher model based on user input commands, and sum all semantic mismatch distances to obtain the teacher mismatch degree; The normalized semantic uncertainty, task risk, privacy sensitivity, network cost, and teacher mismatch are weighted and summed to obtain the collaborative requirement score. S2: Determine the number of teacher models corresponding to the current decoding state based on the collaborative demand score; the teacher models are used to generate a set of candidate collaborative fragments based on structured context representation; S3: When the number of teacher models is 1, the candidate collaborative fragment set generated by the corresponding teacher model is used as the initial collaborative fragment set. When the number of teacher models is two or more, a unified representation mapping and semantic deduplication are performed on the candidate collaborative fragment sets generated by all teacher models to obtain the initial collaborative fragment set. S4: Construct a relation matrix based on the initial set of collaborative fragments and the set of evidence states; calculate the collaborative fragment coverage difference and the evidence coverage difference based on the relation matrix; and complete the initial set of collaborative fragments based on the collaborative fragment coverage difference and the evidence coverage difference to obtain the final set of collaborative fragments. S5: Continue decoding the final collaborative fragment set and inject the resulting final collaborative decoding result into the vehicle-side first language model to generate the final response.
2. The method according to claim 1, characterized in that, The structured context representation also includes an evidence state set; the multi-state coupled representation also includes a semantic state set, a vehicle state set, an environment state set, and a rule state set; wherein, the semantic state set is obtained by performing intent recognition and slot extraction on user input commands and current decoding results; the vehicle state set is obtained by reading vehicle operating status data output by the vehicle bus or diagnostic interface; the environment state set is obtained by information output by positioning, map, and cockpit sensors; the rule state set is obtained by matching with a preset rule base; the evidence state set is obtained by filtering key states from the multi-state coupled representation that are related to the current decoding position and can be used to verify candidate decoding results.
3. The method according to claim 1, characterized in that, The step of determining the number of teacher models corresponding to the current decoding state based on the collaborative demand score includes: The first collaboration threshold and the second collaboration threshold are determined based on the distribution of collaboration demand scores, and the first collaboration threshold is less than the second collaboration threshold. When the collaborative demand score is less than or equal to the first collaborative threshold, the number of teacher models is a first preset value; When the collaborative demand score is greater than the first collaborative threshold and less than or equal to the second collaborative threshold, the number of teacher models is the second preset value; When the collaborative demand score is greater than or equal to the second collaborative threshold, the number of teacher models is the third preset value.
4. The method according to claim 1, characterized in that, The unified representation mapping and semantic deduplication of the candidate collaborative fragment sets generated by all teacher models includes: Each candidate co-contractor in the set of all candidate co-contractors is encoded into a corresponding semantic vector, and the cosine similarity between any two semantic vectors is calculated. When the cosine similarity is greater than or equal to a preset similarity threshold, the candidate collaborative fragment corresponding to either of the two semantic vectors is deleted until all candidate collaborative fragments are traversed to obtain an initial collaborative fragment set.
5. The method according to claim 1, characterized in that, The construction of the relationship matrix based on the initial set of collaborative fragments and the set of evidence states includes: Each initial collaborative fragment in the initial collaborative fragment set is hard-matched or soft-matched with the name dictionary of any evidence state in the evidence state set. If the match is successful, the relationship value between the corresponding initial collaborative fragment and the evidence state is set to 1; otherwise, it is set to 0. The hard matching is to determine whether the name in the name dictionary of the evidence state is contained in the initial collaborative fragment. The soft matching is to determine whether the maximum cosine similarity between the semantic vector of the initial collaborative fragment and the semantic vector of the name in the name dictionary of the evidence state is greater than or equal to the semantic threshold. The relationship matrix is constructed by sequentially traversing each initial collaborative fragment in the initial collaborative fragment set and each evidence state in the evidence state set, based on the relationship value between any initial collaborative fragment and any evidence state.
6. The method according to claim 1, characterized in that, The calculation of collaborative fragment coverage difference and evidence coverage difference based on the relation matrix includes: Sum the relation values corresponding to all evidence states that are successfully matched with any initial collaborative fragment to obtain the number of covered evidence states for the corresponding initial collaborative fragment; Sum the relation values corresponding to all initial collaborative fragments that successfully match any evidence state to obtain the number of collaborative fragments involved in the corresponding evidence state; Calculate the first mean of the number of covered evidence states for all initial collaborative fragments, and calculate the second mean of the number of collaborative fragments involved in all evidence states; The first variance between the number of coverage evidence states of each initial collaborative segment and the first mean is calculated, and the first variance is the collaborative segment coverage difference. The second variance between the number of coordinated segments involved in each evidence state and the second mean is calculated, and the second variance is the evidence coverage difference.
7. The method according to claim 6, characterized in that, The process of completing the initial set of collaborative fragments based on differences in collaborative fragment coverage and differences in evidence coverage includes: A first coefficient of variation is calculated based on the difference in coverage of the cooperating segments and the first mean; A second coefficient of variation is calculated based on the evidence coverage difference and the second mean; The percentage of evidence states involving a number greater than 0 in the set of evidence states is used as the evidence state coverage rate. Evidence states involving fewer than or equal to a preset evidence state coverage threshold are considered undercovered evidence states. Based on the first coefficient of variation, the second coefficient of variation, and the evidence state coverage, the completion trigger command value is determined using the following formula: ; in, This indicates that the trigger command value has been completed. Indicates an indicator function, This represents the first coefficient of variation. This represents the second coefficient of variation. Indicates the coverage rate of evidence status. Indicates the first threshold. This represents the second threshold. This represents the third threshold. Indicates or; For each of the undercovered evidence states, supplementary collaborative fragments are generated in a targeted manner, and the relation matrix is updated until the completion trigger instruction value is not 1 or the maximum number of iterations is reached, thus obtaining the final set of collaborative fragments.
8. The method according to claim 1, characterized in that, The continued decoding of the final co-located fragment set includes: Any final collaborative fragment in the final collaborative fragment set is processed by a teacher model to generate a further decoding result corresponding to the final collaborative fragment. The further decoding result is the final collaborative decoding result corresponding to the final collaborative fragment. Alternatively, any final collaborative fragment in the final collaborative fragment set can be processed by several teacher models to generate several candidate continued decoding results corresponding to the final collaborative fragment. Based on the consistency of evidence and the consistency between models, the final collaborative decoding result corresponding to the final collaborative fragment can be selected from the several candidate continued decoding results.
9. The method according to claim 8, characterized in that, The final collaborative decoding result, selected from several candidate continued decoding results based on evidence consistency and inter-model consistency, includes: Based on the structured context representation, an evidence consistency verification function is constructed, with the expression as follows: ; ; ; in, This represents the evidence consistency verification function. Indicates the first t The final collaborative fragment, Indicates the first t The first final collaborative fragment corresponding to the first k The candidate continues decoding the result. Indicates the candidate continues decoding result Evidence path, Representing a structured context, Indicates an indicator function, Indicates and, This represents the path compliance determination function. Represents any evidence node , Represents the set of compliant evidence paths. This represents the function for determining the consistency of evidence states. This indicates a preset consistency threshold. Represents the security discrimination function, when When not violating preset security rules, ,otherwise ; All evidence states that successfully match the corresponding final collaborative fragment are selected to form a set of evidence states involved in the collaborative fragment, and candidate decoding results are obtained based on the evidence path. The cited evidence states constitute the cited evidence state set; the evidence state consistency ratio is calculated based on the evidence state set involved in the collaborative fragments and the cited evidence state set, using the following formula: ; in, Indicates the candidate continues decoding result The consistency ratio of evidence status This indicates that the collaborative fragment involves a set of evidence states. The number of pieces of evidence in the state of being, Indicates the set of evidence states. The number of pieces of evidence in the state of being, Represents a constant; When the evidence consistency ratio is greater than or equal to a preset ratio threshold, the candidate is determined to continue decoding. Verify the consistency of the evidence status; For candidate decoding results that pass the evidence state consistency check, further evidence consistency check is performed. Based on the candidate decoding results that pass the evidence consistency check, the inter-model consistency score is calculated, and the calculation formula is as follows: ; ; ; in, Indicates the candidate continues decoding result The inter-model consistency score This represents the first set of continued decoding results, consisting of candidate continued decoding results that have passed the evidence consistency check. This indicates the number of candidate results for further decoding in the first set of results. This indicates the total number of candidate results that need to be decoded. Indicates the first t The first final collaborative fragment corresponding to the first l The candidate continues decoding the result. This indicates the set of results from the candidate's continued decoding. This represents the semantic implication discriminant function. Indicates semantic similarity. This indicates a preset semantic implication threshold; A comprehensive score is calculated based on the evidence consistency score, the evidence status consistency ratio, and the inter-model consistency score. The calculation formula is as follows: ; ; in, Indicates the candidate continues decoding result Overall score Indicates the first weight. Indicates the second weight. Indicates the third weight. Indicates the fourth weight. Indicates the confidence score. Indicates the first t The first final collaborative fragment corresponding to the first k The original confidence scores of the candidate decoding results. Indicates the first t The first final collaborative fragment corresponding to the first l The original confidence scores of the candidate decoding results; The candidate with the highest comprehensive score is selected as the final collaborative decoding result corresponding to the final collaborative segment.
Citation Information
Patent Citations
Vehicle-road cloud dynamic collaborative automatic driving method based on space-ground integrated communication network
CN120335353A
Intelligent cockpit vehicle cloud collaborative intelligent decision-making system and method based on multi-modal perception
CN121929079A