Vehicle-mounted intelligent entity data generation method and device, equipment, storage medium and product
Patent Information
- Application Number
- CN202610948043.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明提供了一种车载智能体数据生成方法、装置、设备、存储介质及产品,以解决现有技术中难以针对性的生成高质量数据的问题
[0009]根据本发明的另一方面,提供了一种计算机程序产品,所述计算机程序产品包括计算机程序,所述计算机程序在被处理器执行时实现本发明任一实施例所述的车载智能体数据生成方法。
Smart Images

Figure CN122819299A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to methods, apparatus, devices, storage media, and products for generating data for in-vehicle intelligent agents. Background Technology
[0002] With the development of large-scale model technology, large-scale models are gradually being applied to the field of automotive intelligent cockpits. In order to improve the generalization ability of in-vehicle intelligent agents in navigation, air conditioning, multimedia, vehicle control, and multi-intent mixed commands, and enable in-vehicle intelligent agents to more accurately understand user intentions, it is usually necessary to build large-scale, high-quality training data, evaluation data, and fine-tuning data.
[0003] When generating data for in-vehicle intelligent agents, existing technologies often rely on manual coding or large-scale model generation. However, these methods are typically based on fixed templates or historical log annotations, resulting in high similarity of generated user commands, low utilization value, and difficulty in targeted data generation. For example, questions arise regarding which data should be generated in the next batch, how much data should be generated, and which data should be prioritized. Therefore, a more efficient data generation method for in-vehicle intelligent agents is urgently needed, capable of completing targeted data generation tasks. Summary of the Invention
[0004] This invention provides a method, apparatus, device, storage medium, and product for generating data for in-vehicle intelligent agents, in order to solve the problem of difficulty in generating high-quality data in the prior art.
[0005] According to one aspect of the present invention, a method for generating data for an in-vehicle intelligent agent is provided, comprising: The historical corpus data in the historical intelligent cockpit dataset is divided into several semantic units according to the preset semantic dimensions, and the data distribution information corresponding to each semantic unit is statistically analyzed. Based on the data distribution information corresponding to the semantic unit, the gap information of the semantic unit is determined; wherein, the gap information is used to characterize the degree of absence of historical corpus data currently contained in the semantic unit relative to a preset coverage standard; Based on the gap information, target semantic units are selected from the plurality of semantic units, and the data generation priority of the target semantic units is determined according to the data distribution information of the target semantic units; Based on the data generation priority, a target data generation task corresponding to the target semantic unit is generated, and target data is generated based on the target data generation task; wherein, the target data is used to train the vehicle-mounted intelligent agent.
[0006] According to another aspect of the present invention, an in-vehicle intelligent agent data generation device is provided, comprising: The semantic segmentation module is used to segment the historical corpus data in the historical intelligent cockpit dataset according to the preset semantic dimensions, obtain several semantic units, and statistically analyze the data distribution information corresponding to each semantic unit. The gap information determination module is used to determine the gap information of the semantic unit based on the data distribution information corresponding to the semantic unit; wherein, the gap information is used to characterize the degree of absence of historical corpus data currently contained in the semantic unit relative to a preset coverage standard; The priority calculation module is used to filter out target semantic units from the plurality of semantic units based on the gap information, and determine the data generation priority of the target semantic units according to the data distribution information of the target semantic units; The task generation module is used to generate a target data generation task corresponding to the target semantic unit based on the data generation priority, and generate target data based on the target data generation task; wherein, the target data is used to train the vehicle-mounted intelligent agent.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the vehicle-mounted intelligent agent data generation method according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the vehicle-mounted intelligent agent data generation method according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the vehicle-mounted intelligent agent data generation method according to any embodiment of the present invention.
[0010] The technical solution of this invention divides historical corpus data in a historical intelligent cockpit dataset into several semantic units according to a preset semantic dimension, thus achieving the division of the semantic dimension to which the historical corpus data belongs. By statistically analyzing the data distribution information of each semantic unit, the gap information corresponding to each semantic unit can be determined. This gap information can characterize the degree of absence of the historical corpus data currently contained in the semantic unit relative to a preset coverage standard, realizing the quantification of the semantic coverage and semantic absence of the historical corpus data, providing guidance for subsequent data generation, and ensuring that the target data generated later can accurately fill the semantic gaps of the semantic units. Based on the gap information, target semantic units can be screened, and by determining the data generation priority of the target semantic unit according to its data distribution information, the importance of the target data generation task can be accurately determined, realizing the quantification of the target data generation task, and providing a reference for subsequent target data generation. This invention can accurately identify the gap information of historical corpus data, thereby accurately generating the target data required by the in-vehicle intelligent agent, avoiding the blindness of data generation and the generation of a large number of highly similar redundant samples, and improving the efficiency of data generation.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a method for generating data for an in-vehicle intelligent agent according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of semantic units involved in the embodiments of the present invention; Figure 3 This is a flowchart of a method for generating data for an in-vehicle intelligent agent according to Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of an in-vehicle intelligent agent data generation system applicable to embodiments of the present invention; Figure 5 This is a schematic diagram of the structure of an in-vehicle intelligent agent data generation device according to Embodiment 3 of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device that implements the vehicle-mounted intelligent agent data generation method of Embodiment 4 of the present invention. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0015] It should be noted that the terms "first," "second," "initial," and "target," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 Figure 1 This is a flowchart of a method for generating vehicle-mounted intelligent agent data, provided in Embodiment 1 of the present invention. This embodiment is applicable to data generation scenarios. The method can be executed by a vehicle-mounted intelligent agent data generation device, which can be implemented in hardware and / or software and can be configured in an electronic device. For example... Figure 1 As shown, the method includes: S101. Divide the historical corpus data in the historical intelligent cockpit dataset into several semantic units according to the preset semantic dimensions, and count the data distribution information corresponding to each semantic unit.
[0017] In this embodiment, the preset semantic dimension can be pre-set in conjunction with the interaction scenario of the in-vehicle intelligent agent, and it can serve as a standard for dividing historical corpus data in the historical intelligent cockpit dataset. For example, the preset semantic dimension can be set in conjunction with vehicle information, environmental information, and the expression method of the corpus data.
[0018] Historical intelligent cockpit datasets are used to train in-vehicle intelligent agents, ensuring that the agents can understand user-issued verbal commands and control the vehicle's equipment to operate according to the user's intentions. Historical corpus data can be manually compiled or generated using large models.
[0019] A semantic unit can be understood as a collection of historical corpus data with the same preset semantic dimensions, which can serve as labels for the semantic units. For example, preset semantic dimensions include semantic intent, environmental state, vehicle state, expression mode, task structure, and safety level. If a piece of historical corpus data is "It's a little hard to see ahead, don't take the highway to the company," then this historical corpus data can be divided according to the preset semantic dimensions and mapping relationships: semantic intent maps to air conditioning defogging and navigation route preference, environmental state maps to rainy weather, vehicle state maps to driving, expression mode maps to implicit expression, task structure maps to multiple intents, and safety level maps to executable, etc. Therefore, this historical corpus data can be divided into semantic units corresponding to air conditioning defogging + navigation route preference + rainy weather + driving + implicit expression + multiple intents + executable. The mapping relationships can be preset or learned through a large model. If the first semantic unit corresponding to a piece of historical corpus data is different from all existing semantic units, a new first semantic unit can be created.
[0020] Optionally, to avoid generating too many semantic units, semantic units can also be merged. For example, the semantic similarity between the corresponding tags of semantic units can be calculated, and semantic units with semantic similarity exceeding the similarity threshold can be merged.
[0021] Data distribution information can be determined based on historical corpus data within semantic units. For example, this includes calculating the total number of samples in the historical corpus data and the number of valid samples. The number of valid samples can be understood as the number of historical corpus data samples that meet preset valid filtering criteria. These preset valid filtering criteria can be formulated based on indicators such as the completeness of sentences and semantic clarity in the corpus data.
[0022] Data distribution information can also include evaluation information associated with semantic units. Evaluation information can include the intention recognition error rate of the in-vehicle intelligent agent on historical corpus data within the semantic unit. For example, if the semantic intents corresponding to a semantic unit include air conditioning defogging and navigation route preference, and one piece of historical corpus data for this semantic unit is "It's a bit hard to see ahead, don't take the highway to the company," then it can be determined whether the in-vehicle intelligent agent completely recognized both the air conditioning defogging and navigation route preference semantic intents, thereby determining the intention recognition error rate.
[0023] S102. Determine the gap information of the semantic unit based on the data distribution information corresponding to the semantic unit; wherein, the gap information is used to characterize the degree of absence of the historical corpus data currently contained in the semantic unit relative to the preset coverage standard.
[0024] In this embodiment, the preset coverage criteria can be formulated based on the sample size and evaluation information. For example, thresholds can be set for the total number of samples in the corpus data included in a semantic unit, the number of effective samples, and the intention recognition error rate.
[0025] The data distribution information corresponding to a semantic unit can be compared with the corresponding threshold in a preset coverage standard to obtain the degree of missing historical corpus data currently contained in the semantic unit, thereby determining the gap information. For example, if a semantic unit currently includes 42 valid samples with an intent recognition error rate of 38%, while the threshold for the number of valid samples is 500 and the threshold for the intent recognition error rate is 10%, then the valid samples and the intent recognition error rate can be compared with their corresponding thresholds to quantify the degree of missing data and thus determine the gap information.
[0026] S103. Based on the gap information, select the target semantic unit from the plurality of semantic units, and determine the data generation priority of the target semantic unit according to the data distribution information of the target semantic unit.
[0027] In this embodiment, corresponding filtering conditions can be set for the gap information to filter target semantic units. For example, the filtering condition can be that the degree of absence exceeds the corresponding degree of absence threshold. Alternatively, several semantic units can be sorted based on the gap information to filter target semantic units. For example, several semantic units can be sorted in descending order based on the degree of absence, and the top N semantic units can be used as target semantic units. Here, the several semantic units are the several semantic units obtained by dividing the historical corpus data in the historical intelligent cockpit dataset in the aforementioned steps.
[0028] The data distribution information of the target semantic unit can include the total number of samples in the historical corpus data within the unit, the number of valid samples, and the evaluation information associated with the unit. The data distribution information can be weighted and summed to obtain the data generation priority corresponding to the target semantic unit.
[0029] S104. Based on the data generation priority, generate the target data generation task corresponding to the target semantic unit, and generate target data based on the target data generation task; wherein, the target data is used to train the vehicle-mounted intelligent agent.
[0030] In this embodiment, task generation resources (e.g., computing power resources and storage resources) can be allocated to each target semantic unit based on data generation priority, thereby converting the target semantic unit into an executable target data generation task. Then, target data can be generated based on the target generation task using methods such as large model generation. This target data can be used to expand the historical intelligent cockpit dataset to train the in-vehicle intelligent agent.
[0031] The target data generation task can be associated with task generation requirements, which may include generation quantity indicators, vehicle state constraint indicators, environmental constraint indicators, intent combination requirement indicators, expression method requirement indicators, positive sample ratio, negative sample ratio, clarification sample ratio, and output field requirement indicators. The generation quantity indicator constrains the number of samples in the generated target data. The vehicle state constraint indicator constrains the generated target data to carry vehicle state information, such as whether the vehicle is moving or stationary. The environmental constraint indicator constrains the generated target data to carry environmental state information, such as rainy, sunny, or foggy weather. The intent combination requirement indicator constrains the generated target data to cover multiple semantic intents. The expression method requirement indicator constrains the generated target data to include multiple expression methods, such as veiled, ellipsis, and ambiguity. The positive sample ratio, negative sample ratio, and clarification sample ratio constrain the proportion of positive, negative, and clarification samples in the generated target data; where positive samples can be understood as semantically clear samples, negative samples as semantically unclear samples, and clarification samples as samples that provide semantic clarification for semantically unclear samples. The output field requires that the indicators used to constrain the generated target data must meet the corresponding data format, such as JSON format.
[0032] This invention provides a method for generating data for an in-vehicle intelligent agent. By dividing historical corpus data in a historical intelligent cockpit dataset into several semantic units according to a preset semantic dimension, the method achieves the division of the semantic dimension to which the historical corpus data belongs. By statistically analyzing the data distribution information of each semantic unit, the method can determine the gap information corresponding to each semantic unit. This gap information characterizes the degree of absence of historical corpus data currently contained in the semantic unit relative to a preset coverage standard, quantifying the semantic coverage and semantic absence of the historical corpus data. This provides guidance for subsequent data generation, ensuring that the target data generated later can accurately fill the semantic gaps in the semantic units. Based on the gap information, target semantic units can be selected. By accurately determining the corresponding data generation priority based on the data distribution information of the target semantic unit, the method quantifies the importance of the target data generation task, providing a reference for subsequent target data generation. This invention can accurately identify the gap information of historical corpus data, thereby accurately generating the target data required by the in-vehicle intelligent agent, avoiding blind data generation and generating a large number of highly similar redundant samples, thus improving the efficiency of data generation.
[0033] In some embodiments, the preset semantic dimensions include at least two of the following: semantic intent, semantic parameters, vehicle state, expression mode, environmental state, task structure, safety level, and vehicle type capabilities; each semantic unit corresponds to at least two of the preset semantic dimensions. This setting ensures that the semantic unit can cover multiple semantic dimensions, enabling the subsequently generated target data to contain mixed semantics with multiple intents, making the target data more closely resemble the corpus data in real application scenarios.
[0034] In this embodiment, semantic intent can be understood as information related to intent in the corpus data, such as destination, defogging, and route preference. Semantic parameters can be understood as information about parameters in the corpus data, such as temperature parameters and vehicle speed parameters. Vehicle status can be understood as information related to the vehicle's driving status in the corpus data, such as driving and parked. Expression mode can be understood as information about the expression mode of the corpus data, such as implicit, ambiguous, and ellipsis. Environmental status can be understood as information about the environment, such as rain, high humidity, and fog. Task structure can be understood as information about the semantic intent structure, such as single intent, multiple intents, and intents with conditions. Security level can be understood as identification information about the risk level of the control operation corresponding to the corpus data, such as executable, requiring confirmation, and requiring rejection. Vehicle capability can be understood as information about vehicle functions, such as seat massage.
[0035] like Figure 2 As shown, taking "It's a bit hard to see ahead, don't take the highway to the company" as an example, the semantic unit it corresponds to covers the semantic intent of defogging + navigation route preference, the environmental state is rainy, the vehicle state is driving, the expression method is implicit, and the task structure is multi-intent, etc.
[0036] In some embodiments, the data distribution information includes at least one of semantic coverage information and model inference error information; the gap information includes gap type; wherein, the statistical analysis of the data distribution information corresponding to each semantic unit includes at least one of the following steps A1 to A2: A1. Determine the semantic coverage information of each semantic unit based on the semantic dimensions covered by the historical corpus data in each semantic unit.
[0037] Semantic coverage information can characterize the completeness of historical corpus data in a semantic unit. For example, semantic coverage information can include the number of preset semantic dimensions covered by the semantic unit, and the coverage rate of the historical corpus data in the semantic unit to the preset semantic dimensions.
[0038] A2. Based on the semantic understanding results of the vehicle-mounted intelligent agent on the historical corpus data of each semantic unit, determine the model inference error information of each semantic unit.
[0039] For each semantic unit, the semantic understanding results of the vehicle-mounted intelligent agent on the historical corpus data of the current semantic unit within the historical period can be used to count the number of samples with correct or incorrect semantic understanding, and then determine the model inference error information of the current semantic unit, such as the total number of samples with incorrect semantic understanding and / or the semantic understanding error rate.
[0040] Specifically, determining the gap information of the semantic unit based on the data distribution information corresponding to the semantic unit includes at least one of the following steps B1 to B3: B1. Compare the semantic coverage information corresponding to the semantic unit with the preset target coverage information to obtain a first comparison result, and determine the gap type of the semantic unit based on the first comparison result.
[0041] The first comparison result reflects the degree of missing historical corpus data in the semantic unit. For example, the number of preset semantic dimensions covered by the semantic unit and the coverage rate of the historical corpus data in the semantic unit to the preset semantic dimensions are compared with the corresponding thresholds in the preset target coverage information to obtain the first comparison result. The degree of missing data is quantified using the first comparison result to determine whether the gap type is a high gap type or a low gap type. Among them, the degree of missing data represented by the high gap type is higher than that represented by the low gap type.
[0042] B2. Compare the model inference error information corresponding to the semantic unit with the preset target inference error information to obtain a second comparison result, and determine the gap type of the semantic unit based on the second comparison result.
[0043] Correspondingly, the second comparison result can reflect the degree of missing historical corpus data in the semantic unit. For example, the total number of semantic understanding errors and the semantic understanding error rate are compared with the threshold corresponding to the preset target inference error information to obtain the second comparison result. The degree of missing data is quantified using the second comparison result, thereby determining whether the gap type is a high gap type or a low gap type.
[0044] B3. Compare the semantic coverage information and model inference error information corresponding to the semantic unit with the corresponding preset target coverage information and preset target inference error information to obtain a third comparison result, and determine the gap type of the semantic unit based on the third comparison result.
[0045] Correspondingly, the third comparison result can reflect the degree of missing historical corpus data in the semantic unit. For example, semantic coverage information and model inference error information can be compared with the corresponding thresholds in the preset target coverage information and preset target inference error information, respectively, to obtain the third comparison result. The degree of missing data can be quantified using the third comparison result, thereby determining whether the gap type is a high gap type or a low gap type.
[0046] This invention, through statistical analysis of semantic coverage information corresponding to semantic units, can accurately quantify the coverage of historical corpus data within a semantic unit on its semantic dimensions, providing a basis for subsequent targeted generation of data that covers missing semantics. Furthermore, by statistically analyzing model inference error information corresponding to semantic units, it can accurately identify which corpus data the in-vehicle agent has a weaker understanding of, providing a basis for subsequent targeted generation of data that improves the in-vehicle agent's understanding capabilities. By determining the gap type of a semantic unit based on semantic coverage information and model inference error information, it provides a foundation for determining the data generation priority for that semantic unit, ensuring that data generation tasks can be prioritized based on semantic units with higher degrees of missing semantics.
[0047] In some embodiments, the data distribution information further includes business-related information; the business-related information includes at least one of driving safety level, vehicle type importance, and similarity to existing data generation tasks; the gap type includes a high gap type and a low gap type, wherein the degree of missing information represented by the high gap type is higher than the degree of missing information represented by the low gap type.
[0048] In this embodiment, the driving safety level is related to the driving state; for example, driving while in motion has a higher driving safety level than when stopped. Vehicle model importance can be a pre-set parameter, with different vehicle models corresponding to different importance parameters. The similarity to an existing data generation task can be determined by the semantic similarity of the corpus data corresponding to that task. For example, if the current semantic unit includes first historical corpus data, and the semantic unit corresponding to an existing data generation task includes second historical corpus data, the semantic similarity between the first and second historical corpus data can be calculated. If the semantic similarity is higher than the corresponding semantic similarity threshold, it indicates that the first data generation task corresponding to the current semantic unit is similar to the existing data generation task, and a lower data generation priority can be assigned to the first data generation task subsequently.
[0049] The step of selecting target semantic units from the plurality of semantic units based on the gap information, and determining the data generation priority of the target semantic units according to the data distribution information of the target semantic units, includes the following steps C1 to C3: C1. Select semantic units with high gap type from the plurality of semantic units and determine the semantic units as target semantic units.
[0050] A high gap type indicates a significant lack of semantic unit information, and this semantic unit can be marked as a target semantic unit for determining the data generation task in subsequent steps. A low gap type indicates a relatively low lack of semantic unit information, and the current state can be maintained.
[0051] C2. Determine the data generation cost of the target semantic unit based on the semantic coverage information in the data distribution information.
[0052] The data generation cost of the target semantic unit is quantified based on semantic coverage information. The higher the completeness of the historical corpus data in the target semantic unit represented by semantic coverage information, the lower the data generation cost, and vice versa.
[0053] C3. Based on the data generation cost and the business-related information in the data distribution information, determine the data generation priority of the target semantic unit.
[0054] Different weights can be assigned to various indicators such as data generation cost, driving safety level, vehicle type importance, and similarity to existing data generation tasks. These indicators are then weighted and summed to determine the data generation priority of the target semantic unit. Specifically, indicators related to driving safety can be given the highest weight; for example, the driving safety level can be given the highest weight.
[0055] For example, the preset semantic dimension corresponding to the first target semantic unit is "driving in the rain + defogging + navigation route preference + implicit expression", and the preset semantic dimension corresponding to the second target semantic unit is "parking status + playing music + display expression". Although the data generation cost of the first target semantic unit is higher, it is related to driving safety and has a higher weight. Therefore, the data generation priority of the first target semantic unit is higher than that of the second target semantic unit.
[0056] Example 2 Figure 3 This is a flowchart of a method for generating data for an in-vehicle intelligent agent according to Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiments. In this embodiment, the gap information may include the gap type. For example... Figure 3 As shown, the method includes: S301. Divide the historical corpus data in the historical intelligent cockpit dataset into several semantic units according to the preset semantic dimensions, and count the data distribution information corresponding to each semantic unit.
[0057] For example, such as Figure 4 As shown, the steps in this embodiment can be executed by an in-vehicle intelligent agent data generation system, which includes a semantic unit generation module, a gap assessment module, a priority decision module, a generation task scheduling module, and a feedback update module.
[0058] The historical corpus data in the historical intelligent cockpit dataset can include training data, evaluation data, and online error cases. The historical corpus data is input into the semantic unit generation module, which executes step S301 and outputs the identifier information of each semantic unit, the current number of samples in the semantic unit, the number of valid samples, and associated evaluation information.
[0059] S302. Determine the gap information of the semantic unit based on the data distribution information corresponding to the semantic unit; wherein, the gap information is used to characterize the degree of absence of the historical corpus data currently contained in the semantic unit relative to the preset coverage standard.
[0060] The target coverage configuration can be used as standard information for each indicator in the data distribution information, and input into the gap assessment module to execute S302, thereby determining the gap type of the semantic unit. For example, the target coverage configuration can be the preset target coverage information and preset target inference error information in the aforementioned embodiments.
[0061] S303. Based on the gap information, select the target semantic unit from the plurality of semantic units, and determine the data generation priority of the target semantic unit according to the data distribution information of the target semantic unit.
[0062] For example, the priority decision module determines the data generation priority by combining the security risks, business weights, and data generation costs corresponding to the data distribution information. Security risks can correspond to the driving safety levels in the aforementioned embodiments, and business weights can correspond to the weights of various indicators in the data generation cost and business-related information in the aforementioned embodiments.
[0063] S304. Based on the data generation priority and remaining generation resources, generate the target data generation task corresponding to the target semantic unit; wherein, the remaining generation resources are used to represent the remaining generation resources of the data generation task in this round.
[0064] The remaining generation resources may include the computing power and storage resources currently remaining when the vehicle-mounted intelligent agent data generation system is executing the data generation task for this round.
[0065] For example, the task scheduling module can output a task queue based on the target semantic unit. The task queue can include executable target data generation tasks and task generation requirements corresponding to the target data generation tasks.
[0066] For example, if the plan for this round is to generate 10,000 long-tail data points (i.e., data that appears infrequently and has a small number of effective samples in the historical intelligent cockpit dataset), the system will assign 300 generation tasks to the target semantic units. The requirements for these target data generation tasks are: to generate multi-intent samples showing a user's implicit or fuzzy expression of a windshield defogging request under rainy or high-humidity conditions, while the vehicle is in motion, and simultaneously providing navigation route preferences. The content of these target data generation tasks includes user commands, environmental states, semantic intents, semantic parameters, expected action sequences, and system responses.
[0067] S305. Under the constraints of diversity constraints, generate target data based on the target data generation task; wherein, the diversity constraints are related to the preset semantic dimension corresponding to the target semantic unit.
[0068] Diverse constraints can be set based on the preset semantic dimensions corresponding to the target semantic unit. For example, the preset semantic dimensions can be expanded in a positive or negative direction based on semantic similarity. For example, similar or opposite semantic intentions, similar or opposite route preferences, similar or opposite expression methods, similar or opposite vehicle speed states (e.g., acceleration or deceleration), and single-round or multi-round interactions with users can be generated.
[0069] S306. Verify the target data using preset verification conditions to obtain the qualified information corresponding to the target data.
[0070] The qualified information includes the number of qualified data and / or the qualified rate; the number of qualified data is the number of corpus data in the target data that passes the preset verification conditions; the qualified rate is the proportion of corpus data in the target data that passes the preset verification conditions.
[0071] Preset validation conditions can be formulated based on the number and / or percentage of format-compliant target data, as well as the number and / or percentage of valid samples. For example, they can be formulated based on thresholds for the number of format-compliant data, the percentage of format-compliant data, the number of valid samples, and the percentage of valid samples.
[0072] S307. Based on the qualification information and the training results of the vehicle-mounted intelligent agent, determine the available status information of the target data.
[0073] The training results include the semantic understanding error rate of the vehicle-mounted intelligent agent on the historical corpus data in the target semantic unit after training with the target data.
[0074] Available status information may include the number of qualified items and / or the pass rate, as well as the semantic understanding error rate.
[0075] S308. Based on the availability status information of the target data, update the data generation priority, data generation cost, and diversity constraints corresponding to the target semantic unit according to the update strategy.
[0076] For example, the feedback update module can execute S306 to S308, and after executing S308, when starting the next round of data generation task, it can backtrack to S301.
[0077] The update strategy includes: if the available status information meets the corresponding preset availability conditions, then reduce the data generation priority, reduce the data generation cost, and maintain or relax the constraint range of the diversity constraints; if the available status information does not meet the corresponding preset availability conditions, then increase the data generation priority, increase the data generation cost, and narrow the constraint range of the diversity constraints.
[0078] The preset availability conditions corresponding to qualified information may include preset verification conditions. The preset availability conditions corresponding to training results may be formulated based on the semantic understanding error rate, for example, based on a threshold corresponding to the semantic understanding error rate, or based on a decrease threshold corresponding to the semantic understanding error rate, etc.
[0079] For example, if the semantic understanding error rate of the vehicle-mounted intelligent agent decreases from 38% to 21% after being trained with target data on historical corpus data of the target semantic unit, it indicates that the target data can effectively improve the semantic understanding ability of the vehicle-mounted intelligent agent, and the data generation priority corresponding to the target semantic unit can be reduced in the next round. However, if 300 target data points are generated based on another target semantic unit, but only 30 are valid samples, and the semantic understanding error rate does not decrease significantly, then the data generation cost corresponding to that target semantic unit can be increased in the next round, and the scope of the diversity constraints can be narrowed.
[0080] This invention incorporates diversity constraints to ensure the generation of higher-quality target data. This allows the target data to cover scenarios with multiple intents, implicit expressions, driving safety concerns, and cross-vehicle capability differences, reducing the generation of highly similar redundant samples. Furthermore, based on the availability of the target data, an update strategy is used to update the data generation priority, data generation cost, and diversity constraints corresponding to the target semantic units. This enables the data generation process to adaptively adjust, ensuring the generation of more targeted data and achieving a closed loop between data generation and in-vehicle intelligent agent training.
[0081] Example 3 Figure 5 This is a schematic diagram of the structure of an in-vehicle intelligent agent data generation device provided in Embodiment 3 of the present invention. Figure 5As shown, the device includes: a semantic segmentation module 501, a gap information determination module 502, a priority calculation module 503, and a task generation module 504.
[0082] The semantic segmentation module is used to segment the historical corpus data in the historical intelligent cockpit dataset according to the preset semantic dimensions, obtain several semantic units, and statistically analyze the data distribution information corresponding to each semantic unit. The gap information determination module is used to determine the gap information of the semantic unit based on the data distribution information corresponding to the semantic unit; wherein, the gap information is used to characterize the degree of absence of historical corpus data currently contained in the semantic unit relative to a preset coverage standard; The priority calculation module is used to filter out target semantic units from the plurality of semantic units based on the gap information, and determine the data generation priority of the target semantic units according to the data distribution information of the target semantic units; The task generation module is used to generate a target data generation task corresponding to the target semantic unit based on the data generation priority, and generate target data based on the target data generation task; wherein, the target data is used to train the vehicle-mounted intelligent agent.
[0083] This invention provides a vehicle-mounted intelligent agent data generation device. By dividing historical corpus data in a historical intelligent cockpit dataset into several semantic units according to a preset semantic dimension, it achieves the division of the semantic dimension to which the historical corpus data belongs. By statistically analyzing the data distribution information of each semantic unit, the gap information corresponding to each semantic unit can be determined. This gap information characterizes the degree of absence of the historical corpus data currently contained in the semantic unit relative to a preset coverage standard, quantifying the semantic coverage and semantic absence of the historical corpus data. This provides guidance for subsequent data generation, ensuring that the subsequently generated target data can accurately fill the semantic gaps in the semantic units. Based on the gap information, target semantic units can be selected. By accurately determining the corresponding data generation priority based on the data distribution information of the target semantic unit, the importance of the target data generation task can be quantified, providing a reference for subsequent target data generation. This invention can accurately identify the gap information of historical corpus data, thereby accurately generating the target data required by the vehicle-mounted intelligent agent, avoiding blind data generation and generating a large number of highly similar redundant samples, thus improving the efficiency of data generation.
[0084] Optionally, the preset semantic dimensions include at least two of the following: semantic intent, semantic parameters, vehicle status, expression method, environmental status, task structure, safety level, and vehicle type capabilities; each semantic unit corresponds to at least two of the preset semantic dimensions.
[0085] Optionally, the data distribution information includes at least one of semantic coverage information and model inference error information; the gap information includes gap type; the semantic segmentation module includes: a semantic segmentation unit and at least one of the following statistical units; Among them, the semantic segmentation unit is used to divide the historical corpus data in the historical intelligent cockpit dataset according to the preset semantic dimension to obtain several semantic units; The first statistical unit is used to determine the semantic coverage information of each semantic unit based on the semantic dimensions covered by the historical corpus data in each semantic unit. The second statistical unit is used to determine the model inference error information of each semantic unit based on the semantic understanding results of the vehicle-mounted intelligent agent on the historical corpus data in each semantic unit. Optionally, the gap information determination module includes at least one of the following gap type determination units: The first gap type determination unit is used to compare the semantic coverage information corresponding to the semantic unit with the preset target coverage information to obtain a first comparison result, and determine the gap type of the semantic unit based on the first comparison result; The second gap type determination unit is used to compare the model inference error information corresponding to the semantic unit with the preset target inference error information to obtain a second comparison result, and determine the gap type of the semantic unit based on the second comparison result; The third gap type determination unit is used to compare the semantic coverage information and model inference error information corresponding to the semantic unit with the corresponding preset target coverage information and preset target inference error information, respectively, to obtain a third comparison result, and to determine the gap type of the semantic unit based on the third comparison result.
[0086] Optionally, the data distribution information further includes business-related information; the business-related information includes at least one of driving safety level, vehicle type importance, and similarity to existing data generation tasks; the gap type includes high gap type and low gap type, wherein the degree of missing information represented by the high gap type is higher than the degree of missing information represented by the low gap type; the priority calculation module includes: A filtering unit is used to filter out semantic units with a high gap type from the plurality of semantic units and determine the semantic units as target semantic units; A data generation cost determination unit is used to determine the data generation cost of the target semantic unit based on the semantic coverage information in the data distribution information; The data generation priority determination unit is used to determine the data generation priority of the target semantic unit based on the data generation cost and the business-related information in the data distribution information.
[0087] Optionally, the task generation module includes: The task generation unit is used to generate a target data generation task corresponding to the target semantic unit based on the data generation priority and the remaining generation resources; wherein, the remaining generation resources are used to represent the remaining generation resources of the data generation task in this round; The target data generation unit generates target data based on the target data generation task under the constraints of diversity constraints; wherein the diversity constraints are related to the preset semantic dimension corresponding to the target semantic unit.
[0088] Optionally, the device further includes: The qualification information determination module is used to verify the target data using preset verification conditions to obtain qualification information corresponding to the target data; wherein, the qualification information includes the number of qualified data and / or the qualification rate; the number of qualified data is the number of corpus data in the target data that passes the preset verification conditions; the qualification rate is the proportion of corpus data in the target data that passes the preset verification conditions; The available status information determination module is used to determine the available status information of the target data based on the qualification information and the training results of the vehicle-mounted intelligent agent; wherein, the training results include the semantic understanding error rate of the vehicle-mounted intelligent agent on the historical corpus data in the target semantic unit after training with the target data; An update module is used to update the data generation priority, data generation cost, and diversity constraints corresponding to the target semantic unit based on the availability status information of the target data and an update strategy. The update strategy includes: if the availability status information meets the corresponding preset availability conditions, then reducing the data generation priority, reducing the data generation cost, and maintaining or relaxing the constraint range of the diversity constraints; if the availability status information does not meet the corresponding preset availability conditions, then increasing the data generation priority, increasing the data generation cost, and narrowing the constraint range of the diversity constraints.
[0089] The vehicle-mounted intelligent agent data generation device provided in the embodiments of the present invention can execute the vehicle-mounted intelligent agent data generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0090] Example 4 Figure 6A schematic diagram of an electronic device 600 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0091] like Figure 6 As shown, the electronic device 600 includes at least one processor 601 and a memory, such as a read-only memory (ROM) 602 or a random access memory (RAM) 603, communicatively connected to the at least one processor 601. The memory stores computer programs executable by the at least one processor. The processor 601 can perform various appropriate actions and processes based on the computer program stored in the ROM 602 or loaded into the RAM 603 from storage unit 608. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0092] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0093] Processor 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 601 performs the various methods and processes described above, such as in-vehicle intelligent agent data generation methods.
[0094] In some embodiments, the in-vehicle intelligent agent data generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by processor 601, one or more steps of the in-vehicle intelligent agent data generation method described above may be performed. Alternatively, in other embodiments, processor 601 may be configured to perform the in-vehicle intelligent agent data generation method by any other suitable means (e.g., by means of firmware).
[0095] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0096] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0097] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0099] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0100] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0101] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the vehicle-mounted intelligent agent data generation method provided in the above embodiments.
[0102] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0103] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for generating data for an in-vehicle intelligent agent, characterized in that, include: The historical corpus data in the historical intelligent cockpit dataset is divided into several semantic units according to the preset semantic dimensions, and the data distribution information corresponding to each semantic unit is statistically analyzed. Based on the data distribution information corresponding to the semantic unit, the gap information of the semantic unit is determined; wherein, the gap information is used to characterize the degree of absence of historical corpus data currently contained in the semantic unit relative to a preset coverage standard; Based on the gap information, target semantic units are selected from the plurality of semantic units, and the data generation priority of the target semantic units is determined according to the data distribution information of the target semantic units; Based on the data generation priority, a target data generation task corresponding to the target semantic unit is generated, and target data is generated based on the target data generation task; wherein, the target data is used to train the vehicle-mounted intelligent agent.
2. The method for generating vehicle-mounted intelligent agent data according to claim 1, characterized in that, The preset semantic dimensions include at least two of the following: semantic intent, semantic parameters, vehicle status, expression mode, environmental status, task structure, safety level, and vehicle type capabilities; each semantic unit corresponds to at least two of the preset semantic dimensions.
3. The method for generating vehicle-mounted intelligent agent data according to claim 1, characterized in that, The data distribution information includes at least one of semantic coverage information and model inference error information; the gap information includes gap type; The statistical data distribution information corresponding to each semantic unit includes at least one of the following: Based on the semantic dimensions covered by historical corpus data in each semantic unit, determine the semantic coverage information of each semantic unit; Based on the semantic understanding results of the vehicle-mounted intelligent agent on the historical corpus data of each semantic unit, the model inference error information of each semantic unit is determined; Wherein, determining the gap information of the semantic unit based on the data distribution information corresponding to the semantic unit includes at least one of the following: The semantic coverage information corresponding to the semantic unit is compared with the preset target coverage information to obtain a first comparison result, and the gap type of the semantic unit is determined based on the first comparison result; The model inference error information corresponding to the semantic unit is compared with the preset target inference error information to obtain a second comparison result, and the gap type of the semantic unit is determined based on the second comparison result; The semantic coverage information and model inference error information corresponding to the semantic unit are compared with the corresponding preset target coverage information and preset target inference error information to obtain a third comparison result. The gap type of the semantic unit is determined based on the third comparison result.
4. The method for generating vehicle-mounted intelligent agent data according to claim 3, characterized in that, The data distribution information also includes business-related information; the business-related information includes at least one of driving safety level, vehicle type importance, and similarity to existing data generation tasks; the gap type includes high gap type and low gap type, the degree of missing information represented by the high gap type is higher than the degree of missing information represented by the low gap type; The step of selecting target semantic units from the plurality of semantic units based on the gap information, and determining the data generation priority of the target semantic units according to the data distribution information of the target semantic units, includes: From the aforementioned semantic units, semantic units with a high gap type are selected, and these semantic units are identified as target semantic units; The data generation cost of the target semantic unit is determined based on the semantic coverage information in the data distribution information; Based on the data generation cost and the business-related information in the data distribution information, the data generation priority of the target semantic unit is determined.
5. The method for generating vehicle-mounted intelligent agent data according to claim 1, characterized in that, The step of generating a target data generation task corresponding to the target semantic unit based on the data generation priority, and generating target data based on the target data generation task, includes: Based on the data generation priority and remaining generation resources, a target data generation task corresponding to the target semantic unit is generated; wherein, the remaining generation resources are used to represent the remaining generation resources of the data generation task in this round; Under the constraints of diversity conditions, target data is generated based on the target data generation task; wherein, the diversity constraints are related to the preset semantic dimension corresponding to the target semantic unit.
6. The method for generating vehicle-mounted intelligent agent data according to claim 5, characterized in that, The method further includes: The target data is verified using preset verification conditions to obtain the corresponding qualification information; wherein, the qualification information includes the number of qualified data and / or the qualification rate; the number of qualified data is the number of corpus data in the target data that passes the preset verification conditions; the qualification rate is the proportion of corpus data in the target data that passes the preset verification conditions; Based on the qualified information and the training results of the vehicle-mounted intelligent agent, the usability status information of the target data is determined; wherein, the training results include the semantic understanding error rate of the vehicle-mounted intelligent agent on the historical corpus data in the target semantic unit after training with the target data; Based on the availability status information of the target data, the data generation priority, data generation cost, and diversity constraints corresponding to the target semantic unit are updated according to an update strategy. The update strategy includes: if the availability status information meets the corresponding preset availability conditions, then the data generation priority is reduced, the data generation cost is reduced, and the constraint range of the diversity constraints is maintained or relaxed; if the availability status information does not meet the corresponding preset availability conditions, then the data generation priority is increased, the data generation cost is increased, and the constraint range of the diversity constraints is narrowed.
7. A vehicle-mounted intelligent agent data generation device, characterized in that, include: The semantic segmentation module is used to segment the historical corpus data in the historical intelligent cockpit dataset according to the preset semantic dimensions, obtain several semantic units, and statistically analyze the data distribution information corresponding to each semantic unit. The gap information determination module is used to determine the gap information of the semantic unit based on the data distribution information corresponding to the semantic unit; wherein, the gap information is used to characterize the degree of absence of historical corpus data currently contained in the semantic unit relative to a preset coverage standard; The priority calculation module is used to filter out target semantic units from the plurality of semantic units based on the gap information, and determine the data generation priority of the target semantic units according to the data distribution information of the target semantic units; The task generation module is used to generate a target data generation task corresponding to the target semantic unit based on the data generation priority, and generate target data based on the target data generation task; wherein, the target data is used to train the vehicle-mounted intelligent agent.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the vehicle-mounted intelligent agent data generation method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the vehicle-mounted intelligent agent data generation method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the vehicle-mounted intelligent agent data generation method according to any one of claims 1-6.