Data set generation method and device, electronic equipment, storage medium and vehicle
By analyzing and adjusting the user interaction record information of the on-board system, a high coverage model data set is generated, which solves the problem of low diversity in data set generation in the prior art and improves the model training efficiency.
Patent Information
- Application Number
- CN202411928175.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, there is a problem of low diversity in data set generation, resulting in inefficient model training.
By obtaining log information of the on-board system, extracting user interaction record information, analyzing target coverage, adjusting user interaction record information, and generating a model data set of the on-board system.
This improves the diversity of data sets, improves the efficiency of model training, and has a high coverage of the generated model data set.
Smart Images

Figure CN119988965A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automobile technology, and in particular to a data set generation method, device, electronic equipment, storage medium and vehicle. Background Art
[0002] With the rapid development of the automotive industry and intelligent interaction technology, the combination of vehicles and intelligence enables the use of intelligent voice interaction systems in vehicles, so that users in the vehicle can control related operations of the vehicle through voice. Specifically, the basis for realizing voice interaction is the voice model of the voice interaction system. Due to the particularity of the vehicle scenario, a large number of data sets related to the vehicle application scenario are required in the voice model construction stage to use the data set for model training and parameter adjustment. The data set plays a vital role in the model training of the voice interaction system. The traditional data set generation method usually uses the collected data as the data in the training data set after simple filtering, so that the data in the training data set can be used to train the voice model later. However, the collection process of existing data sets usually does not have a standard and comprehensive collection method, which leads to the problem of low diversity in the data sets collected in the existing related technologies. Summary of the invention
[0003] The present application provides a data set generation method, device, electronic device, storage medium and vehicle to solve the problem of low diversity in data set generation in the existing related technology.
[0004] In a first aspect, the present application provides a method for generating a data set, comprising:
[0005] Get the log information of the vehicle system;
[0006] Extracting user interaction record information of the vehicle-mounted system from the log information;
[0007] Analyze the user interaction record information to obtain the target coverage of the vehicle vertical domain;
[0008] According to the target coverage, the user interaction record information is adjusted to obtain target interaction record information;
[0009] Based on the target interaction record information, a model data set of the in-vehicle system is generated.
[0010] Optionally, analyzing the user interaction record information to obtain the target coverage of the vehicle-mounted vertical domain includes:
[0011] Classifying the user interaction record information by field to obtain record field label information;
[0012] Performing intent recognition on the user interaction record information to obtain record intent label information;
[0013] Extracting the user interaction record information by slot to obtain the record slot label information;
[0014] The vehicle-mounted vertical domain coverage analysis is performed based on the recording domain label information, the recording intent label information and the recording slot label information to obtain the target coverage.
[0015] Optionally, performing the vehicle-mounted vertical domain coverage analysis based on the recording domain label information, the recording intent label information, and the recording slot label information to obtain the target coverage includes:
[0016] Obtaining a vehicle-mounted vertical domain set corresponding to the vehicle-mounted system;
[0017] Based on the recording domain label information, the recording intent label information and the recording slot label information, coverage calculation is performed in combination with the vehicle-mounted vertical domain set to obtain the target coverage.
[0018] Optionally, the calculation of coverage based on the recording domain label information, the recording intent label information, and the recording slot label information in combination with the vehicle-mounted vertical domain set to obtain the target coverage includes:
[0019] Based on the vehicle-mounted vertical domain set, determine preset domain label information, preset intent label information, and preset slot label information;
[0020] Comparing the recorded domain label information with the preset domain label information to obtain domain coverage;
[0021] Comparing the recorded intent label information with the preset intent label information to obtain intent coverage;
[0022] Comparing the recorded slot label information with the preset slot label information to obtain slot coverage;
[0023] The target coverage is obtained based on the domain coverage, the intention coverage and the slot coverage.
[0024] Optionally, adjusting the user interaction record information according to the target coverage to obtain target interaction record information includes:
[0025] Get the preset coverage range;
[0026] When the target coverage is not within the coverage range, comparing the target coverage with the coverage range to obtain a coverage difference;
[0027] The user interaction record information is adjusted according to the coverage difference to obtain the target interaction record information.
[0028] Optionally, after obtaining the preset coverage range, the method further includes:
[0029] Determining whether the target coverage is within the coverage range;
[0030] When the target coverage is within the coverage range, the user interaction record information is determined as the target interaction record information.
[0031] Optionally, the adjusting the user interaction record information according to the coverage difference to obtain the target interaction record information includes:
[0032] When the coverage difference is a positive difference, the user interaction record information is deleted to obtain deleted user interaction record information, and the positive difference indicates that the target coverage is greater than the maximum coverage value within the coverage range;
[0033] The deleted user interaction record information is determined as the target interaction record information.
[0034] Optionally, the adjusting the user interaction record information according to the coverage difference to obtain the target interaction record information includes:
[0035] In the case where the coverage difference is a negative difference, the user interaction record information is added to obtain the added user interaction record information, and the negative difference indicates that the target coverage is less than the minimum coverage value within the coverage range;
[0036] The added user interaction record information is determined as the target interaction record information.
[0037] Optionally, the adding the user interaction record information to obtain the target interaction record information includes:
[0038] Extracting domain coverage, intent coverage, and slot coverage from the target coverage;
[0039] Determining data labels to be supplemented according to the field coverage, the intent coverage, and the slot coverage;
[0040] The to-be-supplemented data corresponding to the to-be-supplemented data tag is added to the user interaction record information to obtain the target interaction record information.
[0041] Optionally, generating a model data set of the vehicle-mounted system based on the target interaction record information includes:
[0042] Extracting sentence patterns from the target interaction record information to obtain a target sentence pattern set;
[0043] Extracting an initial multi-intention sentence from the target sentence set;
[0044] Adding punctuation marks to the initial multi-intent sentence to obtain a target multi-intent sentence;
[0045] Generate a target multi-intent sentence set according to the target multi-intent sentence and the initial multi-intent sentence;
[0046] The target multi-intention sentence set and the target sentence pattern set are combined to obtain the model data set.
[0047] In a second aspect, the present application provides a data set generation device, characterized by comprising:
[0048] The acquisition module is used to obtain the log information of the vehicle system;
[0049] An extraction module, used to extract user interaction record information of the vehicle-mounted system from the log information;
[0050] An analysis module, used to analyze the user interaction record information to obtain a target coverage of the vehicle vertical domain;
[0051] An adjustment module, configured to adjust the user interaction record information according to the target coverage to obtain target interaction record information;
[0052] A generation module is used to generate a model data set of the vehicle-mounted system based on the target interaction record information.
[0053] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0054] Memory, used to store computer programs;
[0055] The processor is used to implement the data set generation method described in any one of the first aspects when executing the program stored in the memory.
[0056] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for generating a data set as described in any one of the first aspects is implemented.
[0057] In a fifth aspect, a vehicle is provided, the vehicle comprising the data set generating device described in the fourth aspect.
[0058] The embodiment of the present application obtains the log information of the vehicle-mounted system, extracts the user interaction record information of the vehicle-mounted system from the log information, analyzes the user interaction record information, obtains the target coverage of the vehicle-mounted vertical domain, and adjusts the user interaction record information according to the target coverage to obtain the target interaction record information, and then generates a model data set corresponding to the vehicle-mounted system based on the target interaction record information; thereby, a model data set with a higher vehicle-mounted vertical domain coverage can be obtained, thereby solving the problem of low diversity in data set generation in the existing related technologies, and can effectively improve the diversity of the data set. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A flowchart of a method for generating a data set provided in an embodiment of the present application;
[0060] Figure 2 Another schematic diagram of a flow chart of a method for generating a data set provided in an embodiment of the present application;
[0061] Figure 3 A schematic diagram of an application scenario of a data set generation method provided in an embodiment of the present application;
[0062] Figure 4 A schematic diagram of the structure of a data set generating device provided in an embodiment of the present application;
[0063] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.
[0065] In the related art, the collected data is usually used to generate a data set, and the data set is used for model training. In the existing data set generation method, the collected data is usually simply filtered to use the data generated after filtering as the data in the data set, and then the model is trained through the data set. During the training process, the data composition of the data set can be adjusted according to the specific training situation, and this is repeated many times until the model training is completed. It can be seen that in the process of model training, multiple training and adjustment of the data set are required, the steps are cumbersome and the efficiency is low. Therefore, the low diversity of the data set will lead to the problem of low efficiency of model training.
[0066] In order to solve the problem of low diversity in data set generation in related technologies, and the problem of low model training efficiency caused by low data set diversity, the present application provides a data set generation method, device, electronic device, storage medium and vehicle, by obtaining log information of the vehicle-mounted system, and extracting user interaction record information of the vehicle-mounted system from the log information, analyzing the user interaction record information, and obtaining the target coverage of the vehicle-mounted vertical domain, and adjusting the user interaction record information according to the target coverage to obtain the target interaction record information, and then generating a model data set corresponding to the vehicle-mounted system based on the target interaction record information; thereby, a model data set with higher vehicle-mounted vertical domain coverage can be obtained, thereby solving the problem of low diversity in data set generation in existing related technologies, and can effectively improve the diversity of the data set, thereby improving the efficiency of subsequent model training.
[0067] Figure 1 A flow chart of a data set generation method provided for an embodiment of the present application. This method can be applied to one or more electronic devices such as vehicles, clients, and servers. In addition, the execution subject of this method can be hardware or software. When the above-mentioned execution subject is hardware, the execution subject can be one or more of the above-mentioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above-mentioned execution subject is software, this method can be implemented as multiple software or software modules, or as a single software or software module, which is not specifically limited here.
[0068] like Figure 1 As shown, a data set generation method provided in an embodiment of the present application may specifically include the following steps:
[0069] Step S110: Obtain log information of the vehicle-mounted system.
[0070] Among them, the vehicle-mounted system represents the vehicle-mounted system whose corresponding data set needs to be generated at present. The vehicle-mounted system can be the vehicle's on-board system, voice interaction system, etc. The log information of the vehicle-mounted system represents the log in the vehicle-mounted system, specifically, it can be the voice conversation log in the vehicle-mounted system, that is, the log information can include the interaction record between the user and the vehicle-mounted system, the vehicle-mounted system operation data, etc.
[0071] Step S120: extracting user interaction record information of the vehicle-mounted system from the log information.
[0072] Specifically, after obtaining the log information of the vehicle-mounted system, user interaction record information of the vehicle-mounted system can be extracted from the log information. The user interaction record information represents recorded data of the user's interaction with the vehicle-mounted system, such as the user's high-frequency voice commands, voice text and other information.
[0073] Step S130: Analyze the user interaction record information to obtain the target coverage of the vehicle vertical domain.
[0074] Specifically, after extracting the user interaction record information, the user interaction record information can be subjected to a vehicle-mounted vertical domain analysis to determine the target coverage of the user interaction record information covering the vehicle-mounted vertical domain corresponding to the vehicle-mounted system. The target coverage can indicate the extent to which the fields, intentions, and slots contained in the user interaction record information cover all fields, all intentions, and all slots corresponding to the vehicle-mounted system.
[0075] It should be noted that the vehicle-mounted vertical domain is also the vehicle-mounted vertical field, which includes three concepts: field, intention and slot.
[0076] Domain: In the in-vehicle vertical, "domain" refers to the specific business scope or functional module served by the dialogue system. For example, the in-vehicle system may include multiple domains such as music playback, navigation, and telephone communication. Each domain corresponds to a specific set of intents and slots for understanding and processing user requests in this domain.
[0077] Intent: "Intent" refers to the intention or purpose behind a user's statement. In an in-vehicle system, intent recognition is the classification of a user's natural language input into predefined intent categories. For example, if a user says "navigate to the nearest gas station", the intent is "navigate".
[0078] Slot: "Slot" refers to the specific parameters or pieces of information required to complete an intent. Slot filling is a task in the dialogue system to extract the values of these parameters from the user's statement. For example, for the intent "navigate to the nearest gas station", the slot may include "destination" and so on.
[0079] These three concepts together form the basis for the dialogue system to understand user requests, enabling the system to identify the task (intent) the user wants to perform and collect the specific information (slots) required to perform the task, ultimately providing users with accurate services in the in-vehicle vertical field.
[0080] It can be seen that different vehicle-mounted systems may have different vehicle-mounted vertical fields corresponding to different fields, intentions and slots, and the fields, intentions and slots included in the vehicle-mounted vertical fields of each specific vehicle-mounted system can be predetermined. For example, the vehicle-mounted system of type A vehicle has fields A, B and C, and the vehicle-mounted system of type B vehicle has fields A, B, C and D. It can be predetermined that the vehicle-mounted vertical fields of vehicle-mounted system A include fields A, B and C, and the vehicle-mounted vertical fields of vehicle-mounted system B include fields A, B, C and D. In the embodiment of the present application, the target coverage can be used to determine the extent to which the user interaction record information covers fields, intentions and slots in the corresponding vehicle-mounted vertical field of the vehicle-mounted system, that is, the target coverage can play a role in judging the diversity or comprehensiveness of the user interaction record information.
[0081] Step S140: adjusting the user interaction record information according to the target coverage to obtain target interaction record information.
[0082] Specifically, after determining the target coverage, the user interaction record information can be adjusted based on the target coverage to obtain target interaction record information. The target interaction record information can represent the data obtained after adjusting the specific data in the user interaction record information, wherein the adjustment can include at least one or more of the following: rewriting, adding, deleting, modifying and other adjustment operations, which are not specifically limited in this embodiment.
[0083] Step S150: Generate a model data set corresponding to the vehicle-mounted system based on the target interaction record information.
[0084] Specifically, after obtaining the target interaction record information, the target interaction record information can be integrated to generate a model data set corresponding to the vehicle system, where the integration method may include but is not limited to sentence adjustment, data cleaning and other operations, and the model data set may represent a data set used to train the corresponding model of the vehicle system.
[0085] It can be seen that this embodiment obtains the log information of the vehicle-mounted system, extracts the user interaction record information of the vehicle-mounted system from the log information, analyzes the user interaction record information, and obtains the target coverage of the vehicle-mounted vertical domain, so as to determine the diversity of the user interaction record information through the target coverage, and adjusts the user interaction record information according to the target coverage to obtain the target interaction record information, and then generates a model data set corresponding to the vehicle-mounted system based on the target interaction record information; thereby, a model data set with higher vehicle-mounted vertical domain coverage can be obtained, thereby solving the problem of low diversity in data set generation in the existing related technologies, and can effectively improve the diversity of the data set, thereby improving the efficiency of subsequent model training.
[0086] In an optional embodiment of the present application, step S130 analyzes the user interaction record information to obtain the target coverage of the vehicle-mounted vertical domain, which may specifically include the following steps: performing field classification on the user interaction record information to obtain record field label information; performing intent recognition on the user interaction record information to obtain record intent label information; performing slot extraction on the user interaction record information to obtain record slot label information; performing vehicle-mounted vertical domain coverage analysis based on the record field label information, the record intent label information and the record slot label information to obtain the target coverage.
[0087] After obtaining the user interaction record information, this embodiment can perform field classification on the user interaction record information to obtain record field label information; the field classification method can be to input the user interaction record information into the field classification model corresponding to the vehicle-mounted vertical domain, and obtain the record field label information corresponding to the user interaction record information based on the model output, wherein the record field label information can represent the label of the field to which the user interaction record information corresponds. Thus, the user interaction record information can be subjected to intent recognition to obtain record intention label information; wherein the intent recognition method can be to input the user interaction record information into the intent recognition model corresponding to the vehicle-mounted vertical domain, and obtain the record intention label information corresponding to the user interaction record information based on the model output, wherein the record intention label information can represent the label of the intent to which the user interaction record information corresponds. Then, the user interaction record information can be subjected to slot extraction to obtain record slot label information; wherein the slot extraction method can be to input the user interaction record information into the slot extraction model corresponding to the vehicle-mounted vertical domain, and obtain the record slot label information corresponding to the user interaction record information based on the model output, wherein the record slot label information can represent the label of the slot to which the user interaction record information corresponds. After obtaining the recording field label information, the recording intention label information and the recording slot label information, the vehicle-mounted vertical domain coverage analysis can be performed based on the recording field label information, the recording intention label information and the recording slot label information to obtain the target coverage; wherein, the vehicle-mounted vertical domain coverage analysis can be to determine the recording field label information, the recording intention label information and the recording slot label information, and analyze the coverage degree of all labels corresponding to the vehicle-mounted vertical domain. The analysis method can be a preset formula, algorithm, network model, etc., and this embodiment does not make any specific limitations on this.
[0088] It should be noted that, since the user interaction record information may contain a large amount of user interaction information, that is, domain classification, intent recognition and slot extraction may be performed for each user interaction information in the user interaction record information; for example, the user interaction information contained in the user interaction record information includes "turn on the air conditioner" and "open the car window", and the recorded domain label information obtained by domain classification of "turn on the air conditioner" is "vehicle control", the recorded intent label information obtained by intent recognition of "turn on the air conditioner" is "turn on device", and the recorded slot label information obtained by slot extraction of "turn on the air conditioner" is "air conditioning", the recorded domain label information obtained by domain classification of "open the car window" is "vehicle control", the recorded intent label information obtained by intent recognition of "open the car window" is "turn on device", and the recorded slot label information obtained by slot extraction of "open the car window" is "car window", so that the vehicle-mounted vertical domain coverage analysis can be performed based on the recorded domain label information, recorded intent label information and recorded slot label information obtained by domain classification, intent recognition and slot extraction of all user interaction information in the user interaction record information to obtain the target coverage. Of course, the above is only an example to illustrate the effect, and this embodiment does not make any specific limitation to this.
[0089] In an optional embodiment of the present application, a vehicle-mounted vertical domain coverage analysis is performed based on the recording domain label information, the recording intent label information and the recording slot label information to obtain a target coverage, which may specifically include the following steps: obtaining a vehicle-mounted vertical domain set corresponding to the vehicle-mounted system; performing coverage calculation based on the recording domain label information, the recording intent label information and the recording slot label information in combination with the vehicle-mounted vertical domain set to obtain a target coverage.
[0090] In the process of performing vehicle vertical domain coverage analysis in this embodiment, the vehicle vertical domain set corresponding to the vehicle system can be obtained. The vehicle vertical domain set can represent the set of all domain labels, intent labels and slot labels contained in the vehicle vertical domain corresponding to the vehicle system. Then, based on the recorded domain label information, the recorded intent label information and the recorded slot label information, the coverage calculation can be performed in combination with the vehicle vertical domain set to obtain the target coverage. The coverage calculation is to calculate the recorded domain label information, the recorded intent label information and the recorded slot label information, which can calculate the coverage of all labels corresponding to the vehicle vertical domain set.
[0091] In an optional embodiment of the present application, based on recording domain label information, recording intent label information and recording slot label information, coverage calculation is performed in combination with the on-board vertical domain set to obtain target coverage, which may specifically include the following steps: based on the on-board vertical domain set, determine the preset domain label information, preset intent label information and preset slot label information; compare the recorded domain label information with the preset domain label information to obtain domain coverage; compare the recorded intent label information with the preset intent label information to obtain intent coverage; compare the recorded slot label information with the preset slot label information to obtain slot coverage; obtain target coverage based on the domain coverage, intent coverage and slot coverage.
[0092] In the specific process of coverage calculation in this embodiment, the preset domain label information, preset intention label information and preset slot label information can be determined based on the vehicle-mounted vertical domain set, wherein the preset domain label information can indicate all domain labels contained in the vehicle-mounted vertical domain corresponding to the vehicle-mounted system, the preset intention label information can indicate all intention labels contained in the vehicle-mounted vertical domain corresponding to the vehicle-mounted system, and the preset slot label information can indicate all slot labels contained in the vehicle-mounted vertical domain corresponding to the vehicle-mounted system; thereby, the recorded domain label information can be compared with the preset domain label information to obtain the domain coverage, and the domain coverage can indicate the recorded domain label information. The coverage degree of all domain labels corresponding to the recorded domain label information and the preset domain label information is compared; the recorded intent label information is compared with the preset intent label information to obtain the intent coverage, which can indicate the coverage degree of all intent labels corresponding to the recorded intent label information and the preset intent label information; the recorded slot label information is compared with the preset slot label information to obtain the slot coverage, which can indicate the coverage degree of all domain labels corresponding to the recorded slot label information and the preset slot label information; then the domain coverage, intent coverage and slot coverage can be integrated to obtain the target coverage.
[0093] In one example, the preset domain tag information includes domain tag A, domain tag B, and domain tag C, the preset intent tag information includes intent tag A, intent tag B, and intent tag C, and the preset slot tag information includes slot tag A, slot tag B, and slot tag C; while the recorded domain tag information includes domain tag A and domain tag B, the recorded intent tag information includes intent tag A and intent tag B, and the recorded slot tag information includes slot tag A and slot tag B; at this time, the recorded domain tag information is compared with the preset domain tag information to determine the coverage of all domain tags corresponding to the recorded domain tag information and the preset domain tag information. The coverage of the domain is 2 / 3, that is, the obtained domain coverage is 2 / 3; by comparing the recorded intent label information with the preset intent label information, it can be determined that the coverage of the recorded intent label information and the preset intent label information corresponding to all intent labels is 2 / 3, that is, the obtained intent coverage is 2 / 3; by comparing the recorded slot label information with the preset slot label information, it can be determined that the coverage of the recorded slot label information and the preset slot label information corresponding to all slot labels is 2 / 3, that is, the obtained slot coverage is 2 / 3; then the domain coverage of 2 / 3, the intent coverage of 2 / 3 and the slot coverage of 2 / 3 can be integrated to obtain the target coverage. Of course, the above is only an example, and this embodiment does not make any specific limitation on this.
[0094] It should be noted that in the process of integrating domain coverage, intent coverage and slot coverage to obtain target coverage, different integration methods can be used according to needs. For example, the set of domain coverage, intent coverage and slot coverage can be used as target coverage, or the domain coverage, intent coverage and slot coverage can be calculated to obtain a target coverage representing the overall coverage.
[0095] In one example, after calculating the domain coverage, intent coverage, and slot coverage to obtain a target coverage representing the overall coverage, the target coverage can be calculated using the following formula:
[0096]
[0097] Where D is the set of all domain labels corresponding to the preset domain label information, I d is the set of all intentions in the domain d∈D, that is, I d The preset intent label information may correspond to the set of all intent labels; S i,d is the set of all slot types under domain d and intention i∈I, that is, S i,d The preset slot label information may correspond to the set of all slot labels; C s,i,d For domain d intention i and slot type s∈Si,d In the example of recording the intent tag information; T s,i,d For domain d intention i and slot type s∈S i,d The vehicle-mounted system corresponds to all instances contained in the vehicle-mounted vertical domain. Therefore, Coverage is used as the target coverage to determine the coverage of the record intention label information for all instances contained in the vehicle-mounted vertical domain, that is, the purpose of determining the diversity of the record intention label information through the target coverage is achieved.
[0098] In an optional embodiment of the present application, step S140 adjusts the user interaction record information according to the target coverage to obtain the target interaction record information, which may specifically include the following steps: obtaining a preset coverage range; when the target coverage is not within the coverage range, comparing the target coverage with the coverage range to obtain the coverage difference; adjusting the user interaction record information according to the coverage difference to obtain the target interaction record information.
[0099] After obtaining the target coverage, this embodiment can obtain a preset coverage range, where the coverage range represents a pre-configured range that meets the required coverage, and determine whether the target coverage is within the coverage range. When the target coverage is not within the coverage range, it means that the interaction data contained in the current user interaction record information does not meet the preset requirements. At this time, the target coverage can be compared with the coverage range to obtain the coverage difference, which can represent the difference between the target coverage and the maximum coverage or minimum coverage contained in the coverage range; thereby, the user interaction record information can be adaptively adjusted according to the coverage difference to obtain the target interaction record information.
[0100] In one example, since the target coverage can be a set of domain coverage, intention coverage and slot coverage, or a target coverage representing the overall coverage is obtained after calculating the domain coverage, intention coverage and slot coverage; when the target coverage is a set of domain coverage, intention coverage and slot coverage, the preset coverage range can include a set of domain coverage range corresponding to the domain coverage, intention coverage range corresponding to the intention coverage and slot coverage range corresponding to the slot coverage, that is, when judging whether the target coverage is within the coverage range, it can be judged in turn whether the domain coverage is within the domain coverage range, whether the intention coverage is within the intention coverage range and whether the slot coverage is within the slot coverage range. In any case where the domain coverage is not within the domain coverage range, the intention coverage is not within the intention coverage range or the slot coverage is not within the slot coverage range, it is determined that the target coverage is not within the coverage range. If the target coverage is a coverage rate representing the whole, the preset coverage range is also a range of the overall coverage rate, and it can be directly judged whether the target coverage is within the coverage range. Of course, the above is only an example to illustrate the effect, and this embodiment does not make any specific limitation to this.
[0101] In an optional embodiment of the present application, after obtaining the preset coverage range, the following steps may be specifically included: determining whether the target coverage is within the coverage range; if the target coverage is within the coverage range, determining the user interaction record information as the target interaction record information.
[0102] After obtaining the preset coverage range, this embodiment can also determine whether the target coverage is within the coverage range. The determination of whether the target coverage is within the coverage range in this embodiment is similar to the aforementioned determination method and will not be repeated here. When the target coverage is within the coverage range, it means that the interaction data contained in the current user interaction record information meets the preset requirements. At this time, the user interaction record information can be directly determined as the target interaction record information.
[0103] In an optional embodiment of the present application, the user interaction record information is adjusted according to the coverage difference to obtain the target interaction record information, which may specifically include the following steps: when the coverage difference is a positive difference, the user interaction record information is deleted to obtain the deleted user interaction record information, the positive difference indicates that the target coverage is greater than the maximum coverage value within the coverage range; the deleted user interaction record information is determined as the target interaction record information.
[0104] In the process of adjusting the user interaction record information according to the coverage difference in this embodiment, it can be determined whether the coverage difference is a positive difference. A positive difference can indicate that the target coverage is greater than the maximum coverage value within the coverage range. When the coverage difference is a positive difference, it means that the record data contained in the user interaction record information is relatively wide, that is, there may be a problem of bloated and cumbersome data. At this time, the user interaction record information can be deleted to obtain the deleted user interaction record information, and the deleted user interaction record information can be determined as the target interaction record information. The specific deletion method can be to delete data with too many data types, to delete data with incomplete record information, etc. This embodiment does not make specific limitations on this.
[0105] In an optional embodiment of the present application, the user interaction record information is adjusted according to the coverage difference to obtain the target interaction record information, which may specifically include the following steps: when the coverage difference is a negative difference, the user interaction record information is added to obtain the added user interaction record information, the negative difference indicating that the target coverage is less than the minimum coverage value within the coverage range; the added user interaction record information is determined as the target interaction record information.
[0106] In the process of adjusting the user interaction record information according to the coverage difference in this embodiment, it can be determined whether the coverage difference is a negative difference. The negative difference can indicate that the target coverage is less than the minimum coverage value within the coverage range. When the coverage difference is a negative difference, it means that the user interaction record information contains less record data, that is, there is a problem of low diversity. At this time, the user interaction record information can be added to obtain the added user interaction record information, and the added user interaction record information can be determined as the target interaction record information. The specific method of adding can be to obtain other interaction record information again and add it to the user interaction record information, or to manually generate new interaction record information and add it to the user interaction record information. Of course, it can also be other methods, and this embodiment does not make specific limitations on this.
[0107] In an optional embodiment of the present application, user interaction record information is supplemented to obtain target interaction record information, which may specifically include the following steps: extracting domain coverage, intent coverage and slot coverage from target coverage; determining data labels to be supplemented based on domain coverage, intent coverage and slot coverage; adding the data to be supplemented corresponding to the data labels to be supplemented to the user interaction record information to obtain target interaction record information.
[0108] In the process of adding user interaction record information in this embodiment, domain coverage, intention coverage and slot coverage can be extracted from the target coverage, and then the domain coverage range, intention coverage range and slot coverage range corresponding to the coverage range can be determined, and whether the domain coverage is within the domain coverage range, whether the intention coverage is within the intention coverage range and whether the slot coverage is within the slot coverage range can be determined in turn. When the domain coverage is not within the domain coverage range, the recorded domain label information corresponding to the user interaction record information and the preset domain label information corresponding to the on-board vertical domain set can be compared to obtain supplementary domain label information. The supplementary domain label information can represent a domain label that exists in the preset domain label information but does not exist in the recorded domain label information, and the supplementary domain label information is determined as the data label to be supplemented. When the domain coverage is within the domain coverage range, but the intention coverage is not within the intention coverage range, the recorded intention label information corresponding to the user interaction record information and the preset domain label information corresponding to the on-board vertical domain set can be compared. The preset intention label information is used to obtain supplementary intention label information. The supplementary intention label information can indicate an intention label that exists in the preset intention label information but does not exist in the recorded intention label information, and the supplementary intention label information is determined as the data label to be supplemented; and when the domain coverage is within the domain coverage range, the intention coverage is within the intention coverage range, but the slot coverage is not within the slot coverage range, the recorded slot label information corresponding to the user interaction record information and the preset slot label information corresponding to the vehicle-mounted vertical domain set can be compared to obtain supplementary slot label information. The supplementary slot label information can indicate a slot label that exists in the preset slot label information but does not exist in the recorded slot label information, and the supplementary slot label information is determined as the data label to be supplemented; then the data to be supplemented corresponding to the data label to be supplemented can be obtained, and the data to be supplemented can be added to the user interaction record information to obtain the target interaction record information, so that the user interaction record information can be accurately and effectively supplemented, and then the target interaction record information with higher vehicle-mounted vertical domain coverage can be obtained.
[0109] like Figure 2 As shown, in an optional embodiment of the present application, step S150 generates a model data set corresponding to the vehicle-mounted system based on the target interaction record information, which may specifically include the following steps:
[0110] Step S151: extracting sentence patterns from the target interaction record information to obtain a target sentence pattern set;
[0111] Step S152: extracting an initial multi-intention sentence from the target sentence set;
[0112] Step S153: adding punctuation marks to the initial multi-intention sentence to obtain a target multi-intention sentence;
[0113] Step S154: Generate a target multi-intention sentence set based on the target multi-intention sentence and the initial multi-intention sentence;
[0114] Step S155: Combine the target multi-intention sentence set and the target sentence pattern set to obtain a model data set.
[0115] After obtaining the target interaction record information, this embodiment can perform sentence extraction on the target interaction record information to obtain a target sentence set, wherein the sentence extraction can represent the process of extracting multiple different sentence patterns for each intent. For example, if the target interaction record information includes the intent "navigate to destination", regular expressions can be used to match sentence patterns such as "navigate to [destination]" and "go to [destination]", so that the target sentence set can include different sentence patterns corresponding to each intent, thereby ensuring the diversity of the data set. Then, an initial multi-intent sentence can be extracted from the target sentence set. The initial multi-intent sentence can represent a sentence containing at least two intents, such as "I want to listen to music and navigate to the phone at the same time." Field", "Adjust the temperature to 22 degrees and close the windows", etc., and the specific extraction process can be to determine whether each sentence in the target sentence set is a multi-intention sentence in turn; thereby, the initial multi-intention sentence can be processed with punctuation marks to obtain a target multi-intention sentence, and the target multi-intention sentence can represent a multi-intention sentence with punctuation marks added, wherein the punctuation mark addition process can be added by manual addition, neural network model, algorithm, etc., which is not specifically limited in this embodiment; then the target multi-intention sentence and the initial multi-intention sentence can be integrated to generate a target multi-intention sentence set, and the target multi-intention sentence set and the target sentence set can be combined to obtain a model data set.
[0116] It should be noted that, in the process of generating a target multi-intent sentence set based on the target multi-intent sentence and the initial multi-intent sentence, the present embodiment can also establish contact label information between the target multi-intent sentence and the initial multi-intent sentence, and the contact label information is used to represent the correspondence between the target multi-intent sentence and the initial multi-intent sentence, so that when the model data set containing the target multi-intent sentence and the initial multi-intent sentence is subsequently used for model training, the model can be trained and recognized based on the target multi-intent sentence, the initial multi-intent sentence and the contact label information to train the model's ability to add punctuation marks, thereby solving the problem that the sentences after speech recognition in the existing in-vehicle system smart cockpit do not have punctuation marks, resulting in the inability of the speech model to accurately and efficiently analyze the multi-intent sentences. The model trained with the model data set of the present embodiment can accurately and efficiently add punctuation marks to multi-intent sentences without punctuation marks, which greatly improves the completeness of the model training during intent recognition and analysis.
[0117] In addition, after extracting sentences from the target interaction record information and obtaining the target sentence set, the target sentence set can also be cleaned to obtain the target sentence set after data cleaning. The subsequent steps are performed on the target sentence set after data cleaning; wherein, data cleaning may include but is not limited to: checking and deleting repeated sentences to ensure that each data in the data set is unique, correcting spelling errors or grammatical errors in the data to ensure the accuracy of the data, and standardizing text containing special characters or marks such as date format, number format, etc. to ensure data consistency.
[0118] It can be seen that the model data set obtained in the embodiment of the present application can have diversity while ensuring the accuracy and uniformity of the data in the model data set, so that the model training efficiency can be greatly improved when the model data set is used to train the model. For example, after the model is trained through the model data set to obtain the target large language model, in order to improve the large language model's understanding and execution capabilities of the smart cockpit user command rewriting task, a structured normalized mapping learning framework can be used to adjust the large language model's output. Specifically, the user's spontaneous, non-standardized voice commands can be used as the input of the large model, and standard commands can be used as the output. A prompt mechanism is introduced to guide the large language model to generate more accurate and expected standard commands; such as Figure 3 As shown, the user input "Hahaha, turn on the music for me, it's okay, turn on the silent mode" is standardized and output as "command: turn on the music i turn on the silent mode".
[0119] like Figure 4 As shown, the present application also discloses an embodiment, providing a data set generating device, including:
[0120] The acquisition module 410 is used to acquire log information of the vehicle-mounted system;
[0121] An extraction module 420, configured to extract user interaction record information of the vehicle-mounted system from the log information;
[0122] An analysis module 430 is used to analyze the user interaction record information to obtain a target coverage of the vehicle vertical domain;
[0123] An adjustment module 440, configured to adjust the user interaction record information according to the target coverage to obtain target interaction record information;
[0124] The generation module 450 is used to generate a model data set of the vehicle-mounted system based on the target interaction record information.
[0125] In an optional embodiment of the present application, the analysis module 430 may include:
[0126] A domain classification unit, used to classify the user interaction record information by domain to obtain record domain label information;
[0127] An intention recognition unit, used to perform intention recognition on the user interaction record information to obtain record intention label information;
[0128] A slot extraction unit, used to extract the slot of the user interaction record information to obtain the record slot label information;
[0129] The vehicle-mounted vertical domain coverage analysis unit is used to perform the vehicle-mounted vertical domain coverage analysis based on the recording field label information, the recording intent label information and the recording slot label information to obtain the target coverage.
[0130] In an optional embodiment of the present application, the vehicle-mounted vertical coverage analysis unit may include:
[0131] An acquisition subunit, used to acquire a vehicle-mounted vertical domain set corresponding to the vehicle-mounted system;
[0132] The coverage calculation subunit is used to perform coverage calculation based on the recording field label information, the recording intent label information and the recording slot label information in combination with the vehicle-mounted vertical domain set to obtain the target coverage.
[0133] In an optional embodiment of the present application, the coverage calculation subunit may include:
[0134] A first determination subunit is used to determine preset domain label information, preset intent label information and preset slot label information based on the vehicle-mounted vertical domain set;
[0135] A first comparison subunit, configured to compare the recorded domain label information with the preset domain label information to obtain domain coverage;
[0136] A second comparison subunit is used to compare the recorded intention label information with the preset intention label information to obtain intention coverage;
[0137] A third comparison subunit, used to compare the recorded slot label information with the preset slot label information to obtain slot coverage;
[0138] The second determination subunit is used to obtain the target coverage according to the field coverage, the intention coverage and the slot coverage.
[0139] In an optional embodiment of the present application, the adjustment module 440 may include:
[0140] An acquisition unit, used for acquiring a preset coverage range;
[0141] A first comparison unit is used to compare the target coverage with the coverage range to obtain a coverage difference when the target coverage is not within the coverage range;
[0142] An adjustment unit is used to adjust the user interaction record information according to the coverage difference to obtain the target interaction record information.
[0143] In an optional embodiment of the present application, the adjustment module 440 may further include:
[0144] A first determining unit, configured to determine whether the target coverage is within the coverage range;
[0145] A second determining unit is configured to determine the user interaction record information as the target interaction record information when the target coverage is within the coverage range.
[0146] In an optional embodiment of the present application, the adjustment unit may include:
[0147] a deletion subunit, configured to delete the user interaction record information to obtain deleted user interaction record information when the coverage difference is a positive difference, wherein the positive difference indicates that the target coverage is greater than a maximum coverage value within the coverage range;
[0148] The third determining subunit is configured to determine the deleted user interaction record information as the target interaction record information.
[0149] In an optional embodiment of the present application, the adjustment unit may include:
[0150] an adding subunit, configured to add the user interaction record information to obtain the added user interaction record information when the coverage difference is a negative difference, wherein the negative difference indicates that the target coverage is less than a minimum coverage value within the coverage range;
[0151] The fourth determining subunit is configured to determine the added user interaction record information as the target interaction record information.
[0152] In an optional embodiment of the present application, the adding subunit may include:
[0153] An extraction subunit, used to extract domain coverage, intent coverage and slot coverage from the target coverage;
[0154] A fifth determination subunit, configured to determine a label of data to be supplemented according to the domain coverage, the intention coverage, and the slot coverage;
[0155] The target subunit is used to add the to-be-supplemented data corresponding to the to-be-supplemented data tag to the user interaction record information to obtain the target interaction record information.
[0156] In an optional embodiment of the present application, the generating module 450 may include:
[0157] A first extraction unit, configured to extract sentence patterns from the target interaction record information to obtain a target sentence pattern set;
[0158] A second extraction unit, configured to extract an initial multi-intention sentence from the target sentence set;
[0159] a punctuation mark adding processing unit, configured to perform punctuation mark adding processing on the initial multi-intention sentence to obtain a target multi-intention sentence;
[0160] A generating unit, configured to generate a target multi-intention sentence set according to the target multi-intention sentence and the initial multi-intention sentence;
[0161] A combining unit is used to combine the target multi-intention sentence set and the target sentence set to obtain the model data set.
[0162] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0163] like Figure 5 As shown, an embodiment of the present application provides an electronic device, including a processor 510, a communication interface 520, a memory 550 and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 550 communicate with each other through the communication bus 540.
[0164] Memory 550, for storing computer programs;
[0165] In one embodiment of the present application, the processor 510, when executing the program stored in the memory 550, implements the data set generation method provided by any of the aforementioned method embodiments, by obtaining the log information of the vehicle-mounted system, and extracting the user interaction record information of the vehicle-mounted system from the log information, analyzing the user interaction record information, and obtaining the target coverage of the vehicle-mounted vertical domain, thereby determining the diversity of the user interaction record information through the target coverage, and adjusting the user interaction record information according to the target coverage to obtain the target interaction record information, and then generating a model data set corresponding to the vehicle-mounted system based on the target interaction record information; thereby, a model data set with higher vehicle-mounted vertical domain coverage can be obtained, thereby solving the problem of low diversity in data set generation in the existing related technologies, and can effectively improve the diversity of the data set, thereby improving the efficiency of subsequent model training.
[0166] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of a data set generation method provided in any of the aforementioned method embodiments are implemented, by obtaining log information of the vehicle-mounted system, and extracting user interaction record information of the vehicle-mounted system from the log information, analyzing the user interaction record information, and obtaining a target coverage of the vehicle-mounted vertical domain, thereby determining the diversity of the user interaction record information through the target coverage, adjusting the user interaction record information according to the target coverage, and obtaining the target interaction record information, and then generating a model data set corresponding to the vehicle-mounted system based on the target interaction record information; thereby, a model data set with a higher vehicle-mounted vertical domain coverage can be obtained, thereby solving the problem of low diversity in data set generation in the existing related technologies, and being able to effectively improve the diversity of the data set, thereby improving the efficiency of subsequent model training.
[0167] An embodiment of the present application also provides a vehicle, the vehicle comprising the data set generating device described in any of the aforementioned embodiments.
[0168] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0170] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "include", "comprise", "contain", and "have" are inclusive, and therefore specify the existence of stated features, steps, operations, elements and / or parts, but do not exclude the existence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not interpreted as necessarily requiring them to be performed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0171] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for generating a data set, characterized in that: include: Get the log information of the vehicle system; Extracting user interaction record information of the vehicle-mounted system from the log information; Analyze the user interaction record information to obtain the target coverage of the vehicle vertical domain; According to the target coverage, the user interaction record information is adjusted to obtain target interaction record information; Based on the target interaction record information, a model data set of the in-vehicle system is generated.
2. The method for generating a data set according to claim 1, characterized in that: The analyzing the user interaction record information to obtain the target coverage of the vehicle vertical domain includes: Classifying the user interaction record information by field to obtain record field label information; Performing intent recognition on the user interaction record information to obtain record intent label information; Extracting the user interaction record information by slot to obtain the record slot label information; Based on the recording domain label information, the recording intent label information and the recording slot label information, a vehicle-mounted vertical domain coverage analysis is performed to obtain the target coverage.
3. The method for generating a data set according to claim 2, characterized in that: Based on the recording domain label information, the recording intent label information, and the recording slot label information, a vehicle vertical domain coverage analysis is performed to obtain the target coverage, including: Obtaining a vehicle-mounted vertical domain set corresponding to the vehicle-mounted system; Based on the recording domain label information, the recording intent label information and the recording slot label information, coverage calculation is performed in combination with the vehicle-mounted vertical domain set to obtain the target coverage.
4. The method for generating a data set according to claim 3, characterized in that: The calculation of coverage based on the recording domain label information, the recording intent label information, and the recording slot label information in combination with the vehicle-mounted vertical domain set to obtain the target coverage includes: Based on the vehicle-mounted vertical domain set, determine preset domain label information, preset intent label information, and preset slot label information; Comparing the recorded domain label information with the preset domain label information to obtain domain coverage; Comparing the recorded intent label information with the preset intent label information to obtain intent coverage; Comparing the recorded slot label information with the preset slot label information to obtain slot coverage; The target coverage is determined based on the domain coverage, the intention coverage and the slot coverage.
5. The method for generating a data set according to claim 1, characterized in that: The step of adjusting the user interaction record information according to the target coverage to obtain target interaction record information includes: Get the preset coverage range; When the target coverage is not within the coverage range, comparing the target coverage with the coverage range to obtain a coverage difference; The user interaction record information is adjusted according to the coverage difference to obtain the target interaction record information.
6. The method for generating a data set according to claim 5, characterized in that: The adjusting the user interaction record information according to the coverage difference to obtain the target interaction record information includes: When the coverage difference is a positive difference, the user interaction record information is deleted to obtain deleted user interaction record information, and the positive difference indicates that the target coverage is greater than the maximum coverage value within the coverage range; The deleted user interaction record information is determined as the target interaction record information.
7. The method for generating a data set according to claim 5, characterized in that: The adjusting the user interaction record information according to the coverage difference to obtain the target interaction record information includes: In the case where the coverage difference is a negative difference, the user interaction record information is added to obtain the added user interaction record information, and the negative difference indicates that the target coverage is less than the minimum coverage value within the coverage range; The added user interaction record information is determined as the target interaction record information.
8. The method for generating a data set according to claim 7, characterized in that: The adding of the user interaction record information to obtain the added user interaction record information includes: Extracting domain coverage, intent coverage, and slot coverage from the target coverage; Determining data labels to be supplemented according to the field coverage, the intent coverage, and the slot coverage; The data to be supplemented corresponding to the data tag to be supplemented is added to the user interaction record information to obtain the added user interaction record information.
9. The method for generating a data set according to any one of claims 1 to 8, characterized in that: The generating a model data set of the vehicle-mounted system based on the target interaction record information includes: Extracting sentence patterns from the target interaction record information to obtain a target sentence pattern set; Extracting an initial multi-intention sentence from the target sentence set; Adding punctuation marks to the initial multi-intent sentence to obtain a target multi-intent sentence; Generate a target multi-intent sentence set according to the target multi-intent sentence and the initial multi-intent sentence; The target multi-intention sentence set and the target sentence pattern set are combined to obtain the model data set.
10. A data set generating device, characterized in that: include: The acquisition module is used to obtain the log information of the vehicle system; An extraction module, used to extract user interaction record information of the vehicle-mounted system from the log information; An analysis module, used to analyze the user interaction record information to obtain a target coverage of the vehicle vertical domain; An adjustment module, configured to adjust the user interaction record information according to the target coverage to obtain target interaction record information; A generation module is used to generate a model data set of the vehicle-mounted system based on the target interaction record information.
11. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor is used to implement the data set generation method described in any one of claims 1 to 9 when executing a program stored in a memory.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data set generation method according to any one of claims 1 to 9 is implemented.
13. A vehicle, characterized in that: The vehicle comprises the data set generating device according to claim 10.
Citation Information
Cited By
Multi-round dialogue sample data generation method, device, equipment and product
CN120541528A
A method, device, equipment and product for generating multi-turn dialogue sample data
CN120541528B