A data processing method, device, equipment and storage medium
By acquiring multimodal data in the vehicle, determining the on-board scenarios and generating structured data sets, the high cost problem during multimodal data processing is solved, and efficient data fusion and computing speed are achieved.
Patent Information
- Application Number
- CN202210077976.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-01-24
AI Technical Summary
In the prior art, a large amount of data calculation is required for multimodal data processing, which consumes a lot of manpower and material costs, and it is difficult to effectively integrate the various modes.
By obtaining the multimodal data collected by the vehicle adaptation layer, determining the on-board scenario, combining the label data to generate descriptive statements, and determining the target feature parameters based on the statements, generating a structured data set, which is directly used for network model training, reducing the individual feature extraction between modals.
It improves data computing speed, reduces the memory usage and CPU usage of computing devices, and reduces data processing costs.
Smart Images

Figure CN114416996B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing technology, and in particular to a data processing method, apparatus, device, and storage medium. Background Art
[0002] In today's society, with a growing population and rapid economic development, the number of cars on the road is increasing year by year, and the demand for driving comfort is also increasing. Currently, the data collected by mass-produced vehicles mostly comes from users performing single-mode operations, such as voice, gestures, and screen touches. The data collected is often limited in dimension and has its own limitations.
[0003] To better understand user driving intent and provide better services, a growing number of automakers are looking to leverage multimodal data to address these challenges. However, multimodal data presents numerous challenges, including large data volumes, high dimensionality, significant differences between modalities, making it difficult to rationalize and complement each other, and modal redundancy. To achieve better integration between modalities, it's often necessary to train models corresponding to each modality and synthesize the outputs of multiple models to determine the user's driving intent. This requires extensive data processing and consumes significant human and material resources. Summary of the Invention
[0004] The present invention provides a data processing method, apparatus, device and storage medium for fusing multimodal data collected in a vehicle, thereby reducing the amount of data calculation required for multimodal data processing, lowering memory occupancy and CPU usage, improving data calculation speed and reducing data processing costs.
[0005] In a first aspect, an embodiment of the present invention provides a data processing method, including:
[0006] Acquire multimodal data collected by the vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information, and behavior information;
[0007] Determining at least one vehicle-mounted scene based on the multimodal data, and determining a label data set corresponding to each vehicle-mounted scene;
[0008] Combining the label data in each label data set to determine at least one descriptive statement corresponding to the label data set;
[0009] The purpose feature parameters corresponding to each label data set are determined according to each descriptive sentence, each descriptive sentence is annotated according to each purpose feature parameter, and a structured data set is generated according to each annotated descriptive sentence.
[0010] In a second aspect, an embodiment of the present invention further provides a data processing device, the data processing device comprising:
[0011] A data acquisition module, configured to acquire multimodal data collected by the vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information, and behavior information;
[0012] A data set determination module, configured to determine at least one vehicle-mounted scene based on the multimodal data, and determine a label data set corresponding to each vehicle-mounted scene;
[0013] A statement determination module, configured to combine the label data in each label data set and determine at least one descriptive statement corresponding to the label data set;
[0014] The set generation module is used to determine the purpose feature parameters corresponding to each label data set based on each descriptive statement, annotate each descriptive statement through each purpose feature parameter, and generate a structured data set based on each annotated descriptive statement.
[0015] In a third aspect, an embodiment of the present invention further provides a data processing device, including:
[0016] a storage device and one or more processors;
[0017] a storage device for storing one or more programs;
[0018] When one or more programs are executed by one or more processors, the one or more processors implement the data processing method of the first aspect described above.
[0019] In a fourth aspect, an embodiment of the present invention further provides a storage medium comprising computer-executable instructions, wherein the computer-executable instructions are used to execute the data processing method of the first aspect when executed by a computer processor.
[0020] A data processing method, apparatus, device and storage medium provided by an embodiment of the present invention obtain multimodal data collected by a vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information and behavior information; at least one vehicle-mounted scene is determined based on the multimodal data, and a label data set corresponding to each vehicle-mounted scene is determined; each label data in each label data set is combined to determine at least one descriptive statement corresponding to the label data set; a purpose feature parameter corresponding to each label data set is determined based on each descriptive statement, each descriptive statement is annotated with each purpose feature parameter, and a structured data set is generated based on each annotated descriptive statement. By adopting the above technical solution, corresponding vehicle information, cloud information, user information and behavior information are extracted from the multimodal data obtained in the vehicle, and at least one vehicle-mounted scenario is determined based on the information. For each vehicle-mounted scenario, the label data set determined based on the multimodal data is combined to obtain a corresponding descriptive statement, and then the purpose feature parameters of the corresponding label data set are determined based on each descriptive statement, so as to directly fuse the multimodal data and determine its corresponding purpose feature parameters, thereby obtaining a structured data set that can be directly used to train the network model. The trained network model can directly determine the purpose features based on the input multimodal data, without the need to extract separate purpose features for different modal data. This solves the problem of requiring a large amount of data calculations and consuming a lot of manpower and material costs when processing multimodal data, improves the data calculation speed, reduces the memory occupancy rate and CPU usage rate of the computing device, and reduces the data processing cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 is a flow chart of a data processing method in embodiment 1 of the present invention;
[0023] Figure 2 is a flow chart of a data processing method in embodiment 2 of the present invention;
[0024] Figure 3 This is an example diagram of a process for dividing multimodal data into scene label data sets corresponding to each scene label information according to each scene label information in the second embodiment of the present invention;
[0025] Figure 4This is an example diagram of a process for determining the confidence of each semantic feature according to a preset confidence setting rule in the second embodiment of the present invention;
[0026] Figure 5 This is a schematic structural diagram of a data processing device in Embodiment 3 of the present invention;
[0027] Figure 6 It is a structural diagram of a data processing device in embodiment 4 of the present invention. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0029] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0030] In the description of the present invention, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. In addition, in the description of the present invention, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0031] Example 1
[0032] Figure 1This is a flowchart of a data processing method provided in Example 1 of the present invention. This embodiment is applicable to determining the purpose characteristics of multimodal data acquired in a vehicle and generating a corresponding structured data set. The method can be executed by a data processing device, which can be implemented by software and / or hardware. The data processing device can be configured on a data processing device, which can be a computer device. The computer device can be composed of two or more physical entities or a single physical entity. Generally speaking, the computer device can be a laptop, desktop computer, smart tablet, etc.
[0033] like Figure 1 As shown, the data processing method provided in this embodiment includes the following steps:
[0034] S101. Acquire multimodal data collected by a vehicle adaptation layer.
[0035] Among them, multimodal data includes at least vehicle information, cloud information, user information and behavior information.
[0036] In this embodiment, the vehicle adaptation layer can be understood as all hardware or software in the vehicle that can interact with the user, adapt to user services, and receive data. Optionally, sensors in the vehicle that directly capture images, video, and audio, as well as controllers that receive user input, can also serve as the vehicle adaptation layer.
[0037] In this embodiment, modality can be specifically understood as a source or form of information. For example, touch, hearing, vision, smell, voice, video, text, radar, infrared, and accelerometer can all be referred to as a modality. Multimodal data can be specifically understood as a collection of data from multiple sources and forms obtained from the vehicle. Vehicle information can be specifically understood as a data dimension used to characterize the driving status of the vehicle, which may include multiple types of vehicle driving status information. For example, vehicle information may include vehicle speed, vehicle range, internal and external temperature, vehicle controller or application status, and user location information. Cloud information can be specifically understood as a data dimension obtained from the cloud to characterize external factors that may affect vehicle driving. For example, cloud information may include holiday information and weather information. User information can be specifically understood as a data dimension used to characterize the characteristics of the vehicle driver. For example, user information may include user ID, user age, user gender, anniversary, and personal portrait information. Behavior information can be specifically understood as a data dimension used to characterize user operation behavior during vehicle driving, which may include multiple types of user behavior information. For example, the behavior information may include user behavior time, user behavior, line of sight direction, gesture posture, and body posture information.
[0038] Specifically, information of various data dimensions is obtained from the controller and sensors serving as the vehicle adaptation layer and is used as multimodal data. The information initially obtained by the vehicle adaptation layer may include visual data, audio data, and text data, etc. Behavioral information used to characterize user behavior, user information used to characterize user identity, vehicle information used to characterize vehicle driving status, and cloud information used to characterize external environmental influences are extracted from the above visual data, audio data, and text data, and the extracted vehicle information, cloud information, user information, and behavioral information are determined as the collected multimodal data.
[0039] S102: Determine at least one vehicle-mounted scene based on the multimodal data, and determine a label data set corresponding to each vehicle-mounted scene.
[0040] In this embodiment, the vehicle scene can be specifically understood as the state information of the vehicle and its driver at a certain moment in the vehicle driving process. The same vehicle scene can contain multiple different scene tags and the states corresponding to the scene tags. For example, scene tags may include windows, seats, user age, vehicle speed, etc., while the states corresponding to the scene tags may be windows open, seats heated, age 20, etc. The label data set can be specifically understood as the collection of data obtained by associating scene tags with the states corresponding to the scene tags.
[0041] Specifically, since multimodal data contains information of multiple dimensions, and the information of each dimension can still be divided into finer-grained information, the multimodal data is divided according to the finest-grained information type, and one data is extracted from each finest-grained information type. The state of the vehicle and its driver corresponding to each set of extracted data is determined as a vehicle-mounted scene, and each extracted data is represented in the form of a scene label plus scene state data, and the data in this form is determined as label data to generate a label data set corresponding to each vehicle-mounted scene.
[0042] S103: Combine the label data in each label data set to determine at least one descriptive sentence corresponding to the label data set.
[0043] In this embodiment, the descriptive sentence can be specifically understood as a sentence generated by describing the information contained in the vehicle scene in a linguistic way, which can also be understood as converting incoherent label data into a logical and understandable sentence.
[0044] Specifically, since each label data set contains multiple label data, and when the order of the label data is different, the focus of the logical descriptive statements converted from them is different, so the label data in the label data set corresponding to each vehicle scene can be arranged and combined, and each arrangement and combination can be expanded to determine a corresponding descriptive statement, and then at least one descriptive statement can be obtained based on each label data set.
[0045] S104 , determining the purpose feature parameters corresponding to each label data set according to each descriptive sentence, annotating each descriptive sentence according to each purpose feature parameter, and generating a structured data set according to each annotated descriptive sentence.
[0046] In this embodiment, the purpose characteristic parameter can be specifically understood as a characteristic parameter used to characterize the user's driving intention or driving purpose. Structured data can be specifically understood as data with a regular and complete data structure, that is, data with complete generation rules.
[0047] Specifically, since each descriptive statement has its corresponding meaning, the purpose to be achieved by the descriptive statement can also be determined based on the meaning, and the purpose can be determined as the purpose feature of the descriptive statement. For the same label data set, it corresponds to multiple descriptive statements with different expression methods. The purpose features corresponding to different descriptive statements expanded from the same label data should be similar to each other. Therefore, based on the purpose features of multiple descriptive statements corresponding to the same label data, the purpose feature with the highest similarity can be determined as the purpose feature parameter corresponding to the label data set. It can be further considered that the purpose of all descriptive statements corresponding to the label data set can be expressed by the determined purpose feature parameter. At this time, the determined purpose feature parameter is annotated in all descriptive statements corresponding to the label data set, and the set after all the descriptive statements corresponding to the label data set are annotated can be determined as a structured data set.
[0048] The embodiment of the present invention obtains multimodal data collected by a vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information and behavior information; determines at least one vehicle-mounted scene based on the multimodal data, and determines a label data set corresponding to each vehicle-mounted scene; combines each label data in each label data set to determine at least one descriptive statement corresponding to the label data set; determines the purpose feature parameters corresponding to each label data set based on each descriptive statement, labels each descriptive statement using each purpose feature parameter, and generates a structured data set based on each labeled descriptive statement. By adopting the above technical solution, corresponding vehicle information, cloud information, user information and behavior information are extracted from the multimodal data obtained in the vehicle, and at least one vehicle-mounted scenario is determined based on the information. For each vehicle-mounted scenario, the label data set determined based on the multimodal data is combined to obtain a corresponding descriptive statement, and then the purpose feature parameters of the corresponding label data set are determined based on each descriptive statement, so as to directly fuse the multimodal data and determine its corresponding purpose feature parameters, thereby obtaining a structured data set that can be directly used to train the network model. The trained network model can directly determine the purpose features based on the input multimodal data, without the need to extract separate purpose features for different modal data. This solves the problem of requiring a large amount of data calculations and consuming a lot of manpower and material costs when processing multimodal data, improves the data calculation speed, reduces the memory occupancy rate and CPU usage rate of the computing device, and reduces the data processing cost.
[0049] Example 2
[0050] Figure 2A flow chart of a data processing method provided for the second embodiment of the present invention. The technical solution of the embodiment of the present invention is further optimized on the basis of the above-mentioned optional technical solutions. By clustering the multimodal data, the scene label information is determined, and then at least one vehicle-mounted scene and a label data set corresponding to the vehicle-mounted scene are determined according to the scene label data corresponding to each scene label information. Then, according to the different arrangement orders of the label data in each label data set, the descriptive sentences corresponding to each label data set are determined, and then according to the preset confidence setting rules, the confidence of the semantic features of each descriptive sentence in the same label data set is determined, and according to the confidence of each semantic feature, the purpose feature parameters corresponding to the label data set are determined, and then the descriptive sentences are labeled using the purpose feature parameters to generate a structured data set. Furthermore, the structured data set and the unlabeled descriptive sentences can be used to train the initial purpose feature determination network model, and then a purpose feature determination network model that can directly determine the purpose feature according to the input multimodal data is obtained. By directly extracting and fusing the features of multimodal data, and directly determining the semantic features based on the descriptive sentences corresponding to the fused multimodal data, direct labeling of the target feature parameters of the multimodal data is achieved, which reduces the amount of data calculations during multimodal data processing and improves the data calculation speed. As a result, the target feature determination network model trained based on the structured data set obtained after labeling can directly determine the corresponding target features based on the input multimodal data, thereby improving the simplicity of data processing and reducing the data processing cost.
[0051] like Figure 2 As shown, a data processing method provided in the second embodiment of the present invention specifically includes the following steps:
[0052] S201. Acquire multimodal data collected by the vehicle adaptation layer.
[0053] Among them, multimodal data includes at least vehicle information, cloud information, user information and behavior information.
[0054] S202: Cluster the multimodal data to determine at least one type of scene label information.
[0055] In this embodiment, clustering can be specifically understood as a process of dividing a collection of objects or abstract objects into multiple classes consisting of similar objects. Scene label information can be understood as the information corresponding to each part of the scene. The same scene label information has the same characteristics. For example, for an in-vehicle scene, the scene label information may include the window status, seat status, vehicle speed, user age, user gender, and the temperature inside and outside the vehicle. It can also be understood as the type information obtained by classifying the information that can be extracted from the in-vehicle scene at a certain moment according to the included content and the extraction subject.
[0056] Specifically, since people's first intuition about data is often to group the data meaningfully so that similar objects are classified into one category and dissimilar objects are classified into different categories, the acquired multimodal data can be clustered. For the data belonging to the same category after clustering, the scene label information corresponding to the data of this category is determined based on the type of information contained therein and the source of the information contained therein. For example, if the data belonging to the same category are all provided for the controller of the car window, and the content contained in the data is the open and closed status of the car window, the car window status can be used as the scene label information corresponding to the data of this category. The above is only a method for determining part of the scene label information provided in the embodiment of the present invention. The embodiment of the present invention does not limit the number and type of the determined scene label information.
[0057] S203: Divide the multimodal data into scene label data sets corresponding to each scene label information according to each scene label information.
[0058] Specifically, since the scene label information is determined based on the clustered multimodal data, it can be known that one scene label information corresponds to at least one data in the multimodal data. The scene label information is combined with the corresponding data to obtain the scene label data corresponding to the scene label information. The set consisting of all scene label data corresponding to the same scene label information is determined as the scene label data set.
[0059] Furthermore, Figure 3 This is a flow chart illustrating an example of dividing multimodal data into scene label data sets corresponding to each scene label information according to each scene label information provided in the second embodiment of the present invention, as shown in FIG. Figure 3 As shown, the specific steps include:
[0060] S2031. For each piece of scene label information, extract at least one original data corresponding to the scene label information from the multimodal data.
[0061] In this embodiment, the original data can be specifically understood as data directly collected and obtained by the vehicle adaptation layer, that is, one of the multimodal data.
[0062] Specifically, for each determined scene label information, it can represent a type of data information in the multimodal data, so the data information of this type in the multimodal data can be extracted as the original data corresponding to the scene label information.
[0063] S2032: transcribe and combine the original data and the scene label information according to a preset label transcription rule, and determine the combined original data and scene label information as scene label data.
[0064] In this embodiment, the preset label transcription rule can be specifically understood as a preset rule for combining original data with scene label information.
[0065] Specifically, based on the extracted original data and the scene label information corresponding to the original data, the two are associated according to the preset label writing rules to form scene label data in the format of "scene label information + original data". The content of the scene label data must include necessary elements such as behavior information, vehicle information, cloud information and user information.
[0066] For example, based on the type of the extracted raw data and the scene label information, the scene label data can be represented as: Window-on, UserAge-18, Speed-60km / h, etc. Furthermore, when the scene label data is in English, the above scene label data can be transcribed into Chinese. The corresponding relationship can be represented by the following table:
[0067] Window-on Open the car window SeatHeat-on Turn on seat heating UserAge-18 Age 18 Weather-sunny it's clear Speed-60km / h Speed 60km / h UserGender-F female UserImage-music Music Master
[0068] S2033: Determine the set of each scene label data as a scene label data set corresponding to the scene label information.
[0069] Specifically, the scene label data corresponding to each scene label information is grouped into a set, and the set is used as the scene label data set corresponding to the scene label information.
[0070] S204 , extracting one scene label data from each scene label data set, determining the scene corresponding to each scene label data set as a vehicle-mounted scene, and determining the set of each scene label data set as a label data set corresponding to the vehicle-mounted scene.
[0071] Specifically, each scene label data in a scene label dataset can represent a scene state corresponding to that scene label. For example, if the scene label is a car window, the corresponding scene label data can be open or closed. Therefore, a scene label data can be extracted from each scene label dataset to represent the scene state of that scene label. The scene states corresponding to all extracted scene label data are then combined, and the resultant combination is determined as a vehicle scene. The set of scene label data used to constitute the vehicle scene is then determined as the corresponding label dataset.
[0072] S205: Arrange the label data in each label data set in different orders to generate label data combinations greater than a first preset number.
[0073] In this embodiment, the label data can be specifically understood as scene label data in the label data set.
[0074] Specifically, since a label data set necessarily includes multiple different label data, and the label data in the label data set are sorted in different orders, the meaning of the expanded sentences will also be different. Therefore, the label data in each label data set can be arranged in different orders, thereby generating multiple label data combinations with different arrangement orders. The number of each label data combination should be greater than the first preset number to meet the extraction of different expressions in the same label data set. Optionally, the first preset number can be pre-set and adjusted according to actual conditions, such as 10, 15 or 20, etc., and the embodiment of the present invention does not limit this.
[0075] S206 , expanding each tag data combination according to a preset data expansion rule, generating a descriptive sentence corresponding to each tag data combination, and determining the descriptive sentence as a descriptive sentence corresponding to the tag data set.
[0076] In this embodiment, the preset data expansion rule can be specifically understood as a pre-set rule for generating a sentence describing a scene according to a plurality of sequentially arranged phrases. The embodiment of the present invention does not limit the type of the data expansion rule.
[0077] Specifically, each label data combination is expanded according to the preset data expansion rules to obtain corresponding multiple descriptive statements. At the same time, since multiple label data combinations can be determined based on a label data set, the multiple descriptive statements determined by each label data combination can be used as descriptive statements corresponding to the label data set.
[0078] For example, if a tag data combination contains the following tags: speed 60 km / h, female, sunny, seat ventilation on, music playing, and windows open, then the descriptive sentence generated based on this tag data combination may be: speed 60 km / h, sunny, female driver, music playing, windows open, and seat ventilation on.
[0079] S207: Determine the semantic features of each descriptive sentence in the same label data set.
[0080] In this embodiment, the semantic feature can be specifically understood as the meaning contained in the descriptive sentence. The semantic feature of the descriptive sentence determined based on the multimodal data obtained in the vehicle can represent the user's expected intention in the corresponding vehicle scenario.
[0081] Continuing with the above example, based on the descriptive sentence generated in the above example, it can be determined that the driver is in a relaxed state based on the statement "the female driver plays music", and the temperature in the car is high based on the statements "the weather is sunny" and "the windows and seats are opened for ventilation". It can also be determined that the semantic features of the descriptive sentence are: relaxed, hot.
[0082] S208: Determine the confidence of each semantic feature according to a preset confidence setting rule, and sort the semantic features according to the confidence.
[0083] In this embodiment, confidence can be specifically understood as the degree of trustworthiness of a measured parameter value, that is, the probability that the true value of the measured parameter falls within the range of the measurement result. The preset confidence setting rule can be specifically understood as a pre-set rule for determining the probability that the meaning represented by each semantic feature is the true meaning of a descriptive statement based on the number of occurrences of each semantic feature and the corresponding meaning of the semantic feature.
[0084] Further, Figure 4 This is an example diagram of a process for determining the confidence of each semantic feature according to a preset confidence setting rule provided in the second embodiment of the present invention, such as Figure 4 As shown, the specific steps include:
[0085] S2081. Cluster each semantic feature to determine a semantic feature group.
[0086] Specifically, the semantic features corresponding to the same label data set are clustered, and the semantic features belonging to the same category after clustering are determined as a semantic feature group. In other words, multiple semantic feature groups can be determined based on the same label data set, and the semantic emphasis corresponding to each semantic feature group is different.
[0087] S2082: Determine the proportion of each semantic feature group in all semantic features, and determine the semantic feature group whose proportion is less than a preset threshold as the semantic feature group to be processed.
[0088] Specifically, the ratio of the number of semantic features in each semantic feature group to the total number of semantic features is determined as the proportion of each semantic feature group in all semantic features. When the proportion of a semantic feature group is less than a preset threshold, it can be considered that the semantic focus corresponding to the semantic feature group is not the semantic focus corresponding to the label data set. At this time, the semantic feature group with a proportion less than the preset threshold can be determined as the semantic feature group to be processed, so as to subsequently determine whether the semantic features in the semantic feature group need to be discarded.
[0089] Optionally, the preset threshold value may be a preset ratio value according to actual conditions, such as 2%, etc., which is not limited in the embodiment of the present invention.
[0090] S2083: If the semantic feature meaning of the semantic feature group to be processed is opposite to the semantic feature group whose proportion is greater than the preset threshold, the semantic feature corresponding to the semantic feature group to be processed is deleted.
[0091] Specifically, the semantic feature meaning corresponding to the semantic feature group to be processed is compared with the semantic feature meaning corresponding to the semantic feature group whose proportion in each semantic feature is greater than a preset threshold. If the two are opposite, it can be considered that the semantic feature corresponding to the semantic feature group to be processed is not the target feature parameter of the label data set. At this time, the semantic feature corresponding to the semantic feature group to be processed can be deleted to reduce the difficulty of determining the target feature parameter of the label data set.
[0092] S2084. The remaining semantic features are weighted and sorted according to the word frequency and the inverse text frequency index, and the confidence level of each semantic feature is determined according to the weighted sorting result.
[0093] In this embodiment, word frequency can be specifically understood as the number of times a given word appears in a document or sentence. Inverse document frequency can be specifically understood as an indicator for evaluating word importance, and its ranking is inversely proportional to the number of documents in which the word appears.
[0094] Specifically, the remaining semantic features after removing the semantic features in the semantic feature group to be processed can be input into the Term Frequency-Inverse Document Frequency (TF-IDF) model. Since a certain word or phrase appears frequently in an article and rarely appears in the article, it can be considered that the word has good category discrimination ability. Therefore, the discrimination ability of each semantic feature can be determined based on the output of the model, and the remaining semantic features can be weighted and sorted. The confidence of each semantic feature can then be determined based on the determined weight sorting results. Semantic features with high weights have high confidence, while semantic features with low weights have low confidence.
[0095] S209: Determine the semantic features ranked in the front by a second preset number as target feature parameters corresponding to the label dataset.
[0096] Specifically, since semantic features with higher weights are more able to reflect the purpose and intent of their corresponding label datasets, for each semantic feature corresponding to the same label dataset, the semantic features ranked by weight in the first second preset number can be determined as the purpose feature parameters corresponding to the label dataset. In other words, it can be assumed that the purpose and intent of all descriptive sentences in the label dataset can be represented by these semantic features. Optionally, the second preset number can be pre-set and adjusted according to actual circumstances, such as 2, 3, or 5, etc., and this is not limited in the present embodiment.
[0097] S210 , annotating each descriptive sentence using characteristic parameters of each purpose, and generating a structured data set based on the annotated descriptive sentences.
[0098] Furthermore, after generating a structured data set according to the annotated descriptive sentences, the method further includes:
[0099] The initial target feature determination network model is trained according to the structured data set and the unlabeled descriptive sentences until the preset convergence conditions are met to obtain the target feature determination network model.
[0100] In this embodiment, the preset convergence condition can be specifically understood as a condition used to determine whether the network model has entered a converged state based on the initial target features of the training. Optionally, the preset convergence condition may include the change in weight parameters between two iterations of model training being less than a preset parameter change threshold, the number of iterations exceeding a set maximum number of iterations, or the completion of training of all training samples, etc., which are not limited in this embodiment of the present invention.
[0101] Specifically, a set of structured data and a set of unlabeled descriptive sentences can be determined as a training sample set. The training sample set is input into the initial purpose feature determination network model for training. During the training process, multiple different intermediate results of the initial purpose feature determination network model are extracted. Then, the loss function used to train the initial purpose feature determination network model can be determined based on the multiple different intermediate results and preset weighting rules. Then, the initial purpose feature determination network model is back-propagated using the determined loss function, so that the weight parameters used to constitute the initial purpose feature determination network model can be adjusted according to the determined loss function, until the preset convergence condition is met and the trained initial purpose feature determination network model is determined as the purpose feature determination network model.
[0102] Optionally, the initial purpose feature determination network model may be an intelligent speech semantic model, which is not limited in the embodiment of the present invention.
[0103] Furthermore, the trained purpose feature determination network model can directly obtain the driver's expected intention at the time of data acquisition based on the multimodal data collected by the vehicle adaptation layer input into it, and then control the vehicle to switch to the corresponding mode based on the expected intention, or predict the driver's next operation, and adjust the hardware and software parameters in the vehicle accordingly to provide more targeted services for the driver.
[0104] The technical solution of this embodiment clusters multimodal data acquired from a vehicle to determine at least one vehicle-mounted scene and a label dataset corresponding to the vehicle-mounted scene. Corresponding descriptive statements are then generated based on the different orderings of the label data in each label dataset. The target feature parameters corresponding to the label dataset are determined based on the confidence level of the semantic features of each descriptive statement. Each descriptive statement in the label dataset is then calibrated using the target feature parameters to generate a structured data set. The structured data set and the unlabeled descriptive statements can then be used as a training sample set to train an initial target feature determination network model, thereby obtaining a network model that can directly determine target features based on the input multimodal data. By directly extracting and fusing features from the multimodal data and directly determining semantic features based on the descriptive statements corresponding to the fused multimodal data, direct labeling of target feature parameters for the multimodal data is achieved, reducing the amount of data computation required for multimodal data processing and improving data computation speed. This allows the target feature determination network model trained on the labeled structured data set to directly determine the corresponding target features based on the input multimodal data, thereby improving the simplicity and cost of data processing.
[0105] Example 3
[0106] Figure 5 This is a structural diagram of a data processing device provided in the third embodiment of the present invention. The data processing device includes: a data acquisition module 31, a data set determination module 32, a statement determination module 33 and a set generation module 34.
[0107] Among them, the data acquisition module 31 is used to obtain multimodal data collected by the vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information and behavior information; the data set determination module 32 is used to determine at least one vehicle-mounted scene based on the multimodal data, and determine the label data set corresponding to each vehicle-mounted scene; the statement determination module 33 is used to combine each label data in each label data set, and determine at least one descriptive statement corresponding to the label data set; the set generation module 34 is used to determine the purpose feature parameters corresponding to each label data set according to each descriptive statement, annotate each descriptive statement by each purpose feature parameter, and generate a structured data set according to each annotated descriptive statement.
[0108] The technical solution of this embodiment extracts corresponding vehicle information, cloud information, user information and behavior information from the multimodal data obtained in the vehicle, and determines at least one vehicle-mounted scenario based on the information. For each vehicle-mounted scenario, the label data set determined based on the multimodal data is combined to obtain a corresponding descriptive statement, and then the purpose feature parameters of the corresponding label data set are determined based on each descriptive statement, so as to directly fuse the multimodal data and determine its corresponding purpose feature parameters, thereby obtaining a structured data set that can be directly used to train the network model. The trained network model can directly determine the purpose features based on the input multimodal data, without the need to extract separate purpose features for different modal data. This solves the problem of requiring a large amount of data calculations and consuming a lot of manpower and material costs when processing multimodal data, improves the data calculation speed, reduces the memory occupancy and CPU usage of the computing device, and reduces the data processing cost.
[0109] Optionally, the data set determination module 32 includes:
[0110] a label information determination unit, configured to cluster the multimodal data and determine at least one type of scene label information;
[0111] A scene label set determining unit, configured to divide the multimodal data into scene label data sets corresponding to each scene label information according to each scene label information;
[0112] The label data set determination unit is used to extract a scene label data from each scene label data set, determine the scene corresponding to each scene label data as a vehicle scene, and determine the set of each scene label data as a label data set corresponding to the vehicle scene.
[0113] Optionally, the scene label set determination unit is specifically configured to:
[0114] For each scene label information, extract at least one original data corresponding to the scene label information from the multimodal data;
[0115] The original data and the scene label information are transcribed and combined according to a preset label transcription rule, and the combined original data and scene label information are determined as scene label data;
[0116] A set of each scene label data is determined as a scene label data set corresponding to the scene label information.
[0117] Optionally, the statement determination module 33 includes:
[0118] a combination generating unit, configured to arrange the label data in each label data set in different orders to generate a label data combination greater than a first preset number;
[0119] The sentence determination unit is used to expand each label data combination according to a preset data expansion rule, generate a descriptive sentence corresponding to each label data combination, and determine the descriptive sentence as a descriptive sentence corresponding to the label data set.
[0120] Optionally, the set generation module 34 includes:
[0121] A semantic feature determination unit, used to determine the semantic features of each descriptive sentence in the same label data set;
[0122] A feature sorting unit, configured to determine the confidence of each semantic feature according to a preset confidence setting rule, and sort the semantic features according to the confidence;
[0123] A feature parameter determination unit, configured to determine the semantic features ranked in the top second preset number as target feature parameters corresponding to the label data set;
[0124] The set generation unit is used to mark each descriptive sentence by using characteristic parameters of each purpose, and generate a structured data set according to each marked descriptive sentence.
[0125] Optional feature sorting unit, specifically used to:
[0126] Clustering each semantic feature to determine the semantic feature group;
[0127] Determine the proportion of each semantic feature group in all semantic features, and determine the semantic feature group whose proportion is less than a preset threshold as the semantic feature group to be processed;
[0128] If the semantic feature meaning of the semantic feature group to be processed is opposite to the semantic feature group whose proportion is greater than the preset threshold, the semantic feature corresponding to the semantic feature group to be processed is deleted;
[0129] The remaining semantic features are weighted and sorted according to word frequency and inverse text frequency index, and the confidence of each semantic feature is determined according to the weighted sorting result.
[0130] Optionally, the data processing device further includes:
[0131] The network model training module is used to train the initial target feature determination network model based on the structured data set and unlabeled descriptive statements until the target feature determination network model is obtained by meeting the preset convergence conditions.
[0132] The data processing device provided by the embodiment of the present invention can execute the data processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0133] Example 4
[0134] Figure 6 This is a structural diagram of a data processing device provided in the fourth embodiment of the present invention. The data processing device includes: a processor 40, a storage device 41, a display screen 42, an input device 43, and an output device 44. The number of processors 40 in the data processing device can be one or more. Figure 6 In the example, a processor 40 is used. The number of storage devices 41 in the data processing device can be one or more. Figure 6 A storage device 41 is taken as an example. The processor 40, storage device 41, display screen 42, input device 43 and output device 44 of the data processing device can be connected by a bus or other means. Figure 6 In the embodiment, the data processing device can be a computer, a notebook or a smart tablet.
[0135] The storage device 41 is a computer-readable storage medium that can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the data processing device described in any embodiment of the present application (for example, the data acquisition module 31, the data set determination module 32, the statement determination module 33 and the set generation module 34). The storage device 41 may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the storage device 41 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the storage device 41 may further include a memory remotely located relative to the processor 40, and these remote memories can be connected to the device via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0136] The display screen 42 may be a touch screen 42, which may be a capacitive screen, an electromagnetic screen, or an infrared screen. Generally speaking, the display screen 42 is used to display data according to the instructions of the processor 40, and is also used to receive touch operations acting on the display screen 42 and send corresponding signals to the processor 40 or other devices.
[0137] Input device 43 can be used to receive input digital or character information and generate key input signals related to user settings and function control of the display device. It can also be a camera for capturing images and a sound pickup device for capturing audio data. Output device 44 can include audio equipment such as speakers. It should be noted that the specific composition of input device 43 and output device 44 can be set according to actual circumstances.
[0138] The processor 40 executes the software programs, instructions and modules stored in the storage device 41 to perform various functional applications and data processing of the device, that is, to implement the above-mentioned data processing method.
[0139] The computer device provided above can be used to execute the data processing method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0140] Example 5
[0141] A fifth embodiment of the present invention further provides a storage medium containing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, the computer-executable instructions are used to perform a data processing method, the method comprising:
[0142] Acquire multimodal data collected by the vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information, and behavior information;
[0143] Determining at least one vehicle-mounted scene based on the multimodal data, and determining a label data set corresponding to each vehicle-mounted scene;
[0144] Combining the label data in each label data set to determine at least one descriptive statement corresponding to the label data set;
[0145] The purpose feature parameters corresponding to each label data set are determined according to each descriptive sentence, each descriptive sentence is annotated according to each purpose feature parameter, and a structured data set is generated according to each annotated descriptive sentence.
[0146] Of course, the computer executable instructions of a storage medium containing computer executable instructions provided by an embodiment of the present invention are not limited to the method operations described above, and can also execute related operations in the data processing method provided by any embodiment of the present invention.
[0147] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0148] It is worth noting that in the embodiment of the above-mentioned search device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0149] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.
Claims
1. A data processing method, characterized in that: include: Acquire multimodal data collected by the vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information, and behavior information; Determining at least one vehicle-mounted scene according to the multimodal data, and determining a label data set corresponding to each of the vehicle-mounted scenes; Combining each of the label data in each of the label data sets to determine at least one descriptive statement corresponding to the label data set; Determining the purpose feature parameters corresponding to each of the label data sets according to each of the descriptive sentences, annotating each of the descriptive sentences using each of the purpose feature parameters, and generating a structured data set according to each of the annotated descriptive sentences; The combining of the label data in each of the label data sets to determine at least one descriptive statement corresponding to the label data set includes: Arranging the label data in each of the label data sets in different orders to generate label data combinations greater than a first preset number; Expanding each of the label data combinations according to a preset data expansion rule, generating a descriptive statement corresponding to each of the label data combinations, and determining the descriptive statement as a descriptive statement corresponding to the label data set; The step of determining the target characteristic parameters corresponding to each of the label data sets according to each of the descriptive statements includes: Determining the semantic features of each of the descriptive sentences in the same label data set; Determining the confidence of each of the semantic features according to a preset confidence setting rule, and sorting the semantic features according to the confidence; Determine the semantic features ranked in the top second preset number as target feature parameters corresponding to the label data set; The step of labeling each of the descriptive statements by using each of the purpose characteristic parameters includes: For each of the target feature parameters, the target feature parameter is annotated in all descriptive sentences of the label data set corresponding to the target feature parameter.
2. The method according to claim 1, characterized in that The determining of at least one vehicle-mounted scene according to the multimodal data and determining a label data set corresponding to each of the vehicle-mounted scenes includes: Clustering the multimodal data to determine at least one scene label information; Dividing the multimodal data into scene label data sets corresponding to the scene label information according to the scene label information; A scene label data set is extracted from each of the scene label data sets, the scene corresponding to each of the scene label data sets is determined as a vehicle-mounted scene, and the set of each of the scene label data sets is determined as a label data set corresponding to the vehicle-mounted scene.
3. The method according to claim 2, characterized in that The dividing the multimodal data into scene label data sets corresponding to the scene label information according to the scene label information includes: For each piece of scene label information, extracting at least one original data corresponding to the scene label information from the multimodal data; Transcribing and combining the original data and the scene label information according to a preset label transcription rule, and determining the combined original data and the scene label information as scene label data; The sets of each of the scene label data are determined as scene label data sets corresponding to the scene label information.
4. The method according to claim 1, wherein The step of determining the confidence of each semantic feature according to a preset confidence setting rule includes: Clustering the semantic features to determine a semantic feature group; Determining the proportion of each of the semantic feature groups in all semantic features, and determining the semantic feature group whose proportion is less than a preset threshold as the semantic feature group to be processed; If the semantic feature meaning of the semantic feature group to be processed is opposite to the semantic feature group whose proportion is greater than a preset threshold, deleting the semantic feature corresponding to the semantic feature group to be processed; The remaining semantic features are weighted and sorted according to the word frequency and the inverse text frequency index, and the confidence level of each semantic feature is determined according to the weighted sorting result.
5. The method according to claim 1, wherein After generating a structured data set according to the annotated descriptive sentences, the method further includes: The initial target feature determination network model is trained according to the structured data set and the unlabeled descriptive sentences until a preset convergence condition is met to obtain the target feature determination network model.
6. A data processing device, characterized in that: include: A data acquisition module, configured to acquire multimodal data collected by the vehicle adaptation layer; wherein the multimodal data includes at least vehicle information, cloud information, user information, and behavior information; a data set determination module, configured to determine at least one vehicle-mounted scene based on the multimodal data, and determine a label data set corresponding to each of the vehicle-mounted scenes; a statement determination module, configured to combine the label data in each of the label data sets to determine at least one descriptive statement corresponding to the label data set; A set generation module is used to determine the purpose feature parameters corresponding to each of the label data sets according to each of the descriptive statements, annotate each of the descriptive statements according to each of the purpose feature parameters, and generate a structured data set according to each of the annotated descriptive statements; The statement determination module includes: a combination generating unit, configured to arrange the label data in each of the label data sets in different orders to generate a label data combination greater than a first preset number; a statement determination unit, configured to expand each of the label data combinations according to a preset data expansion rule, generate a descriptive statement corresponding to each of the label data combinations, and determine the descriptive statement as a descriptive statement corresponding to the label data set; Wherein, the set generation module includes: A semantic feature determination unit, configured to determine the semantic features of each of the descriptive sentences in the same label data set; A feature sorting unit, configured to determine the confidence of each of the semantic features according to a preset confidence setting rule, and sort the semantic features according to the confidence; A feature parameter determination unit, configured to determine the semantic features ranked in the top second preset number as target feature parameters corresponding to the label data set; The step of labeling each of the descriptive statements by using each of the purpose characteristic parameters includes: For each of the target feature parameters, the target feature parameter is annotated in all descriptive sentences of the label data set corresponding to the target feature parameter.
7. A data processing device, characterized in that: include: a storage device and one or more processors; The storage device is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to any one of claims 1 to 5.
8. A storage medium containing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, they are used to perform the data processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data processing method, device and equipment based on unmanned vehicle and storage medium
CN109829395A
Simulation scenario generation method and unmanned driving system test method
CN110597086A