Data analysis method, device and equipment in unmanned driving field and storage medium
By utilizing large language models and similarity matching of historical analysis libraries in the field of autonomous driving, the problem of high data analysis resource consumption is solved, and efficient and accurate data analysis results are achieved to meet user needs.
Patent Information
- Application Number
- CN202510886787.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
Data analysis in autonomous driving technology consumes a lot of resources, affecting service stability.
By obtaining the task requirement association information and requirement description text of the current analysis task, using the large language model to extract features, and combining the task results with the highest similarity in the historical analysis library for data analysis, we avoid directly processing the current task.
It improves data analysis efficiency and resource utilization, ensures that analysis results meet user needs, reduces resource consumption, and improves response speed.
Smart Images

Figure CN120804653A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of unmanned driving technology, and particularly relates to a data analysis method, device, equipment and storage medium in the field of unmanned driving. BACKGROUND
[0002] With the rapid development of unmanned driving technology, unmanned vehicles are increasingly widely used, and a large amount of data is generated during the operation of unmanned vehicles. In order to understand the operation of unmanned vehicles, demanders such as researchers and operators have many data analysis needs.
[0003] In the prior art, various demands proposed by demanders generally need to be calculated and analyzed based on target data, resulting in a large amount of resource consumption. Moreover, once the resource usage increases, the stability of the service will also be affected. SUMMARY
[0004] To solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a data analysis method, device, equipment and storage medium in the field of unmanned driving, which solves the problem of large amount of resource consumption in data analysis in related technologies, thereby ensuring the stability of data analysis service.
[0005] In a first aspect, the embodiments of the present disclosure provide a data analysis method in the field of unmanned driving, which comprises:
[0006] obtaining task demand association information of a current analysis task, and obtaining a demand description text of the demander for the current analysis task;
[0007] determining demand association features according to the task demand association information, and extracting demand description features from the demand description text using a large language model;
[0008] determining the similarity between the current analysis task and each historical analysis task in the historical analysis library according to the demand association features and the demand description features of the current analysis task, and the demand association features and the demand description features of each historical analysis task in the preset historical analysis library;
[0009] determining a historical analysis task with the highest similarity and greater than a preset first threshold as an associated analysis task corresponding to the current analysis task;
[0010] determining a data analysis result of the associated analysis task from the historical analysis library, and determining a data analysis result of the current analysis task according to the data analysis result of the associated analysis task.
[0011] In a second aspect, the embodiments of the present disclosure further provide a data analysis device in the field of unmanned driving, which comprises:
[0012] an information obtaining module, configured to obtain task requirement association information of a current analysis task, and obtain requirement description text of a requirement party for the current analysis task;
[0013] a feature extraction module, configured to determine requirement association features according to the task requirement association information, and extract requirement description features from the requirement description text using a large language model;
[0014] a feature comparison module, configured to determine similarities between the current analysis task and each of historical analysis tasks in a preset historical analysis library according to the requirement association features and the requirement description features of the current analysis task, and the requirement association features and the requirement description features of each of the historical analysis tasks in the historical analysis library;
[0015] an associated task determination module, configured to determine a historical analysis task with the highest similarity and greater than a preset first threshold as an associated analysis task corresponding to the current analysis task;
[0016] a result determination module, configured to determine a data analysis result of the associated analysis task from the historical analysis library, and determine a data analysis result of the current analysis task according to the data analysis result of the associated analysis task.
[0017] In a third aspect, an electronic device is provided, and the electronic device includes one or more processors, and a memory configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the data analysis method in the field of unmanned driving.
[0018] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements the data analysis method in the field of unmanned driving.
[0019] The disclosed embodiment provides a data analysis method in the field of unmanned driving, which obtains task requirement association information of the current analysis task and the demand description text of the demander for the current analysis task, determines the requirement association features according to the task requirement association information, and uses a large language model to extract the requirement description features from the requirement description text, and then determines the similarity between the current analysis task and each historical analysis task according to the requirement association features and the requirement description features of the current analysis task, as well as the requirement association features and the requirement description features of each historical analysis task in a preset historical analysis library, and determines the historical analysis task with the highest similarity and a similarity greater than a preset first threshold as the associated analysis task, thereby determining the data analysis results of the associated analysis task in the historical analysis library, and determining the current analysis task through the data analysis results of the associated analysis task. Data analysis results, realize data analysis, this method can combine the demand characteristics of the current analysis task, use the historical analysis tasks to obtain analysis results, avoid directly processing the current analysis task, and improve the efficiency of task analysis while improving the utilization rate of existing resources, thereby improving the task response speed, and the historical analysis tasks cover richer result outputs, compared with directly processing the current analysis task, it can obtain results that are more in line with user expectations. In addition, this method extracts the demand characteristics of the analysis task from the task demand association information and the demand description text respectively, and can predict the user's needs through other information related to the task, and extract the user's described needs through the text of the user's description of the task, so as to achieve the purpose of analyzing user needs from different angles, so that the final data analysis results can better meet user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 This is a flow chart of a data analysis method in the field of unmanned driving according to an embodiment of the present disclosure;
[0022] Figure 2 A schematic diagram of setting a time window in an embodiment of the present disclosure;
[0023] Figure 3 is a schematic diagram of continuing to correct a first correction result based on road information acquired within a set time window in an embodiment of the present disclosure;
[0024] Figure 4 A schematic diagram of further optimizing a second correction result using a third posture obtained by repositioning in an embodiment of the present disclosure;
[0025] Figure 5 FIG. 1 is a schematic diagram of a positioning architecture according to an embodiment of the present disclosure;
[0026] Figure 6 FIG. 1 is a schematic diagram of a positioning architecture according to an embodiment of the present disclosure;
[0027] Figure 7 FIG. 1 is a schematic diagram of a positioning architecture according to an embodiment of the present disclosure; DETAILED DESCRIPTION
[0028] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0029] It should be noted that the terms "first", "second", and the like used in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0030] Figure 1 FIG. 1 is a schematic diagram of a positioning architecture according to an embodiment of the present disclosure; Figure 1 As shown in FIG. 1, the method can specifically include the following steps:
[0031] S110, obtaining task demand association information of a current analysis task, and obtaining a demand description text of a demand party for the current analysis task.
[0032] The current analysis task can be a data analysis task to be processed. For example, the current analysis task can be to count the top 10 vehicles with the highest automatic driving mileage, to analyze the vehicle model with the highest engine failure rate, or to count the service life of a vehicle-mounted sensor.
[0033] In the embodiments of the present disclosure, the task demand association information can be information associated with the task demand of the current analysis task, i.e., information that can reflect the demand of the demand party for the current analysis task.
[0034] Exemplarily, the task demand association information can include demand proposer information, which can include role information, a department to which the demand proposer belongs, and an existing analysis task, considering the role of the demand side (i.e., the task proposer), the department to which the demand side belongs, and the existing analysis task, which can reflect the current task demand of the demand side to some extent.
[0035] Alternatively, the task demand association information can include associated fault problems and system modules, considering that the demand side may perform data analysis to solve a fault problem of a vehicle or analyze a system module in the vehicle, which can reflect the current task demand of the demand side to some extent. In addition, the task demand association information can include an analysis view, considering that a view type constructed by the demand side according to a data analysis result can also reflect the task demand of the demand side.
[0036] In the embodiments of the present application, the demand description text can be a task description text of the current analysis task by the demand side, which can directly reflect the task demand of the demand side. For example, the demand description text can be obtained by recognizing a voice instruction initiated by the demand side, or the text input by the demand side in the interface for the current analysis task.
[0037] In S120, demand association features are determined according to the task demand association information, and demand description features are extracted from the demand description text using a large language model.
[0038] After obtaining the task demand association information and the demand description text, demand association features can be further extracted from the task demand association information, and demand description features can be further extracted from the demand description text. The demand association features can be features associated with the task demand of the demand side, and the demand description features can be features in the task demand described by the demand side.
[0039] Exemplarily, the task demand association information can be input into a pre-trained feature extractor to obtain demand association features output by the feature extractor. Alternatively, the role information of the demand side, the task keywords of the existing analysis task, the associated fault problems, and the analysis view in the task demand association information can be converted into feature vectors to obtain the demand association features.
[0040] In the embodiments of the present application, a large language model (LLM) can be used for extracting demand description features. The large language model can be trained in advance through a large amount of text data. The large language model can adopt a Transformer architecture and use a self-attention mechanism to capture long-distance semantic dependencies.
[0041] Specifically, for the demand description text of the demand side, the demand description text can be input into the large language model, and the intention of the demand side is understood and the demand description features are extracted from the large language model. For example, the large language model can first perform word segmentation processing on the demand description text, and then extract task keywords from the word segmentation result, and convert the extracted task keywords into vectors to obtain demand description features.
[0042] In S130, similarity between the current analysis task and each historical analysis task is determined according to the demand association features and the demand description features of the current analysis task, and the demand association features and the demand description features of each historical analysis task in the preset historical analysis library.
[0043] The historical analysis library can store the demand association features and the demand description features of each historical analysis task. The historical analysis task can be understood as a data analysis task that has been executed and obtained a result. The historical analysis library can also store the data analysis results of each historical analysis task.
[0044] For example, Figure 2 is a schematic diagram of a historical analysis library provided by an embodiment of the present application, as Figure 2 As shown, the demand association features and the demand description features of each historical analysis task (historical analysis task 1 to historical analysis task n) can be stored in the historical analysis library, and the data analysis results of each historical analysis task can be stored in the historical analysis library in the form of result links. The result link can be used to link to the data analysis result of the historical analysis task.
[0045] In the embodiments of the present application, in order to quickly obtain the data analysis result of the current analysis task, whether there is a task identical or similar to the current analysis task can be found in the historical analysis library according to the demand association features and the demand description features of the current analysis task, so that the data analysis result of the historical analysis task can be used subsequently, without the need to process the current analysis task according to the complete task analysis process.
[0046] Specifically, the similarity between the current analysis task and each historical analysis task can be determined according to the demand association features and the demand description features of the current analysis task, and the demand association features and the demand description features of each historical analysis task. For example, the similarity between the current analysis task and each historical analysis task can be determined by calculating the task similarity.
[0047] In a specific implementation, the historical analysis library is used to store a feature set of each historical analysis task, and the feature set includes demand association features and demand description features.
[0048] According to the requirement association features and requirement description features of the current analysis task and the requirement association features and requirement description features of each historical analysis task in the historical analysis library, the similarity between the current analysis task and each historical analysis task is determined, including the following steps:
[0049] Step 11, constructing a feature set of the current analysis task according to the requirement association features and requirement description features of the current analysis task;
[0050] Step 12, for each historical analysis task in the historical analysis library, determining the similarity between the feature set of the historical analysis task and the feature set of the current analysis task, and determining the similarity as the similarity between the historical analysis task and the current analysis task.
[0051] Among them, for each historical analysis task, the feature set can be constructed in advance according to the requirement association features and requirement description features of the historical analysis task, and then the requirement association features and requirement description features are stored in the historical analysis library in the form of the feature set.
[0052] Exemplarily, Figure 3 is a feature set construction process of a historical analysis task provided by the embodiments of the present application, as shown in Figure 3 , the requirement association features can be extracted according to the role of the demander, the module in the vehicle involved, the associated fault problem and the view to be analyzed in the requirement association information; and the requirement description features can be extracted by performing LLM semantic analysis on the requirement document. The requirement document can be a task description text generated by the demander for the historical analysis task. After obtaining the requirement association features and the requirement description features, the feature set of the historical analysis task can be constructed according to the requirement association features and the requirement description features.
[0053] Specifically, in step 11, for the current analysis task, the feature set of the current analysis task can be constructed according to the requirement association features and the requirement description features.
[0054] Exemplarily, Figure 4 is a feature extraction schematic diagram of a current analysis task provided by the embodiments of the present application, as shown in Figure 4 , the requirement association features can be extracted from the role of the demander and the associated fault problem in the requirement association information, and the requirement description features can be obtained by performing LLM semantic analysis on the requirement description text, and then the feature set of the current analysis task can be constructed according to the requirement association features and the requirement description features.
[0055] Further, in step 12, for each historical analysis task in the historical analysis library, the similarity between the feature set of the historical analysis task and the feature set of the current analysis task can be calculated, and the similarity is taken as the similarity between the historical analysis task and the current analysis task.
[0056] Through the above embodiments, the task matching based on the feature set can be realized, the efficiency of searching for a similar or repeated task to the current analysis task in the historical analysis library can be greatly improved, and the response efficiency of the current analysis task is improved.
[0057] In another specific embodiment, the similarity between the current analysis task and each historical analysis task in the historical analysis library is determined according to the demand association features of the current analysis task and the demand association features of each historical analysis task in the historical analysis library, and the demand description features of the current analysis task and the demand description features of each historical analysis task in the historical analysis library, including the following steps:
[0058] Step 21, for each historical analysis task in the historical analysis library, determining a first feature similarity between the demand association features of the historical analysis task and the demand association features of the current analysis task, and determining a second feature similarity between the demand description features of the historical analysis task and the demand description features of the current analysis task.
[0059] Step 22, weighting the first feature similarity and the second feature similarity according to the preset weights corresponding to the demand association features and the demand description features, to obtain the similarity between the historical analysis task and the current analysis task.
[0060] In step 21, for each historical analysis task in the historical analysis library, the similarity between the demand association features of the historical analysis task and the demand association features of the current analysis task can be calculated as the first feature similarity. And the similarity between the demand description features of the historical analysis task and the demand description features of the current analysis task is calculated as the second feature similarity.
[0061] Further, in step 22, considering that the demand description features reflect the demander's description information of the task, which can more directly reflect the demander's task analysis demand, and the demand association features reflect other information associated with the task demand, which can be used to predict the demander's task analysis demand. Therefore, the corresponding weights of the demand description features and the demand association features can be set respectively, and the similarity between the historical analysis task and the current analysis task is obtained by weighting.
[0062] Specifically, the first feature similarity and the second feature similarity can be weighted using the preset weights corresponding to the demand association features and the demand description features. Since the demand description features are obtained from the demand description text of the demander, they can directly reflect the demander's task analysis demand, and therefore the preset weight corresponding to the demand description features can be greater than the preset weight corresponding to the demand association features.
[0063] It should be noted that the preset weight of the demand association feature and the preset weight of the demand description feature can be set according to actual business scenarios, and the sum of the two should satisfy the condition of equaling 1.
[0064] Through the above implementation, considering that the demand association feature and the demand description feature can reflect the task analysis demand of the demand side to different degrees, by respectively determining the similarity of the demand association feature and the demand description feature, and then obtaining the similarity between the historical analysis task and the current analysis task through the weighted manner, the accuracy of the similarity between the tasks is ensured, thereby the accuracy of the subsequent matched associated analysis task can be further ensured.
[0065] S140, the historical analysis task with the highest similarity and greater than the preset first threshold is determined as the associated analysis task corresponding to the current analysis task.
[0066] After determining the similarity between the current analysis task and each historical analysis task, each historical analysis task can be sorted in descending or ascending order of similarity, so as to select the historical analysis task with the highest similarity.
[0067] Further, it can be judged whether the historical analysis task with the highest similarity is greater than the preset first threshold. The preset first threshold can be a pre-set similarity threshold representing that two tasks are at least similar.
[0068] If there is a historical analysis task with the highest similarity and greater than the preset first threshold, the historical analysis task can be used as the associated analysis task corresponding to the current analysis task, so as to subsequently determine the data analysis result of the current analysis task by referring to the data analysis result of the associated analysis task.
[0069] S150, the data analysis result of the associated analysis task is determined from the historical analysis library, and the data analysis result of the current analysis task is determined according to the data analysis result of the associated analysis task.
[0070] Specifically, after the associated analysis task is determined, the data analysis result of the associated analysis task can be obtained from the historical analysis library. For example, the result link of the associated analysis task can be extracted from the historical analysis library, and then the data analysis result of the associated analysis task can be obtained through the result link.
[0071] Further, since the associated analysis task and the current analysis task are at least similar, i.e. similar or repeated, the data analysis result of the current analysis task can be determined on the basis of the data analysis result of the associated analysis task.
[0072] In a specific embodiment, the data analysis result of the current analysis task is determined according to the data analysis result of the correlation analysis task, including:
[0073] If the similarity between the correlation analysis task and the current analysis task is greater than or equal to a preset second threshold, the data analysis result of the correlation analysis task is determined as the data analysis result of the current analysis task, wherein the preset second threshold is greater than the preset first threshold.
[0074] If the similarity between the correlation analysis task and the current analysis task is less than the preset second threshold, the data analysis result of the correlation analysis task is adjusted to obtain the data analysis result of the current analysis task.
[0075] The preset second threshold can be a pre-set similarity threshold representing repetition between two tasks.
[0076] Specifically, if the similarity between the correlation analysis task and the current analysis task is greater than or equal to the preset second threshold, it means that the correlation analysis task and the current analysis task are repeated, i.e., the two tasks are the same, and the data analysis result of the correlation analysis task can be directly determined as the data analysis result of the current analysis task. The result link of the data analysis result of the correlation analysis task can be fed back as the result link of the current analysis task, and the data analysis result of the current analysis task can be obtained through the result link in the future.
[0077] If the similarity between the correlation analysis task and the current analysis task is less than the preset second threshold, it means that the correlation analysis task and the current analysis task are similar, i.e., there is a certain difference, and the data analysis result of the current analysis task can be obtained by adjusting the data analysis result of the correlation analysis task.
[0078] For example, the existing historical analysis task in the historical analysis library is to statistically and summarize the fields of automatic driving mileage, automatic driving time, manual driving mileage, and manual driving time of vehicles in each dimension of week, month, and year. The demand description text of the current analysis task is: "Want to know the top 10 vehicles with the highest automatic driving mileage since 2025", and the above embodiment can be used to determine that the two tasks are repeated, and the existing historical analysis task in the historical analysis library can meet the demand of the current analysis task, and the data analysis result of the historical analysis task can be directly output.
[0079] For example, the requirement description text of the current analysis task is: "Want to know the top 10 cars with the highest total mileage of automatic driving and manual driving since 2025", and the historical analysis task in the historical analysis library is the analysis of automatic driving mileage and manual driving mileage, which lacks the analysis of the sum of automatic driving mileage and manual driving mileage. At this time, it can be judged that the two tasks are similar through the above implementation, and the data analysis result of the current analysis task can be obtained by simply adding and subtracting the data analysis result of the historical analysis task and sorting.
[0080] Through the above implementation, the data analysis result of the current analysis task can be determined by further judging whether the associated analysis task and the current analysis task meet the similar or repetitive relationship, and the data analysis result of the associated analysis task can be directly used in the repetitive case, and the data analysis result of the associated analysis task can be adjusted in the similar case, to ensure the accuracy of the data analysis result of the current analysis result.
[0081] In the above implementation, for the case that the associated analysis task is similar to the current analysis task, the data analysis result of the associated analysis task is adjusted, and specifically, the data analysis result of the associated analysis task can be processed in combination with the difference between the associated analysis task and the current analysis task, to obtain the data analysis result of the current analysis result.
[0082] In an example, the data analysis result of the associated analysis task is adjusted to obtain the data analysis result of the current analysis task, including the following steps:
[0083] Step 31, determining the requirement difference information between the associated analysis task and the current analysis task according to the requirement association features and the requirement description features of the current analysis task, and the requirement association features and the requirement description features of the associated analysis task, and determining the analysis supplement task according to the requirement difference information;
[0084] Step 32, performing the analysis supplement task on the basis of the data analysis result of the associated analysis task to obtain the data analysis result of the current analysis task.
[0085] In step 31, the requirement difference information between the two tasks can be determined according to the requirement association features and the requirement description features of the current analysis task, and the requirement association features and the requirement description features of the associated analysis task.
[0086] For example, in the requirement association features of the current analysis task, the features not covered by the requirement association features of the associated analysis task can be determined, and in the requirement description features of the current analysis task, the features not covered by the requirement description features of the associated analysis task can be determined, to obtain the requirement difference information.
[0087] After obtaining the demand difference information, further, the analysis supplement task can be determined according to the demand difference information. The analysis supplement task can be a supplement processing task performed on the data analysis result of the correlation analysis task.
[0088] Continuing with the above example, the demand difference information can be that the sum of the autonomous driving mileage and the manual driving mileage is the highest, and therefore, it can be determined that the analysis supplement task is to sum the autonomous driving mileage and the manual driving mileage, and to sort after the summation.
[0089] After obtaining the analysis supplement task, further, in step 32, the data analysis result of the correlation analysis task can be processed according to the analysis supplement task to execute the analysis supplement task, and obtain the data analysis result of the current analysis task.
[0090] Through the above example, in the case of two similar tasks, the post-processing of the data analysis result of the correlation analysis task is realized by determining the demand difference between the two tasks, thereby obtaining the data analysis result of the current analysis task, avoiding executing the complete processing flow of the current analysis task, improving the response efficiency of the current analysis task, and reducing the resources occupied by processing the current analysis task.
[0091] In the embodiments of the present application, considering the case where the similarity between all historical analysis tasks and the current analysis task is less than the preset first threshold, i.e., no correlation analysis task can be searched in the historical analysis library, for such cases, the current analysis task can be processed in combination with a large language model to improve the processing efficiency of the current analysis task.
[0092] In some embodiments, after determining the similarity between the current analysis task and each historical analysis task, the following steps are further included:
[0093] Step 41, if the similarity between each historical analysis task and the current analysis task is less than the preset first threshold, the demand description text is identified using a large language model to obtain task key parameters;
[0094] Step 42, an analysis statement of the current analysis task is constructed according to the task key parameters, and the analysis statement is executed in the target database to obtain the data analysis result of the current analysis task.
[0095] In step 41, if the similarity between all historical analysis tasks and the current analysis task is less than the preset first threshold, it can be determined that no historical analysis task similar or repetitive to the current analysis task can be found in the historical analysis library.
[0096] For example, the requirement description text of the current analysis task is: "Want to know the top 10 cars with the highest power consumption since 2025", and there is no historical analysis task related to power consumption in the historical analysis library, so the current analysis task can be determined as a new requirement task different from the historical analysis library.
[0097] Specifically, if the similarity between each historical analysis task and the current analysis task is less than a preset first threshold, the requirement description text of the current analysis task can be input into the large language model. The requirement description text can be entered through a text interaction interface or through a voice interaction system.
[0098] In step 41, the large language model can analyze the user's intention in the requirement description text to extract the parameter values corresponding to the task key fields, i.e., the task key parameters, from the requirement description text. The task key field can be a related field required to generate an analysis statement, and the task key parameter can be a task structured element required to generate an analysis statement.
[0099] For example, the task key field can include operation type, target field, sorting condition, limit result, and search condition. For example, the requirement description text is: "Want to know the top 10 cars with the highest autonomous mileage since 2025", and the task key parameters can be determined as:
[0100] Operation type: query (SELECT);
[0101] Target field: vehicle name (vehicle);
[0102] Sorting condition: descending order of autonomous mileage (aoto_odometer DESC);
[0103] Limit result: top 10 (LIMIT 10);
[0104] Search condition: 2025 (date>='2025-01-01');
[0105] After obtaining the task key parameters, further, in step 42, an analysis statement of the current analysis task can be constructed according to the task key parameters. The analysis statement can be a SQL (Structured Query Language) statement.
[0106] Exemplarily, the validity of the task-critical field corresponding to each task-critical parameter can be verified in combination with the target database, for example, whether the target database has the task-critical field is determined; if the verification is passed, the analysis statement can be generated in combination with the task-critical parameter. Wherein, the metadata service provided to the target database is requested for table structure, and the existence of each task-critical field is verified by table, for example, table name: odometers, field existence: aoto_odometer, vehicle, etc.
[0107] Following the above example, the generated analysis statement can be: SELECT vehicle FROM odometers WHERE date >= '2025-01-01' ORDER BY aoto_odometer DESC LIMIT 10.
[0108] After obtaining the analysis statement, the target database can be connected, and the analysis statement can be executed in the target database to obtain the data analysis result of the current analysis task; the data analysis result can be in the form of a table, and specifically includes structured data results such as JSON, DataFrame, etc.
[0109] In the embodiments of the present application, during the execution of the analysis statement, a security mechanism such as SQL injection detection and query timeout control can also be introduced. Wherein, the SQL injection detection refers to the technology and process of identifying and preventing SQL injection attacks; the query timeout control refers to a mechanism for setting a limit on the execution time of a database query to prevent long-running queries from affecting system performance.
[0110] After obtaining the data analysis result, a visual view can also be provided according to the data analysis result to achieve the purpose of automatic visualization. In the embodiments of the present application, the Dify framework can be used to output the intelligent chart corresponding to the data analysis result through the BI (Business Intelligence, Business Intelligence) analysis tool.
[0111] Exemplarily, the chart type can be automatically matched according to the data characteristics in the data analysis result, for example, a line chart is matched for time series data, and a pie chart is matched for proportion data. Further, the intelligent chart corresponding to the data analysis result can be output through an interactive visualization component.
[0112] In the embodiments of the present application, during the display of the data analysis result, drilling or filtering of the data analysis result can also be supported. Drilling can refer to when a user views the data analysis result, the user can drill down to a more detailed data level (drill down) or return to a more macro summary view (drill up) through interaction such as clicking. Filtering can refer to extracting records that meet the conditions from the data set and excluding records that do not meet the conditions according to specific conditions.
[0113] Exemplary, Figure 5 is a processing flowchart of a current analysis task provided by an embodiment of the present application, as Figure 5 indicated, the demand side can first perform natural language input, i.e., enter a demand description text, and then perform LLM semantic understanding on the demand description text, automatically generate an SQL statement according to the understanding content, perform a query according to the SQL statement, and then automatically visualize the obtained data analysis result, which can be used for business decision-making.
[0114] Through the above implementation, in the case that an associated analysis task cannot be searched from the historical analysis library, the current analysis task can be processed in combination with a large language model, an analysis statement is generated through a task key parameter, the processing efficiency of the current analysis task is improved, and the processing accuracy of the current analysis task is also guaranteed.
[0115] In addition, in the case that the current analysis task is a new demand task, in order to facilitate subsequent improvement of the response efficiency of other similar new demand tasks, the data analysis result and the related features of the current analysis task can also be stored in the historical analysis library.
[0116] In an example, after the analysis statement is executed in the target database to obtain the data analysis result of the current analysis task, the method further includes:
[0117] storing the data analysis result of the current analysis task, the demand association features, and the demand description features in the historical analysis library.
[0118] Specifically, in the case that the similarity between all historical analysis tasks in the historical analysis library and the current analysis task is less than a preset first threshold, after the data analysis result of the current analysis task is obtained, the data analysis result of the current analysis task, the demand association features, and the demand description features can be associated and stored in the historical analysis library, so as to be used as a historical analysis task for similarity calculation with other new analysis tasks in the future.
[0119] Through the above example, dynamic updating of the historical analysis library can be achieved, new demand tasks are supplemented to the historical analysis library in a timely manner, and the response efficiency of subsequent other tasks is improved.
[0120] The data analysis method in the field of unmanned driving provided by the embodiment of the present application obtains the task requirement association information of the current analysis task and obtains the demand description text of the demander for the current analysis task, determines the demand association features according to the task requirement association information, and uses a large language model to extract the demand description features from the demand description text, and then determines the similarity between the current analysis task and each historical analysis task according to the demand association features and demand description features of the current analysis task, as well as the demand association features and demand description features of each historical analysis task in a preset historical analysis library, and determines the historical analysis task with the highest similarity and a similarity greater than a preset first threshold as the associated analysis task, thereby determining the data analysis results of the associated analysis task in the historical analysis library, and determining the data of the current analysis task through the data analysis results of the associated analysis task. According to the analysis results, data analysis is realized. This method can combine the demand characteristics of the current analysis task and use the historical analysis tasks to obtain the analysis results, avoiding direct processing of the current analysis task. While improving the efficiency of task analysis, it can also improve the utilization rate of existing resources, thereby improving the task response speed. In addition, the historical analysis tasks cover richer result outputs. Compared with directly processing the current analysis task, it can obtain results that are more in line with user expectations. In addition, this method extracts the demand characteristics of the analysis task from the task demand association information and the demand description text respectively, and can predict the user's needs through other information related to the task, and extract the user's described needs through the text of the user's description of the task, so as to achieve the purpose of analyzing user needs from different angles, so that the final data analysis results can better meet user needs.
[0121] The data analysis method provided in the embodiment of the present application takes into account and gives full play to the utilization value of existing historical analysis tasks when building a data analysis link for a large language model. Most of the existing historical analysis tasks are relatively mature and general analysis tasks with multiple dimensions and wide coverage. They can cover most new needs input by users and can better serve unmanned vehicle data analysis users.
[0122] Furthermore, by matching related analysis tasks, there's no need to use SQL to query raw data and generate visual reports, saving resources, improving existing resource utilization, and increasing response speed. The demander may provide a description of the requirements for the current analysis task, but the actual description may not be accurate or comprehensive enough. By leveraging existing analysis tasks, richer output can be obtained, resulting in data analysis results that better meet expectations compared to directly executing the current analysis task.
[0123] And, in constructing the historical analysis library, the demand-related features and demand documents of the historical analysis task are combined, semantic analysis is performed by using the LLM, the historical analysis library is further optimized in depth, and relevant feature attributes of the demander are added to improve the analysis and inspection accuracy. In addition, the method can realize task processing by using the large language model for new demand tasks, thereby greatly improving the task processing efficiency and accuracy.
[0124] Figure 6 FIG. 1 is a structural schematic diagram of a data analysis device in the field of unmanned driving according to an embodiment of the present disclosure. As shown in the figure, the device comprises an information acquisition module 610, a feature extraction module 620, a feature comparison module 630, an associated task determination module 640, and a result determination module 650. Figure 6 The information acquisition module 610 is configured to acquire task demand-related information of a current analysis task and acquire demand description text of the demander for the current analysis task.
[0125] The feature extraction module 620 is configured to determine demand-related features according to the task demand-related information and extract demand description features from the demand description text by using a large language model.
[0126] The feature comparison module 630 is configured to determine the similarity between the current analysis task and each historical analysis task according to the demand-related features and the demand description features of the current analysis task and the demand-related features and the demand description features of each historical analysis task in a preset historical analysis library.
[0127] The associated task determination module 640 is configured to determine a historical analysis task with the highest similarity and greater than a preset first threshold as an associated analysis task corresponding to the current analysis task.
[0128] The result determination module 650 is configured to determine a data analysis result of the associated analysis task from the historical analysis library and determine a data analysis result of the current analysis task according to the data analysis result of the associated analysis task.
[0129] Optionally, the historical analysis library is configured to store a feature set of each historical analysis task, and the feature set comprises demand-related features and demand description features. The feature comparison module 630 comprises a feature set comparison unit, which is configured to:
[0130] construct a feature set of the current analysis task according to the demand-related features and the demand description features of the current analysis task;
[0131]
[0132] For each historical analysis task in the historical analysis library, a feature set of the historical analysis task is determined to be similar to a feature set of the current analysis task, and the similarity is determined as the similarity between the historical analysis task and the current analysis task.
[0133] Optionally, the feature comparison module 630 comprises a feature separate comparison unit, configured to:
[0134] For each historical analysis task in the historical analysis library, a first feature similarity of a requirement association feature of the historical analysis task to a requirement association feature of the current analysis task is determined, and a second feature similarity of a requirement description feature of the historical analysis task to a requirement description feature of the current analysis task is determined.
[0135] According to preset weights corresponding to the requirement association feature and the requirement description feature respectively, the first feature similarity and the second feature similarity are weighted to obtain the similarity between the historical analysis task and the current analysis task.
[0136] Optionally, the result determination module 650 is specifically configured to:
[0137] If the similarity between the associated analysis task and the current analysis task is greater than or equal to a preset second threshold, the data analysis result of the associated analysis task is determined as the data analysis result of the current analysis task, wherein the preset second threshold is greater than the preset first threshold.
[0138] If the similarity between the associated analysis task and the current analysis task is less than the preset second threshold, the data analysis result of the associated analysis task is adjusted to obtain the data analysis result of the current analysis task.
[0139] Optionally, the result determination module 650 is further configured to:
[0140] According to the requirement association feature and the requirement description feature of the current analysis task, and the requirement association feature and the requirement description feature of the associated analysis task, requirement difference information between the associated analysis task and the current analysis task is determined, and an analysis supplementary task is determined according to the requirement difference information.
[0141] On the basis of the data analysis result of the associated analysis task, the analysis supplementary task is executed to obtain the data analysis result of the current analysis task.
[0142] Optionally, the result determination module 650 is further configured to:
[0143] If the similarity between each of the historical analysis tasks and the current analysis task is less than the preset first threshold, then using a large language model to recognize the requirement description text to obtain task key parameters;
[0144] An analysis statement of the current analysis task is constructed according to the key parameters of the task, and the analysis statement is executed in the target database to obtain a data analysis result of the current analysis task.
[0145] Optionally, the result determination module 650 is further configured to:
[0146] The data analysis results, requirement association features and requirement description features of the current analysis task are stored in the historical analysis library.
[0147] The data analysis device in the unmanned driving field provided by the embodiment of the present disclosure can execute the steps of the data analysis method in the unmanned driving field provided by the method embodiment of the present disclosure. The execution steps and beneficial effects are no longer repeated here.
[0148] Figure 7 This is a schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 7 , which shows a structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. Figure 7 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0149] like Figure 7 As shown, the electronic device 500 may include a processing device 501, a ROM 502, a RAM 503, a bus 504, an input / output (I / O) interface 505, an input device 506, an output device 507, a storage device 508, and a communication device 509. The processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501 can perform various appropriate actions and processes to implement the method of the embodiment as described in the present disclosure according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via the bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0150] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts, thereby implementing the data analysis method for the unmanned field as described above. In such embodiments, the computer program can be downloaded and installed from a network by the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
[0151] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take on many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, transmit, or propagate program code for use by or in connection with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted using any suitable medium, including but not limited to a wire, an optical fiber, an RF (radio frequency) or the like, or any suitable combination thereof.
[0152] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and not be assembled into the electronic device. The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform:
[0153] obtain task requirement association information of a current analysis task, and obtain requirement description text of the current analysis task from a requirement party;
[0154] determine requirement association features according to the task requirement association information, and extract requirement description features from the requirement description text using a large language model;
[0155] determine similarity between the current analysis task and each of the historical analysis tasks according to the requirement association features and the requirement description features of the current analysis task, and the requirement association features and the requirement description features of each of the historical analysis tasks in the preset historical analysis library;
[0156] determine a historical analysis task with the highest similarity and greater than a preset first threshold as an associated analysis task corresponding to the current analysis task;
[0157] determine a data analysis result of the associated analysis task from the historical analysis library, and determine a data analysis result of the current analysis task according to the data analysis result of the associated analysis task.
[0158] Optionally, when the one or more programs are executed by the electronic device, the electronic device can further perform other steps described in the above embodiments.
[0159] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium can include one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical storage devices, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0160] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the disclosed concept. For example, the above features can be replaced with technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
Claims
1. A data analysis method in the field of unmanned driving, characterized in that: The method comprises: Obtaining task requirement related information of the current analysis task, and obtaining the demand description text of the demander for the current analysis task; Determine requirement association features based on the task requirement association information, and extract requirement description features from the requirement description text using a large language model; Determining the similarity between the current analysis task and each of the historical analysis tasks based on the requirement association characteristics and requirement description characteristics of the current analysis task and the requirement association characteristics and requirement description characteristics of each of the historical analysis tasks in a preset historical analysis library; Determine the historical analysis task with the highest similarity and a similarity greater than a preset first threshold as the associated analysis task corresponding to the current analysis task; The data analysis result of the associated analysis task is determined from the historical analysis library, and the data analysis result of the current analysis task is determined based on the data analysis result of the associated analysis task.
2. The method according to claim 1, characterized in that The historical analysis library is used to store the feature set of each historical analysis task, wherein the feature set includes demand association features and demand description features; Determining the similarity between the current analysis task and each of the historical analysis tasks based on the requirement association features and requirement description features of the current analysis task and the requirement association features and requirement description features of each of the historical analysis tasks in the historical analysis library includes: Constructing a feature set of the current analysis task based on the requirement association features and requirement description features of the current analysis task; For each historical analysis task in the historical analysis library, the similarity between the feature set of the historical analysis task and the feature set of the current analysis task is determined, and the similarity is determined as the similarity between the historical analysis task and the current analysis task.
3. The method according to claim 1, characterized in that Determining the similarity between the current analysis task and each of the historical analysis tasks based on the requirement association features and requirement description features of the current analysis task and the requirement association features and requirement description features of each of the historical analysis tasks in the historical analysis library includes: For each historical analysis task in the historical analysis library, determining a first feature similarity between a requirement-related feature of the historical analysis task and a requirement-related feature of the current analysis task, and determining a second feature similarity between a requirement description feature of the historical analysis task and a requirement description feature of the current analysis task; The first feature similarity and the second feature similarity are weighted according to preset weights corresponding to the requirement association feature and the requirement description feature, so as to obtain the similarity between the historical analysis task and the current analysis task.
4. The method according to claim 1, wherein Determining the data analysis result of the current analysis task according to the data analysis result of the associated analysis task includes: If the similarity between the association analysis task and the current analysis task is greater than or equal to a preset second threshold, determining the data analysis result of the association analysis task as the data analysis result of the current analysis task, wherein the preset second threshold is greater than the preset first threshold; If the similarity between the associated analysis task and the current analysis task is less than the preset second threshold, the data analysis result of the associated analysis task is adjusted to obtain the data analysis result of the current analysis task.
5. The method according to claim 4, characterized in that Adjusting the data analysis result of the association analysis task to obtain the data analysis result of the current analysis task includes: Determining requirement difference information between the associated analysis task and the current analysis task based on the requirement association characteristics and requirement description characteristics of the current analysis task and the requirement association characteristics and requirement description characteristics of the associated analysis task, and determining an analysis supplementary task based on the requirement difference information; Based on the data analysis result of the associated analysis task, the analysis supplement task is executed to obtain the data analysis result of the current analysis task.
6. The method according to claim 1, characterized in that After determining the similarity between the current analysis task and each of the historical analysis tasks, the method further includes: If the similarity between each of the historical analysis tasks and the current analysis task is less than the preset first threshold, then using a large language model to recognize the requirement description text to obtain task key parameters; An analysis statement of the current analysis task is constructed according to the key parameters of the task, and the analysis statement is executed in the target database to obtain a data analysis result of the current analysis task.
7. The method according to claim 6, characterized in that After executing the analysis statement in the target database to obtain the data analysis result of the current analysis task, the method further includes: The data analysis results, requirement association features and requirement description features of the current analysis task are stored in the historical analysis library.
8. A data analysis device for unmanned driving, characterized in that: The device comprises: An information acquisition module is used to obtain task requirement related information of the current analysis task and obtain the demand description text of the demander for the current analysis task; A feature extraction module is used to determine requirement-related features based on the task requirement-related information, and to extract requirement description features from the requirement description text using a large language model; a feature comparison module, configured to determine the similarity between the current analysis task and each of the historical analysis tasks based on the requirement association features and requirement description features of the current analysis task and the requirement association features and requirement description features of each of the historical analysis tasks in a preset historical analysis library; an associated task determining module, configured to determine the historical analysis task with the highest similarity and a similarity greater than a preset first threshold as the associated analysis task corresponding to the current analysis task; The result determination module is used to determine the data analysis result of the associated analysis task from the historical analysis library, and determine the data analysis result of the current analysis task based on the data analysis result of the associated analysis task.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data analysis method in the unmanned driving field as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the data analysis method in the field of unmanned driving as described in any one of claims 1 to 7 is implemented.