High-speed traffic acquisition method based on fused large model and related equipment

By adopting a parallel processing method based on a fusion large model in high-speed traffic scenarios, and using the prompt word model and the plug-in model to obtain preliminary and accurate traffic data respectively, the problem of poor accuracy of high-speed traffic query results is solved, and more stable and accurate query results are achieved.

CN120086248APending Publication Date: 2025-06-03NANJING MICROVIDEO TECH

Patent Information

Application Number
CN202510102862.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-15
Filing Date
2025-01-22
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art user query statements are not standardized in high-speed traffic scenarios, resulting in poor accuracy of query results and complex data query between business systems, which increases the query time and may cause data statistical deviations.

Method used

The high-speed traffic acquisition method based on the fusion large model is adopted, and the prompt word model and the plug-in model are processed in parallel. The prompt word model quickly obtains preliminary traffic data based on intent parameters, and the plug-in model obtains more accurate traffic data through dynamic orchestration and execution.

Benefits of technology

It improves the accuracy of high-speed traffic query results and the stability of system response, realizes rapid intent recognition and parameter extraction, and ensures data accuracy and consistency in complex query scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086248A_ABST
    Figure CN120086248A_ABST
Patent Text Reader

Abstract

The invention provides a high-speed traffic acquisition method based on a fused large model and related equipment, and relates to the technical field of natural language processing. The system utilizes the two models to start query processing at the same time, the cue word model directly and quickly obtains first traffic data based on intention parameters, and the plug-in model obtains more accurate second traffic data through dynamic arrangement and execution. When the second flow data is obtained, the system uses a more accurate query result; when the second flow data is not obtained, the system can still respond in time based on the first flow data. The two-channel parallel mechanism can obtain more accurate data in a more flexible question and answer mode during high-speed traffic, and ensures the accuracy of a query result and the stability of system response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a high-speed traffic acquisition method and related devices based on a fused large model. Background Art

[0002] With the acceleration of global economic integration and urbanization, the transportation system, as an important part of the national infrastructure, plays a key role in connecting cities, promoting economic development, and facilitating social communication. Expressways are a type of specially designed and constructed road infrastructure aimed at providing efficient, safe, and fast transportation services for vehicles. Expressways significantly improve traffic efficiency in terms of function, are suitable for long-distance rapid transportation, and reduce the incidence of traffic accidents by separating lanes and reducing level crossings. In addition, expressways also promote regional economic development, improve connections between regions, and effectively relieve traffic pressure within cities. The acquisition of high-speed traffic can enhance traffic management and optimize infrastructure to cope with the increasing vehicle flow and improve road safety. The acquisition of high-speed traffic can enhance traffic management efficiency, improve road safety, optimize infrastructure planning and maintenance, enhance travel experience, support intelligent transportation systems, and promote economic development.

[0003] Currently, operators mainly use template matching systems to handle users' traffic query requirements. This system extracts information such as time, location, and service type from the user input through a preset keyword list and query template, and calls the corresponding data interface to obtain traffic information according to the matching results. For complex query scenarios that the system cannot match, it is transferred to manual customer service to manually query multiple business systems and gradually summarize the query results.

[0004] However, in actual applications, especially in high-speed traffic scenarios, users' query expressions are often not standardized and vary greatly from the preset keywords and templates. At the same time, due to the relative independence of business systems, data queries need to be switched and transferred between multiple systems, which not only increases the query time but also easily causes deviations in data statistics and affects the accuracy of query results.

[0005] Therefore, the query results of related technologies for high-speed traffic are less accurate. Summary of the Invention

[0006] This application provides a high-speed traffic acquisition method and related devices based on a fused large model to improve the accuracy of query results for high-speed traffic.

[0007] In a first aspect, the present application provides a high-speed traffic acquisition method based on a fusion large model, including: receiving a user question and inputting it into a prompt model and a plugin model respectively for parallel processing; wherein, the prompt model calls a traffic query function based on intent parameters to obtain first traffic data; the intent parameters are business type parameters, time parameters, and location parameters; the plugin model arranges the calling order of plugins according to the user question; the plugin model executes the calling of plugins and processes the result transmission between plugins to obtain second traffic data of the plugin call; when the plugin model obtains the second traffic data, a final answer is generated based on the second traffic data; when the plugin model does not obtain the second traffic data, a final answer is generated based on the first traffic data.

[0008] By adopting the above technical solution, the system uses two models to start query processing simultaneously. The prompt model directly obtains the first traffic data quickly based on the intent parameters, and the plugin model obtains more accurate second traffic data through dynamic arrangement and execution. When the second traffic data is obtained, the system uses the more accurate query result; when the second traffic data is not obtained, the system can still respond in a timely manner based on the first traffic data. This dual-channel parallel mechanism can obtain more accurate data in a more flexible question-and-answer manner at high speed traffic, ensuring the accuracy of the query result and the stability of the system response.

[0009] Combined with some embodiments of the first aspect, in some embodiments, the step of the prompt model calling a traffic query function based on intent parameters to obtain first traffic data specifically includes: determining whether the user question is a common traffic question; if the user question is a common traffic question, the prompt model extracts the first intent parameter in the user question; the prompt model calls a traffic query function based on the first intent parameter to obtain first traffic data; if the user question is not a common traffic question, the prompt model calculates the similarity between the user question and the common traffic questions respectively; the prompt model selects the common traffic question with the highest similarity as the target common traffic question, and then the prompt model extracts the second intent parameter in the target common traffic question; the prompt model calls a traffic query function based on the second intent parameter to obtain first traffic data.

[0010] By adopting the above technical solution, for questions with standard expressions, the system directly extracts intent parameters for accurate query; for questions with non-standard expressions, the system finds the closest standard question template through similarity calculation and reuses its parameter extraction logic. This hierarchical recognition mechanism avoids the complex semantic parsing process, realizes fast intent recognition and parameter extraction, and greatly improves the speed of the system to process various queries.

[0011] In some embodiments in combination with some embodiments of the first aspect, before the step of the prompt model calling the traffic query function based on the first intention parameter to obtain the first traffic data, the method further includes: sorting business types according to high-speed traffic; based on the sorted high-speed traffic, establishing a traffic query method library for common traffic problems, and the traffic query method library performs fast query of traffic data based on time parameters, location parameters, and business type parameters.

[0012] By adopting the above technical solution, the discrete business data is systematically sorted out. Data indexes are established through three core dimensions of time, location, and business type, realizing the rapid positioning and extraction of data, standardizing the data organization method, and corresponding traffic data can be directly obtained based on time parameters, location parameters, and business type parameters.

[0013] In a second aspect, the present application provides a training method for a plugin model, including: constructing a fine-tuning data set based on artificial business processes; adding plugin call tool information and corresponding plugin call result information to the constructed fine-tuning data set; training a basic large model according to the fine-tuning data set to obtain a plugin model, and the training parameters include: batch size parameter, learning rate parameter, number of training steps parameter, rank parameter, and optimizer parameter; configuring the system prompt word of the plugin model, and the system prompt word includes identity configuration information, scenario configuration information, business type processing method, and example information.

[0014] By adopting the above technical solution, the fine-tuning data set constructed based on artificial business processes accurately reflects the actual business scenario and processing rules, avoiding the model's incorrect understanding of business logic; by presetting plugin call information and corresponding call results in the training data, the model can accurately grasp the usage boundary and data processing requirements of the plugin; the precise configuration of the system prompt word further standardizes the behavior constraints of the model, enabling it to strictly follow business rules when processing queries. This targeted training solution significantly improves the processing accuracy of the model in actual business scenarios. Especially in complex query scenarios, the model can accurately select and arrange the plugin call order based on the deeply understood business rules, ensuring the accuracy and consistency of the final data.

[0015] In some embodiments in combination with some embodiments of the second aspect, in some embodiments, a plug-in model is obtained by training a basic large model according to a fine-tuning data set. The training parameters include the steps of batch size parameter, learning rate parameter, number of training steps parameter, rank parameter, and optimizer parameter, which specifically include: setting fine-tuning training parameters, where the fine-tuning training parameters include pre-trained model path, fine-tuning data set path, fine-tuning phase identifier, fine-tuning template configuration, target layer configuration, model saving path, batch processing parameter, maximum sequence length parameter, gradient accumulation step parameter, learning rate scheduler parameter, logging step parameter, weight saving interval step parameter, number of training epochs parameter, maximum number of samples parameter, and quantization bit width parameter; obtaining the weight matrix of the basic large model, and generating a first low-rank matrix and a second low-rank matrix by means of low-rank decomposition; updating the first low-rank matrix and the second low-rank matrix based on the fine-tuning data set; updating the weight matrix based on the product of the first low-rank matrix and the second low-rank matrix; obtaining the training set loss value and the validation set loss value during the training process; and ending the training process when the training set loss value shows a downward trend and the validation set loss value tends to be stable.

[0016] By adopting the above technical solution, through a comprehensive training parameter configuration and optimization mechanism, it is ensured that the model can fully learn business rules and data associations. The system automatically adjusts the training process by monitoring the training effect in real time, avoiding the problem of overfitting. This training method ensures the stability and accuracy of the model when processing complex queries.

[0017] In some embodiments in combination with some embodiments of the second aspect, in some embodiments, the calculation formula for updating the weight matrix based on the product of the first low-rank matrix and the second low-rank matrix is: In the formula, is the updated weight matrix, is the weight matrix, is in its transposed form, is in its transposed form, is the first low-rank matrix, is the second low-rank matrix, is the rank number.

[0018] By adopting the above technical solution, through the method of low-rank decomposition, the original high-dimensional weight matrix is decomposed into the product of two low-rank matrices, reducing the number of parameters that need to be updated and stored. Compared with the related technology that needs to update all model parameters, this solution only needs to optimize the parameters of the low-dimensional matrices, greatly reducing the computational complexity and storage overhead while maintaining the model performance. Specifically, when the dimension of the original weight matrix is m×n, by choosing a rank k that is much smaller than min(m, n), the number of parameters can be reduced from m×n to k×(m + n), and at the same time, an effective approximation of the original weight matrix is maintained through matrix multiplication.

[0019] In some embodiments in combination with the second aspect, in some embodiments, a fine-tuning data set is constructed based on artificial business process language; the steps of adding plug-in call tool information and corresponding plug-in call result information to the constructed fine-tuning data set specifically include: establishing a set of traffic business processing tools, where the set of traffic business processing tools includes a traffic query tool; obtaining artificial dialogue corpus, where the artificial dialogue corpus includes question information and corresponding answer information; processing the artificial dialogue corpus to obtain dialogue training samples, where: identifying questions in the question information that need to call the set of traffic business processing tools; using the tool call result corresponding to the question as an independent dialogue turn; generating a fine-tuning data set based on the dialogue training samples.

[0020] By adopting the above technical solutions, a structured training data set is constructed, closely associating business tool calls with query scenarios. By using the tool call result as an independent dialogue turn, not only is the query intent recognition learned, but also the rules for cross-system data integration are mastered. This organization method of training data ensures that the model can accurately understand complex query requirements and provide accurate query results.

[0021] In a third aspect, the present application provides a high-speed traffic acquisition system based on a fusion large model. The high-speed traffic acquisition system based on a fusion large model includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the high-speed traffic acquisition system based on a fusion large model to execute the methods described in the first aspect and any possible implementation manner in the first aspect, and the second aspect and any possible implementation manner in the second aspect.

[0022] In a fourth aspect, the present application provides a computer program product containing instructions. When the computer program product runs on a high-speed traffic acquisition system based on a fusion large model, it enables the high-speed traffic acquisition system based on a fusion large model to execute the methods described in the first aspect and any possible implementation manner in the first aspect, and the second aspect and any possible implementation manner in the second aspect.

[0023] In a fifth aspect, the present application provides a computer-readable storage medium including instructions. When the instructions run on a high-speed traffic acquisition system based on a fusion large model, it enables the high-speed traffic acquisition system based on a fusion large model to execute the methods described in the first aspect and any possible implementation manner in the first aspect, and the second aspect and any possible implementation manner in the second aspect.

[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The system starts query processing using two models simultaneously. The prompt model directly obtains the first traffic data quickly based on the intent parameters, and the plugin model obtains more accurate second traffic data through dynamic orchestration and execution. When the second traffic data is obtained, the system uses the more accurate query results; when the second traffic data is not obtained, the system can still respond in a timely manner based on the first traffic data. This dual-channel parallel mechanism can obtain more accurate data in a more flexible Q&A manner at high traffic speeds, ensuring the accuracy of query results and the stability of system responses.

[0025] 2. For problems with standard expressions, the system directly extracts intent parameters for accurate query; for problems with non-standard expressions, the system finds the closest standard question template through similarity calculation and reuses its parameter extraction logic. This hierarchical recognition mechanism avoids complex semantic parsing processes, realizes fast intent recognition and parameter extraction, and greatly improves the speed of the system to process various queries.

[0026] 3. The fine-tuning dataset constructed based on the artificial business process accurately reflects the actual business scenario and processing rules, avoiding the model's misinterpretation of business logic; by presetting plugin call information and corresponding call results in the training data, the model can accurately grasp the usage boundaries and data processing requirements of the plugin; the precise configuration of the system prompt further standardizes the behavior constraints of the model, enabling it to strictly follow business rules when processing queries. This targeted training scheme significantly improves the processing accuracy of the model in actual business scenarios. Especially in complex query scenarios, the model can accurately select and orchestrate the plugin call order based on the deeply understood business rules, ensuring the accuracy and consistency of the final data. Description of the Drawings

[0027] Figure 1 is a schematic flowchart of a method for obtaining high-speed traffic based on a fusion large model in an embodiment of the present application; Figure 2 is a specific schematic flowchart of step S301 in an embodiment of the present application; Figure 3 is a schematic flowchart of a training method for a plugin model in an embodiment of the present application; Figure 4 is an exemplary hardware structure schematic diagram of a system for obtaining high-speed traffic based on a fusion large model in an embodiment of the present application. Detailed Embodiments

[0028] The terms used in the following embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and appended claims of this application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in this application refers to and includes any or all possible combinations of one or more of the listed items.

[0029] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0030] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a high-speed traffic acquisition method based on a fusion large model in the embodiments of this application; S101. Receive a user question and input it into a prompt model and a plugin model respectively for parallel processing; Among them, the user question represents the content of the query request submitted by the user through the system interface; parallel processing represents a mechanism for simultaneously inputting the user question into two models for processing.

[0031] In some embodiments, both the prompt model and the plugin model are deployed in an API manner to achieve the distributed architecture and service decoupling of the system. Specifically, each model is encapsulated as an independent microservice and provides a service interface to the outside through the RESTful API or gRPC protocol. After receiving the user question, the system will call the API services of these two models in parallel for question-answering processing. The prompt model API mainly processes standardized traffic query requests and can quickly return basic query results; the plugin model API is responsible for processing complex query scenarios and obtaining more accurate query results by calling multiple plugins.

[0032] In some embodiments, before inputting the user question into the prompt model and the plugin model respectively for parallel processing, it is also necessary to preprocess the received user question, including removing special characters, unifying the format, etc.; details are not described herein again.

[0033] It should be noted that this embodiment mainly describes the usage process of the prompt model and the plugin model, and focuses on introducing the complete process of how the system uses these two models to process user questions and obtain traffic data. As for the specific implementation principles, internal architectures, training methods and other technical details of the prompt model and the plugin model, they will be described in detail in the following embodiments and will not be elaborated here.

[0034] S102. Among them, the prompt model calls the traffic query function based on the intent parameters to obtain the first traffic data; the intent parameters are the business type parameter, the time parameter, and the location parameter. Among them, the intent parameter represents the key query information extracted from the user question; the business type parameter is used to represent the specific business field of the query, such as airports, stations, etc.; the time parameter refers to the query time range or specific time point; the location parameter represents the spatial location information of the query; the traffic query function is a system interface for quickly obtaining basic traffic data; the first traffic data is used to represent the preliminary query result quickly obtained through the prompt model. For example, from "the passenger flow of Shanghai Airport yesterday", the time "yesterday", the location "Shanghai Airport", and the business type "passenger flow" are extracted.

[0035] In some embodiments, the prompt model first identifies and extracts the intent parameters through a preset parameter extraction rule; then standardizes the extracted parameters to ensure compliance with the input requirements of the query function; then calls the corresponding traffic query function to obtain data; finally, formats the obtained data to form the first traffic data.

[0036] S103. The plugin model arranges the call order of the plugins according to the user question. Among them, the plugin represents an independent processing module with specific functions; the call order refers to the execution sequence relationship between multiple plugins; arranging means the process of determining the optimal plugin combination and execution order according to the problem characteristics. For example, for a complex query such as "compare the passenger flow trends of each terminal in the past three days", it is necessary to first call the data acquisition plugin and then call the data analysis plugin.

[0037] This step is executed after the plugin model completes the problem analysis. In some embodiments, the plugin model first analyzes the query intent and complexity of the user question; then selects the required plugin combination from the plugin library according to the problem characteristics; then constructs the call link based on the dependency relationship between the plugins; finally, generates a specific execution plan to determine the input-output relationship and execution timing of each plugin.

[0038] S104. The plugin model executes the call of the plugins and processes the result transfer between the plugins to obtain the second traffic data of the plugin call. Among them, the plug-in call represents the process of executing each plug-in in a predetermined order; the result transfer refers to the data flow process of using the output of the previous plug-in as the input of the subsequent plug-in; the second traffic data is used to represent the accurate query result obtained through the complete plug-in call chain. In some embodiments, the system first initializes the plug-in running environment; then calls each plug-in in sequence according to the orchestration order and ensures the correct transfer of data; at the same time, monitors the execution status of the plug-ins, including execution time, resource occupancy, etc.; finally, summarizes the processing results of each plug-in to form the complete second traffic data.

[0039] S105. When the plug-in model obtains the second traffic data, generate a final answer based on the second traffic data; when the plug-in model fails to obtain the second traffic data, generate a final answer based on the first traffic data.

[0040] It should be noted that in actual applications, the plug-in model may fail to obtain the second traffic data for various reasons. For example, during the plug-in call process, problems such as failed initialization of the plug-in running environment, abnormal data transfer between plug-ins, or interruption of the call chain due to an error in the execution of a certain plug-in may occur; in addition, reasons such as the temporary unavailability of the required data source, unexpected data format, or excessive data volume resulting in processing timeout may also cause the situation of failing to obtain the second traffic data.

[0041] It is precisely considering the above possible abnormal situations that the system adopts a mechanism of parallel processing of the prompt model and the plug-in model. When the plug-in model cannot normally obtain the second traffic data, the system can use the first traffic data obtained by the prompt model to generate an answer.

[0042] It can be seen that the system uses two models to start query processing simultaneously. The prompt model directly obtains the first traffic data based on the intent parameters, and the plug-in model obtains more accurate second traffic data through dynamic orchestration and execution. When the second traffic data is obtained, the system uses the more accurate query result; when the second traffic data is not obtained, the system can still respond in a timely manner based on the first traffic data. This dual-channel parallel mechanism can obtain more accurate data in a more flexible Q&A manner at high traffic speeds, ensuring the accuracy of the query result and the stability of the system response.

[0043] Please refer to Figure 2 , Figure 2 which is a specific process schematic diagram of step S301 in the embodiments of the present application; In some embodiments, step S103 specifically includes; S1031. Determine whether the user's question is a common traffic problem; Among them, the common traffic problem represents a standardized traffic query problem template predefined by the system; In some embodiments, the user's question is matched with a preset common traffic problem template; it is determined whether it belongs to a common traffic problem by means of rule matching or pattern recognition; finally, the judgment result is output to determine the subsequent processing path.

[0044] In some other embodiments, when the vector similarity calculation method is adopted, the determination criterion of "common traffic problem" changes from simple pattern matching to judgment based on a similarity threshold. If there is a template problem with a similarity greater than the preset threshold, it is determined that the current user's question is a common traffic problem. This judgment method is more flexible and can handle query requests with slightly different expression forms but the same essence.

[0045] Correspondingly, step S1035 also needs to be adjusted: change the original "select the common traffic problem with the highest similarity" to "select the common traffic problem with a similarity greater than the preset threshold and the highest similarity" as the target common traffic problem.

[0046] S1032. If the user's question is a common traffic problem; then the prompt word model extracts the first intention parameter in the user's question; in some embodiments, the prompt word model first performs semantic analysis on the user's question to identify the key information in the question; then according to the preset parameter extraction rules, locates and extracts various intention parameters; then performs standardization processing on the extracted parameters to ensure unified format; finally, organizes the processed parameters into a standard query parameter set.

[0047] In some specific embodiments, named entity recognition technology is used to identify parameters. First, the input text is segmented and part-of-speech tagged to identify key entity information such as time, location, and business type; then parameter value extraction is performed based on a preset rule template, and the identified entities are mapped to parameter types; then the extracted parameters are verified for validity, checking the integrity and reasonableness of the parameters, and eliminating invalid or abnormal parameters; finally, the verified parameters are uniformly formatted to ensure that the output result conforms to the system standard specification, which is not limited here.

[0048] In some specific embodiments, parameters are located based on dependency syntactic analysis. First, a syntactic tree of the sentence is constructed to analyze the dependency relationships between words; then semantic role labeling technology is used to identify the semantic components in the sentence to determine the roles of each parameter at the semantic level; then normalization processing is performed on the identified parameters to uniformly convert different expression forms into a standard format; finally, the processed parameters are organized into a structured parameter set for subsequent query operations, which is not limited here.

[0049] In some specific embodiments, step S1032 specifically includes: The prompt word model performs intent recognition and parameter extraction based on the Qwen-max large model. Through the preset prompt word templates, the model can accurately identify the major intent categories in the user's question and extract key parameter information, including parameters such as operation type, time range, geographical location, etc.

[0050] The prompt word structure includes the following key components: Definition of the model role: Clearly define the identity and responsibilities of the model as a traffic data query assistant; Definition of the business scope: Specify the business types that can be processed and their specific meanings; Definition of parameter types: Standardize the value ranges of parameters such as location, time, and operation; Definition of output format: Specify the output structure and field requirements in JSON format; The specific configuration of the prompt words includes: Business type: Divided into five major categories: traffic, provincial boundary traffic, events, accidents, and traffic control; Location parameter: Supports the entire G2 line and its subdivided sections (such as the Xuzhou-Suqian section, Huai'an section, Yangzhou section, etc.). When only G2 is mentioned, it is defaulted to the entire G2 line; Operation parameter: Corresponding to the business type, serving as the core basis for intent recognition; Time parameter: Supports three time dimensions: today's cumulative, yesterday's cumulative, and hourly cumulative; Output parameter: Includes two types: month-on-month and empty; Output format: Adopts the standard JSON format, including two main fields: intent and parameters, and requires the model to strictly follow the preset format to output the results.

[0051] S1033. The prompt word model calls the traffic query function based on the first intent parameter to obtain the first traffic data; Among them, the traffic query function represents a system interface for accessing the traffic database and obtaining data.

[0052] In some embodiments, the system selects the corresponding query function according to the parameter type; then performs a data query operation and obtains the raw data; finally, formats the query result to form the standardized first traffic data.

[0053] S1034. If the user's question is not a common traffic question, the prompt word model calculates the similarity between the user's question and the common traffic questions respectively; Among them, the similarity represents a numerical index for measuring the degree of proximity between two questions in terms of semantics or structure; The formula for vector similarity matching is: In the formula, cos(θ) is the cosine similarity value between two vectors (the user's question and the common traffic question), x iis the value of the i-th component in the first vector, y i is the value of the i-th component in the second vector, and n is the dimension of the vector.

[0054] In some embodiments, the system first performs feature extraction and vectorization processing on the user's question; then traverses each question in the common traffic problem library; then uses a preset similarity algorithm to calculate the matching degree between questions; and finally generates an ordered list of all similarity calculation results.

[0055] S1035. The prompt word model selects the common traffic problem with the highest similarity as the target common traffic problem, and the prompt word model extracts the second intention parameter in the target common traffic problem; It should be noted that the principle and process of this step are similar to those of step S1032. The relevant principles and processes can be referred to step S1032, and are not limited here.

[0056] S1036. The prompt word model calls the traffic query function based on the second intention parameter to obtain the first traffic data.

[0057] It should be noted that the principle and process of this step are similar to those of step S1033. The relevant principles and processes can be referred to step S1033, and are not limited here.

[0058] It should be noted that the advantages of this embodiment can be seen here. For example: First, the vector similarity calculation method has the characteristic of high calculation efficiency. Compared with complex semantic understanding and deep learning methods, the vectorization calculation process is direct and efficient, and can quickly complete the similarity comparison; Second, since each user question can calculate a definite similarity value with the preset template and select the template with the highest similarity as the matching result, the query process will surely obtain a definite result, avoiding the situation of query failure; Third, by introducing a similarity threshold judgment mechanism, while ensuring the query speed, a certain matching accuracy is also maintained, achieving a balance between efficiency and accuracy.

[0059] It can be seen that for questions with standard expressions, the system directly extracts the intention parameters for accurate query; for questions with non-standard expressions, the system finds the closest standard question template through similarity calculation and reuses its parameter extraction logic. This hierarchical recognition mechanism avoids complex semantic parsing processes, realizes fast intention recognition and parameter extraction, and greatly improves the speed of the system to process various queries.

[0060] In some other embodiments, before step S1033, it further includes: S1037. Sort the business types according to high-speed traffic; In some embodiments, the system first collects the original traffic data at each monitoring point, then divides the data into different business types according to business scenarios and application requirements, then cleans and standardizes the data for each business type, and finally establishes a unified data index and storage structure to ensure the queryability and consistency of the data.

[0061] In some embodiments, the major business categories are as follows: Dispatch category: Involves traffic command and emergency response; Operation category: Includes revenue analysis and operation management; Toll collection category: Responsible for business related to toll collection; Supervision category: Conducts road condition monitoring and safety management; User profile category: Analyzes the characteristics of road users; Vehicle profile category: Studies the traffic patterns of vehicles; The traffic data is classified by monitoring location as: Section traffic: Measures the vehicle flow of the road cross-section; Vehicle type traffic: Counts the traffic volume of different types of vehicles; Entrance / exit traffic: Records the vehicle data entering and leaving the toll station; Provincial boundary traffic: Counts the vehicle data passing across provinces; The question-and-answer types are divided into: Numeric type: Directly queries specific traffic data; Index type: Calculates relevant statistical indexes; S1038. Based on the sorted highway traffic, a traffic query method library is established for common traffic problems. The traffic query method library performs fast query of traffic data based on time parameters, location parameters, and business type parameters.

[0062] Among them, the traffic query method library represents a system component containing various query functions and interfaces; the time parameter represents the query time range or time point; the location parameter represents the query spatial location or road section; the business type parameter represents the specific query business category.

[0063] In some specific embodiments, the query mode and access characteristics are first analyzed, and then a multi-dimensional index structure is established, so that the query process can quickly locate and extract the target data block, avoiding full-scale data scanning, thereby realizing efficient retrieval of traffic data.

[0064] It can be understood that other methods can also be used to establish the query method library, such as query optimization methods based on distributed computing frameworks, etc., which are not limited here.

[0065] Continuing with the previous example: For high-frequency traffic query requirements, a standardized query method system has been established. This method enables fast query through three core parameters: Time parameter: Determine the time range of the query; Location parameter: Specify the spatial location of the query; Operation parameter: Define the specific query type.

[0066] It can be seen that the discrete business data has been systematically organized. By establishing data indexes through the three core dimensions of time, location, and business type, the rapid positioning and extraction of data have been achieved, the data organization method has been standardized, and the corresponding traffic data can be directly obtained based on the time parameter, location parameter, and business type parameter.

[0067] Please refer to Figure 3 , Figure 3 which is a flowchart of the training method of the plug-in model in the embodiment of the present application; S301. Construct a fine-tuning data set based on artificial business process language; add plug-in call tool information and corresponding plug-in call result information to the constructed fine-tuning data set; Among them, the artificial business process corpus refers to a collection of representative business conversations and operation instructions written or collected manually; the fine-tuning data set represents a standardized data set for model training; the plug-in call tool information is used to represent the call methods and parameter specifications of various business plug-ins; the plug-in call result information refers to the standardized result data returned after the plug-in call is executed.

[0068] The main purpose of constructing the fine-tuning data set based on the artificial business process corpus is to provide high-quality training samples for model training. In some embodiments, first, typical business scenario conversations are collected and sorted out, then these conversations are converted into a standard training sample format, then the trigger points and call methods of the plug-in call are marked in the samples, and finally the corresponding call result information is supplemented to form a complete training data set.

[0069] In some embodiments, step S301 specifically includes: S3011. Establish a traffic business processing tool set, and the traffic business processing tool set includes a traffic query tool; Among them, the traffic business processing tool set refers to a function component library for processing highway traffic-related businesses; the traffic query tool represents a dedicated program interface for obtaining and analyzing traffic data; the business processing tool represents various function modules for traffic data operation and analysis; the tool set refers to a tool library organized by function classification.

[0070] In some embodiments, first, the core functional requirements in highway traffic services are sorted out, then standardized tool interface specifications are designed, then the query and processing functions of various traffic data are implemented, and finally these tools are integrated into a unified tool set to provide necessary functional support for subsequent model training.

[0071] S3012. Obtain artificial dialogue corpus, where the artificial dialogue corpus includes question information and corresponding answer information; Among them, the artificial dialogue corpus refers to the real business dialogue content compiled or collected manually; the question information represents the specific consultation or request of the user in the traffic service scenario; the answer information is the standard processing response corresponding to the question; the dialogue corpus is used to represent the complete question-answer interaction record.

[0072] In some embodiments, real traffic service dialogue records are collected through multiple channels, including customer service logs, business consultation records, and standard Q&A sets, etc., to ensure that the collected dialogue corpus covers various typical business scenarios and contains complete question descriptions and corresponding processing answers.

[0073] S3013. Process the artificial dialogue corpus to obtain dialogue training samples, where: identify the questions in the question information that need to call the traffic service processing tool set; use the tool call result corresponding to the question as an independent dialogue turn; In some embodiments, the collected original dialogue is analyzed and processed to identify the types of questions that need to call tools, and the tool call process and results are used as independent dialogue links to form a standardized training sample format, ensuring that the model can learn the timing and method of tool calls.

[0074] S3014. Generate a fine-tuning data set based on the dialogue training samples.

[0075] It can be seen that a structured training data set is constructed, closely associating business tool calls with query scenarios. By using the tool call result as an independent dialogue turn, not only the query intent recognition is learned, but also the rules of cross-system data integration are mastered. This organization method of training data ensures that the model can accurately understand complex query requirements and provide accurate query results.

[0076] S302. Train the basic large model according to the fine-tuning data set to obtain a plug-in model, and the training parameters include: batch size parameter, learning rate parameter, training step parameter, rank parameter, and optimizer parameter; In some embodiments, based on the prepared fine-tuning data set, appropriate training parameters are set, and the pre-trained basic model is used for directional training. Through multiple rounds of iterative optimization, the model's understanding and usage ability of plug-in calls are gradually improved, and finally a plug-in model adapted to specific business scenarios is obtained.

[0077] Step S302 specifically includes: S3021. Set fine-tuning training parameters, which include the pre-trained model path, fine-tuning dataset path, fine-tuning stage identifier, fine-tuning template configuration, target layer configuration, model saving path, batch processing parameters, sequence maximum length parameter, gradient accumulation step parameter, learning rate scheduler parameter, logging step parameter, weight saving interval step parameter, training epoch parameter, maximum sample number parameter, and quantization bit width parameter; In some embodiments, various training parameters are reasonably configured according to business requirements and hardware resource conditions, including model path setting, training control parameters, optimizer parameters, resource limit parameters, etc.

[0078] In some specific embodiments, when the stage is sft, it is necessary to specify the pre-trained local model path, fine-tuning dataset path and name, and at the same time configure the fine-tuning template for setting the configuration of specific tasks or formats. The fine-tuning type is selected as LORA, the target layer for fine-tuning application is W_pack, and the result model saving path is set. In addition, it is also necessary to configure whether to train from scratch, set the batch size, set the maximum length of tokens to 1024, the gradient accumulation step to 8, select cosine as the learning rate scheduler, record logs every 10 steps, and save weights every 100 steps. The overall training control parameters include: the total number of training epochs is 500, the maximum sample number is 3000, and the quantization bit width is 4 bits.

[0079] S3022. Obtain the weight matrix of the basic large model, and generate the first low-rank matrix and the second low-rank matrix by means of low-rank decomposition; S3023. Update the first low-rank matrix and the second low-rank matrix based on the fine-tuning dataset; S3024. Update the weight matrix based on the product of the first low-rank matrix and the second low-rank matrix; The calculation formula for updating the weight matrix based on the product of the first low-rank matrix and the second low-rank matrix is: In the formula, is the updated weight matrix, is the weight matrix, is the transposed form of, is the transposed form of, is the first low-rank matrix, is the second low-rank matrix, is the rank number.

[0080] Continuing with the above example, in the LoRA fine-tuning technique, it is mainly to perform low-rank approximation on the weight matrix of the large model to save storage and computing resources, rather than updating all parameters.

[0081] For the weight matrix Instead of directly updating the entire matrix, LoRA introduces two smaller matrices where is a lower rank number. During the fine-tuning process, only these two low-rank matrices are updated, and then the effect of the original weight matrix is updated through matrix multiplication The update formula during fine-tuning can be expressed as: In practical applications and are small-scale parameter matrices learned through backpropagation, used to fit the weight changes related to the task. In the context of LoRA, and are constrained by regularization to ensure the low-rank property, and they usually maintain a small rank during training thus achieving efficient and resource-saving fine-tuning.

[0082] It can be seen that through the method of low-rank decomposition, the original high-dimensional weight matrix is decomposed into the product of two low-rank matrices, reducing the number of parameters that need to be updated and stored. Compared with the related technology that needs to update all model parameters, this scheme only needs to optimize the parameters of the low-dimensional matrices, greatly reducing the computational complexity and storage overhead while maintaining the model performance. Specifically, when the dimension of the original weight matrix is m×n, by choosing a rank k much smaller than min(m, n), the number of parameters can be reduced from m×n to k×(m + n), and at the same time, an effective approximation of the original weight matrix is maintained through matrix multiplication.

[0083] S3025. Obtain the training set loss value and the validation set loss value during the training process; S3026. When the training set loss value shows a downward trend and the validation set loss value tends to be stable, end the training process.

[0084] Among them, the training set loss value refers to the error index of the model on the training data; the validation set loss value represents the performance of the model on the validation data; the loss trend is used to represent the convergence situation of the model training; the training termination condition refers to the judgment criterion for deciding to stop training.

[0085] In some embodiments, during the training process, continuously monitor the performance of the model on the training set and the validation set, judge the learning state of the model through the change trend of the loss value. When the training loss continuously decreases and the validation loss tends to be stable, it indicates that the model has achieved a good training effect, and at this time, the training process can be terminated.

[0086] It can be seen that through comprehensive training parameter configuration and optimization mechanisms, the model is ensured to fully learn business rules and data associations. By monitoring the training effect in real time and automatically adjusting the training process, the system avoids the problem of overfitting. This training method guarantees the stability and accuracy of the model when processing complex queries.

[0087] S303. Configure the system prompt words of the plug-in model. The system prompt words include identity configuration information, scenario configuration information, business type processing methods, and example information.

[0088] Among them, the system prompt words represent the configuration information used to guide the behavior of the model; the identity configuration information is used to define the role characteristics of the model; the scenario configuration information represents the business environment applicable to the model; the business type processing method refers to the standard processing process for different business scenarios; and the example information is used to display the standard input and output forms.

[0089] In some embodiments, according to specific business requirements, design a complete system of prompt words, including the identity positioning of the model, application scenarios, processing procedures, and example explanations. Through these configuration information, guide the behavior performance of the model in actual applications to ensure that the model output meets business expectations.

[0090] In some embodiments, when configuring the system prompt words for the plug-in model, it is necessary to design a concise and complete system of prompt words. This configuration mainly includes the following core elements: Clarify the identity positioning of the model: Configure the role characteristics of the model as a professional AI assistant to maintain professional and friendly behavioral guidelines during the answering process. Standardize the plug-in usage process: Guide the model to obtain information through the plug-in first; require the model to clearly state the plug-ins used when answering; stipulate alternative solutions when the plug-ins cannot handle; provide standard example explanations: show the processing procedures for common problems; illustrate the input and output formats through specific cases; cover the processing methods for different types of business scenarios. These prompt word configurations together constitute the behavior guidelines of the model, ensuring that the model can accurately understand user needs and provide services through appropriate plug-ins. Through scientific prompt word design, the model shows consistent and expected behavior performance in actual applications and outputs standardized high-quality answers. At the same time, this configuration system also provides clear behavior boundaries for the model, helping to maintain the stability of service quality.

[0091] It can be seen that the fine-tuning dataset constructed based on the manual business process accurately reflects the actual business scenarios and processing rules, avoiding the model's misinterpretation of business logic. By presetting the plugin call information and corresponding call results in the training data, the model can accurately grasp the usage boundaries and data processing requirements of the plugins. The precise configuration of the system prompt words further standardizes the behavior constraints of the model, enabling it to strictly follow business rules when processing queries. This targeted training scheme significantly improves the processing accuracy of the model in actual business scenarios. Especially in complex query scenarios, the model can accurately select and arrange the plugin call order based on the deeply understood business rules, ensuring the accuracy and consistency of the final data.

[0092] The following introduces the exemplary high-speed traffic acquisition system 400 based on the integrated large model provided by the embodiments of the present application. Figure 4 FIG. 5 is an exemplary hardware structure diagram of the high-speed traffic acquisition system 400 based on the integrated large model provided by the embodiments of the present application.

[0093] In some embodiments, the high-speed traffic acquisition system 400 based on the integrated large model is a computer device or the high-speed traffic acquisition system 400 based on the integrated large model includes a computer device. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with other external terminals or servers through a network connection. In some embodiments, the network interface can be a wired network interface, and in some embodiments, the network interface can also be a wireless network interface. The computer program, when executed by the processor, implements the method in the embodiments of the present application.

[0094] Those skilled in the art can understand that Figure 4 the structure shown in FIG. 5 is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0095] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present application.

[0096] In the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

[0097] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc.

[0098] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. These processes can be completed by relevant hardware instructed by a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage media include: ROM or random access memory RAM, magnetic disks, or optical disks and other media that can store program codes.

Claims

1. A high-speed traffic acquisition method based on a fusion large model, characterized in that: include: Receive user questions and input them into the prompt word model and plug-in model for parallel processing; The prompt word model calls the traffic query function based on the intention parameter to obtain the first traffic data; the intention parameter is a business type parameter, a time parameter and a location parameter; The plug-in model arranges the calling order of the plug-ins according to the user question; The plug-in model executes the call of the plug-in and processes the result transfer between plug-ins, and obtains the second flow data of the plug-in call; When the plug-in model obtains the second traffic data, a final answer is generated based on the second traffic data; when the plug-in model does not obtain the second traffic data, a final answer is generated based on the first traffic data.

2. The method according to claim 1, characterized in that The prompt word model calls the traffic query function based on the intention parameter to obtain the first traffic data, specifically including: Determine whether the user problem is a common traffic problem; If the user question is a common traffic problem, the prompt word model extracts the first intention parameter in the user question; The prompt word model calls a traffic query function based on the first intention parameter to obtain the first traffic data; If the user problem is not a common traffic problem, the prompt word model calculates the similarity between the user problem and the common traffic problem respectively; The prompt word model selects the common traffic problem with the highest similarity as the target common traffic problem, and the prompt word model extracts the second intention parameter in the target common traffic problem; The prompt word model calls a traffic query function based on the second intent parameter to obtain the first traffic data.

3. The method according to claim 2, characterized in that Before the step of the prompt word model calling a traffic query function based on the first intention parameter to obtain the first traffic data, the method further includes: Sort business types according to high-speed traffic; Based on the sorted high-speed traffic, a traffic query method library is established for the common traffic problems, and the traffic query method library performs a quick query on the traffic data based on the time parameter, the location parameter and the business type parameter.

4. A method for training a plug-in model, characterized in that: A method as claimed in any one of claims 1 to 3, comprising: Constructing a fine-tuning data set based on the artificial business process language; adding plug-in calling tool information and corresponding plug-in calling result information to the constructed fine-tuning data set; The plug-in model is obtained by training the basic large model according to the fine-tuning data set, and the training parameters include: batch size parameter, learning rate parameter, training step parameter, rank parameter and optimizer parameter; The system prompt words of the plug-in model are configured, wherein the system prompt words include identity configuration information, scenario configuration information, business type processing method and sample information.

5. The method according to claim 4, characterized in that The step of training the basic large model according to the fine-tuning data set to obtain the plug-in model, wherein the training parameters include: batch size parameter, learning rate parameter, training step parameter, rank parameter and optimizer parameter, specifically includes: Set fine-tuning training parameters, which include pre-training model path, fine-tuning dataset path, fine-tuning stage identifier, fine-tuning template configuration, target layer configuration, model save path, batch processing parameters, sequence maximum length parameters, gradient accumulation step parameters, learning rate scheduler parameters, logging step parameters, weight save interval step parameters, training round number parameters, maximum sample number parameters, and quantization bit width parameters; Obtaining a weight matrix of the basic large model, and generating a first low-rank matrix and a second low-rank matrix by using a low-rank decomposition method; updating the first low-rank matrix and the second low-rank matrix based on the fine-tuning dataset; Updating the weight matrix based on the product of the first low-rank matrix and the second low-rank matrix; Get the training set loss value and the validation set loss value during the training process; When the loss value of the training set shows a downward trend and the loss value of the validation set tends to be stable, the training process is terminated.

6. The method according to claim 5, characterized in that The calculation formula for updating the weight matrix based on the product of the first low-rank matrix and the second low-rank matrix is: 𝑊′=𝑊+𝑈⋅𝑉 𝑇 𝑊∈𝑅 𝑚×𝑛 𝐸∈𝑅 𝑚×𝑘 𝑘≪𝑚𝑖𝑛(𝑚,𝑛) 𝐹∈𝑅 𝑘×𝑛 𝑘≪𝑚𝑖𝑛(𝑚,𝑛) In the formula, 𝑊′ is the updated weight matrix, 𝑊 is the weight matrix, 𝑈 is the transposed form of 𝐸, 𝑉 is the transposed form of 𝐹, 𝐸 is the first low-rank matrix, 𝐹 is the second low-rank matrix, and 𝑘 is the rank number.

7. The method according to claim 4, characterized in that The fine-tuning dataset is constructed based on the artificial business process language; The step of adding the plug-in calling tool information and the corresponding plug-in calling result information to the fine-tuning data set specifically includes: Establishing a flow service processing tool set, wherein the flow service processing tool set includes a flow query tool; Acquire an artificial dialogue corpus, wherein the artificial dialogue corpus includes question information and corresponding answer information; The artificial dialogue corpus is processed to obtain a dialogue training sample, wherein: a problem in the problem information that requires calling the traffic service processing tool set is identified; and a tool calling result corresponding to the problem is used as an independent dialogue round; The fine-tuning dataset is generated based on the dialogue training sample.

8. A high-speed traffic acquisition system based on a fusion large model, characterized in that: The high-speed traffic acquisition system based on the fused large model includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the high-speed traffic acquisition system based on the fused large model to execute the method described in any one of claims 1-7.

9. A computer program product comprising instructions, characterized in that When the computer program product runs on a high-speed traffic acquisition system based on a fused large model, the high-speed traffic acquisition system based on a fused large model executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on a high-speed traffic acquisition system based on a fused large model, the high-speed traffic acquisition system based on a fused large model executes the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Intelligent question answering system, method and device for offline environment

    CN117194647A

  • Service plug-in calling method based on large language model

    CN118193174A

  • Charging scene intention acquisition and business processing method and system based on large model

    CN118708683A

  • Data processing method and device, electronic equipment and readable storage medium

    CN118820408A

  • Network-based communication session copilot

    WO2024182148A1

Cited By

  • Large model business usage monitoring method and device based on eBPF, and medium

    CN121119160A