Traffic comprehensive query optimization method and system applied to intelligent human-computer interaction

Through multi-modal feature analysis and dynamic interactive feature fusion, multi-dimensional intent features are generated, which solves the problem of real-time environment and equipment adaptation in intelligent traffic query, and realizes efficient and reliable traffic query services.

CN120336372AActive Publication Date: 2025-07-18GUIZHOU JIAOTOU HIGH TECH CO LTD

Patent Information

Application Number
CN202510833178.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing intelligent traffic query method cannot be understood in combination with real-time environmental parameters for dynamic intention, resulting in deviations from feedback strategies and user needs and physical scenarios. The path optimization algorithm lacks adaptability considerations for the execution capabilities of heterogeneous terminal devices, resulting in the scheme being easily failed when executed on the device.

Method used

By obtaining the user's traffic query request, multi-modal feature analysis generates semantic understanding features and scene-related features, using preset language models to fusion dynamic interactive features, generate multi-dimensional intention features, and combine real-time environmental information to optimize service paths to generate comprehensive response strategies to ensure the dynamic adaptation of feedback information and device parameters.

Benefits of technology

It improves the real-time response, rational decision-making and reliability of the intelligent traffic query service, ensuring the natural interactive experience of feedback information and good equipment compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336372A_ABST
    Figure CN120336372A_ABST
Patent Text Reader

Abstract

The invention provides a traffic comprehensive query optimization method and system applied to intelligent human-computer interaction, and the method comprises the steps: obtaining a traffic query request input by a target user in an interaction scene, carrying out the multi-modal feature analysis of the traffic query request, generating a semantic understanding feature and a scene correlation feature, and obtaining a semantic understanding feature and a scene correlation feature based on a preset language large model; performing dynamic interaction feature fusion processing on the semantic understanding feature and the scene association feature to generate a multi-dimensional intention feature, and performing service path optimization on the traffic query request of the target user according to the multi-dimensional intention feature to generate a comprehensive response strategy; and converting the comprehensive response strategy into interactive feedback information, dynamically adapting to equipment parameters in the real-time environment association information, and executing feedback operation to complete optimized response of the traffic query request. According to the invention, the response real-time performance, decision reasonability and execution reliability of the intelligent traffic query service can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and more particularly, to a traffic comprehensive query optimization method and system applied to intelligent human-computer interaction. Background Art

[0002] With the deep integration of intelligent transportation systems and human-computer interaction technologies, current mainstream intelligent transportation query methods usually generate traffic route suggestions by parsing natural language texts input by users, relying on pre-set databases to provide fixed-mode feedback information, such as route planning based on keyword matching or real-time traffic condition announcements in a single dimension. However, such methods have significant limitations: on the one hand, traditional natural language processing technologies can only extract explicit user requirements and cannot perform dynamic intention understanding by combining real-time environmental parameters (such as device response latency, terminal compatibility), resulting in a deviation between the feedback strategy and the user's true needs and physical scenarios; on the other hand, existing route optimization algorithms mostly focus on calculating theoretically optimal routes and lack consideration of the adaptability to the execution capabilities of heterogeneous terminal devices, making the generated route plans prone to failure due to protocol mismatch or resource overrun when executed on terminal devices. Summary of the Invention

[0003] The present invention provides a traffic comprehensive query optimization method and system applied to intelligent human-computer interaction.

[0004] In a first aspect, an embodiment of the present invention provides a traffic comprehensive query optimization method applied to intelligent human-computer interaction, including: Obtaining a traffic query request input by a target user in an interaction scenario, where the traffic query request includes a natural language text and real-time environment association information; Performing multi-modal feature parsing on the traffic query request to generate semantic understanding features and scenario association features, where the scenario association features are used to indicate the real-time traffic environment state related to the traffic query request; Based on a pre-set language large model, performing dynamic interaction feature fusion processing on the semantic understanding features and the scenario association features to generate multi-dimensional intention features; According to the multi-dimensional intention features, optimizing the service path of the traffic query request of the target user to generate a comprehensive response strategy; Converting the comprehensive response strategy into interactive feedback information, and dynamically adapting the device parameters in the real-time environment association information, and performing a feedback operation to complete the optimized response to the traffic query request.

[0005] In a second aspect, an embodiment of the present invention provides a human-computer interaction system, including: A memory in which a computer program is stored; A processor for loading the computer program to implement the traffic comprehensive query optimization method applied to intelligent human-computer interaction as described above.

[0006] The traffic comprehensive query optimization method applied to intelligent human-computer interaction provided by the present invention generates a composite data expression with both semantic understanding features and scene association features through multi-modal feature parsing that integrates user natural language text and real-time environment association information, effectively overcoming the limitations of scene perception caused by single-modal data processing in traditional traffic query systems; based on a language large model, it dynamically interacts and fuses multi-source features to generate multi-dimensional intention features, breaking through the one-way parsing limitation of explicit requirements by traditional intention recognition models, realizing multi-level coupling analysis of user needs and real-time traffic environment status, and significantly improving the integrity of intention understanding in complex scenarios; by deeply combining multi-dimensional intention features with a service path optimization algorithm to generate a comprehensive response strategy, it not only meets the core query needs of users, but also incorporates device parameter constraints and resource dynamic distribution as optimization conditions into the decision-making process, effectively solving the problem of infeasible solutions caused by traditional path planning algorithms ignoring the physical device status; finally, through the time synchronization and device adaptation conversion mechanism of multi-channel feedback data, it realizes end-to-end closed-loop optimization from strategy generation to terminal execution, ensuring both the natural interaction experience of feedback information and the precise compatibility of control instructions with heterogeneous devices, comprehensively improving the response real-time performance, decision-making rationality, and execution reliability of intelligent traffic query services. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0008] Figure 1 It is a flowchart of a traffic comprehensive query optimization method applied to intelligent human-computer interaction provided by an embodiment of the present invention.

[0009] Figure 2 It is a schematic diagram of the composition of a human-computer interaction system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0011] Please refer to Figure 1 , Figure 1 which is a flowchart of a traffic comprehensive query optimization method applied to intelligent human-computer interaction provided by an embodiment of the present invention. The traffic comprehensive query optimization method applied to intelligent human-computer interaction can be executed by a human-computer interaction system, and the traffic comprehensive query optimization method applied to intelligent human-computer interaction may include the following steps: Step S100: Obtain a traffic query request input by a target user in an interaction scenario, where the traffic query request includes natural language text and real-time environment association information.

[0012] In the traffic query scenario of intelligent human-computer interaction, the target user is an individual who initiates a traffic query operation, and the interaction scenario covers various environments where users interact with the system, such as in an intelligent vehicle system, a traffic query application on a mobile phone, etc. The traffic query request is the content input by the user to obtain traffic-related information, which consists of two parts: natural language text and real-time environment association information. The natural language text is the query content expressed by the user in daily language, such as "How to get from home to the company" "Where is the nearest subway station" etc. The real-time environment association information is data related to the current real-time environment, such as the geographical location where the user is currently located, the current time, the real-time traffic flow status, etc.

[0013] As an implementation manner, step S100, obtaining a traffic query request input by a target user in an interaction scenario, may specifically include the following steps S110~S150: Step S110: Receive the original query information submitted by the target user through voice input, text input or touch input.

[0014] In different interaction scenarios, the target user can submit the original query information in various ways. Voice input is the way for the user to convey the query content to the system by speaking. For example, in an intelligent vehicle system, the user can say "Query the route to the airport". The system will use speech recognition technology to convert the speech signal into text information. The speech recognition technology can be based on deep learning models, such as recurrent neural network (RNN) and its variants long short-term memory network (LSTM), gated recurrent unit (GRU), etc. These models are trained with a large amount of speech data to learn the mapping relationship between speech features and corresponding texts. In practical applications, the system will perform preprocessing operations such as sampling and feature extraction on the user input speech, and then input the processed features into the trained speech recognition model to obtain the corresponding text information.

[0015] The text input is that the user directly enters text in the input box to express the query content. For example, in the traffic query application on the mobile phone, the user manually enters "Find nearby bus stops". The touch input is that the user performs query operations by touching specific areas or icons on the screen. For example, on the intelligent traffic query tablet, the user clicks on the "Peripheral parking lot query" icon. The system will receive and perform preliminary format processing on the original query information input in these different ways to unify the subsequent processing process.

[0016] Step S120: Perform noise filtering processing on the original query information to remove irrelevant characters or invalid audio segments.

[0017] The original query information may contain some irrelevant characters or invalid audio segments, which will affect the subsequent processing effect. Irrelevant characters may be punctuation marks, garbled characters, etc. mis-entered by the user, and invalid audio segments may be background noise, the user's cough, etc. The purpose of the noise filtering processing is to remove these interfering information and improve the quality of the query information.

[0018] For the original query information in text form, noise filtering can be achieved by means of regular expression matching. Regular expression is a tool for describing string patterns. By defining pattern rules, irrelevant characters can be matched and removed. For example, define a regular expression pattern to match all characters that are not letters, numbers, and common punctuation marks, and then use the regular expression library function in the programming language to replace these characters with an empty string in the original query information.

[0019] For the original query information in audio form, noise filtering can adopt audio processing algorithms such as spectral subtraction, Wiener filtering, etc. Spectral subtraction is a method for noise estimation and removal based on the frequency domain. It estimates the spectral characteristics of the noise, and then subtracts the noise spectrum from the spectrum of the original audio to obtain the audio signal after removing the noise. Wiener filtering is an optimal linear filtering method. It filters the audio signal according to the statistical characteristics of the noise and the prior knowledge of the signal to achieve the purpose of removing the noise.

[0020] Step S130: Perform intent pre-recognition on the filtered original query information to determine whether it contains key semantic units related to traffic.

[0021] Intent pre-recognition is to perform semantic analysis on the filtered original query information to determine whether it contains key semantic units related to traffic. Key semantic units refer to words or phrases that can clearly express the traffic query intent, such as "route", "bus", "subway", "parking lot", etc.

[0022] Intent pre-recognition can adopt a method that combines keyword matching and semantic understanding models in natural language processing technology. Keyword matching is to pre-define a keyword list related to traffic, and then check whether these keywords exist in the filtered original query information. If they exist, it is considered that the key semantic units related to traffic may be included. The semantic understanding model can use deep learning models such as convolutional neural network (CNN), bidirectional long short-term memory network (Bi-LSTM), etc. These models can learn the semantic features of the text, and judge whether it is related to the traffic query intent by encoding and classifying the input query information. In practical applications, keyword matching is first used for preliminary screening, and then the screened query information is input into the trained semantic understanding model for further judgment.

[0023] Step S140: If the key semantic units are included, mark the original query information as a valid traffic query request and trigger the subsequent processing flow.

[0024] When it is determined through intent pre-recognition that the filtered original query information contains key semantic units related to traffic, the original query information is marked as a valid traffic query request. The purpose of marking is to distinguish between valid and invalid query information, so as to perform targeted processing on the valid query information subsequently. Triggering the subsequent processing flow means that the system will start to execute step S200 and subsequent steps to further analyze and process the valid traffic query request to provide an accurate traffic query response.

[0025] Step S150: If the key semantic units are not included, send guiding information to the target user to re-enter or supplement the query content.

[0026] If it is found during the intent pre-recognition process that the filtered original query information does not contain key semantic units related to traffic, it indicates that the query information may be irrelevant to the traffic query or is not clearly expressed. At this time, the system will send guiding information to the target user. The guiding information can be a text prompt, such as "Your query seems to be irrelevant to traffic. Please enter content related to routes, buses, subways, etc."; it can also be a voice prompt, which is converted into voice through text-to-speech technology and played to the user. The text-to-speech technology is usually based on neural network models such as the Tacotron series models, which can learn the mapping relationship between text and speech and convert the input text into a natural and fluent voice signal. By sending guiding information, the target user is guided to re-enter or supplement the query content so that they can accurately express the traffic query intent.

[0027] Step S200: Perform multi-modal feature parsing on the traffic query request to generate semantic understanding features and scene association features. The scene association features are used to indicate the real-time traffic environment state related to the traffic query request.

[0028] Multi-modal feature analysis refers to comprehensively analyzing the natural language text and real-time environment associated information in a traffic query request to extract key features therefrom. The semantic understanding feature is a feature used to represent the core semantic intention of a traffic query request, which can reflect the user's true query needs. The scene association feature is a feature related to the real-time traffic environment state, used to indicate the current traffic conditions, such as the real-time traffic flow distribution, device response delay, the spatial distance between the user's location and the target traffic node, etc.

[0029] When performing multi-modal feature analysis, it is necessary to combine natural language processing technology and data analysis technology. For natural language text, it is necessary to analyze its grammatical structure and semantic meaning and extract key information; for real-time environment associated information, it is necessary to process and analyze various environmental parameters to generate features related to the traffic environment. The semantic understanding features and scene association features generated through multi-modal feature analysis provide an important basis for subsequent feature fusion and intention analysis.

[0030] As an implementation manner, in step S200, perform multi-modal feature analysis on the traffic query request to generate semantic understanding features and scene association features, which may specifically include the following steps S210 to S240: Step S210: Perform word segmentation on the natural language text to obtain multiple semantic units, and extract environmental parameters from the real-time environment associated information to obtain multiple environmental state parameters.

[0031] Word segmentation divides the natural language text into independent semantic units. For example, for the text "How to get from home to the company", after word segmentation, semantic units such as "from", "home", "to", "company", "how", "get" can be obtained. Multiple methods can be used for word segmentation, such as rule-based word segmentation methods, statistic-based word segmentation methods, and deep learning-based word segmentation methods. The rule-based word segmentation method divides the text according to pre-defined word segmentation rules, such as the forward maximum matching method, the backward maximum matching method, etc. The statistic-based word segmentation method performs word segmentation by statistically analyzing a large amount of text data and learning the probability of word occurrence and context relationship. The deep learning-based word segmentation method uses neural network models, such as recurrent neural networks (RNN), convolutional neural networks (CNN), etc., to learn the semantic features of the text and automatically perform word segmentation.

[0032] The real-time environment correlation information includes various data related to the traffic environment, and the environmental parameter extraction is to extract useful environmental state parameters from this data. For example, parameters such as traffic flow and vehicle speed on different road sections are extracted from the real-time traffic flow data; the response delay parameter of the device is extracted from the device information; the current location of the user and the location of the target traffic node are extracted from the geographical location information. The environmental parameter extraction can be achieved through data mining and data analysis techniques. For example, methods such as data filtering and feature selection are used to screen out the environmental state parameters related to the traffic query request from a large amount of real-time environment correlation information.

[0033] Step S220: Perform context semantic encoding on multiple semantic units to generate semantic understanding features; among them, the semantic understanding features are used to represent the core semantic intention of the traffic query request.

[0034] The context semantic encoding is to encode the multiple semantic units obtained by word segmentation, considering their context relationship, to generate features that can represent the core semantic intention of the traffic query request. The context semantic encoding can use deep learning models, such as the BERT (Bidirectional Encoder Representations from Transformers) model. The BERT model is a pre-trained language model based on the Transformer architecture, which can learn the context information of the text through a bidirectional attention mechanism. When performing context semantic encoding, first input the multiple semantic units into the BERT model, and the model will encode each semantic unit to generate the corresponding word vector representation. Then, through further processing of these word vectors, such as pooling operations (average pooling, max pooling, etc.), a comprehensive semantic understanding feature vector is obtained. This feature vector can reflect the core semantic intention of the traffic query request, such as whether the user is querying a route, looking for a bus stop, or understanding the traffic conditions, etc.

[0035] Step S230: Perform spatio-temporal correlation analysis on multiple environmental state parameters to generate an environmental state encoding vector.

[0036] Spatio-temporal correlation analysis is to comprehensively analyze multiple environmental state parameters and consider their correlation relationships in time and space. In the traffic query scenario, the time distribution and spatial topological relationship of environmental state parameters can help understand the real-time traffic environmental state. For example, the traffic flow distribution in different time periods may vary greatly, and the traffic conditions in different geographical locations will also affect each other. Spatio-temporal correlation analysis can be achieved by constructing a spatio-temporal encoding network. The spatio-temporal encoding network includes a time convolutional layer and a spatial graph attention layer. The time convolutional layer is used to extract the time series dependence relationship of environmental state parameters. It can perform convolutional operations on environmental state parameters in the time dimension to capture the variation law of parameters over time. The spatial graph attention layer is used to extract the correlation weights between geographical locations. It can represent geographical location information as a graph structure and learn the correlation degree between different geographical locations through the attention mechanism.

[0037] When performing spatio-temporal correlation analysis, first extract timestamp data and geographical location data from the real-time environmental correlation information, then determine the time distribution characteristics of environmental state parameters according to the timestamp data, and determine the spatial topological relationship of environmental state parameters according to the geographical location data. Then, input the time distribution characteristics and spatial topological relationship into the spatio-temporal encoding network to generate spatio-temporal correlation features. Finally, perform dimensionality reduction processing on the spatio-temporal correlation features to obtain the environmental state encoding vector. Dimensionality reduction processing can use methods such as principal component analysis (PCA) to reduce the dimension of features while retaining the main information.

[0038] As an implementation manner, in step S230, perform spatio-temporal correlation analysis on multiple environmental state parameters to generate an environmental state encoding vector, which may specifically include the following steps S231 to S235: Step S231: Extract timestamp data and geographical location data from the real-time environmental correlation information.

[0039] The timestamp data is information recording the acquisition time of environmental state parameters, which can be accurate to the specific date and time. The geographical location data is information representing the acquisition location of environmental state parameters, usually represented by longitude and latitude coordinates. Extracting timestamp data and geographical location data from the real-time environmental correlation information can be achieved through data parsing and extraction methods. For example, in real-time traffic flow data, each data record may contain a timestamp field and a geographical location field. By parsing these fields, the timestamp data and geographical location data can be extracted.

[0040] Step S232: Determine the time distribution characteristics of environmental state parameters according to the timestamp data.

[0041] The time distribution characteristics reflect the variation law of environmental state parameters over time. Determining the time distribution characteristics of environmental state parameters based on timestamp data can be achieved through statistical analysis and time series analysis methods. For example, statistical quantities such as the average value, maximum value, and minimum value of environmental state parameters within different time periods can be calculated to understand the overall situation of the parameters in different time periods. Time series analysis methods such as the ARIMA (Autoregressive Integrated Moving Average) model and seasonal decomposition methods can also be used to model and analyze the time series of environmental state parameters and predict the future change trend of the parameters. In practical applications, first, the timestamp data is sorted in chronological order, and then the environmental state parameters are grouped by time to calculate the statistical quantities within each time period or use time series analysis methods for modeling.

[0042] Step S233: Determine the spatial topological relationship of environmental state parameters based on geographical location data.

[0043] The spatial topological relationship describes the relative positions and connection relationships between different geographical locations. Determining the spatial topological relationship of environmental state parameters based on geographical location data can be achieved through Geographic Information System (GIS) technology and graph theory methods. First, the geographical location data is converted into points in the geographical space, and then a graph structure is constructed based on the distances and connection relationships between these points. In the graph structure, each node represents a geographical location, and the edge represents the connection relationship between nodes. Through graph theory algorithms such as the shortest path algorithm and centrality analysis algorithm, the distribution and propagation laws of environmental state parameters in space can be analyzed. For example, the Dijkstra algorithm is used to calculate the shortest path between different geographical locations to understand the propagation situation of traffic flow on different road sections. In practical applications, GIS software or open-source geographical information processing libraries such as GeoPandas and NetworkX can be used to process and analyze the geographical location data to determine the spatial topological relationship of environmental state parameters.

[0044] Step S234: Input the time distribution characteristics and spatial topological relationship into the spatio-temporal encoding network to generate spatio-temporal correlation characteristics.

[0045] The spatio-temporal encoding network is a neural network model for processing spatio-temporal data, which can fuse time and space information to generate spatio-temporal correlation characteristics. When inputting the time distribution characteristics and spatial topological relationship into the spatio-temporal encoding network, the input data is first preprocessed to meet the input requirements of the network. For example, the time distribution characteristics and spatial topological relationship are converted into a suitable tensor format and normalized to improve the training effect of the network.

[0046] The temporal convolutional layer of the spatio-temporal encoding network performs a convolutional operation on the temporal distribution features to extract the temporal sequence dependencies. The temporal convolutional layer can use a one-dimensional convolutional kernel that slides along the temporal dimension to extract features from the temporal distribution features. The spatial graph attention layer processes the spatial topological relationships and learns the association weights between different geographical locations through the attention mechanism. The spatial graph attention layer can represent the geographical location information as a graph structure and use the graph attention mechanism to update the features of the nodes in the graph. Finally, the outputs of the temporal convolutional layer and the spatial graph attention layer are fused to generate spatio-temporal association features.

[0047] Step S235: Perform dimensionality reduction on the spatio-temporal association features to obtain an environmental state encoding vector; wherein, the spatio-temporal encoding network includes a temporal convolutional layer and a spatial graph attention layer, the temporal convolutional layer is used to extract temporal sequence dependencies, and the spatial graph attention layer is used to extract the association weights between geographical locations.

[0048] The dimensionality reduction process is to reduce the dimension of the spatio-temporal association features while retaining the main information, improving the efficiency and accuracy of subsequent processing. The principal component analysis (PCA) method can be used to perform dimensionality reduction on the spatio-temporal association features. The PCA method projects the high-dimensional data into a low-dimensional space by finding the principal components of the data. The specific steps are as follows: First, centralize the spatio-temporal association features so that their mean is zero. Then, calculate the covariance matrix of the spatio-temporal association features. Next, solve the eigenvalues and eigenvectors of the covariance matrix, and select the first k eigenvectors with the largest eigenvalues as the principal components. Finally, project the spatio-temporal association features onto these k principal components to obtain the dimensionality-reduced environmental state encoding vector. Through dimensionality reduction, the high-dimensional spatio-temporal association features are converted into low-dimensional environmental state encoding vectors, facilitating subsequent feature fusion and intent analysis.

[0049] Step S240: Construct scene association features based on the environmental state encoding vector and the device type parameter in the real-time environmental association information.

[0050] The scene association features are used to indicate the real-time traffic environmental state related to the traffic query request, and it includes real-time traffic flow distribution parameters, device response delay parameters, spatial distance parameters between the user's location and the target traffic node, etc. Constructing scene association features based on the environmental state encoding vector and the device type parameter in the real-time environmental association information requires comprehensive consideration of the environmental state and device characteristics.

[0051] First, the environmental state encoding vector already reflects the spatio-temporal correlation information of environmental state parameters, and it is used as part of the scene correlation feature. Then, in combination with the device type parameters in the real-time environmental correlation information, according to different device types, determine their impact on traffic queries. For example, different types of devices may have different response delay characteristics, and the corresponding device response delay parameters can be obtained through the device type parameters. The spatial distance parameter between the user location and the target traffic node can be calculated from the geographical location data in the real-time environmental correlation information. Finally, the environmental state encoding vector, the device response delay parameter, the spatial distance parameter between the user location and the target traffic node, etc. are combined and weighted to construct the scene correlation feature. In practical applications, the weights of each parameter can be adjusted according to different traffic query scenarios and requirements to obtain a more accurate scene correlation feature.

[0052] Step S300: Based on a preset large language model, perform dynamic interaction feature fusion processing on the semantic understanding feature and the scene correlation feature to generate multi-dimensional intention features.

[0053] The preset large language model is a pre-trained natural language processing model with strong semantic understanding and feature representation capabilities. The dynamic interaction feature fusion processing refers to fusing the semantic understanding feature and the scene correlation feature on the basis of considering their dynamic relationship to generate more comprehensive and accurate multi-dimensional intention features. The multi-dimensional intention features include user explicit demand features, implicit scene adaptation features, and real-time resource constraint features, which can more comprehensively reflect the user's traffic query intention.

[0054] When performing dynamic interaction feature fusion processing, effectively combine the semantic understanding feature and the scene correlation feature, and consider their interaction in different scenarios. The preset large language model can provide a powerful feature fusion framework, and through the learning ability of the model, automatically adjust the weights of the semantic understanding feature and the scene correlation feature to achieve dynamic interaction feature fusion.

[0055] As an implementation, step S300, based on a preset large language model, perform dynamic interaction feature fusion processing on the semantic understanding feature and the scene correlation feature to generate multi-dimensional intention features, which can specifically include the following steps S310~S360: Step S310: Map the semantic understanding feature to a first feature vector and map the scene correlation feature to a second feature vector.

[0056] Mapping is the process of converting semantic understanding features and scene association features into vector representations. Semantic understanding features are obtained by performing context semantic encoding on natural language texts, and they may be complex feature representations. Mapping it to the first feature vector can be achieved through linear transformation or non-linear transformation. For example, a fully connected layer can be used to linearly map the semantic understanding features to obtain a first feature vector with a fixed dimension.

[0057] Scene association features are constructed based on the environmental state encoding vector and device type parameters, and it is also necessary to map them to the second feature vector. The mapping process is similar to that of semantic understanding features, and a fully connected layer or other neural network layers can be used for mapping. By mapping semantic understanding features and scene association features to vector representations, it is convenient for subsequent feature fusion and calculation.

[0058] Step S320: Perform dynamic weight assignment processing on the first feature vector and the second feature vector to determine the semantic weight coefficient and the environmental weight coefficient.

[0059] Dynamic weight assignment processing is to dynamically adjust the weights of semantic understanding features and scene association features in feature fusion according to different traffic query scenarios and requirements. The semantic weight coefficient represents the importance of semantic understanding features in feature fusion, and the environmental weight coefficient represents the importance of scene association features in feature fusion.

[0060] When performing dynamic weight assignment processing, historical interaction data and the current query situation need to be considered. First, obtain the sample semantic feature vector set and the sample environmental feature vector set in the historical interaction dataset. Each sample semantic feature vector is associated with the influence weight value of the semantic feature in the actual decision-making, and each sample environmental feature vector is associated with the influence weight value of the environmental feature in the actual decision-making. Then, use these historical data to train a weight assignment model to learn the weight assignment rules of semantic features and environmental features in different situations.

[0061] As an implementation method, in step S320, performing dynamic weight assignment processing on the first feature vector and the second feature vector to determine the semantic weight coefficient and the environmental weight coefficient may specifically include the following steps S321~S3210: Step S321: Obtain the sample semantic feature vector set and the sample environmental feature vector set in the historical interaction dataset, where each sample semantic feature vector is associated with the influence weight value of the semantic feature in the actual decision-making, and each sample environmental feature vector is associated with the influence weight value of the environmental feature in the actual decision-making.

[0062] The historical interaction dataset is the data recorded by the system during past traffic query interactions. It contains a large number of sample semantic feature vectors and sample environmental feature vectors. The sample semantic feature vectors are the feature vectors obtained by semantic encoding of the natural language text in the historical queries, and the sample environmental feature vectors are the feature vectors obtained by processing the real-time environmental association information in the historical queries.

[0063] Each sample semantic feature vector and sample environmental feature vector are associated with the influence weight values of the semantic features and environmental features in the actual decision-making. These influence weight values are determined based on the actual results of the historical queries and user feedback, and reflect the importance of semantic features and environmental features to the final decision-making in different traffic query scenarios. For example, in some cases, the user's query intention is mainly determined by semantic information, and at this time the influence weight value of the semantic features is higher; while in other cases, the real-time traffic environment has a greater impact on the query results, and the influence weight value of the environmental features is higher. Obtaining the set of sample semantic feature vectors and the set of sample environmental feature vectors in the historical interaction dataset can be achieved through a data storage and management system. For example, using a database system to store the historical interaction data, and then extracting the required sample feature vectors and influence weight values through query statements.

[0064] Step S322: Input the sample semantic feature vectors and sample environmental feature vectors into the initial weight prediction model to generate corresponding semantic weight prediction sequences and environmental weight prediction sequences.

[0065] The initial weight prediction model is a model for predicting semantic weights and environmental weights. It can use a neural network model such as a multi-layer perceptron (MLP). A multi-layer perceptron is a feedforward neural network composed of an input layer, a hidden layer, and an output layer. In the input layer, the sample semantic feature vectors and sample environmental feature vectors are used as input data. The hidden layer can contain multiple neurons, and the input data is subjected to feature transformation through a non-linear activation function (such as the ReLU function). The output layer outputs the semantic weight prediction sequences and environmental weight prediction sequences.

[0066] When training the initial weight prediction model, it is necessary to normalize the sample semantic feature vectors and sample environmental feature vectors so that they have the same scale and range. Then, the normalized sample feature vectors are input into the initial weight prediction model, and the semantic weight prediction sequences and environmental weight prediction sequences are obtained through forward propagation calculation. These prediction sequences are the weight prediction results of the model for semantic features and environmental features under different samples.

[0067] Step S323: Calculate the error loss between each predicted value in the semantic weight prediction sequence and its associated actual influence weight value to generate a semantic weight error set.

[0068] Error loss calculation is to measure the difference between each predicted value in the semantic weight prediction sequence and the actual influence weight value. The mean squared error (MSE) can be used as the error loss function, and the calculation formula is: , where is the actual influence weight value, is the predicted value in the semantic weight prediction sequence, and n is the number of samples.

[0069] In practical applications, each predicted value in the semantic weight prediction sequence is traversed, compared with the corresponding actual influence weight value, and the error loss is calculated. The error loss values of all samples are collected to generate a semantic weight error set. The semantic weight error set reflects the accuracy of the initial weight prediction model in predicting semantic weights. The smaller the error value, the more accurate the model's prediction.

[0070] Step S324: Calculate the error loss for each predicted value in the environmental weight prediction sequence and its associated actual influence weight value to generate an environmental weight error set.

[0071] Similar to the semantic weight error calculation, calculate the error loss for each predicted value in the environmental weight prediction sequence and its associated actual influence weight value. The mean squared error (MSE) can also be used as the error loss function. Traverse each predicted value in the environmental weight prediction sequence, compare it with the corresponding actual influence weight value, and calculate the error loss. The error loss values of all samples are collected to generate an environmental weight error set. The environmental weight error set reflects the accuracy of the initial weight prediction model in predicting environmental weights. The smaller the error value, the more accurate the model's prediction.

[0072] Step S325: Weightedly fuse the semantic weight error set and the environmental weight error set to generate a total error loss value, and iteratively adjust the parameters of the initial weight prediction model according to the total error loss value until the total error loss value converges to a stable interval, generating a trained weight assignment model.

[0073] Weighted fusion comprehensively considers the semantic weight error set and the environmental weight error set, and generates a total error loss value by weighted summation. A weight coefficient can be set for the semantic weight error set and the environmental weight error set respectively, and then their error loss values are multiplied by the corresponding weight coefficients and added together to obtain the total error loss value. For example, let the weight coefficient of the semantic weight error set be , and the weight coefficient of the environmental weight error set be , the total error loss value , where is the error value in the semantic weight error set, is the number of samples in the semantic weight error set, is the error value in the environmental weight error set, is the sample number of the environmental weight error set.

[0074] The initial weight prediction model can be iteratively adjusted according to the total error loss value, and the gradient descent algorithm can be used. The gradient descent algorithm is an optimization algorithm that calculates the gradient of the total error loss value with respect to the model parameters and then updates the model parameters in the opposite direction of the gradient, making the total error loss value gradually decrease. In each iteration, the gradient of the total error loss value is calculated, and then the model parameters are updated according to the learning rate. This process is repeated until the total error loss value converges to a stable interval, and at this time, the trained weight assignment model is generated.

[0075] Step S326: Input the first feature vector and the second feature vector into the trained weight assignment model, and output the initial semantic weight value and the initial environmental weight value before normalization.

[0076] Input the first feature vector and the second feature vector into the trained weight assignment model. The model will process the first feature vector and the second feature vector according to the learned weight assignment rule and output the initial semantic weight value and the initial environmental weight value before normalization. The trained weight assignment model has learned the weight assignment rules of semantic features and environmental features in historical interaction data and can predict their weights in feature fusion based on the current first feature vector and second feature vector.

[0077] Step S327: Perform dynamic balance constraint processing on the initial semantic weight value and the initial environmental weight value so that the sum of the two satisfies a preset constant relationship, and generate an intermediate semantic weight value and an intermediate environmental weight value.

[0078] The dynamic balance constraint processing is to ensure the balance relationship between the semantic weight value and the environmental weight value. The preset constant relationship can be set according to specific traffic query scenarios and requirements. For example, set the sum of the two to 1. The dynamic balance constraint processing of the initial semantic weight value and the initial environmental weight value can be achieved by normalization. For example, let the initial semantic weight value be , and the initial environmental weight value be , then the intermediate semantic weight value , and the intermediate environmental weight value . Through the dynamic balance constraint processing, the sum of the intermediate semantic weight value and the intermediate environmental weight value satisfies the preset constant relationship, ensuring the balance of the relative importance of semantic features and environmental features in feature fusion.

[0079] Step S328: Perform non-linear activation transformation processing on the intermediate semantic weight value to generate a smooth semantic weight coefficient after eliminating the sudden fluctuations of the weight value.

[0080] Nonlinear activation transformation processing can use functions such as the Sigmoid function or the Tanh function. Taking the Sigmoid function as an example, when the intermediate semantic weight value is input into the Sigmoid function, the function will perform a nonlinear transformation on the intermediate semantic weight value and map it to the range of 0 to 1. Through nonlinear activation transformation processing, sudden fluctuations in the weight value can be eliminated, making the weight value smoother. For example, when there are large fluctuations in the intermediate semantic weight value, the Sigmoid function will compress it into a relatively stable range to generate a smooth semantic weight coefficient.

[0081] Step S329: Perform nonlinear activation transformation processing on the intermediate environmental weight value to eliminate sudden fluctuations in the weight value and generate a smooth environmental weight coefficient.

[0082] Similar to the processing of the intermediate semantic weight value, for the intermediate environmental weight value, nonlinear activation transformation processing can also be performed, and functions such as the Sigmoid function or the Tanh function can be used. When the intermediate environmental weight value is input into the nonlinear activation function, the function will perform a nonlinear transformation on it, eliminate sudden fluctuations in the weight value, and generate a smooth environmental weight coefficient. Through the smoothing process, the environmental weight coefficient becomes more stable, avoiding the impact of drastic changes in the weight value on the feature fusion result.

[0083] Step S3210: Perform joint verification processing on the smooth semantic weight coefficient and the smooth environmental weight coefficient to ensure that their interaction contribution degrees to the first feature vector and the second feature vector match, and generate the finally available semantic weight coefficient and environmental weight coefficient.

[0084] The joint verification processing is to ensure the rationality and effectiveness of the smooth semantic weight coefficient and the smooth environmental weight coefficient in feature fusion. It can be verified by calculating the interaction contribution degrees of the smooth semantic weight coefficient and the smooth environmental weight coefficient with the first feature vector and the second feature vector. The interaction contribution degree can be measured by calculating the sum of the products of the feature vector and the weight coefficient. For example, calculate the sum of the products of the first feature vector and the smooth semantic weight coefficient, and the sum of the products of the second feature vector and the smooth environmental weight coefficient, and compare their magnitudes and proportional relationships. If their interaction contribution degrees do not match, the smooth semantic weight coefficient and the smooth environmental weight coefficient need to be adjusted. Through the joint verification processing, the finally available semantic weight coefficient and environmental weight coefficient are generated, ensuring the reasonable contributions of semantic features and environmental features in feature fusion.

[0085] Step S330: According to the semantic weight coefficient, perform weighted adjustment on the first feature vector to obtain the adjusted semantic feature vector.

[0086] The weighted adjustment is to perform an element-wise multiplication of the semantic weight coefficient and the first feature vector. Through the weighted adjustment, according to the magnitude of the semantic weight coefficient, the importance of each element in the first feature vector is adjusted. If the semantic weight coefficient is large, the elements in the first feature vector have a larger proportion in the adjusted semantic feature vector; conversely, if the semantic weight coefficient is small, the elements in the first feature vector have a smaller proportion in the adjusted semantic feature vector.

[0087] Step S340: According to the environmental weight coefficient, perform a weighted adjustment on the second feature vector to obtain an adjusted environmental feature vector.

[0088] Similar to the weighted adjustment of the first feature vector, according to the environmental weight coefficient, perform a weighted adjustment on the second feature vector. Through the weighted adjustment, according to the magnitude of the environmental weight coefficient, the importance of each element in the second feature vector is adjusted. If the environmental weight coefficient is large, the elements in the second feature vector have a larger proportion in the adjusted environmental feature vector; conversely, if the environmental weight coefficient is small, the elements in the second feature vector have a smaller proportion in the adjusted environmental feature vector.

[0089] Step S350: Concatenate the adjusted semantic feature vector and the adjusted environmental feature vector to generate an initial fusion feature.

[0090] Concatenation is to connect the adjusted semantic feature vector and the adjusted environmental feature vector in sequence to generate a longer vector. Through the concatenation operation, the semantic features and environmental features are fused to obtain an initial fusion feature containing more information. The initial fusion feature combines the information of semantic understanding features and scene association features, providing a basis for subsequent intention dimension expansion.

[0091] Step S360: Invoke the intention recognition layer in the language large model to perform intention dimension expansion on the initial fusion feature to generate multi-dimensional intention features; among them, the multi-dimensional intention features include user explicit demand features, implicit scene adaptation features, and real-time resource constraint features.

[0092] The intention recognition layer can analyze and process the input features to identify the intention information therein. Performing intention dimension expansion on the initial fusion feature is to extract more intention information from the initial fusion feature and generate more comprehensive multi-dimensional intention features.

[0093] The multi-dimensional intent features include user explicit demand features, implicit scenario adaptation features, and real-time resource constraint features. User explicit demand features are demand information directly extracted from the user's query text, such as the route query and bus stop search demands explicitly proposed by the user. Implicit scenario adaptation features are the adaptation information that can be inferred according to the real-time traffic environment and the user's query situation. For example, in the case of traffic congestion, a more suitable travel mode is recommended. Real-time resource constraint features are features that consider the current traffic resource limitations, such as the road capacity and bus transport capacity, to constrain and optimize the user's query.

[0094] In the embodiment of the present invention, the intent recognition layer can be specifically implemented to include a semantic analysis sub-layer, a scenario reasoning sub-layer, a resource evaluation sub-layer, and a feature integration sub-layer.

[0095] The main task of the semantic analysis sub-layer is to deeply analyze the semantic information in the initial fusion features and extract user explicit demand features. It can adopt the multi-head attention mechanism based on the Transformer architecture. The Transformer architecture has powerful parallel computing capabilities and long sequence processing capabilities, and can capture semantic associations at different positions in the text.

[0096] The multi-head attention mechanism allows the model to parallelly focus on different parts of the input sequence in different representation sub-spaces. Specifically, the input initial fusion features are first linearly projected into multiple low-dimensional sub-spaces to form multiple "heads". Each head independently calculates the attention scores, that is, determines the importance of each position by calculating the similarity between the query (Query), key (Key), and value (Value). Then, these attention scores are weighted and summed to obtain the output of each head. Finally, the outputs of all heads are concatenated and passed through a linear transformation to obtain the final semantic analysis result.

[0097] For example, for the user's query "How to take the subway from home to the mall at 10 am tomorrow", the semantic analysis sub-layer will determine the travel time by focusing on "10 am tomorrow" through the multi-head attention mechanism, determine the travel starting point and ending point by "from home to the mall", and determine the travel mode by "take the subway", so as to accurately extract the user's explicit demand features.

[0098] The scene inference sub-layer, based on the scene association information in the initial fusion features, combines real-time traffic data and predefined scene rules to perform dynamic scene adaptation analysis and generate implicit scene adaptation features. This sub-layer can use graph neural networks (GNNs) to process the spatial topology information in the traffic network. A graph neural network can represent the traffic network as a graph, where nodes represent traffic nodes (such as bus stops, subway stations, intersections, etc.), and edges represent the connection relationships between nodes (such as roads, lines, etc.). Through the message passing mechanism, nodes can receive and aggregate information from their neighboring nodes, thereby updating their own feature representations.

[0099] At the same time, the scene inference sub-layer also combines time series analysis methods and considers the changes in traffic conditions over time. For example, long short-term memory networks (LSTMs) or gated recurrent units (GRUs) are used to process the time series data of traffic flow and predict the traffic conditions in the next period of time.

[0100] In practical applications, the scene inference sub-layer infers the dynamic constraint conditions of the traffic resources for the user's current scene based on information such as the current traffic flow, weather conditions, and special events. For example, if it is the peak period on a weekday and the queried route passes through a bustling commercial area, the scene inference sub-layer may infer that the road section is congested, and thus generate implicit scene adaptation features that suggest the user choose other routes or travel modes.

[0101] The resource evaluation sub-layer is responsible for evaluating traffic resources and determining real-time resource constraint features. It combines a real-time traffic database and a machine learning model to perform a quantitative analysis of the availability and usage cost of traffic resources.

[0102] The real-time traffic database contains information such as road capacity, the transport capacity of buses and subways, and the number of vacant spaces in parking lots. The resource evaluation sub-layer queries these databases in real time to obtain the latest resource status. The resource evaluation sub-layer can use support vector regression (SVR) to predict resource consumption and demand. Support vector regression finds an optimal hyperplane to minimize the error between the predicted value and the actual value. In resource evaluation, it can predict the resource consumption of different travel plans, such as time cost, energy consumption, cost, etc., based on historical data and the current traffic conditions.

[0103] The feature integration sub-layer integrates the explicit user requirement features extracted by the semantic analysis sub-layer, the implicit scenario adaptation features generated by the scenario reasoning sub-layer, and the real-time resource constraint features determined by the resource evaluation sub-layer to generate the final multi-dimensional intention features. This sub-layer can use residual connections and layer normalization techniques to ensure the stability of features and the effective transmission of information. Residual connections allow the model to skip some layers during training and directly pass the input features to subsequent layers, avoiding the vanishing gradient problem and enabling the model to learn more complex feature representations. Layer normalization normalizes the input of each layer, making the features have a similar distribution in different dimensions, improving the training efficiency and generalization ability of the model.

[0104] During the feature integration process, the feature integration sub-layer will perform weighted fusion on different features, and the weights are determined based on the statistical analysis of historical data and the training results of the model. For example, for some users with high requirements for travel time, the weight of time in the explicit user requirement features will be relatively high; while for some users who are sensitive to costs, the weight of costs in the real-time resource constraint features will be greater.

[0105] Finally, the feature integration sub-layer outputs multi-dimensional intention features containing explicit requirements, implicit scenario adaptation, and real-time resource constraints, providing comprehensive and accurate information for subsequent service path optimization.

[0106] As an implementation, in step S360, the intention recognition layer in the language large model is called to expand the intention dimension of the initial fusion features to generate multi-dimensional intention features, which can specifically include the following steps S361~S366: Step S361: Decompose the initial fusion features into a semantics-dominated feature branch and an environment-dominated feature branch; among them, the semantics-dominated feature branch is composed of the features after the channel attention screening of the adjusted semantic feature vector, and the environment-dominated feature branch is composed of the features after the spatial attention screening of the adjusted environmental feature vector.

[0107] Channel attention screening processes the adjusted semantic feature vector and selects the feature channels that are more important for semantic understanding through the channel attention mechanism. The channel attention mechanism can use the channel attention module in the convolutional neural network, such as the Squeeze-and-Excitation (SE) module. The SE module compresses the adjusted semantic feature vector in the spatial dimension through global average pooling to obtain the global feature representation of each channel. Then, the global feature representation is processed through a fully connected layer and a non-linear activation function (such as the Sigmoid function) to obtain the attention weight of each channel. Finally, the attention weight is multiplied element-wise with the adjusted semantic feature vector to select the feature channels that are more important for semantic understanding, forming the semantics-dominated feature branch.

[0108] Spatial attention screening processes the adjusted environmental feature vectors and selects the feature spaces that are more important for environmental understanding through a spatial attention mechanism. The spatial attention mechanism can use the spatial attention module in a convolutional neural network, such as the SpatialAttentionModule (SAM). The SAM module processes the adjusted environmental feature vectors through convolutional operations to obtain a spatial attention map. Then, the spatial attention map is multiplied element-wise with the adjusted environmental feature vectors to select the feature spaces that are more important for environmental understanding, forming the environmental dominant feature branch.

[0109] Step S362: Conduct a context demand correlation analysis on the semantic dominant feature branch to extract the explicit user demand features; among them, the context demand correlation analysis is achieved by traversing the matching relationship between the semantic units in the semantic dominant feature branch and the predefined set of traffic service keywords.

[0110] The context demand correlation analysis analyzes the semantic units in the semantic dominant feature branch to find the semantic units that match the predefined set of traffic service keywords, thereby extracting the explicit user demand features. The predefined set of traffic service keywords is a predefined set of keywords related to traffic services, such as "route", "bus", "subway", "parking lot", etc.

[0111] When conducting the context demand correlation analysis, traverse each semantic unit in the semantic dominant feature branch and match it with the predefined set of traffic service keywords. If a semantic unit matches a keyword in the keyword set, it is considered that the semantic unit is related to the traffic service demand. By comprehensively analyzing all the matching semantic units, the explicit user demand features are extracted. For example, if the semantic dominant feature branch contains "the route from home to company", through matching with the predefined set of traffic service keywords, the explicit user demand feature of "route query" is extracted.

[0112] Step S363: Conduct a dynamic scenario adaptation analysis on the environmental dominant feature branch to generate implicit scenario adaptation features; among them, the dynamic scenario adaptation analysis includes: determining the dynamic constraint conditions of the target user's current scenario on traffic resources based on the real-time traffic state parameters in the environmental dominant feature branch.

[0113] The dynamic scenario adaptation analysis analyzes the scenario where the target user is located according to the real-time traffic state parameters in the environmental dominant feature branch, determines the dynamic constraint conditions of this scenario on traffic resources, and thereby generates implicit scenario adaptation features. The real-time traffic state parameters include real-time traffic flow, road congestion conditions, bus operation status, etc.

[0114] When performing dynamic scenario adaptation analysis, first, determine the scenario where the target user is located based on real-time traffic state parameters, such as whether it is the peak traffic period or the off-peak period, whether it is a congested section or a smooth section, etc. Then, according to different scenarios, determine the dynamic constraint conditions for traffic resources. For example, during the peak traffic period, the traffic capacity of the road may be restricted, and the transport capacity of the bus may be insufficient; in a congested section, it may be necessary to choose a detour travel mode. By analyzing these dynamic constraint conditions, implicit scenario adaptation features are generated. For example, if the real-time traffic state parameters show that a certain section is congested, the generated implicit scenario adaptation feature may be to recommend that the user choose other routes or travel modes.

[0115] Step S364: Input the user's explicit demand features and implicit scenario adaptation features into the cross-modal interaction network to generate potential demand compensation features; among them, the cross-modal interaction network realizes feature compensation by alternately performing projection mapping from semantic features to environmental features and feedback correction from environmental features to semantic features.

[0116] The cross-modal interaction network is a neural network model used to process the interaction between different modal features. Input the user's explicit demand features and implicit scenario adaptation features into the cross-modal interaction network, and through alternately performing projection mapping from semantic features to environmental features and feedback correction from environmental features to semantic features, feature compensation is realized to generate potential demand compensation features.

[0117] The projection mapping from semantic features to environmental features projects the user's explicit demand features into the environmental feature space, enabling the environmental features to better understand the user's needs. The feedback correction from environmental features to semantic features corrects the user's explicit demand features according to the environmental features, taking into account the impact of the real-time traffic environment on the user's needs. The cross-modal interaction network can be implemented using a bidirectional recurrent neural network (Bi-RNN) or a Transformer architecture. During the training process of the network, by alternately performing projection mapping and feedback correction operations, the parameters of the network are continuously optimized, so that the generated potential demand compensation features can better supplement the deficiencies of the user's explicit demand features and implicit scenario adaptation features.

[0118] Step S365: Perform feature orthogonalization processing on the user's explicit demand features, implicit scenario adaptation features, and potential demand compensation features, and generate an orthogonal demand feature set after eliminating redundant information.

[0119] Feature orthogonalization is to eliminate redundant information among the explicit demand features of users, implicit scenario adaptation features, and potential demand compensation features, making these features independent of each other. The Gram-Schmidt orthogonalization method can be used for feature orthogonalization. The Gram-Schmidt orthogonalization method is a method of converting a set of linearly independent vectors into a set of orthogonal vectors. The specific steps are as follows: First, regard the explicit demand features of users, implicit scenario adaptation features, and potential demand compensation features as a set of vectors. Then, select one of the vectors as the initial orthogonal vector. Next, orthogonalize the other vectors in turn by subtracting the projection of the vector on the already orthogonalized vectors, so that the newly generated vector is orthogonal to the already orthogonalized vectors. Repeat this process until all vectors are orthogonalized. Finally, an orthogonal demand feature set is obtained. Through feature orthogonalization, the redundant information between features is eliminated, and the independence and effectiveness of features are improved.

[0120] Step S366: Based on each feature vector in the orthogonal demand feature set, construct a multi-dimensional demand distribution matrix, and perform feature dimension expansion on the multi-dimensional demand distribution matrix to generate a multi-dimensional intention feature including explicit demand, implicit scenario adaptation, and real-time resource constraints; among them, the feature dimension expansion is realized by performing tensor splicing on the multi-dimensional demand distribution matrix and a pre-trained resource constraint embedding vector.

[0121] The multi-dimensional demand distribution matrix is a matrix composed of each feature vector in the orthogonal demand feature set, which can intuitively represent the multi-dimensional demand distribution of users. When constructing the multi-dimensional demand distribution matrix, each feature vector in the orthogonal demand feature set is used as a row or a column of the matrix to form a matrix.

[0122] Feature dimension expansion is to add real-time resource constraint information to the multi-dimensional demand distribution matrix to generate a multi-dimensional intention feature including explicit demand, implicit scenario adaptation, and real-time resource constraints. The pre-trained resource constraint embedding vector is a pre-trained vector, which represents the constraints of real-time traffic resources, such as the traffic capacity of roads, the transport capacity of buses, etc. By performing tensor splicing on the multi-dimensional demand distribution matrix and the pre-trained resource constraint embedding vector, the dimension of the matrix is expanded to generate a multi-dimensional intention feature containing more information. For example, the pre-trained resource constraint embedding vector is used as a column or a row of the matrix and spliced with the multi-dimensional demand distribution matrix to obtain the expanded multi-dimensional intention feature.

[0123] Step S400: According to the multi-dimensional intention feature, optimize the service path for the traffic query request of the target user to generate a comprehensive response strategy.

[0124] Service path optimization is to select the optimal transportation service path according to multi-dimensional intention features, considering the user's explicit needs, implicit scenario adaptation, and real-time resource constraints, so as to meet the user's transportation query needs. The comprehensive response strategy is a response plan that includes route guidance information and resource adaptation strategies generated based on the optimized service path and other needs of the user.

[0125] When performing service path optimization, multiple factors need to be considered, such as traffic flow, road conditions, travel time, resource consumption, etc. The multi-dimensional intention features provide rich information, including the user's needs and the real-time traffic environment, which provides a basis for service path optimization. By analyzing and processing the multi-dimensional intention features, the optimal service path is selected, and a comprehensive response strategy is generated based on this path.

[0126] As an implementation method, in step S400, according to the multi-dimensional intention features, the traffic query request of the target user is optimized for the service path, and a comprehensive response strategy is generated, which can specifically include the following steps S410~S460: Step S410: Split the service requirements of the multi-dimensional intention features to obtain the core service requirement features and auxiliary service requirement features.

[0127] Service requirement splitting is to decompose the multi-dimensional intention features and extract the core service requirement features and auxiliary service requirement features. The core service requirement features are the features directly related to the user's main traffic query needs, such as the starting point and ending point of the route queried by the user, the travel mode, etc. The auxiliary service requirement features are other auxiliary requirement features related to the core service requirements, such as the user's requirements for travel time, comfort, etc.

[0128] When performing service requirement splitting, it can be analyzed according to different dimensions and semantic information of the multi-dimensional intention features. For example, extract the features related to route planning from the multi-dimensional intention features as the core service requirement features, and extract the features related to travel time, comfort, etc. as the auxiliary service requirement features. Through service requirement splitting, the complex multi-dimensional intention features are decomposed into more easily processed core service requirement features and auxiliary service requirement features, providing a basis for subsequent service path optimization.

[0129] Step S420: Based on the real-time traffic database, obtain the initial service path set that matches the core service requirement features.

[0130] The real-time traffic database is a database that stores real-time traffic information, which includes road network information, traffic flow information, bus operation information, etc. Based on the real-time traffic database, according to the core service requirement features, search for the initial service path set that matches them.

[0131] For example, if the core service requirement feature is a route query from location A to location B, all possible routes from location A to location B are found through the road network information and traffic flow information in the real-time traffic database, forming an initial set of service paths. In practical applications, graph search algorithms such as Dijkstra's algorithm, A* algorithm, etc. can be used to search for paths that meet the core service requirement feature in the road network of the real-time traffic database.

[0132] Step S430: Match each service path in the initial set of service paths with the constraint conditions to determine the resource consumption parameters and path adaptation scores corresponding to each service path.

[0133] Constraint condition matching is to match each service path in the initial set of service paths with the constraint conditions of the real-time traffic environment and user requirements, and evaluate the feasibility and applicability of each service path. Resource consumption parameters refer to the resources required to be consumed during the execution of each service path, such as time, energy, cost, etc. The path adaptation score is to score the matching degree of each service path with user requirements and the real-time traffic environment. The higher the score, the more in line with user requirements and the real-time traffic environment the service path is. In one implementation, for the resource consumption parameters, the energy consumption of each path can be calculated according to the traffic flow and road conditions of different sections; the estimated travel time of each path can be estimated based on the real-time road conditions as the time consumption; if there are toll sections, etc., the cost is counted as the cost consumption. When calculating the path adaptation score, quantitative standards can be set first for each consideration factor of user requirements and the real-time traffic environment. For the travel time dimension, calculate the absolute value of the difference between the estimated travel time of the service path and the user's expected time, and divide this difference into intervals. For example, if the difference is within 0 - 5 minutes, the score is 8 - 10 points; 5 - 10 minutes, the score is 5 - 7 points; more than 10 minutes, the score is 1 - 4 points. For the travel mode preference dimension, if it fully meets the user preference, the score is 8 - 10 points; partially meets, the score is 4 - 7 points; does not meet at all, the score is 1 - 3 points. For the comfort dimension, grades are divided according to the number of transfers, road conditions, etc. and corresponding scores are given. For the traffic environment dimension, grades are scored according to the traffic flow and congestion conditions. According to actual needs, corresponding weights are assigned to each dimension, and the scores of each dimension are multiplied by the corresponding weights and then added together to obtain the path adaptation score.

[0134] When performing constraint condition matching, multiple factors need to be considered, such as traffic flow, road conditions, travel time, resource consumption, etc. For example, if a service path passes through a congested section, its required travel time will increase, and the resource consumption will also increase accordingly, and the path adaptation score will decrease. The resource consumption parameters and path adaptation scores of each service path can be calculated through the traffic flow information and road condition information in the real-time traffic database.

[0135] Step S440: Construct a dynamic optimization objective function according to the auxiliary service demand characteristics; wherein, the dynamic optimization objective function is used to balance the comprehensive optimization weights of resource consumption parameters and path adaptation scores.

[0136] The dynamic optimization objective function is a function constructed according to the auxiliary service demand characteristics, which is used to balance the comprehensive optimization weights of resource consumption parameters and path adaptation scores. The auxiliary service demand characteristics reflect the requirements of users for aspects such as travel time and comfort. Through the dynamic optimization objective function, the weights of resource consumption parameters and path adaptation scores can be adjusted according to different user needs.

[0137] For example, if the user has a high requirement for travel time, the dynamic optimization objective function can increase the weight of the time factor in the resource consumption parameters and reduce the weights of other factors in the path adaptation score, so as to preferentially select the service path with the shortest travel time. The dynamic optimization objective function can be expressed as: F = w1R + w2S, where F is the value of the dynamic optimization objective function, R is the resource consumption parameter, S is the path adaptation score, and w1 and w2 are weight coefficients, which are dynamically adjusted according to the auxiliary service demand characteristics.

[0138] Step S450: Call the path optimization algorithm, and iteratively screen the initial service path set based on the dynamic optimization objective function to generate the optimal service path.

[0139] The path optimization algorithm is an algorithm used to select the optimal path among multiple service paths. Based on the dynamic optimization objective function, the initial service path set is iteratively screened, and the dynamic optimization objective function values of each service path are continuously compared to select the optimal service path.

[0140] As an implementation, the path optimization algorithm realizes iterative screening through the following steps S451~S455: Step S451: Initialize the path screening queue and add all service paths in the initial service path set to the path screening queue.

[0141] The path screening queue is a queue used to store service paths to be screened. When initializing the path screening queue, all service paths in the initial service path set are added to the queue in sequence. In the subsequent iterative screening process, service paths will be taken out from the queue for processing.

[0142] Step S452: Extract the current path from the path screening queue and calculate the objective function value corresponding to the current path.

[0143] Take a service path from the path screening queue as the current path, and then calculate the objective function value corresponding to this path according to the dynamic optimization objective function. The dynamic optimization objective function comprehensively considers the resource consumption parameters and the path adaptation score. By calculating the objective function value, the quality of the current path can be evaluated.

[0144] Step S453: If the objective function value is greater than the preset threshold, add the current path to the candidate path set.

[0145] The preset threshold is a threshold of the objective function value set in advance, which is used to judge whether the current path meets the preset requirements. If the objective function value of the current path is greater than the preset threshold, it means that this path performs well in terms of resource consumption and path adaptation, and it will be added to the candidate path set. The candidate path set is a set used to store service paths that meet the set requirements.

[0146] Step S454: If the objective function value is less than or equal to the preset threshold, perform path segmentation optimization on the current path, generate an optimized set of sub-paths, and add the set of sub-paths to the path screening queue.

[0147] If the objective function value of the current path is less than or equal to the preset threshold, it means that this path performs poorly in terms of resource consumption and path adaptation, and it needs to be optimized. Path segmentation optimization is to divide the current path into multiple sub-paths, and then optimize each sub-path separately. Local search algorithms such as the greedy algorithm and the simulated annealing algorithm can be used to optimize each sub-path. After generating the optimized set of sub-paths, add it to the path screening queue for continuous screening in subsequent iterations.

[0148] Step S455: Repeat the above extraction, calculation, and processing steps until the path screening queue is empty or the maximum number of iterations is reached, and select the service path with the largest objective function value from the candidate path set as the optimal service path.

[0149] Repeat steps S452 - S454, continuously extract service paths from the path screening queue for processing until the path screening queue is empty or the maximum number of iterations is reached. The maximum number of iterations is an upper limit of the number of iterations set in advance, which is used to control the running time of the algorithm. When the path screening queue is empty or the maximum number of iterations is reached, select the service path with the largest objective function value from the candidate path set as the optimal service path.

[0150] Step S460: Generate a comprehensive response strategy including path guidance information and resource adaptation strategies according to the optimal service path and the characteristics of auxiliary service requirements.

[0151] Generate a comprehensive response strategy based on the optimal service path and the characteristics of auxiliary service requirements. Route guidance information is detailed information on how to travel along the optimal service path, such as the specific direction of the route, the locations passed through, transfer information, etc. The resource adaptation strategy is a strategy for reasonably allocating transportation resources according to the characteristics of auxiliary service requirements, such as selecting appropriate transportation means and arranging appropriate travel times.

[0152] When generating the comprehensive response strategy, combine the optimal service path and the characteristics of auxiliary service requirements, and integrate the route guidance information and the resource adaptation strategy. For example, if the characteristics of auxiliary service requirements indicate that the user has a high requirement for travel time, the resource adaptation strategy can select the fastest transportation means and arrange an appropriate travel time; the route guidance information can provide the specific way of taking this transportation means and the route. By generating a comprehensive response strategy that includes route guidance information and resource adaptation strategy, the traffic query needs of users are met.

[0153] Step S500: Convert the comprehensive response strategy into interactive feedback information, dynamically adapt the device parameters in the real-time environment association information, and perform a feedback operation to complete the optimized response to the traffic query request.

[0154] Converting the comprehensive response strategy into interactive feedback information is to enable the target user to conveniently obtain and understand the results of the traffic query. The interactive feedback information can be in text form, voice form, graphic form, etc. Dynamically adapting the device parameters in the real-time environment association information is to ensure that the feedback information can be normally displayed and interacted on different devices. Perform a feedback operation to send the interactive feedback information to the target user and complete the optimized response to the traffic query request.

[0155] As an implementation method, in step S500, converting the comprehensive response strategy into interactive feedback information, dynamically adapting the device parameters in the real-time environment association information, and performing a feedback operation to complete the optimized response to the traffic query request can specifically include the following steps S510~S580: Step S510: Perform multi-channel coding decomposition on the route guidance information and the resource adaptation strategy in the comprehensive response strategy to generate a text feedback data stream and a device control data stream.

[0156] Multi-channel coding decomposition is to decompose the route guidance information and the resource adaptation strategy in the comprehensive response strategy and encode them into a text feedback data stream and a device control data stream respectively. The text feedback data stream is used to provide the user with feedback information in text form, such as route instructions, transportation means selection suggestions, etc. The device control data stream is the data stream used to control relevant devices, such as controlling the navigation device to display the route and controlling the intelligent transportation card for payment.

[0157] When performing multi-channel coding decomposition, classification coding can be carried out according to the different contents and uses of path guidance information and resource adaptation strategies. For example, the text description part in the path guidance information is encoded into a text feedback data stream, and the instruction part related to device control in the resource adaptation strategy is encoded into a device control data stream.

[0158] Step S520: Reorganize the natural language structure of the text feedback data stream to generate a contextually coherent text sequence that conforms to the user interaction habit.

[0159] The natural language structure reorganization is to process the text feedback data stream so that it is more in line with the user interaction habit in terms of language structure and expression. The contextually coherent text sequence means that the text content is semantically and logically coherent and easy for users to understand. When performing natural language structure reorganization, natural language processing technologies such as syntactic analysis and semantic understanding can be used. First, perform syntactic analysis on the text feedback data stream to understand the structure and grammatical relationship of the sentence. Then, according to the user interaction habit and language expression norms, adjust and optimize the sentence to make the text content smoother and easier to understand. For example, split some complex sentences into simple sentences and adjust the order of the sentences to make the logic of the text clearer.

[0160] Step S530: Perform device type matching processing on the device control data stream, and extract the target device communication protocol rules from the predefined protocol adaptation library based on the device type parameters in the real-time environment association information.

[0161] The device type matching processing is to determine the type of the target device according to the device type parameters in the real-time environment association information. The predefined protocol adaptation library is a library that stores communication protocol rules for different devices, which contains communication protocol information for various devices.

[0162] Based on the device type parameters in the real-time environment association information, extract the target device communication protocol rules from the predefined protocol adaptation library. For example, if the target device is a smart phone, extract the protocol rules related to smart phone communication from the protocol adaptation library, such as Bluetooth protocol, Wi-Fi protocol, etc. Through device type matching processing and protocol rule extraction, ensure that the device control data stream can communicate correctly with the target device.

[0163] Step S540: According to the target device communication protocol rules, perform instruction format conversion on the device control data stream to generate a set of device execution instructions that are compatible with the target device hardware interface.

[0164] Instruction format conversion is to convert the device control data stream into an instruction format compatible with the hardware interface of the target device according to the communication protocol rules of the target device. Different devices may have different hardware interfaces and communication protocols, and the device control data stream needs to be correspondingly converted to be correctly recognized and executed by the target device. When performing instruction format conversion, the instructions in the device control data stream are adjusted in format and encoded according to the communication protocol rules of the target device. For example, if the hardware interface of the target device requires instructions to be in binary encoding format, the instructions in the device control data stream are converted to this binary encoding format. Through instruction format conversion, a set of device execution instructions compatible with the hardware interface of the target device is generated.

[0165] Step S550: Perform real-time alignment processing on the contextually coherent text sequence and the set of device execution instructions, and calculate the synchronization offset between the text display timing and the device control timing based on the device response delay parameter in the real-time environment association information.

[0166] Real-time alignment processing is to ensure the synchronization in time of the display of the contextually coherent text sequence and the execution of the set of device execution instructions. Based on the device response delay parameter in the real-time environment association information, calculate the synchronization offset between the text display timing and the device control timing. The device response delay parameter refers to the response time of the device to instructions, and different devices may have different response delays. When performing real-time alignment processing, first obtain the device response delay parameter in the real-time environment association information. Then, according to the device response delay parameter, calculate the synchronization offset between the text display timing and the device control timing. For example, if the response delay of the device is 1 second, the text display timing needs to be delayed by 1 second accordingly to ensure the synchronization of text display and device control.

[0167] Step S560: Dynamically calibrate the push time nodes of the contextually coherent text sequence according to the synchronization offset to generate a text push queue aligned with the time axis.

[0168] Dynamically calibrate the push time nodes of the contextually coherent text sequence according to the calculated synchronization offset. Dynamic calibration means adjusting the push time of the contextually coherent text sequence according to the synchronization offset to make it synchronized with the device control timing. Generate a text push queue aligned with the time axis, and arrange the calibrated contextually coherent text sequences in the queue in chronological order. In subsequent feedback operations, push text information in sequence according to the text push queue aligned with the time axis to ensure the synchronization of text display and device control.

[0169] Step S570: Perform cross-channel data encapsulation on the text push queue aligned with the time axis and the set of device execution instructions to generate a multi-modal interaction instruction packet.

[0170] Cross-channel data encapsulation is to integrate and encapsulate data from different channels, namely the text data represented by the text push queue aligned with the time axis and the device control data represented by the set of device execution instructions, so as to be orderly and efficiently transmitted to the corresponding device of the target user during the same interaction process. This process needs to consider data format compatibility, transmission stability, and parsability on different devices.

[0171] When performing cross-channel data encapsulation, first determine a unified data structure to accommodate text data and device control data. Formats such as JSON (JavaScript Object Notation) or XML (eXtensible Markup Language) can be used. Taking JSON as an example, create a JSON object with two main fields. One field is used to store information about the text push queue, and the other field is used to store the set of device execution instructions. Inside each field, organize the data according to preset rules. For example, the text push queue can store each text message and its corresponding push time in the form of an array in chronological order, and the set of device execution instructions can store each instruction and its related parameters in the form of an object in an array.

[0172] In practical applications, assume that there are three text messages in the text push queue, namely "Please turn right at the intersection ahead", "The next stop is XX Station", and "There are still 2 kilometers to the destination", with corresponding push times of t1, t2, and t3 respectively; there are two instructions in the set of device execution instructions, one is an instruction to control the navigation device to display the route, and the other is an instruction to control the smart speaker to play a voice prompt.

[0173] Step S580: Parallelly send the multi-modal interaction instruction packet to the corresponding display terminal and device control terminal of the target user through the distributed communication interface, and trigger the display terminal to execute the phased text rendering operation, while triggering the device control terminal to execute the control commands in the set of device execution instructions.

[0174] The distributed communication interface is an interface for data transmission between different devices, which can achieve efficient and stable data transmission. In this step, use the distributed communication interface to parallelly send the multi-modal interaction instruction packet to the corresponding display terminal and device control terminal of the target user. Parallel sending can improve the efficiency of data transmission and ensure that the display terminal and device control terminal can receive the required data simultaneously.

[0175] A display terminal refers to a device used to present information to users, such as a smartphone screen, a vehicle-mounted display, etc. When the display terminal receives a multi-modal interaction instruction packet, it performs a phased text rendering operation according to the text push queue information therein. The phased text rendering operation means that, according to the set push time nodes in the text push queue, the text information is sequentially displayed on the screen. For example, at time t1, "Please turn right at the intersection ahead" is displayed; at time t2, "The next stop is XX Station" is displayed; at time t3, "There are still 2 kilometers to the destination" is displayed.

[0176] The device control terminal is a terminal used to control various devices, such as a smart speaker, a navigation device, etc. When the device control terminal receives a multi-modal interaction instruction packet, it will parse the device execution instruction set therein and execute the corresponding control commands. For example, controlling the navigation device to display a specified route and set an appropriate zoom level; controlling the smart speaker to play a preset voice prompt and adjust the volume size, etc.

[0177] Through the above steps, the comprehensive response strategy is transformed into interactive feedback information, and the device parameters in the real-time environment association information are dynamically adapted, and the feedback operation is performed, thus completing the optimized response to the traffic query request. The entire traffic comprehensive query optimization method, through a series of steps such as obtaining the user's traffic query request, performing multi-modal feature parsing, feature fusion, service path optimization, and feedback information generation, fully considers the user's needs and the real-time traffic environment, and can provide accurate, efficient, and personalized traffic query services for users.

[0178] It can be understood that in the above introductions of the embodiments of the present invention, various algorithms, models, network layers involved, such as the greedy algorithm, the simulated annealing algorithm, etc., can be obtained from the relevant content in the prior art. For the sake of saving space, they are not elaborated in the embodiments of the present application. In addition, those skilled in the art can make detailed supplements according to the common general knowledge in the art when implementing the solution of the present application. For example, according to the general knowledge in the art, normalization can be used to eliminate the dimension conflict before feature fusion, interpolation can be used to eliminate the dimension difference, historical data, experience or business scenario requirements can be combined to reasonably set the threshold, and the model can be trained based on the general model training method, etc. The present application will no longer give redundant introductions to the overly detailed implementation process.

[0179] In addition, those skilled in the art can make optimization modifications to this solution based on their own knowledge for some unstated content. For example, after feature orthogonalization, a feature compatibility verification step can be added to calculate the correlation coefficient between the newly added vector and the orthogonal basis. When the correlation coefficient exceeds the threshold, perform Schur complement orthogonalization processing on the resource constraint embedding vector to eliminate potential conflicts that may exist between the linear independence of orthogonal features and the correlation of the newly added vector; when training the weight model, it can be optimized to adopt a two-stage weight prediction mechanism. In the offline stage, pre-train the basic weight allocation model, and in the online stage, update the weight parameters in real time through Kalman filtering; for some network structures, those skilled in the art can make specific choices according to actual needs. For example, the temporal convolutional layer uses causal dilated convolution, and the spatial graph attention uses a multi-head attention mechanism, etc.

[0180] Please refer to Figure 2 , Figure 2 FIG. is a schematic structural diagram of a human-computer interaction system provided by an embodiment of the present invention. The human-computer interaction system is, for example, a smart phone, a vehicle-mounted terminal, etc., and at least includes a processor 101, a communication interface 102, and a memory 103. Among them, the processor 101, the communication interface 102, and the memory 103 can be connected through a bus or other means. Among them, the processor 101 (or Central Processing Unit (CPU)) is the computing core and control core of the human-computer interaction system, which can parse various instructions in the human-computer interaction system and process various data in the human-computer interaction system. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and under the control of the processor 101, can be used for sending and receiving data; the communication interface 102 can also be used for the transmission and interaction of internal data in the human-computer interaction system. The memory 103 (Memory) is a memory device in the human-computer interaction system, used to store programs and data. It can be understood that the memory 103 here can include both the built-in memory of the human-computer interaction system and, of course, the extended memory supported by the human-computer interaction system. The memory 103 provides a storage space, and the operating system of the human-computer interaction system is stored in this storage space, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc., and the present invention does not make any limitations in this regard.

[0181] In one embodiment, the processor 101 executes the traffic comprehensive query optimization method applied to intelligent human-computer interaction provided by the above embodiments of the present invention by running the computer program in the memory 103.

Claims

1. A traffic comprehensive query optimization method applied to intelligent human-computer interaction, characterized in that Including: Obtain a traffic query request input by a target user in an interaction scenario, where the traffic query request includes natural language text and real-time environment association information; Perform multi-modal feature parsing on the traffic query request to generate semantic understanding features and scenario association features, where the scenario association features are used to indicate the real-time traffic environment state related to the traffic query request; Based on a preset language large model, perform dynamic interaction feature fusion processing on the semantic understanding features and the scenario association features to generate multi-dimensional intention features; According to the multi-dimensional intention features, optimize the service path of the traffic query request of the target user to generate a comprehensive response strategy; Convert the comprehensive response strategy into interactive feedback information, and dynamically adapt the device parameters in the real-time environment association information, and perform a feedback operation to complete the optimized response of the traffic query request.

2. The method according to claim 1, wherein The performing multi-modal feature parsing on the traffic query request to generate semantic understanding features and scenario association features includes: Perform word segmentation processing on the natural language text to obtain multiple semantic units, and extract environmental parameters from the real-time environment association information to obtain multiple environmental state parameters; Perform context semantic encoding on the multiple semantic units to generate the semantic understanding features; where the semantic understanding features are used to represent the core semantic intention of the traffic query request; Perform spatio-temporal association analysis on the multiple environmental state parameters to generate an environmental state encoding vector; Based on the environmental state encoding vector and the device type parameters in the real-time environment association information, construct the scenario association features; Wherein, the scenario association features include at least one of the following: real-time traffic flow distribution parameters, device response delay parameters, and spatial distance parameters between the user location and the target traffic node.

3. The method according to claim 2, wherein The performing dynamic interaction feature fusion processing on the semantic understanding features and the scenario association features based on a preset language large model to generate multi-dimensional intention features includes: Map the semantic understanding features to a first feature vector, and map the scenario association features to a second feature vector; Perform dynamic weight allocation processing on the first feature vector and the second feature vector to determine a semantic weight coefficient and an environmental weight coefficient; According to the semantic weight coefficient, perform weighted adjustment on the first feature vector to obtain an adjusted semantic feature vector; According to the environmental weight coefficient, perform weighted adjustment on the second feature vector to obtain an adjusted environmental feature vector; Concatenate the adjusted semantic feature vector and the adjusted environmental feature vector to generate an initial fusion feature; Call the intention recognition layer in the language large model to perform intention dimension expansion on the initial fusion feature to generate the multi-dimensional intention features; where the multi-dimensional intention features include user explicit demand features, implicit scenario adaptation features, and real-time resource constraint features.

4. The method according to claim 3, characterized in that, The calling the intention recognition layer in the language large model to perform intention dimension expansion on the initial fusion feature to generate the multi-dimensional intention features includes: Decompose the initial fusion feature into a semantics-dominated feature branch and an environment-dominated feature branch; wherein, the semantics-dominated feature branch consists of the features obtained by screening the adjusted semantics feature vector through channel attention, and the environment-dominated feature branch consists of the features obtained by screening the adjusted environment feature vector through spatial attention; Perform context demand association analysis on the semantics-dominated feature branch to extract user explicit demand features; wherein, the context demand association analysis is achieved by traversing the matching relationship between the semantic units in the semantics-dominated feature branch and a predefined set of traffic service keywords; Perform dynamic scene adaptation analysis on the environment-dominated feature branch to generate implicit scene adaptation features; wherein, the dynamic scene adaptation analysis includes: determining the dynamic constraint conditions of the traffic resources for the scene where the target user is located based on the real-time traffic state parameters in the environment-dominated feature branch; Input the user explicit demand features and the implicit scene adaptation features into a cross-modal interaction network to generate potential demand compensation features; wherein, the cross-modal interaction network realizes feature compensation by alternately performing projection mapping from semantic features to environment features and feedback correction from environment features to semantic features; Perform feature orthogonalization processing on the user explicit demand features, the implicit scene adaptation features, and the potential demand compensation features, and generate an orthogonal demand feature set after eliminating redundant information; Based on each feature vector in the orthogonal demand feature set, construct a multi-dimensional demand distribution matrix, and perform feature dimension expansion on the multi-dimensional demand distribution matrix to generate the multi-dimensional intention feature including explicit demand, implicit scene adaptation, and real-time resource constraint; wherein, the feature dimension expansion is achieved by performing tensor splicing on the multi-dimensional demand distribution matrix and a pre-trained resource constraint embedding vector; 5. The method according to claim 3, wherein According to the multi-dimensional intention feature, optimize the service path for the traffic query request of the target user to generate a comprehensive response strategy, including: Split the service demand of the multi-dimensional intention feature to obtain a core service demand feature and an auxiliary service demand feature; Based on a real-time traffic database, obtain an initial set of service paths that match the core service demand feature; Match the constraint conditions for each service path in the initial set of service paths, and determine the resource consumption parameter and path adaptation score corresponding to each service path; Construct a dynamic optimization objective function according to the auxiliary service demand feature; wherein, the dynamic optimization objective function is used to balance the comprehensive optimization weight of the resource consumption parameter and the path adaptation score; Call a path optimization algorithm to iteratively screen the initial set of service paths based on the dynamic optimization objective function to generate an optimal service path; Generate the comprehensive response strategy including path guidance information and resource adaptation strategy according to the optimal service path and the auxiliary service demand feature; Wherein, the path optimization algorithm realizes iterative screening through the following steps: Initialize a path screening queue, and add all service paths in the initial set of service paths to the path screening queue; Extract the current path from the path screening queue and calculate the objective function value corresponding to the current path; If the objective function value is greater than the preset threshold, add the current path to the candidate path set; If the objective function value is less than or equal to the preset threshold, perform path segmentation optimization on the current path to generate an optimized sub-path set, and add the sub-path set to the path screening queue; Repeat the above extraction, calculation, and processing steps until the path screening queue is empty or the maximum number of iterations is reached; Select the service path with the largest objective function value from the candidate path set as the optimal service path.

6. The method according to claim 5, wherein The step of converting the comprehensive response policy into an interactive feedback message, dynamically adapting the device parameters in the real-time environment association information, and performing a feedback operation to complete the optimized response to the traffic query request includes: Perform multi-channel encoding and decomposition on the path guidance information and resource adaptation policy in the comprehensive response policy to generate a text feedback data stream and a device control data stream; Reorganize the natural language structure of the text feedback data stream to generate a contextually coherent text sequence that conforms to the user interaction habit; Perform device type matching on the device control data stream, and based on the device type parameters in the real-time environment association information, extract the target device communication protocol rules from the predefined protocol adaptation library; According to the target device communication protocol rules, perform instruction format conversion on the device control data stream to generate a set of device execution instructions compatible with the target device hardware interface; Perform real-time alignment processing on the contextually coherent text sequence and the set of device execution instructions, and based on the device response delay parameter in the real-time environment association information, calculate the synchronization offset between the text display timing and the device control timing; Dynamically calibrate the push time node of the contextually coherent text sequence according to the synchronization offset to generate a time-axis aligned text push queue; Perform cross-channel data encapsulation on the time-axis aligned text push queue and the set of device execution instructions to generate a multi-modal interaction instruction packet; Parallelly send the multi-modal interaction instruction packet to the display terminal and the device control terminal corresponding to the target user through a distributed communication interface, and trigger the display terminal to perform a phased text rendering operation, and at the same time trigger the device control terminal to execute the control commands in the set of device execution instructions.

7. The method according to claim 2, wherein The step of performing spatio-temporal correlation analysis on the multiple environmental state parameters to generate an environmental state encoding vector includes: Extract the timestamp data and geographical location data from the real-time environment association information; Determine the time distribution characteristics of the environmental state parameters according to the timestamp data; Determine the spatial topological relationship of the environmental state parameters according to the geographical location data; Input the time distribution characteristics and the spatial topological relationship into a spatio-temporal encoding network to generate spatio-temporal correlation characteristics; Perform dimensionality reduction processing on the spatio-temporal correlation features to obtain the environmental state encoding vector; wherein, the spatio-temporal encoding network includes a temporal convolutional layer and a spatial graph attention layer, the temporal convolutional layer is used to extract temporal sequence dependencies, and the spatial graph attention layer is used to extract the correlation weights between geographical locations.

8. The method according to claim 3, wherein The dynamic weight assignment processing of the first feature vector and the second feature vector to determine the semantic weight coefficient and the environmental weight coefficient includes: Obtain the sample semantic feature vector set and the sample environmental feature vector set in the historical interaction dataset, where each sample semantic feature vector is associated with the influence weight value of the semantic feature in the actual decision-making, and each sample environmental feature vector is associated with the influence weight value of the environmental feature in the actual decision-making; Input the sample semantic feature vector and the sample environmental feature vector into the initial weight prediction model to generate the corresponding semantic weight prediction sequence and environmental weight prediction sequence; Calculate the error loss between each predicted value in the semantic weight prediction sequence and its associated actual influence weight value to generate a semantic weight error set; Calculate the error loss between each predicted value in the environmental weight prediction sequence and its associated actual influence weight value to generate an environmental weight error set; Perform weighted fusion on the semantic weight error set and the environmental weight error set to generate a total error loss value, and iteratively adjust the parameters of the initial weight prediction model according to the total error loss value until the total error loss value converges to a stable interval to generate a trained weight assignment model; Input the first feature vector and the second feature vector into the trained weight assignment model to output the initial semantic weight value and the initial environmental weight value before normalization; Perform dynamic balance constraint processing on the initial semantic weight value and the initial environmental weight value so that the sum of the two satisfies a preset constant relationship to generate an intermediate semantic weight value and an intermediate environmental weight value; Perform non-linear activation transformation processing on the intermediate semantic weight value to generate a smooth semantic weight coefficient after eliminating the sudden fluctuations of the weight value; Perform non-linear activation transformation processing on the intermediate environmental weight value to generate a smooth environmental weight coefficient after eliminating the sudden fluctuations of the weight value; Perform joint verification processing on the smooth semantic weight coefficient and the smooth environmental weight coefficient to ensure that their interaction contribution degrees to the first feature vector and the second feature vector match, and generate the finally available semantic weight coefficient and environmental weight coefficient.

9. The method according to claim 1, wherein The obtaining of the traffic query request input by the target user in the interaction scenario includes: Receive the original query information submitted by the target user through voice input, text input or touch input; Perform noise filtering processing on the original query information to remove irrelevant characters or invalid audio segments; Perform intention pre-recognition on the filtered original query information to determine whether it contains key semantic units related to traffic; If the key semantic units are included, mark the original query information as a valid traffic query request and trigger the subsequent processing flow; If the key semantic units are not included, send guiding information to the target user to re-enter or supplement the query content.

10. A human-computer interaction system, characterized in that, Comprising: A memory in which a computer program is stored; A processor for loading the computer program to implement the traffic comprehensive query optimization method applied to intelligent human-computer interaction according to any one of claims 1-9.

Citation Information

Patent Citations

  • Language-driven traffic interaction scene generation system and method

    CN118333035A

  • Man-machine interaction type query optimization method and system for incomplete query

    CN118643060A

  • Multi-modal voice interaction method based on large model, electronic equipment and storage medium

    CN119559946A

  • Traffic special situation optimization processing method and system based on large model

    CN119723899A

  • Multi-modal large model training system and method applied to traffic field

    CN119903887A

Cited By

  • User portrait combined store personalized recommendation interaction method and system

    CN120563214A

  • Operation and maintenance technology service remote guidance interaction method and system combined with agent assistance

    CN120670563A

  • AI voice interaction-based pension service scheduling method and system

    CN120748370A

  • Semantic understanding-based medical data query optimization method and system

    CN121412270A

  • Method and system for optimizing medical data query based on semantic understanding

    CN121412270B