An intelligent reply system and method
By constructing multidimensional feature vectors and prediction models, and combining text and map material with an arrangement and response mode, the problem of existing map query services struggling to understand complex address queries has been solved, resulting in a more accurate and personalized map query service.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI BIG DATA INC
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-31
AI Technical Summary
Existing map search services struggle to understand users' complex address search needs, and cannot meet users' diverse requirements solely through text or map content.
A multidimensional feature vector is constructed, and the semantic feature information of the natural query language input by the user is extracted through the vector construction module. Combined with the prediction model and intent scoring module, the arrangement of the response is determined. A multimodal arrangement response mode is adopted, which combines text description and map materials for the response.
It accurately meets users' address query needs and provides richer and more personalized map query services.
Smart Images

Figure CN122489585A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of map query, and more particularly to an intelligent reply system and method. Background Technology
[0002] In modern society, people cannot live without map search services. Users input natural language to map search software to express their address search needs and receive responses from the software.
[0003] However, existing map search services can only respond with text content or monotonous map images based on fixed rules, making it difficult to understand the user's intent. Complex address search needs often require a combination of text and map content; simply responding with text or map content alone is insufficient to meet users' complex address search requirements. Summary of the Invention
[0004] To address the aforementioned problems in the existing technology, this invention provides an intelligent reply system applied to map query scenarios, comprising: The vector construction module is used to extract semantic feature information from multiple dimensions of the natural query language input by the user, and fuse the semantic feature information to construct a multi-dimensional feature vector; The multidimensional feature vector includes at least a map scene vector associated with the address query requirement; The intent scoring module, connected to the vector construction module, uses a pre-trained prediction model to process the multidimensional feature vector to output a predicted value. The intelligent response module, connected to the intent scoring module, is used to determine the formatted response corresponding to the natural query language based on the predicted value; The proposed response arrangement method includes at least a multimodal response mode that combines map materials with textual descriptions.
[0005] Preferably, the semantic feature information includes at least one of spatial feature information, guidance feature information, action feature information, interaction feature information, and scene feature information; The vector construction module includes: The first assignment unit is used to assign scores to the spatial feature information according to a preset address scoring standard in order to form the map scene vector; The second assignment unit is used to assign scores to the guidance feature information according to a preset guidance keyword library to form a guidance feature vector; The third assignment unit is used to assign scores to the action feature information according to a preset action keyword library to form an action feature vector; The fourth assignment unit is used to assign values to the interaction feature information based on the number of map queries in the user's history, so as to form an interaction feature vector; The fifth assignment unit is used to encode and assign values to the scene feature information according to a preset scene encoding library to form a scene feature vector; The vector construction unit, which is connected to the first assignment unit, the second assignment unit, the third assignment unit, the fourth assignment unit, and the fifth assignment unit respectively, is used to integrate the map scene vector, the guidance feature vector, the action feature vector, the interaction feature vector, and the scene feature vector into the multidimensional feature vector.
[0006] Preferably, the intelligent reply module includes: A threshold determination unit is used to determine a first threshold based on the scene feature vector; A response mode determination unit, connected to the threshold determination unit, is used to compare the predicted value with the first threshold and determine the arrangement response mode based on the comparison result; If the predicted value is greater than the first threshold, the multimodal orchestration response mode is executed; If the predicted value is not greater than the first threshold, execute the text reply mode that only contains textual descriptions; The intelligent reply unit, connected to the reply mode determination unit, is used to reply to the natural query language input by the user according to the arrangement reply method determined by the reply mode determination unit.
[0007] Preferably, multiple call intervals are pre-set for the multimodal orchestration response mode, and the display content of the map material corresponding to each call interval is different; The intelligent reply module further includes a content display determination unit, which is connected to the reply mode determination unit and the intelligent reply unit respectively, and is used to determine the call interval that matches the predicted value, thereby determining the content to be displayed in the map material; The intelligent reply unit uses the displayed content of the determined map material to determine the reply content corresponding to the multimodal arrangement reply mode.
[0008] Preferably, the map materials have different types of display content, and each type of display content has a corresponding display priority; The response content determination unit includes: The priority determination subunit is used to determine that the display priority of the text description is always greater than the display priority of the map material, and The display priority of the map material's content is determined based on the calling interval; A multimodal material library, which stores text materials for different addresses and different display content of the map materials; The calling subunit connects to the multimodal material library and is used to call the display content of the text material and the map material from the multimodal material library according to the calling interval; The text and map materials invoked are respectively matched with the semantic feature information; An arrangement subunit, connected to the priority determination subunit and the invocation subunit, is used to organize the text materials into the textual narrative content, and to arrange the textual narrative content and the map materials based on the display priority. The display content of the map materials is arranged based on the display priority of the corresponding call interval.
[0009] Preferably, the arrangement subunit determines the arrangement method of the text material based on the calling interval, so as to determine the text narrative content.
[0010] Preferably, the calling interval includes a first calling interval, a second calling interval, and a third calling interval, wherein the numerical range of the first calling interval is smaller than the numerical range of the second calling interval, the numerical range of the second calling interval is smaller than the numerical range of the third calling interval, and the first threshold is used as the lower limit of the numerical range of the first calling interval. The display content of the map material corresponding to the first call interval includes thumbnail content; The display content of the map material corresponding to the second call interval includes thumbnail content and map location content; When the calling interval is determined to be the second calling interval, the priority determination subunit determines that the display priority of the map positioning is greater than the display priority of the thumbnail; The display content of the map material corresponding to the third call interval includes thumbnail content, map location content, and associated layer content; When the calling interval is determined to be the third calling interval, the priority determination subunit determines that the display priority of the associated layer content is greater than the display priority of the map positioning content, and the display priority of the map positioning content is greater than the display priority of the thumbnail content.
[0011] This invention provides an intelligent reply method applicable to map query scenarios; The intelligent response method includes: Step S1: Extract semantic feature information from multiple dimensions of the natural query language input by the user, and fuse the semantic feature information to construct a multi-dimensional feature vector; The multidimensional feature vector includes at least a map scene vector associated with the address query requirement; Step S2: The multidimensional feature vector is processed using a pre-trained prediction model to output a predicted value; Step S3: Determine the arrangement and response method corresponding to the natural query language based on the predicted value; The proposed response arrangement method includes at least a multimodal response mode that combines map materials with textual descriptions.
[0012] Preferably, the semantic feature information includes at least one of spatial feature information, guidance feature information, action feature information, interaction feature information, and scene feature information; Step S1 includes: Step S11: Assign values to the spatial feature information according to a preset address scoring standard to form the map scene vector, and The guiding feature information is scored and assigned values based on a preset guiding keyword library to form a guiding feature vector. The action feature information is scored and assigned values based on a preset action-related keyword library to form an action feature vector. The interaction feature information is assigned values based on the user's historical map query count to form an interaction feature vector. The scene feature information is encoded and assigned values according to a preset scene encoding library to form a scene feature vector; Step S12: Integrate the map scene vector, the guidance feature vector, the action feature vector, the interaction feature vector, and the scene feature vector into the multidimensional feature vector.
[0013] Preferably, multiple call intervals are pre-set for the multimodal response mode, and each call interval corresponds to the modal type that needs to be called; Step S3 includes: Step S31: Determine the first threshold based on the scene feature vector; Step S32: Determine whether the predicted value is greater than the first threshold. If so, execute the multimodal orchestration response mode, and then proceed to step S33; If not, execute the text response mode that only contains textual descriptions, and then proceed to step S33; Step S33: Reply to the natural query language input by the user according to the arranged reply method.
[0014] The following beneficial effects can be obtained by using the present invention: By extracting semantic features from natural query language and constructing multidimensional feature vectors, processing these feature vectors and outputting predicted values, determining the appropriate response format based on the predicted values, and then replying to the user with the corresponding content, the system can accurately meet the user's address query needs. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the intelligent response system of the present invention; Figure 2 This is a schematic diagram of the vector construction module of the present invention; Figure 3 This is a schematic diagram of the intelligent response module of the present invention; Figure 4 This is a schematic diagram of the structure of the response content determination unit of the present invention; Figure 5 This is a flowchart illustrating the intelligent response method of the present invention; Figure 6 This is a schematic diagram of the process for constructing multidimensional feature vectors in this invention; Figure 7 This is a flowchart illustrating the process of determining the arrangement and response method in this invention; In the attached image: 1. Vector Construction Module, 11. First Assignment Unit, 12. Second Assignment Unit, 13. Third Assignment Unit, 14. Fourth Assignment Unit, 15. Fifth Assignment Unit, 16. Vector Construction Unit; 2. Intent Scoring Module; 3. Intelligent Response Module, 31. Threshold Determination Unit, 32. Response Mode Determination Unit, 33. Intelligent Response Unit, 34. Response Content Determination Unit, 341. Priority Determination Subunit, 342. Multimodal Material Library, 343. Calling Subunit, 344. Arrangement Subunit Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0018] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.
[0019] This invention provides an intelligent reply system for use in map query scenarios, such as... Figure 1 As shown, it includes: Vector construction module 1 is used to extract semantic feature information from multiple dimensions of the natural query language input by the user, and fuse the semantic feature information to construct a multi-dimensional feature vector; The multidimensional feature vector should include at least a map scene vector related to the address query requirement; Intent scoring module 2 and connection vector construction module 1 use a pre-trained prediction model to process multi-dimensional feature vectors to output predicted values; The intelligent response module 3, connected to the intent scoring module 2, is used to determine the arrangement of responses corresponding to the natural query language based on the predicted value. The reply arrangement methods should include at least a multimodal reply mode that uses map materials and combines them with text descriptions.
[0020] Specifically, users input natural language queries related to address searches on the client side. Input methods can include manually entering natural language queries or converting voice input into natural language queries.
[0021] Vector construction module 1 extracts semantic feature information from natural query language and saves the extracted semantic feature information.
[0022] Furthermore, before using the prediction model, the prediction model is trained using the following method: Beforehand, over 100,000 historical interaction data points were collected. Each historical interaction data point included the natural language of the user's input query and whether the user subsequently opened the map function. Historical interaction data where the user opened the map function was labeled as positive samples, indicating a need for the map function, while historical interaction data where the user did not manually open the map function was labeled as negative samples, indicating a need for only a text reply. A corresponding multidimensional feature vector was extracted from each historical interaction data point, and each multidimensional feature vector and its corresponding label were input into a random forest model for training.
[0023] Specifically, the random forest model is set to have 100 trees, a maximum depth of 10, and a minimum number of splits of 5. Each tree is trained sequentially using random sampling. After each sampling, the sampled data is put back and resampled until the maximum number of samples is reached (default is 100,000 total samples). The root node of each tree contains all sampled data. Starting from the root node, branches are formed, continuously splitting the sampled data according to data features (one feature per dimension, with the total number of dimensions not exceeding the maximum depth). Splitting stops when a node reaches the minimum number of splits, and this node becomes a leaf node. Each leaf node is classified, with leaf nodes having a large number of positive samples marked as positive sample nodes and leaf nodes having a large number of negative samples marked as negative sample nodes. Based on the number of positive and negative samples in the leaf nodes of each dimension, it is determined whether that dimension is a positive or negative sample on the corresponding tree. At this point, a tree that makes a judgment result for each dimension is obtained. The 100 trees are trained repeatedly and then combined to form a prediction model. Upon receiving a new multidimensional feature vector, each tree in the prediction model determines the positive or negative sample for each dimension, classifying positive samples as 1 and negative samples as 0, thus obtaining the judgment value for each dimension. The sum of the judgment values for each dimension is then averaged to obtain the predicted value of the multidimensional feature vector.
[0024] By using cross-validation, new multidimensional feature vectors are continuously input into the prediction model for verification, ultimately ensuring that the F1 score of the prediction model is not less than 0.92.
[0025] The F1 score is the average of the precision and recall of the prediction model. The higher the F1 score, the more accurate the predictions output by the model.
[0026] Furthermore, the intelligent response module 3 determines the response arrangement method based on the predicted value. The larger the predicted value, the more necessary it is to adopt a multimodal response mode that combines text descriptions and map materials.
[0027] In a preferred embodiment of the present invention, the semantic feature information includes at least one of spatial feature information, guidance feature information, action feature information, interaction feature information, and scene feature information; like Figure 2 As shown, vector construction module 1 includes: The first assignment unit 11 is used to assign scores to spatial feature information according to a preset address scoring standard in order to form a map scene vector; The second assignment unit 12 is used to assign scores to the guidance feature information according to the preset guidance keyword library to form a guidance feature vector; The third assignment unit 13 is used to assign scores to action feature information based on a preset action keyword library to form an action feature vector. The fourth assignment unit 14 is used to assign values to the interaction feature information based on the number of map queries in the user's history, so as to form an interaction feature vector; The fifth assignment unit 15 is used to encode and assign values to scene feature information according to a preset scene encoding library to form a scene feature vector; The vector construction unit 16 is connected to the first assignment unit 11, the second assignment unit 12, the third assignment unit 13, the fourth assignment unit 14 and the fifth assignment unit 15 respectively, and is used to integrate the map scene vector, the guidance feature vector, the action feature vector, the interaction feature vector and the scene feature vector into a multi-dimensional feature vector.
[0028] Specifically, on the one hand, the first assignment unit 11 scores the spatial feature information. For example, for semantic feature information including precise road names and location names such as XX Road XX Number, if its spatial feature information score is 0.8, then the corresponding map scene vector is 0.8.
[0029] Similarly, for semantic feature information that includes first-level geographic entities, such as XX province or XX region, the spatial feature information score is 0.3, and the corresponding map scene vector is 0.3.
[0030] Similarly, for semantic feature information that includes distribution needs, such as surrounding schools and restaurants, the spatial feature information score is 0.6, and the corresponding map scene vector is 0.6.
[0031] Furthermore, when semantic feature information includes multiple types of spatial feature information, a comprehensive score is calculated based on the scores of each type of spatial feature information, but the comprehensive score will not be lower than the highest score among the single type of spatial feature information. For example, if a user queries the distribution information of multiple locations around a certain road segment, the output spatial feature information score will necessarily be greater than 0.8.
[0032] On the other hand, the second assignment unit 12 parses the guiding keywords in the semantic feature information according to the preset guiding keyword library. If the semantic feature information includes guiding keywords from the guiding keyword library, such as navigation, distribution, boundary, etc., the guiding feature information is rated as 1 point and the guiding feature vector is 1. If there are no guiding keywords, the score is 0 and the corresponding guiding feature vector is 0.
[0033] On the other hand, the third assignment unit 13 parses the action-related keywords in the semantic feature information according to the preset action-related keyword library. If the semantic feature information includes action-related keywords from the action-related keyword library, such as "view," "plan," or "compare," then the action feature information is rated as 1 point, and the action feature vector is 1. If there are no action-related keywords, the score is 0, and the corresponding action feature vector is 0.
[0034] On the other hand, the fourth assignment unit 14 queries the user ID, counts and calculates the ratio of the number of times the user queries the map within 30 days to the total number of times the user inputs natural language queries, and uses this value as the interaction feature information.
[0035] On the other hand, the fifth assignment unit 15 identifies the query scenario in the semantic feature information and performs one-hot encoding. For example, if the semantic feature information includes the XX Municipal Government, the fifth assignment unit 15 determines it to be a government affairs scenario and encodes it as [1,0,0], and uses this encoding as the scenario feature vector.
[0036] Furthermore, the vector construction unit 16 integrates the map scene vector, guidance feature vector, action feature vector, interaction feature vector, and scene feature vector into a multi-dimensional feature vector, in the form of: [map scene vector, guidance feature vector, action feature vector, interaction feature vector, scene feature vector].
[0037] In a preferred embodiment of the present invention, such as Figure 3 As shown, the intelligent reply module 3 includes: The threshold determination unit 31 is used to determine the first threshold based on the scene feature vector; The response mode determination unit 32 and the connection threshold determination unit 31 are used to compare the predicted value with the first threshold and determine the arrangement of the response based on the comparison result. If the predicted value is greater than the first threshold, execute the multimodal orchestration response mode; If the predicted value is not greater than the first threshold, execute the text reply mode that only contains textual descriptions; The intelligent reply unit 33 is connected to the reply mode determination unit 32 and is used to reply to the natural query language entered by the user according to the reply arrangement determined by the reply mode determination unit 32.
[0038] In one embodiment, the initial first threshold is 0.6.
[0039] Specifically, the threshold determination unit 31 queries the scene feature vector and determines the first threshold according to the preset threshold determination rules. For example, if the scene feature vector corresponds to a government affairs scene, the first threshold is adjusted to 0.5 to increase the probability of map triggering. Or, if the scene feature vector corresponds to a life services scene, the first threshold is adjusted to 0.7 to reduce the probability of map triggering and prevent frequent map triggering.
[0040] Furthermore, the response pattern determination unit 32 compares the predicted value with the magnitude of the first threshold.
[0041] On the one hand, if the predicted value is greater than the first threshold, a multimodal orchestration response mode is executed, and the response content includes textual descriptions and map materials.
[0042] On the other hand, if the predicted value is not greater than the first threshold, a text reply mode containing only textual descriptions will be executed.
[0043] Furthermore, a backup strategy is set in the response mode determination unit 32. When the predicted value is not greater than the first threshold but close to the first threshold, and the corresponding ratio of the interaction feature information is very high, it is determined that the user has potential map needs, and the multimodal orchestration response mode is still entered.
[0044] Furthermore, the intelligent response unit 33 returns the arrangement results to the client for the user to query. When the user finishes querying and exits, the intelligent response unit 33 collects the user's feedback information, including at least whether the user opened the map function, whether the user added the map to their favorites, whether the user immediately closed the map or re-entered the natural language query, etc. The collected feedback information is saved to the database to update historical interaction data and delete data that has been stored for too long (such as 30 days).
[0045] The random forest model retrains the prediction model incrementally or fully every 7 days based on the updated database, replacing the old prediction model with the new one to achieve adaptive optimization of the prediction model.
[0046] In a preferred embodiment of the present invention, multiple call intervals are pre-set for the multimodal orchestration response mode, and the display content of the map material corresponding to each call interval is different; The intelligent response module 3 also includes a response content determination unit 34, which is connected to the response mode determination unit 32 and the intelligent response unit 33 respectively. It is used to determine the call range that matches the predicted value, and then determine the display content of the map material. The intelligent response unit 33 uses the displayed content of the determined map material to determine the response content corresponding to the multimodal arrangement response mode.
[0047] Specifically, map materials include primary geographic entities such as provinces and cities, as well as specific addresses such as road sections and buildings. Each type of map material includes different display content, such as corresponding map coordinates, surrounding dynamic maps, map hierarchy, and thumbnails.
[0048] Each call interval corresponds to the display content of the map assets that need to be called. The display content of the map assets is determined based on the call interval that matches the predicted value.
[0049] Furthermore, the display content of the map materials can be saved in advance. For display content that is not saved, the response content determination unit 34 can also generate the corresponding map material display content based on real-time calculations. For example, if it is necessary to call up the thumbnail and surrounding dynamic map of region A, but the display content of region A only has the thumbnail saved, then the response content determination unit 34 can obtain the dynamic map of the surrounding area of region A through GPS positioning and map it to the called map material in real time for display. This map can also be arranged together with text and other display content.
[0050] In a preferred embodiment of the present invention, the map materials have different types of display content, and each type of display content has a corresponding display priority; The response content confirmation unit 34 includes: Priority determination subunit 341 is used to determine that the display priority of textual descriptions is always greater than the display priority of map elements, and The display priority of map assets is determined based on the call interval; Multimodal material library 342 contains different display contents of text materials and map materials for different addresses; Call subunit 343 to connect to multimodal material library 342, which is used to call the display content of text material and map material from the multimodal material library according to the calling range; The text and map materials invoked are matched with semantic feature information respectively; The arrangement subunit 344, connected to the calling subunit 343, is used to organize text materials into textual narrative content, and to arrange the textual narrative content and map materials based on display priority. The display content of map materials is arranged based on the display priority corresponding to the call interval.
[0051] Specifically, the priority determination subunit 341 determines that the display priority of the text description content is greater than the display priority of the map material, and then determines the display content of the map material to be called and the display priority among the display content based on the calling interval.
[0052] Furthermore, the calling subunit 343 retrieves text and map materials that match the semantic feature information from the multimodal material library 342. For example, if the semantic feature information includes a field for querying region B, then region B is used as a keyword to search the multimodal material library 342, retrieves the corresponding text material, and determines the display content for region B based on the retrieval range, and retrieves it from the multimodal material library 342.
[0053] Furthermore, the arrangement subunit 344 organizes the text materials called by the calling subunit 343 to form textual narrative content, arranges the textual narrative content before the map materials based on display priority, and sends the arrangement result as reply content to the intelligent reply unit 33.
[0054] In a preferred embodiment of the present invention, the arrangement subunit 344 determines the arrangement method of the text material based on the call interval in order to determine the text narrative content.
[0055] Specifically, each call interval corresponds to different user needs. The arrangement subunit 344 selects content corresponding to the user needs from the retrieved text materials and integrates them, and then processes and polishes them to form textual narrative content that is easy for users to understand.
[0056] Furthermore, when the reply format is text-based, there is no need to consider map materials; simply retrieve the corresponding text materials from the multimodal material library 342 and format them into a text narrative.
[0057] In a preferred embodiment of the present invention, the call interval includes a first call interval, a second call interval, and a third call interval. The numerical range of the first call interval is smaller than the numerical range of the second call interval, the numerical range of the second call interval is smaller than the numerical range of the third call interval, and the first threshold is used as the lower limit of the numerical range of the first call interval. The map material displayed in the first call interval includes thumbnail content; The map materials displayed in the second call interval include thumbnail content and map location content; When the calling interval is determined to be the second calling interval, the priority determination subunit 341 determines that the display priority of the map positioning content is greater than the display priority of the thumbnail content; The map material displayed in the third call interval includes thumbnail content, map location content, and related layer content; When the calling interval is determined to be the third calling interval, the priority determination subunit 341 determines that the display priority of the associated layer content is greater than the display priority of the map positioning content, and the display priority of the map positioning content is greater than the display priority of the thumbnail content.
[0058] Specifically, the priority determination subunit 341 determines the display content and priority of map materials based on the call interval. For example, in a government affairs scenario, the predicted value is 0.65, which is greater than the first threshold for government affairs scenarios and falls within the second call interval [0.6, 0.7]. In this case, the thumbnail content and map coordinate content of the corresponding map material need to be called as the display content. Therefore, the thumbnail and map coordinates are called from the multimodal material library 342. Furthermore, the display priority of the map coordinate content is 0.6, and the display priority of the thumbnail content is 0.4 within this call interval. Therefore, the arrangement subunit 344 arranges the different display content of the map material according to the display priority, prioritizing the display of map coordinates and attaching a thumbnail after the map coordinates.
[0059] Those skilled in the art should recognize that the above is merely a preferred embodiment. In practice, maintenance personnel of the intelligent response system can adjust the specific values of the calling intervals, the displayed map content corresponding to each calling interval, and the display priority between different displayed contents according to the actual situation, and are not limited to the above embodiment.
[0060] Furthermore, the range of values in the calling interval is always greater than the first threshold.
[0061] This invention provides an intelligent reply method applicable to map query scenarios; like Figure 5 As shown, the intelligent reply methods include: Step S1: Extract semantic feature information from multiple dimensions of the natural query language input by the user, and fuse the semantic feature information to construct a multi-dimensional feature vector; The multidimensional feature vector should include at least a map scene vector related to the address query requirement; Step S2: The multidimensional feature vector is processed using a pre-trained prediction model to output the predicted value; Step S3: Determine the arrangement and response method corresponding to the natural query language based on the predicted value; The reply arrangement methods should include at least a multimodal reply mode that uses map materials and combines them with text descriptions.
[0062] Specifically, the system receives natural language query input from the user and parses it. It extracts and saves multi-dimensional semantic feature information, scores or encodes each feature dimension separately to obtain the corresponding feature vector. All feature vectors are then integrated into a multi-dimensional feature vector.
[0063] Furthermore, the multidimensional feature vector is input into the prediction model, which then predicts the multidimensional feature vector based on pre-trained prediction logic and outputs the predicted value.
[0064] The predicted value is within the prediction range.
[0065] Furthermore, the arrangement of responses is determined based on the position of the predicted value within the prediction interval (e.g., the predicted value is 0.7 in the prediction interval [0,1].
[0066] The format of the reply can be either a text-based reply mode containing only textual descriptions, or a multimodal reply mode that combines textual descriptions with map materials.
[0067] In a preferred embodiment of the present invention, the semantic feature information includes at least one of spatial feature information, guidance feature information, action feature information, interaction feature information, and scene feature information; like Figure 6 As shown, step S1 includes: Step S11: Assign values to spatial feature information according to preset address scoring criteria to form a map scene vector, and The guiding feature information is scored and assigned values based on a pre-defined guiding keyword library to form a guiding feature vector. Based on a pre-defined keyword library for actions, scores and assigns values to action feature information to form action feature vectors. The interaction feature information is assigned values based on the user's historical map query count to form an interaction feature vector, and The scene feature information is encoded and assigned values according to the preset scene encoding library to form a scene feature vector; Step S12: Integrate the map scene vector, guidance feature vector, action feature vector, interaction feature vector, and scene feature vector into a multi-dimensional feature vector.
[0068] Specifically, in the process of constructing a multidimensional feature vector, the features of each dimension of semantic feature information are first extracted, and then each dimension of features is scored or encoded.
[0069] Specifically, on the one hand, spatial feature information is scored based on the level of detail of the addresses mentioned in the semantic feature information. The more precise the address, the higher the score. The score of spatial feature information is represented as a spatial feature vector, and a label representing the spatial feature information is attached.
[0070] On the other hand, semantic feature information is matched against a keyword library to determine whether it contains navigation, distribution, boundary, or other guiding keywords. If a match is found, the guiding feature information is scored 1 point; otherwise, it is scored 0 points. The score of the guiding feature information is represented as a guiding feature vector, with an appended label representing the guiding feature information.
[0071] On the other hand, semantic feature information is matched against a database of action-related keywords to determine whether it contains action-related keywords such as "view," "plan," or "compare." If a match is found, the action feature information is scored 1 point; otherwise, it is scored 0 points. The score of the action feature information is represented as an action feature vector, with an appended label indicating the action feature information.
[0072] On the other hand, the ratio of the number of times a user has opened the map in history to the total number of queries is used to represent the interaction feature vector, and a label is attached to represent the interaction feature information.
[0073] On the other hand, semantic feature information is matched with scene coding library, and the corresponding coding is determined according to the scene included in the semantic feature information. The coding is represented as scene feature vector and a label is attached to represent scene feature information.
[0074] Furthermore, the spatial feature vector, guidance feature vector, action feature vector, interaction feature vector, and scene feature vector are integrated into a multi-dimensional feature vector. The subsequent prediction model then uses the labels in this multi-dimensional feature vector to predict the features of each dimension.
[0075] In a preferred embodiment of the present invention, multiple calling intervals are pre-set for the multimodal response mode, and each calling interval corresponds to the modal type that needs to be called; like Figure 7 As shown, step S3 includes: Step S31: Determine the first threshold based on the scene feature vector; Step S32: Determine whether the predicted value is greater than the first threshold. If so, execute the multimodal orchestration response mode, and then proceed to step S33; If not, execute the text response mode that only contains textual descriptions, and then proceed to step S33; Step S33: Reply to the natural language query entered by the user according to the arrangement of the reply method.
[0076] In one embodiment, the initial first threshold is 0.6.
[0077] Specifically, the encoding corresponding to the scene feature in the scene feature vector is obtained, and the first threshold is dynamically adjusted and determined based on the encoding.
[0078] Two specific embodiments are provided to help those skilled in the art better understand the content of responses to users in different scenarios.
[0079] Example 1: The natural language search query entered by users on map query software is: "How do I get to the XX District Administrative Service Center, and what documents do I need to bring?" The first assignment unit 11 parsed the map scene vector to 0.9 (containing the specific address "XX District Administrative Service Center"); the second assignment unit 12 parsed the guidance feature vector to 1 (containing "how to get there"); the third assignment unit 13 parsed the action feature vector to 0; and the fourth assignment unit 14 parsed the interaction feature vector to 0.5 (this user has queried 10 times in the past 30 days, triggering the location). Figure 5 The fifth assignment unit 15 parses the scene feature vector as [1,0,0] (government affairs scene, encoded as [1,0,0]). The vector construction unit 16 integrates the feature vector as [0.9, 1, 0, 0.5, 1,0,0]. The feature vector is input into the trained prediction model, and the model outputs a predicted value of 0.85. The first threshold is determined to be 0.5 based on the government affairs scene. Since the predicted value is greater than the first threshold, the multimodal orchestration response mode is entered. The map material display content is determined to be map location and surrounding traffic layer based on the call interval, and the display priority of map location is greater than that of surrounding traffic layer. The call sub-unit 343 calls the text processing material description, map location, and surrounding traffic layer (including subway station, bus station, etc.) of "XX District Administrative Service Center". The orchestration sub-unit 344 orchestrates the called content. According to the display priority of text description and map material, the text processing material description is organized and displayed first, and then the map location and surrounding traffic layer are displayed in sequence. The intelligent reply unit 33 submits the arrangement result as the reply content to the client.
[0080] Example 2: The natural language search query entered by users on map search software is: "Good Sichuan restaurants nearby".
[0081] The first assignment unit 11 resolves the map scene vector to 0.5 (including "nearby"). The second assignment unit 12 resolves the guiding feature vector to 1 (including "nearby"). The third assignment unit 13 resolves the action feature vector to 0. The fourth assignment unit 14 resolves the interaction feature vector to 0.9 (this user frequently uses the map, querying 20 times in the past 30 days, triggering the map 18 times, accounting for 0.9%). The fifth assignment unit 15 resolves the scene feature vector to [0,1,0] (life service scene, encoded [0,1,0]). The feature vector is [0.5, 1, 0, 0.9, 0,1,0]. The feature vector is input into the trained prediction model, and the model outputs a predicted value of 0.68. Based on the life service scene, the first threshold is determined to be 0.7. Since the predicted value is not greater than the first threshold, it should enter the text reply mode. However, because the predicted value is close to the first threshold and the value corresponding to the interaction feature vector is high, it is determined that the user has potential map needs. Therefore, a backup strategy is adopted, and it still enters the multimodal orchestration reply mode. Based on the call interval being (0.6, 0.7), the map material display content is determined to be only thumbnails. Sub-unit 343 calls text materials of nearby Sichuan restaurants with high ratings, along with thumbnails of the user's location. Sub-unit 344, based on the display priority between text descriptions and map materials, organizes the Sichuan restaurant reviews into a list format and displays it as text descriptions, followed by thumbnails. Smart reply unit 33 submits the arrangement result as a reply to the client.
[0082] The above description is merely a preferred embodiment of the present invention and does not limit the implementation and protection scope of the present invention. Those skilled in the art should realize that any equivalent substitutions and obvious changes made based on the description and illustrations of the present invention should be included within the protection scope of the present invention.
Claims
1. An intelligent reply system, applied in map query scenarios, characterized in that, include: The vector construction module is used to extract semantic feature information from multiple dimensions of the natural query language input by the user, and fuse the semantic feature information to construct a multi-dimensional feature vector; The multidimensional feature vector includes at least a map scene vector associated with the address query requirement; The intent scoring module, connected to the vector construction module, uses a pre-trained prediction model to process the multidimensional feature vector to output a predicted value. The intelligent response module, connected to the intent scoring module, is used to determine the formatted response corresponding to the natural query language based on the predicted value; The proposed response arrangement method includes at least a multimodal response mode that combines map materials with textual descriptions.
2. The intelligent response system according to claim 1, characterized in that, The semantic feature information includes at least one of spatial feature information, guidance feature information, action feature information, interaction feature information, and scene feature information; The vector construction module includes: The first assignment unit is used to assign scores to the spatial feature information according to a preset address scoring standard in order to form the map scene vector; The second assignment unit is used to assign scores to the guidance feature information according to a preset guidance keyword library to form a guidance feature vector; The third assignment unit is used to assign scores to the action feature information according to a preset action keyword library to form an action feature vector; The fourth assignment unit is used to assign values to the interaction feature information based on the number of map queries in the user's history, so as to form an interaction feature vector; The fifth assignment unit is used to encode and assign values to the scene feature information according to a preset scene encoding library to form a scene feature vector; The vector construction unit, which is connected to the first assignment unit, the second assignment unit, the third assignment unit, the fourth assignment unit, and the fifth assignment unit respectively, is used to integrate the map scene vector, the guidance feature vector, the action feature vector, the interaction feature vector, and the scene feature vector into the multidimensional feature vector.
3. The intelligent response system according to claim 2, characterized in that, The intelligent reply module includes: A threshold determination unit is used to determine a first threshold based on the scene feature vector; A response mode determination unit, connected to the threshold determination unit, is used to compare the predicted value with the first threshold and determine the arrangement response mode based on the comparison result; If the predicted value is greater than the first threshold, the multimodal orchestration response mode is executed; If the predicted value is not greater than the first threshold, execute the text reply mode that only contains textual descriptions; The intelligent reply unit, connected to the reply mode determination unit, is used to reply to the natural query language input by the user according to the arrangement reply method determined by the reply mode determination unit.
4. The intelligent response system according to claim 3, characterized in that, Multiple call intervals are pre-set for the multimodal orchestration response mode, and the display content of the map material corresponding to each call interval is different; The intelligent reply module further includes a reply content determination unit, which is connected to the reply mode determination unit and the intelligent reply unit respectively, and is used to determine the call interval that matches the predicted value, thereby determining the display content of the map material; The intelligent reply unit uses the displayed content of the determined map material to determine the reply content corresponding to the multimodal arrangement reply mode.
5. The intelligent response system according to claim 4, characterized in that, The map materials have different types of display content, and each type of display content has a corresponding display priority; The response content determination unit includes: The priority determination subunit is used to determine that the display priority of the text description is always greater than the display priority of the map material, and The display priority of the map material's content is determined based on the calling interval; A multimodal material library, which stores text materials for different addresses and different display content of the map materials; The calling subunit connects to the multimodal material library and is used to call the display content of the text material and the map material from the multimodal material library according to the calling interval; The text and map materials invoked are respectively matched with the semantic feature information; An arrangement subunit, connected to the priority determination subunit and the invocation subunit, is used to organize the text materials into the textual narrative content, and to arrange the textual narrative content and the map materials based on the display priority. The display content of the map materials is arranged based on the display priority of the corresponding call interval.
6. The intelligent response system according to claim 5, characterized in that, The arrangement subunit determines the organization method of the text material based on the calling interval, so as to determine the text narrative content.
7. The intelligent response system according to claim 5, characterized in that, The call interval includes a first call interval, a second call interval, and a third call interval. The numerical range of the first call interval is smaller than the numerical range of the second call interval, the numerical range of the second call interval is smaller than the numerical range of the third call interval, and the first threshold is used as the lower limit of the numerical range of the first call interval. The display content of the map material corresponding to the first call interval includes thumbnail content; The display content of the map material corresponding to the second call interval includes thumbnail content and map location content; When the calling interval is determined to be the second calling interval, the priority determination subunit determines that the display priority of the map positioning is greater than the display priority of the thumbnail; The display content of the map material corresponding to the third call interval includes thumbnail content, map location content, and associated layer content; When the calling interval is determined to be the third calling interval, the priority determination subunit determines that the display priority of the associated layer content is greater than the display priority of the map positioning content, and the display priority of the map positioning content is greater than the display priority of the thumbnail content.
8. An intelligent reply method applied to map query scenarios, characterized in that, The intelligent response system described in any one of claims 1-7 is adopted; The intelligent response method includes: Step S1: Extract semantic feature information from multiple dimensions of the natural query language input by the user, and fuse the semantic feature information to construct a multi-dimensional feature vector; The multidimensional feature vector includes at least a map scene vector associated with the address query requirement; Step S2: The multidimensional feature vector is processed using a pre-trained prediction model to output a predicted value; Step S3: Determine the arrangement and response method corresponding to the natural query language based on the predicted value; The proposed response arrangement method includes at least a multimodal response mode that combines map materials with textual descriptions.
9. The intelligent response method according to claim 8, characterized in that, The semantic feature information includes at least one of spatial feature information, guidance feature information, action feature information, interaction feature information, and scene feature information; Step S1 includes: Step S11: Assign values to the spatial feature information according to a preset address scoring standard to form the map scene vector, and The guiding feature information is scored and assigned values based on a preset guiding keyword library to form a guiding feature vector. The action feature information is scored and assigned values based on a preset action-related keyword library to form an action feature vector. The interaction feature information is assigned values based on the user's historical map query count to form an interaction feature vector. The scene feature information is encoded and assigned values according to a preset scene encoding library to form a scene feature vector; Step S12: Integrate the map scene vector, the guidance feature vector, the action feature vector, the interaction feature vector, and the scene feature vector into the multidimensional feature vector.
10. The intelligent response method according to claim 9, characterized in that, Multiple call intervals are pre-set for the multimodal response mode, and each call interval corresponds to the modal type that needs to be called; Step S3 includes: Step S31: Determine the first threshold based on the scene feature vector; Step S32: Determine whether the predicted value is greater than the first threshold. If so, execute the multimodal orchestration response mode, and then proceed to step S33; If not, execute the text response mode that only contains textual descriptions, and then proceed to step S33; Step S33: Reply to the natural query language input by the user according to the arranged reply method.