An intelligent question and answer implementation method and system based on a large model and a semantic graph

By constructing an intelligent question-answering system based on large models and semantic graphs, the problem of high error rate in spatiotemporal dynamic modeling in the retail supply chain was solved, and accurate mapping of stores on multiple platforms and real-time business fluctuation analysis were achieved, improving the modeling accuracy of the causal relationship of Chenfeng's stockout.

CN120723876BActive Publication Date: 2025-11-28ZHEJIANG PISTACHIO SHUZHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511174146.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-28
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Existing intelligent retail supply chain solutions have bottlenecks in spatiotemporal dynamic modeling, resulting in high error rates in multi-platform store entity alignment methods. Static maps fail to accurately capture real-time business fluctuations and cannot effectively analyze the root causes of Chenfeng's stockouts.

Method used

By constructing an intelligent question-answering system based on large models and semantic graphs, structured and unstructured data are acquired, cleaned and normalized, and then trained using graph neural networks to generate dynamic semantic graphs. These graphs integrate spatiotemporal features, identify spatiotemporal patterns, and generate structured answers.

Benefits of technology

It reduced the multi-platform store mapping error rate, improved the accuracy of the causal relationship modeling of Chenfeng stockouts, and enhanced the accuracy of the spatiotemporal relationship modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723876B_ABST
    Figure CN120723876B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent question and answer implementation method and system based on a large model and a semantic graph, relates to the technical field of supply chain intelligentization, and comprises the following steps: obtaining structured distribution data tables and unstructured text data streams, performing data cleaning and normalization processing, constructing a unified knowledge base, performing semantic understanding and knowledge extraction on the unified knowledge base, extracting key entities and semantic relationships through a named entity recognition and relationship extraction task, and constructing a preliminary semantic graph by fusing structured distribution data features; integrating the preliminary semantic graph with spatial dimension features obtained through a data interface and store inspection assistant real-time task execution logs, giving nodes space-time attributes through graph neural network training, and outputting a dynamic semantic graph integrating space-time dynamic mode features. The application uses Gaode API to inject geographic fence coordinates and crowd density features, superimposes store inspection assistant real-time replenishment state logs, and improves the modeling accuracy of causal correlation strength through ST-GNN training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of supply chain intelligence, and in particular to an intelligent question and answer implementation method and system based on a large model and a semantic graph. BACKGROUND

[0002] The current retail supply chain intelligence field evolves relying on a multi-source data fusion technology framework, and a standardized technical route has been formed: a mainstream general scheme adopts a rule engine to realize structured cleaning and pipeline transmission of cross-platform API data, supports integration of inventory and order data of an O2O platform, realizes entity recognition and relationship extraction of commodity description text in combination with a pre-trained language model, and constructs a static network topology based on a knowledge graph for lack-of-stock problem attribution analysis.

[0003] However, the existing scheme has significant bottlenecks in the time-space dynamic modeling aspect: firstly, a multi-platform store entity alignment method relying on a single feature does not fuse the inverse weight of geographic coordinates and the cross-platform name word frequency feature, resulting in a high platform ID mapping error rate; secondly, the static graph does not access the geographic fence data of the Gaode API and the minute-level state update of the store patrol log, so that the time-space correlation strength modeling accuracy is insufficient and real-time business fluctuations cannot be captured. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides an intelligent question and answer implementation method based on a large model and a semantic graph to solve the problem of excessively high morning peak lack-of-stock root cause attribution error rate caused by insufficient time-space fusion of multi-source retail data.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides an intelligent question and answer implementation method based on a large model and a semantic graph, which includes obtaining a structured distribution data table and an unstructured text data stream, performing data cleaning and normalization processing, and constructing a unified knowledge base.

[0008] The unified knowledge base is subjected to semantic understanding and knowledge extraction, key entities and semantic relationships are extracted through a named entity recognition and relationship extraction task, and structured distribution data features are fused to construct a preliminary semantic graph.

[0009] The preliminary semantic graph is integrated with spatial dimension features and store patrol assistant real-time task execution logs obtained through a data interface, time-space attributes are given to nodes through graph neural network training, and a dynamic semantic graph integrating time-space dynamic mode features is output.

[0010] Based on the natural language question of user input, the large language model is called to analyze the intention and entity, and the dynamic semantic graph retrieval store multi-dimensional dynamic sales data, task execution record and user feedback aggregation data are combined with the integrated space-time dynamic mode characteristics to generate a structured query vector;

[0011] The structured query vector is input into the dynamic semantic graph with integrated space-time dynamic mode characteristics for multi-hop reasoning, traversing the store node analysis state and edge attribute, identifying the space-time mode and associating the user feedback data node to generate a structured answer.

[0012] As a preferred scheme of the intelligent question and answer implementation method based on the large model and semantic graph, wherein: the structured distribution data table includes brand commodity field, real-time on-shelf state field, accurate inventory field, historical sales trend field and prediction data field;

[0013] The unstructured text data stream includes user evaluation text content, product question and answer interaction text and text store record.

[0014] As a preferred scheme of the intelligent question and answer implementation method based on the large model and semantic graph, wherein: the construction of the unified knowledge base includes the following steps,

[0015] The structured distribution data table and the unstructured text data stream are cleaned, and the cleaned structured distribution data table and the cleaned unstructured text data stream are output;

[0016] The similarity of the cleaned structured distribution data is calculated and weighted summed to generate a comprehensive similarity;

[0017] The K-medoids clustering algorithm is executed on the store entity whose comprehensive similarity exceeds the similarity threshold, and the cross-platform entity clustering group is output;

[0018] A globally unique standardized store ID code is assigned to each clustering group, replacing the original store code field in the cleaned structured distribution data table, and a normalized processed structured distribution data table is generated;

[0019] The normalized processed structured distribution data table is associated with the cleaned unstructured text data stream to construct a unified knowledge base.

[0020] As a preferred scheme of the intelligent question and answer implementation method based on the large model and semantic graph, wherein: the construction of the preliminary semantic graph includes the following steps,

[0021] The pre-trained large language model parameters are loaded and the entity recognition task and relationship extraction task are initialized, and the user evaluation content and store text record of the unified knowledge base are extracted, and the standardized text analysis input stream is generated after segmentation and coding processing;

[0022] The standardized text input stream is processed by the embedding layer and the Transformer encoding layer of the large language model, and a structured entity set is generated through entity recognition function analysis and label mapping;

[0023] A semantic relationship set is generated by a relationship extraction function based on the large language model, and the correlation strength is calculated through a self-attention mechanism after fusing the commodity distribution data features to construct a preliminary semantic graph.

[0024] As a preferred scheme of the intelligent question answering implementation method based on the large model and the semantic graph according to the application, the dynamic semantic graph integrating the time-space dynamic mode features comprises the following steps,

[0025] Based on the preliminary semantic graph of the store geographic location, the shopping district label, the geographic fence coordinates and the crowd density data are obtained, and the data are bound to the global store ID field primary key through data analysis and output, and the spatial enhanced semantic graph is output;

[0026] Taking the spatial enhanced semantic graph as input, the global store ID field of the store node is extracted to generate a data request list, a real-time task log is obtained by calling a store inspection assistant interface, a node state attribute is updated through field mapping and a timestamp field is added, and a time-space enhanced dynamic semantic graph is output;

[0027] The store node attribute field is converted into a node feature matrix, an adjacency matrix is generated in combination with the existing problem relationship edges, and a graph convolutional neural network model is iteratively trained to output the dynamic semantic graph integrating the time-space dynamic mode features.

[0028] As a preferred scheme of the intelligent question answering implementation method based on the large model and the semantic graph according to the application, the generation of the structured query vector comprises the following steps,

[0029] After text normalization of the natural language input by the user, tokenization, intent classification and named entity recognition are performed by the large language model to generate an analysis result tuple;

[0030] Taking the regional entity field value in the analysis result tuple as an index key, the corresponding shopping district store node in the dynamic semantic graph integrating the time-space dynamic mode features is matched, all problem type nodes are traversed, and a structured search result is generated;

[0031] Based on the associated store node set in the structured search result, GMV trend, restocking efficiency and user feedback data are extracted to construct a three-dimensional feature vector, and a structured query vector is assembled.

[0032] As a preferred scheme of the intelligent question answering implementation method based on the large model and the semantic graph according to the application, the generation of the structured answer comprises the following steps,

[0033] Extract the corresponding field values ​​according to the structure of the structured query vector field, match the store nodes of the business district in the dynamic semantic graph that integrates spatiotemporal dynamic pattern features and filter the node types, and output a set of high-potential stores.

[0034] Based on a set of high-potential stores, the system traverses the three-level links of products, tasks, and issues, integrates relationship weights and status data, and constructs a complete traversal path tree.

[0035] Based on the traversal path tree, the task completion rate and complaint outbreak coefficient are calculated, and a spatiotemporal pattern matching report is generated.

[0036] Based on the task completion rate and complaint outbreak coefficient, the root cause is output, the estimated GMV loss value is calculated, and a structured answer is generated.

[0037] Secondly, the present invention provides an intelligent question-answering system based on a large model and semantic graph, including a knowledge base construction module, which acquires structured distribution data tables and unstructured text data streams, performs data cleaning and normalization processing, and constructs a unified knowledge base;

[0038] The graph construction module performs semantic understanding and knowledge extraction on the unified knowledge base. It extracts key entities and semantic relationships through named entity recognition and relation extraction tasks, and integrates structured distribution data features to construct a preliminary semantic graph.

[0039] The spatiotemporal enhancement module integrates the preliminary semantic graph with the spatial dimension features obtained from the data interface and the real-time task execution logs of the store patrol assistant. It then uses graph neural network training to assign spatiotemporal attributes to the nodes and outputs a dynamic semantic graph that integrates spatiotemporal dynamic pattern features.

[0040] The vector generation module, based on the natural language questions input by the user, calls a large language model to parse the intent and entities, and combines dynamic semantic graphs that integrate spatiotemporal dynamic pattern features to retrieve multi-dimensional sales data of stores, task execution records and aggregated user feedback data to generate structured query vectors.

[0041] The answer generation module takes the structured query vector as input and integrates it with the dynamic semantic graph of spatiotemporal dynamic pattern features for multi-hop reasoning. It traverses the store nodes to analyze the status and edge attributes, identifies the spatiotemporal patterns and associates them with user feedback data nodes to generate structured answers.

[0042] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the intelligent question answering method based on large models and semantic graphs as described in the first aspect of the present invention.

[0043] In a fourth aspect, the present application provides a computer readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements any step of the method for implementing intelligent question answering based on a large model and a semantic graph according to the first aspect of the present application.

[0044] The present application has the beneficial effects that: the global store ID is generated by K-medoids clustering based on the comprehensive similarity, so that the multi-platform store mapping error rate is reduced; the geographic fence coordinates and crowd density features are injected by using the Gaode API, the real-time replenishment state log of the store inspection assistant is superimposed, and the modeling accuracy of the causal association strength of "breakfast milk shortage -> delivery delay" is improved through ST-GNN training. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0046] Fig. 1 The flowchart of the method for implementing intelligent question answering based on a large model and a semantic graph.

[0047] Fig. 2 The schematic diagram for constructing a unified knowledge base.

[0048] Fig. 3 The schematic diagram for constructing a preliminary semantic graph.

[0049] Fig. 4 The schematic diagram for generating a structured query vector and an answer. DETAILED DESCRIPTION

[0050] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings.

[0051] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the concept of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0052] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.

[0053] Referring to Figs. 1-4 For an embodiment of the present application, the embodiment provides an intelligent question and answer implementation method based on a large model and a semantic graph, comprising the following steps:

[0054] S1. Obtain structured distribution data table and unstructured text data stream, perform data cleaning and normalization processing, and construct a unified knowledge base.

[0055] S1.1. Clean the structured distribution data table and the unstructured text data stream, and output the cleaned structured distribution data table and the cleaned unstructured text data stream.

[0056] Specifically, for multiple O2O platforms such as Meituan, Ele.me and Jingdong to home, API interface is called respectively to initiate programmatic request, and original API response data stream of each O2O platform is output; brand commodity field, real-time on-shelf state field, accurate inventory quantity field, historical sales trend field and prediction data field are extracted from the original API response data stream of each O2O platform, mapped to each store code field, converted to unified store coding rule, and a data set with standardized store code is generated; data table structure is constructed, and the data set with standardized store code is reorganized (such as brand commodity column, real-time on-shelf state column, accurate inventory quantity column, historical sales trend column, and prediction data column), and a structured distribution data table is generated;

[0057] The user evaluation page URL list and the commodity question and answer detail page URL list of multiple O2O platforms such as Meituan, Ele.me and Jingdong to home are configured as the crawler entry point, the request queue of the network crawler framework is initialized, the crawler thread is dispatched to access the corresponding user evaluation page and commodity question and answer detail page at a preset frequency, the web page parsing technology is applied to extract the user evaluation text content and the commodity question and answer interactive text, the text tour records submitted by the business personnel are obtained through the tour assistant applet storage interface, and the continuous user evaluation text content, the commodity question and answer interactive text and the text tour records are respectively organized according to the multiple O2O platforms such as Meituan, Ele.me and Jingdong to home, and a complete unstructured text data stream is formed;

[0058] Load the structured distribution data table into the rule engine of the B1 data cleaning product, enable the regular expression matching function to remove invalid characters in the brand commodity field and the real-time on-shelf state field, standardize the field format, filter out negative values in the accurate inventory quantity field, and abnormal fluctuation values (such as more than 500% of the same period growth) in the historical sales trend field and the prediction data field through the abnormal value filtering rule, and output the cleaned structured distribution data table;

[0059] The user evaluation text content, commodity question and answer interaction text and text store inspection record in the unstructured text data stream are cleaned of HTML tags, advertisement codes and special control symbols, and abnormal value filtering rules are used to remove evaluation records with more than 20 continuous repeated characters and question and answer / store inspection records with more than 30% of random code characters, and the cleaned unstructured text data stream is output.

[0060] S1.2. The cleaned structured distribution data is respectively subjected to similarity calculation and then weighted summation to generate a comprehensive similarity.

[0061] Specifically, the core module of the One-ID algorithm is executed on the store name field, address field and longitude and latitude field of the cleaned structured distribution data table, and the name text similarity, address edit distance score and longitude and latitude spherical distance difference between stores of different O2O platforms such as Meituan, Ele.me and Jingdong to home are respectively calculated by a similarity function.

[0062] Further, the calculation of the name text similarity between stores of different platforms refers to performing Chinese word segmentation and stop word filtering processing on each store name string to generate a standardized store name word vector; the standardized store name word vector is input into a cosine similarity algorithm to calculate the term frequency-inverse document frequency weight value between store names of different O2O platforms such as Meituan, Ele.me and Jingdong to home, and the output is a name text similarity score in the interval of 0-1.

[0063] The calculation of the address edit distance score refers to extracting the address field string value, performing province-city-district administrative unit standardization mapping and road name normalization processing on each address string to generate a formatted address string; the formatted address strings of different O2O platforms are combined into pairs and input into an edit distance algorithm to calculate the minimum single character operation count value between address strings of platforms such as Meituan, Ele.me and Jingdong to home, and the value is converted into an address edit distance score in the interval of 0-1.

[0064] The calculation of the longitude and latitude spherical distance refers to extracting the longitude and latitude field numerical value and converting the decimal system coordinate into spherical coordinate system radian value; the longitude and latitude radian values of different O2O platforms are input into the Haversine formula to calculate the shortest path distance on the ground.

[0065] The inverse function is applied to the longitude and latitude spherical distance to convert it into a 0-1 similarity score (1 / (1+longitude and latitude spherical distance)), and the standardized longitude and latitude similarity is output.

[0066] The name text similarity score (weight 40%), address edit distance score (weight 30%) and standardized longitude and latitude similarity (weight 30%) are weighted and summed to obtain a comprehensive similarity.

[0067] S1.3. Perform K-medoids clustering algorithm on store entities with integrated similarity exceeding similarity threshold, output cross-platform entity clustering groups of Meituan, Ele.me, Jingdong to home and other O2O platform stores.

[0068] Further explanation, performing K-medoids clustering algorithm means scanning the integrated similarity scores of all store entity pairs, retaining store entity pairs with integrated similarity exceeding similarity threshold, generating high-confidence matching entity pair set; randomly select K store entities from the high-confidence matching entity pair set as initial clustering centers (K value is determined adaptively according to the total number of entities), where K is not less than the integer value of the total number of platform stores of Meituan, Ele.me and Jingdong to home divided by 5; assign each store entity to the nearest cluster center (distance is defined as 1-integrated similarity score); reselect the entity with the highest average similarity to other entities in the cluster as the new cluster center in each cluster; repeat the assignment and center selection until the cluster center does not change for three consecutive iterations, output the final stable clustering group structure, each clustering group contains Meituan, Ele.me and Jingdong to home and other O2O platform store entities, and the entities in the clustering group share the same real world store identifier.

[0069] It should be noted that the similarity threshold is set based on the accuracy requirement of cross-platform entity strict matching, aiming to filter out extremely high-confidence entity pairs to ensure the accuracy of multi-platform store clustering.

[0070] S1.4. Assign a globally unique standardized store ID code to each clustering group, replace the original store code field in the cleaned structured distribution data table, and generate a normalized structured distribution data table containing all fields such as brand commodity field, real-time on-shelf state field and accurate inventory field.

[0071] S1.5. Take the global store ID field in the normalized structured distribution data table as the primary key, associate it with the store attribute label (derived from the address field + longitude and latitude field circle type label) and the user evaluation text content and commodity question and answer interaction text in the cleaned unstructured text data stream, and build a unified knowledge base.

[0072] Further explanation, constructing a unified knowledge base refers to mapping brand commodity fields, real-time on-shelf state fields, accurate inventory fields, historical sales trend fields and prediction data fields into columns, constructing a commodity distribution state dimension table; mapping address fields into STRING type columns, longitude and latitude fields into GEO POINT type columns, and commercial district type label fields into STRING type columns, constructing a store basic information dimension table; mapping user evaluation text content fields into TEXT type columns, commodity question and answer interaction text fields into TEXT type columns, and text store inspection record fields into TEXT type columns, constructing a user voice dimension table; creating a commodity distribution state dimension table, a store basic information dimension table and a user voice dimension table in a distributed storage engine, and completing the construction of a unified knowledge base.

[0073] S2. Perform semantic understanding and knowledge extraction on the unified knowledge base, extract key entities and semantic relationships through named entity recognition and relationship extraction tasks, and fuse structured distribution data features to construct a preliminary semantic graph.

[0074] S2.1. Load pre-trained large language model parameters and initialize entity recognition tasks and relationship extraction tasks, while extracting user evaluation content and store text records from the unified knowledge base, and generate standardized text analysis input stream after segmentation and coding processing.

[0075] Specifically, the model parameter file storage path of the pre-trained large language model (such as BERT or similar variants) is obtained; the model parameter file is loaded into the memory, and the inference calculation graph of the basic semantic understanding engine is instantiated; the entity category label set of the named entity recognition task is initialized in the basic semantic understanding engine;

[0076] Further explanation, the entity category label set includes four categories of specific commodity SKU name, store ID, question type and time point.

[0077] It should be noted that the training of the pre-trained large language model is as follows: First, basic training is performed on a massive general corpus using masked language modeling and next-sentence prediction tasks, with the cross-entropy loss function optimized through multiple rounds of iteration using a BERT-like Transformer architecture. Second, retail-specific training is performed, inputting vertical corpora such as cross-platform product descriptions, store inspection reports, and user complaint texts from a unified knowledge base. Entity boundary detection and relation type classification tasks are added on top of the MLM task, and entity position offset loss and relation triple confidence loss are jointly optimized. Finally, spatiotemporal dynamic pattern features are incorporated for graph adaptive training, transforming product distribution data (distribution rate, GMV curve) and user feedback time series into feature vectors and concatenating them to the text embedding layer. Graph topology regularization constraints and relation edge weight modeling are used to enhance the modeling ability of semantic topology in retail scenarios such as "store-product-problem". Finally, a trained large language model that supports complex spatiotemporal pattern parsing is output.

[0078] The process involves: synchronously initializing the relation type template library for relation extraction (including predefined relation patterns such as "existing problem" and "occurred at"); extracting all text records of the user review content field from the user voice dimension table of the unified knowledge base; extracting all text content of the text store visit records from the log storage area of ​​the store visit assistant mini-program associated with the unified knowledge base; merging the text records of the user review content field with the text content of the text store visit record data stream to form an unstructured text data input stream; dividing the unstructured text data input stream into batches of text segments with a fixed length of 512 characters; inputting each batch of text segments into the text parsing interface of the basic semantic understanding engine, performing UTF-8 encoding conversion and special character escaping processing on each text segment to generate a standardized text parsing input stream.

[0079] S2.2. The standardized text input stream is processed through the embedding layer and Transformer encoding layer of the large language model, and a structured entity set is generated by parsing with entity recognition function and label mapping.

[0080] Specifically, each text segment of the standardized text parsing input stream is input into the embedding layer of a pre-trained large language model, converting the text segments into 768-dimensional dense vector representations to generate a text vector sequence. The text vector sequence is then input into the Transformer encoding layer of the pre-trained large language model, where context-aware feature vectors are calculated using a multi-head self-attention mechanism, outputting an enhanced text feature vector sequence. This enhanced text feature vector sequence is then input into the named entity recognition function of the pre-trained large language model, generating a position label sequence through character-level position label prediction (B / I / O label set) based on a Softmax classifier.

[0081] Perform entity classification mapping operation on the location label sequence:

[0082] The continuous character sequence of the label category "BRAND_SKU" is mapped to the specific commodity SKU name entity field and associated with the brand commodity field of the unified knowledge base;

[0083] The label category "STORE_ID" is mapped to the store ID entity field and associated with the global store ID field;

[0084] The label category "ISSUE_TYPE" is mapped to the issue type entity field (such as "commodity damage");

[0085] The label category "TIMESTAMP" is parsed into the minute-level time point entity field through the regular expression "\d{4}-\d{2}-\d{2} \d{2}:\d{2}";

[0086] The structured entity set containing the specific commodity SKU name entity field, the store ID entity field, the issue type entity field, and the time point entity field is output.

[0087] S2.3. The relationship extraction function based on the large language model generates a semantic relationship set, and the associated strength is calculated through the self-attention mechanism after fusing the commodity distribution data features to construct a preliminary semantic graph.

[0088] Specifically, the structured entity set is input into the relationship extraction function of the pre-trained large language model, and the "specific commodity SKU name-store ID-issue type" format is matched through the relationship extraction function to generate triples and verify the matching with the unified knowledge base field, and the "issue type-time point" relationship is identified to generate a relationship pair, and the semantic relationship set is output;

[0089] The semantic relationship set is merged with the store average SKU number field value and the specific period of goods completion rate field value extracted from the unified knowledge base commodity distribution state dimension table to form a composite feature vector, which is input into the feature fusion layer of the pre-trained large language model for weighted integration, and an enhanced feature vector that fuses the text semantic relationship and the structured data feature is output;

[0090] The enhanced feature vector is input into the self-attention mechanism layer of the pre-trained large language model to calculate the semantic correlation score between entity nodes (such as the correlation strength between the "breakfast milk" node and the "Chaoyang district store" node) and the confidence score of the relationship edge (0-1 floating point value), and a semantic correlation score set containing entity correlation strength and relationship confidence is output;

[0091] The specific commodity SKU name entity field, the store ID entity field, the question type entity field, and the time point entity field based on the semantic relationship set are used to create four types of nodes, namely commodity nodes, store nodes, question type nodes, and solution nodes. The commodity node attribute maps the specific commodity SKU name entity field and the unified knowledge base brand commodity field. The store node attribute maps the store ID entity field and the unified knowledge base commercial district type tag field. The question type node attribute maps the question type entity field. The solution node attribute is automatically generated based on a predefined relationship pattern. The time point entity field is mapped to the timestamp attribute, and the relationship confidence score of the semantic correlation degree score set is mapped to the confidence attribute and attached to the relationship edge. The preliminary graph structure containing node attributes and relationship edge metadata is output.

[0092] The preliminary graph structure connects the relationship edge through the graph layout function of the pre-trained large language model: the commodity node is connected to the question type node through the "problem exists" relationship edge, the question type node is connected to the time point attribute through the "occurs in" relationship edge, the store node is connected to the commercial district attribute through the "location belongs to" relationship edge, and the entity correlation strength in the semantic correlation degree score set is used to adjust the edge weight. Finally, the multi-dimensional metadata preliminary semantic graph with timestamp attribute, confidence attribute, and spatial relationship is output.

[0093] S3. The multi-dimensional metadata preliminary semantic graph is integrated with the spatial dimension features obtained through the data interface and the real-time task execution log of the store patrol assistant. The graph neural network is trained to give the node spatio-temporal attributes, and the spatio-temporal enhanced dynamic semantic graph is output.

[0094] S3.1. Based on the store geographic location of the preliminary semantic graph, the commercial district label, geographic fence coordinate, and crowd density data are obtained. After data analysis and global store ID field primary key binding, the spatial enhanced semantic graph is output.

[0095] Specifically, based on the latitude and longitude field value and the store name field value of the store node in the preliminary semantic graph, the store geographic location request parameter list is extracted and combined. The store geographic location request parameter list requests the offline store label field through the Geocoding API of the Gaode map API, requests the geographic fence boundary coordinate field through the GeoFencing API, and requests the commercial district real-time heat map data field through the Heatmap API, and outputs the original geographic data set.

[0096] performing JSON format parsing on the original geographic data set: extracting the "business_area" field value from the geocoding service response as the offline sales force label field, extracting the "fence_vertices" field value from the geofence service response as the geofence boundary coordinate field, and extracting the "density_index" field value from the heat map service response as the commercial district population density field to generate a structured spatial feature data table;

[0097] Taking the global store ID field of the store node in the preliminary semantic graph as the primary key for association mapping, the offline sales force label field of the structured spatial feature data table is bound as the "belonging commercial district" attribute field of the store node, the geofence boundary coordinate field is bound as the "geofence coordinate" attribute field, and the commercial district population density field is bound as the "surrounding population density" attribute field, and a spatially enhanced semantic graph is output.

[0098] S3.2. Taking the spatially enhanced semantic graph as input, extracting the global store ID field of the store node to generate a data request list, calling the store inspection assistant interface to obtain real-time task logs, updating node state attributes through field mapping and adding a timestamp field, and outputting a spatio-temporally enhanced dynamic semantic graph.

[0099] Specifically, the global store ID field values of all store nodes in the spatially enhanced semantic graph are extracted to generate a real-time data request ID list, and the current timestamp field value is obtained as a data query cutoff time point field; the real-time data request ID list and the current timestamp field value request the latest data through the real-time task execution log interface of the store inspection assistant applet: calling the OCR task result interface to obtain the recognition result text field (including the product name field and the on-shelf time field) of the on-shelf complete photo field, calling the inventory inspection interface to obtain the digital inventory status update data field (including the accurate inventory quantity field and the inspection time field), and outputting an original real-time task log data set;

[0100] performing JSON format parsing on the original real-time task log data set: extracting the "product_status" field value from the OCR task result interface response to map it to the on-shelf completion status field, and extracting the "current_stock" field value from the inventory inspection interface response to map it to the latest inventory status field, to generate a structured real-time status data table containing the global store ID field, the on-shelf completion status field, and the latest inventory status field;

[0101] Taking the global store ID field of the store node in the spatially enhanced semantic graph as the primary key, the latest inventory status field value of the structured real-time status data table is matched to update the inventory status attribute field of the store node, the on-shelf completion status field value is matched to update the task completion status attribute field of the store node, and the original state attribute field value is retained for the unmatched nodes.

[0102] After the update, add a "last update time" attribute field to all store nodes (taken from the current timestamp field value), keep the original spatial enhancement semantic graph's belonging to the shopping mall field, geographic fence coordinate field, and surrounding crowd density field unchanged, and output the spatio-temporal enhanced dynamic semantic graph.

[0103] S3.3. Convert the store node attribute field into a node feature matrix, combine the existing problem relationship edge to generate an adjacency matrix, and output the dynamic semantic graph integrated with spatio-temporal dynamic pattern features through iterative training of the graph convolutional neural network model.

[0104] Specifically, the store node attribute fields (belonging to the shopping mall field, surrounding crowd density field, and latest inventory status field) in the spatio-temporal enhanced dynamic semantic graph are converted into numerical vectors to form row vectors of the node feature matrix, and the connection relationship with the "existing problem" relationship edge in the spatio-temporal enhanced semantic graph is generated to generate an adjacency matrix (e.g., the matrix element value of the connected node pair is 1, and the matrix element value of the unconnected node pair is 0);

[0105] Load the graph convolutional layer standard parameter initialization file of the graph convolutional neural network model, input the node feature matrix and the adjacency matrix, and perform iterative training:

[0106] Perform neighbor node feature aggregation operations through the graph convolutional layer function, perform weighted summation on the belonging to the shopping mall field, surrounding crowd density field, and directly connected problem type node features of each store node, and output a new node representation matrix containing spatial correlation features; input the new node representation matrix into the next round of training iteration, repeat the execution, update the node feature matrix with ReLU activation function processing convolution result each time, and update the connection strength of the "existing problem" relationship edge according to the feature propagation result, terminate the training when the loss function change value converges for 10 consecutive iterations, and output the trained dynamic semantic graph;

[0107] In the trained dynamic semantic graph, the store node retains the spatial position attribute field (geographic fence boundary coordinate field), the problem type node adds the time dynamic attribute field (latest inventory status field + task completion status field), and the edge relationship increases the spatio-temporal correlation weight field (e.g., the correlation strength between "breakfast milk shortage problem" and "office building lunch order peak" is 0.92), and outputs the dynamic semantic graph integrated with spatio-temporal dynamic pattern features.

[0108] S4. Based on the user input natural language, call the large language model to parse the intent and entity, and combine the dynamic semantic graph integrated with spatio-temporal dynamic pattern features to retrieve store multi-dimensional dynamic sales data, task execution records, and user feedback aggregated data, and generate a structured query vector.

[0109] S4.1. After text normalization of the natural language input by the user, the big language model is used for word segmentation, intent classification and named entity recognition to generate an analysis result tuple.

[0110] Specifically, the natural language question string input by the user on the front-end interface (such as "Why is there often a shortage of breakfast milk in Chaoyang District?") is processed by removing the leading and trailing spaces and control characters to generate a normalized question string; the normalized question string is input into the word segmentation function of the pre-trained big language model for sub-word segmentation processing to generate a word token ID sequence vector with special markers added, output a word token embedding vector sequence, and perform intent classification; a pooling feature vector is extracted at the special marker position and the probability distribution of the intent category is calculated by a Softmax classifier; the "inventory attribution analysis" intent classification label field with a probability distribution > intent classification probability threshold is selected as the effective output.

[0111] It should be noted that the setting of the intent classification probability threshold is a decision boundary determined based on the ROC curve analysis of the big language model on the retail domain intent classification validation set.

[0112] At the same time, the word token embedding vector sequence is subjected to named entity recognition, and a BIO tagging scheme is used to output an optimized label sequence through a CRF layer to generate a product category entity field = "breakfast milk" and a region entity field = "Chaoyang District" by merging the continuous entity label field; the intent classification label field is mapped to a fixed format string, and the entity field is converted into a key-value pair dictionary structure to generate an analysis result tuple.

[0113] S4.2. Take the region entity field value in the analysis result tuple as the index key to match the associated commercial district store nodes in the dynamic semantic graph integrated with spatiotemporal dynamic pattern features, and traverse all question type nodes to generate a structured search result.

[0114] Specifically, the region entity field value (such as "Chaoyang District") in the analysis result tuple is input as an index key to traverse the "associated commercial district attribute field" of all store nodes in the dynamic semantic graph integrated with spatiotemporal dynamic pattern features, and perform attribute value exact matching to filter out store nodes with an attribute field value equal to "Chaoyang District" to generate a "Chaoyang District associated store node set";

[0115] Based on the global store ID field of all store nodes in the Chaoyang District associated store node set, topological traversal is performed along the "problem existence" relationship edge in the dynamic semantic graph integrated with spatiotemporal dynamic pattern features to find all directly associated question type nodes, and the nodes with a product name attribute field value containing "breakfast milk" in the question type nodes are filtered out to generate a "breakfast milk shortage problem node set";

[0116] The mapping relationship between each store node ID field and corresponding problem node ID field is recorded through the newly created association index table, the independent storage structures of the "Chaoyang District associated store node set" and the "breakfast milk out-of-stock problem node set" are maintained, and the structured retrieval result of the associated store node set containing the global store ID attribute field and the geographic fence coordinate attribute field and the associated problem node set containing the problem type attribute field and the occurrence timestamp attribute field is output.

[0117] S4.3. Based on the associated store node set in the structured retrieval result, extract GMV trend, replenishment efficiency and user feedback data, construct a three-dimensional feature vector, and assemble it into a structured query vector.

[0118] Specifically, taking the associated store node set in the structured retrieval result as input, all global store ID attribute field values are extracted to form a store ID index list; the store ID index list is used to query the commodity distribution state dimension table, the GMV contribution rate trend curve field of the breakfast commodity is retrieved, and the records matching "breakfast milk" in the commodity name field are filtered, and a multi-dimensional dynamic sales data field block is output; the store ID index list is used to call the store patrol assistant task record interface to request the average completion delay rate field of the historical replenishment task and limit the time range, and output the task execution record field block;

[0119] The store ID index list is used to retrieve the user feedback dashboard system aggregation database to extract the "breakfast milk delivery exception" complaint period statistics field and calculate the complaint peak percentage value in the morning period (such as 07:00-09:00), and output the user feedback statistics field block; merge the multi-dimensional dynamic sales data field block, the task execution record field block, and the user feedback statistics field block, and align them by store ID field grouping to form a multi-modal data block;

[0120] The complaint peak percentage value in the morning period is extracted from the user feedback statistics field block of the multi-modal data block as a time dimension feature; the geographic fence coordinate attribute field value of the first store node is extracted from the associated store node set as a spatial dimension feature; the average completion delay rate field value of the historical replenishment task is extracted from the task execution record field block of the multi-modal data block as an execution link feature; and the time dimension feature, the spatial dimension feature, and the execution link feature are encapsulated as a three-dimensional feature vector;

[0121] The time dimension field in the three-dimensional feature vector is renamed as "morning period", the spatial dimension field is renamed as "Chaoyang District commercial circle", and the execution link dimension field is renamed as "replenishment task efficiency", and a structured query vector containing an intent label field (such as "inventory attribution analysis"), a key entity field (such as a commodity category: "breakfast milk", a region: "Chaoyang District"), a time dimension field, a spatial dimension field, and an execution link field is assembled according to a preset field structure.

[0122] It should be noted that the user feedback refers to unstructured text information for capturing real-time feedback, questions and records of users on goods or services, and the user feedback mainly includes three types of unstructured text:

[0123] User evaluation content: User comments or service feedback on O2O platforms (such as Meituan and Eleme), such as evaluation of the quality of "breakfast milk" or the timeliness of delivery.

[0124] Product question and answer interaction text: User and merchant question and answer records on the product detail page, such as asking about the inventory status or complaining about the lack of goods.

[0125] Text store inspection records: On-site inspection reports submitted by business personnel through the store inspection assistant applet, including product on-shelf status or problem records.

[0126] S5. Input the structured query vector into the integrated spatiotemporal dynamic pattern feature dynamic semantic graph for multi-hop reasoning, traverse the store node analysis state and edge attribute, identify the spatiotemporal pattern and associate the user feedback data node, and generate a structured answer.

[0127] S5.1. According to the field structure of the structured query vector, extract the corresponding field value, match the corresponding store node in the integrated spatiotemporal dynamic pattern feature dynamic semantic graph, and filter the node type, and output the high-potential store set.

[0128] Specifically, according to the field structure of the structured query vector, extract the region entity field value as the spatial index key, the time dimension field value as the time window identifier, and the product category entity field value as the target product identifier; In the integrated spatiotemporal dynamic pattern feature dynamic semantic graph, scan the attribute field value of the store node belonging to the commercial circle, and filter out the store node whose attribute field value is equal to the spatial index key through string complete matching, to generate a set of region matching store nodes; Perform attribute filtering on the set of region matching store nodes, check the node type attribute field value of each store node, and output a set of high-potential store nodes.

[0129] S5.2. Based on the high-potential store set, traverse the three-level link of goods, tasks and questions, integrate the relationship weight and state data, and build a complete traversal path tree.

[0130] Specifically, take the high-potential store node set as the input starting point, traverse the outgoing relationship edge of each store node, filter the relationship edge with the relationship type of "commodity-store" and connect the target commodity node along the edge, and output the commodity node traversal set; then take the high-potential store node set as the starting point to traverse the outgoing relationship edge, filter the relationship edge with the relationship type of "store-task" and connect the store task node along the edge, and output the task node traversal set; then take the task node traversal set as the starting point to traverse the outgoing relationship edge, filter the relationship edge with the relationship type of "task-question" and connect the distribution question node along the edge, and output the question node traversal set;

[0131] In the traversal process, the space-time correlation weight field values of the "commodity-store", "store-task", and "task-question" relationship edges are recorded synchronously, the inventory state attribute field values of the commodity nodes on the connection path and the task completion state attribute field values of the task nodes are copied, and a path record containing the relationship edge weight list, the inventory state list, and the task state list is generated; finally, all path records are grouped according to the high-potential store node ID field, a tree structure is constructed with the high-potential store node as the root node, the target commodity node as the first-level child node (connected through the "commodity-store" edge), the store task node as the second-level child node (connected through the "store-task" edge), and the distribution question node as the third-level child node (connected through the "task-question" edge), and a complete traversal path tree carrying path record attribute data is output.

[0132] S5.3. Calculate the task completion rate and the complaint outbreak coefficient based on the traversal path tree, and generate a space-time pattern matching report.

[0133] Specifically, the task completion state attribute field values and the occurrence timestamp attribute field values of all second-level child nodes in the complete traversal path tree are extracted, the task nodes with the occurrence timestamp attribute field values in the "07:00-09:00" interval are filtered, and a morning period task node set is generated; the number of nodes with the task completion state attribute field value of "True" in the morning period task node set is counted as the completion number, the total number of nodes in the set is counted as the total number, and the task completion rate field value is calculated.

[0134] The task completion rate field value = (completion number / total number x 100%);

[0135] The corresponding first-level child node with a label task completion rate field value lower than a task completion rate threshold value is a "promotion goods not timely on-shelf" gap node, and a gap node list and a completion rate value are output; problem type attribute field values and occurrence timestamp attribute field values of all third-level child nodes in the complete traversal path tree are extracted, problem nodes with timestamp field values in the "07:00-09:00" interval are screened, and a morning period problem node set is generated; the number of nodes with a problem type attribute field value equal to "breakfast period delivery delay" in the morning period problem node set is counted as a complaint node number, the total node number is counted as a total problem node number, and a complaint outbreak coefficient is calculated;

[0136] The complaint outbreak coefficient = (complaint node number / total problem node number);

[0137] It should be noted that the task completion rate threshold value is set based on the business standard value of the retail industry, such as 85%, which is a rigid baseline to ensure that the on-shelf time of promotion goods meets the breakfast consumption demand peak.

[0138] The "promotion goods not timely on-shelf" gap node field, the morning period task completion rate field, and the "breakfast period delivery delay" complaint outbreak coefficient field are integrated to generate a spatiotemporal pattern matching report.

[0139] S5.4. Based on the task completion rate and the complaint outbreak coefficient, the root cause is output, the estimated GMV loss value is calculated, and a structured answer is generated.

[0140] Specifically, based on the spatiotemporal pattern matching report, the morning period task completion rate field value and the complaint outbreak coefficient field value are extracted, and when the morning period task completion rate < task completion rate threshold value and the complaint outbreak coefficient > complaint outbreak coefficient threshold value, the root cause field string "insufficient morning inspection coverage of business personnel" is output; the "breakfast milk" GMV contribution rate change curve field is retrieved from the commodity distribution state dimension table through the global store ID field corresponding to the gap node, the estimated GMV loss value is calculated, and the quantitative impact field string "estimated GMV loss" is output.

[0141] The estimated GMV loss value = (historical same period morning period average - current value) / historical same period morning period average x 100%;

[0142] It should be noted that the complaint outbreak coefficient threshold value is a causal strong correlation critical point determined through historical event attribution analysis, such as 0.7, and when the proportion of breakfast period delivery delay complaints is > 0.7, it is determined as a systematic service failure.

[0143] Retrieve the global store ID field for all gap nodes and generate the solution field string "Store patrol schedule to be forcibly adjusted to start at 7:00 AM"; combine the root cause field, quantified impact field, and solution field into a structured answer framework for output; use the global store ID field from the gap node list as input to call the store list database interface to extract the relevant store name list field for output, call the store patrol task dashboard interface to obtain the real-time screenshot field of the task execution efficiency dashboard for output, and simultaneously use the morning time period and the global store ID field of the gap nodes as input to obtain the morning time period complaint distribution heat map field from the user feedback dashboard system for output; combine the store name list field, task efficiency screenshot field, and complaint heat map field into a supporting attachment field dictionary for output; merge the structured answer framework and the supporting attachment field dictionary into a structured answer.

[0144] This embodiment also provides an intelligent question-answering system based on a large model and semantic graph, including:

[0145] The knowledge base construction module acquires structured distribution data tables and unstructured text data streams, performs data cleaning and normalization processing, and builds a unified knowledge base.

[0146] The graph construction module performs semantic understanding and knowledge extraction on the unified knowledge base. It extracts key entities and semantic relationships through named entity recognition and relation extraction tasks, and integrates structured distribution data features to construct a preliminary semantic graph.

[0147] The spatiotemporal enhancement module integrates the preliminary semantic graph with the spatial dimension features obtained from the data interface and the real-time task execution logs of the store patrol assistant. It then uses graph neural network training to assign spatiotemporal attributes to the nodes and outputs a dynamic semantic graph that integrates spatiotemporal dynamic pattern features.

[0148] The vector generation module, based on the natural language questions input by the user, calls a large language model to parse the intent and entities, and combines dynamic semantic graphs that integrate spatiotemporal dynamic pattern features to retrieve multi-dimensional sales data of stores, task execution records and aggregated user feedback data to generate structured query vectors.

[0149] The answer generation module takes the structured query vector as input and integrates it with the dynamic semantic graph of spatiotemporal dynamic pattern features for multi-hop reasoning. It traverses the store nodes to analyze the status and edge attributes, identifies the spatiotemporal patterns and associates them with user feedback data nodes to generate structured answers.

[0150] This embodiment also provides a computer device applicable to the implementation of intelligent question answering methods based on large models and semantic graphs, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the intelligent question answering method based on large models and semantic graphs as proposed in the above embodiment.

[0151] The computer device can be a terminal, which includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved by WIFI, operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0152] The embodiment also provides a storage medium having a computer program stored thereon, the program being executed by a processor to implement the method for implementing intelligent question answering based on a large model and a semantic graph as described above. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.

[0153] To sum up, the present application: by synthesizing the similarity, generating global store ID through K-medoids clustering, reducing the multi-platform store mapping error rate; using Gaode API to inject geographic fence coordinates and crowd density features, superimposing real-time replenishment state logs of store inspection assistants, and training ST-GNN to improve the modeling accuracy of the causal association strength of "breakfast milk shortage → delayed distribution".

[0154] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced, without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.

Claims

1. A method for implementing intelligent question answering based on large models and semantic graphs, characterized in that: include, Acquire structured distribution data tables and unstructured text data streams, perform data cleaning and normalization, and build a unified knowledge base; Semantic understanding and knowledge extraction are performed on the unified knowledge base. Key entities and semantic relationships are extracted through named entity recognition and relation extraction tasks, and the features of structured distribution data are integrated to construct a preliminary semantic graph. The initial semantic graph is integrated with spatial dimension features obtained from the data interface and real-time task execution logs from the store patrol assistant. Nodes are then trained using a graph neural network to assign spatiotemporal attributes, outputting a dynamic semantic graph that integrates spatiotemporal dynamic pattern features. This process includes the following steps: Based on the store's geographical location from the preliminary semantic graph, we obtain business district tags, geofence coordinates, and crowd density data. After data parsing and binding with the primary key of the global store ID field, we output a spatially enhanced semantic graph. Using a spatially enhanced semantic graph as input, the global store ID field of the store node is extracted to generate a data request list. The store patrol assistant interface is called to obtain real-time task logs. The node status attributes are updated through field mapping and a timestamp field is added. The spatiotemporally enhanced dynamic semantic graph is output. The store node attribute fields are transformed into node feature matrices, and adjacency matrices are generated by combining problematic relationship edges. Through iterative training of a graph convolutional neural network model, a dynamic semantic graph integrating spatiotemporal dynamic pattern features is output. Based on the natural language questions input by the user, the large language model is invoked to parse the intent and entities, and combined with the dynamic semantic graph that integrates spatiotemporal dynamic pattern features to retrieve multi-dimensional sales data of stores, task execution records and aggregated user feedback data, and generate structured query vectors. The structured query vector is input into a dynamic semantic graph that integrates spatiotemporal dynamic pattern features for multi-hop reasoning. The state and edge attributes of store nodes are analyzed, spatiotemporal patterns are identified, and user feedback data nodes are associated to generate structured answers.

2. The intelligent question answering method based on large models and semantic graphs as described in claim 1, characterized in that: The structured distribution data table includes fields for brand products, real-time listing status, precise inventory, historical sales trends, and forecast data. The unstructured text data stream includes user review content, product Q&A interaction text, and text store visit records.

3. The intelligent question answering method based on large models and semantic graphs as described in claim 1, characterized in that: The construction of the unified knowledge base includes the following steps. Clean the structured distribution data table and the unstructured text data stream, and output the cleaned structured distribution data table and the cleaned unstructured text data stream. The similarity of the cleaned structured distribution data is calculated and weighted summed to generate a comprehensive similarity score. For store entities whose overall similarity exceeds the similarity threshold, perform the K-medoids clustering algorithm to output cross-platform entity cluster groups; Assign a globally unique standardized store ID code to each cluster group, replace the original store code field in the cleaned structured distribution data table, and generate a normalized structured distribution data table. By associating the normalized structured distribution data table with the cleaned unstructured text data stream, a unified knowledge base is constructed.

4. The intelligent question answering method based on large models and semantic graphs as described in claim 1, characterized in that: The construction of the preliminary semantic graph includes the following steps. Load the parameters of the pre-trained large language model and initialize the entity recognition and relation extraction tasks. At the same time, extract user evaluation content and store visit text records from the unified knowledge base, and generate a standardized text parsing input stream through segmentation and encoding. The standardized text input stream is processed through the embedding layer and Transformer encoding layer of the large language model, and a structured entity set is generated by parsing through entity recognition function and label mapping. A semantic relation set is generated based on the relation extraction function of the large language model. After integrating the features of commodity distribution data, the association strength is calculated through a self-attention mechanism to construct a preliminary semantic graph.

5. The intelligent question answering method based on large models and semantic graphs as described in claim 1, characterized in that: The generation of structured query vectors includes the following steps. After normalizing the natural language input by the user, the parsed result tuple is generated through word segmentation, intent classification, and named entity recognition using a large language model. Using the regional entity field value in the parsed result tuple as the index key, the store node in the business district is matched in the dynamic semantic graph that integrates spatiotemporal dynamic pattern features, and all question type nodes are traversed to generate structured search results. Based on the set of associated store nodes in the structured search results, GMV trends, replenishment efficiency, and user feedback data are extracted, a three-dimensional feature vector is constructed, and then assembled into a structured query vector.

6. The intelligent question answering method based on large models and semantic graphs as described in claim 1, characterized in that: The generation of structured answers includes the following steps. Extract the corresponding field values ​​according to the structure of the structured query vector field, match the store nodes of the business district in the dynamic semantic graph that integrates spatiotemporal dynamic pattern features and filter the node types, and output a set of high-potential stores. Based on a set of high-potential stores, the system traverses the three-level links of products, tasks, and issues, integrates relationship weights and status data, and constructs a complete traversal path tree. Based on the traversal path tree, the task completion rate and complaint outbreak coefficient are calculated, and a spatiotemporal pattern matching report is generated. Based on the task completion rate and complaint outbreak coefficient, the root cause is output, the estimated GMV loss value is calculated, and a structured answer is generated.

7. An intelligent question-answering system based on large models and semantic graphs, based on the intelligent question-answering method based on large models and semantic graphs as described in any one of claims 1 to 6, characterized in that: include, The knowledge base construction module acquires structured distribution data tables and unstructured text data streams, performs data cleaning and normalization processing, and builds a unified knowledge base. The graph construction module performs semantic understanding and knowledge extraction on the unified knowledge base. It extracts key entities and semantic relationships through named entity recognition and relation extraction tasks, and integrates structured distribution data features to construct a preliminary semantic graph. The spatiotemporal enhancement module integrates the preliminary semantic graph with the spatial dimension features obtained from the data interface and the real-time task execution logs of the store patrol assistant. It then uses graph neural network training to assign spatiotemporal attributes to the nodes and outputs a dynamic semantic graph that integrates spatiotemporal dynamic pattern features. The vector generation module, based on the natural language questions input by the user, calls a large language model to parse the intent and entities, and combines dynamic semantic graphs that integrate spatiotemporal dynamic pattern features to retrieve multi-dimensional sales data of stores, task execution records and aggregated user feedback data to generate structured query vectors. The answer generation module takes the structured query vector as input and integrates it with the dynamic semantic graph of spatiotemporal dynamic pattern features for multi-hop reasoning. It traverses the store nodes to analyze the status and edge attributes, identifies the spatiotemporal patterns and associates them with user feedback data nodes to generate structured answers.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent question answering method based on large models and semantic graphs as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent question answering method based on large models and semantic graphs as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Knowledge graph-based medicine supply chain question-answering method and system

    CN117251543A

  • Knowledge-driven underground space information retrieval method, system and equipment

    CN120216612A