A large model-oriented multi-source data capability hierarchical discovery and joint calling method, system, device and medium
Patent Information
- Application Number
- CN202610717465.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-15
Smart Images

Figure CN122759136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and in particular to a method, system, device, and medium for hierarchical discovery and joint invocation of multi-source data capabilities for large models. Background Technology
[0002] As large language models are increasingly used in scenarios such as personal data question answering, intelligent assistants, multi-source information integration, consumption analysis, content preference analysis, and to-do list identification, models relying solely on parameter memorization can no longer meet the demands for real-time, accurate, and verifiable data acquisition. Therefore, more and more systems are opening up external data access capabilities, business interface capabilities, and retrieval tool capabilities to large models through MCP (Model Context Protocol) or similar tool call protocols. In such systems, the large model typically decides whether to call a tool, which tool to call, and how to construct parameters based on the user's question. The server then executes the corresponding tool and returns structured results, which the model ultimately uses to generate an answer.
[0003] Current multi-source data platforms generally adopt the following organization method: 1. Various types of data are stored in corresponding data tables or corresponding storage units; 2. Each type of data is accessed by one or more corresponding query interfaces; 3. The model indirectly accesses the underlying data records by calling the interfaces. This method is simple and direct in engineering implementation and is also easy to extend functionality according to data sources, therefore it is one of the mainstream implementation methods.
[0004] However, as the types of data, data sources, and number of interfaces that are accessed continue to grow, at least the following new technical contradictions will be exposed in the large model calling scenario: the number of interfaces expands rapidly with the expansion of data sources, the tool decision space of the model is too large and the model selection complexity is high, the prompt context becomes longer, the call error rate increases, and ultimately the accuracy of the large model's answer is low, which affects the user's experience of using the large model on the user terminal. Summary of the Invention
[0005] Based on the above-mentioned technical problems, the present invention provides a method, system, device and medium for hierarchical discovery and joint invocation of multi-source data capabilities for large models, aiming to overcome the above problems or at least partially solve the above problems.
[0006] The first aspect of this invention provides a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models, deployed in a personal data management application on a user terminal, the method comprising: The current round of question text input by the user on the user terminal is obtained through the user interaction page of the personal data management application; Each topic in the topic layer of the personal data management application is exposed to the large model. The large model then matches the semantics of each topic with the semantics of the current round of question text, and identifies the target topic that semantically matches the current round of question text from each topic. If the number of candidate interfaces that semantically match the target topic is greater than the number of interfaces, and the number of candidate tags that semantically match the target topic is greater than the number of tags, then each candidate interface in the interface layer of the personal data management application is exposed to the large model, and each candidate tag in the tag layer of the personal data management application is exposed to the large model. The large model retrieves data records with at least one candidate label from the multi-source data layer of the personal data management application based on the candidate labels. The large model also calls various candidate interfaces to read data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. Using the large model, based on data records with at least one candidate label and data records that semantically match the current round of question text, the answer text corresponding to the current round of question text is generated, and the answer text corresponding to the current round of question text is displayed on the user interaction page of the personal data management application.
[0007] A second aspect of the present invention provides a system for hierarchical discovery and joint invocation of multi-source data capabilities for large models, the system comprising: The text acquisition module is used to acquire the current round of question text input by the user on the user terminal through the user interaction page of the personal data management application; The first matching module is used to expose each topic in the topic layer of the personal data management application to the large model, and then match the semantics of each topic with the semantics of the current round of question text through the large model, and determine the target topic that matches the semantics of the current round of question text from each topic. The data exposure module is used to expose each candidate interface in the interface layer of the personal data management application to the large model, and to expose each candidate tag in the tag layer of the personal data management application to the large model when the number of each candidate interface that semantically matches the target topic is greater than the number of interfaces and the number of each candidate tag that semantically matches the target topic is greater than the number of tags. The record reading module is used to retrieve data records with at least one candidate label from the multi-source data layer of the personal data management application based on the candidate labels using the large model, and to call the candidate interfaces through the large model to read data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. The text generation module is used to generate the answer text corresponding to the current round of questions by using the large model, based on data records with at least one candidate label and data records that semantically match the current round of questions, and to display the answer text corresponding to the current round of questions through the user interaction page of the personal data management application.
[0008] A third aspect of the present invention provides an electronic device comprising a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method for hierarchical discovery and joint invocation of multi-source data capabilities for large models as described in the first aspect of the present invention.
[0009] The fourth aspect of the present invention provides a readable storage medium on which a program or instruction is stored, which, when executed by a processor, implements the steps of the multi-source data capability hierarchical discovery and joint invocation method for large models as described in the first aspect of the present invention.
[0010] In the proposed method for multi-source data capability hierarchical discovery and joint invocation for large models, for personal data management applications on user terminals, the current round of question text input by the user is obtained through the user interaction page. The current round of question text first flows into the topic layer of the personal data management application, exposing each topic in the topic layer to the large model. The large model matches the semantics of each topic with the semantics of the current round of question text to determine the target topic that semantically matches the current round of question text. Next, the number of candidate interfaces and candidate tags that match the target topic is judged to determine the exposure mechanism of the interface layer and / or tag layer of the personal data management application. Specifically, in each candidate interface... If the number of candidate interfaces exceeds the threshold and the number of candidate tags exceeds the threshold for the number of tags, then each candidate interface in the interface layer and each candidate tag in the tag layer is exposed to the large model. The large model then searches the multi-source data layer of the personal data management application based on each candidate tag to obtain data records with at least one candidate tag, and calls each candidate interface to read data records from the multi-source data layer that semantically match the current round of question text. Finally, based on the data records with at least one candidate tag and the data records that semantically match the current round of question text, the large model generates the answer text corresponding to the current round of question text and displays the answer text through the user interaction page.
[0011] Thus, this invention addresses the scenario of large models calling multi-source data tools in user terminals. Through a hierarchical capability organization structure built within a personal data management application—comprising a topic layer, interface layer, tag layer, and multi-source data layer—the problem text first flows into the topic layer to determine the target topic. Then, based on the number of candidate interfaces and / or candidate tags corresponding to the target topic, the visibility range of the large model in the interface and / or tag layers is controlled. This allows the large model to make controllable selections, joint searches, and / or unified calls between interface and tag paths, without having to face all interfaces and tags initially. Finally, in the multi-source data layer, the same data record can be hit through multiple interfaces and / or multiple usage scenario entry points, avoiding data omissions that could lead to biased answers from the large model. This solves the problem of model call complexity caused by an excessive number of interfaces and the problem of missed detections due to multiple usage scenarios for a single data entry, improving the accuracy of the large model's answers and the user experience of using the large model on the user terminal. This achieves personal data management and intelligent recommendation on the user terminal. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating the steps of a multi-source data capability hierarchical discovery and joint invocation method for large models, as shown in an embodiment of the present invention. Figure 2 This is a timing diagram illustrating a shared topic-driven joint discovery method according to an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating a data retrieval mode according to an embodiment of the present invention; Figure 4 This is a structural block diagram of a multi-source data capability hierarchical discovery and joint invocation system for large models provided by an embodiment of the present invention; Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] Under the organizational method described above, as the number of data sources, data tables, and business domains increases, the number of interfaces exposed by the server to the large model will increase rapidly. If all these interfaces are provided to the model at once, the model needs to filter and construct parameters among a large number of interfaces with different names, descriptions, and parameter structures (for example, in a system where interfaces are designed by table, source, or business module, the server can easily create a large number of independent API interfaces. If all these API interface tools are exposed to the large model at once, the model must simultaneously understand: what problem each tool is suitable for, how to fill in the parameters of each tool, how to choose between multiple similar tools, and which tools come from different sources but can jointly support the same problem, etc.). This results in an excessively large tool decision space for the model, longer prompt contexts, higher call error rates, and a lack of effective exposure convergence mechanism on the server side. Therefore, this direct interface exposure method that only organizes data calls from the interface path will produce the "large toolface" problem.
[0016] In addition to the issues mentioned above, this invention also found that current systems typically assume that the "data source semantics" and "data usage scenario semantics" are basically consistent. However, in real-world use, the same data often has multiple semantic uses. For example, a data record might contain "I want to go to place A recently and need to buy refrigerator magnets," corresponding to shopping needs, to-do items, and corresponding travel arrangements. From the source (i.e., interface) perspective, this record belongs to chat data from a client application; but from the usage scenario perspective, this record has three uses simultaneously: shopping, to-do items, and travel. If the system only allows large models to access this record through the "social data interface," when users ask questions like "What do I need to buy recently?", "What are my to-do items recently?", or "What are my recent travel plans?", the large model might only select paths related to shopping, to-do items, or travel, without actively associating it with the social interface, thus causing highly relevant data to be missed. In real-world business scenarios, users do not ask questions based on the "data source module," but rather on the "task objective" or "usage scenario." For example: What do I need to buy recently? What are my to-do items recently? What content am I mainly focusing on recently? Where might I travel to recently? However, the source semantics of a record are not the same as its usage scenario semantics. A chat record, though originating from chat data, may be used for shopping, to-do lists, or travel scenarios; a browsing record, though originating from a video platform, may reflect travel intentions, health concerns, or consumption preferences. If the system still categorizes data rigidly based on its source interface, the model will struggle to retrieve the data from different scenario entry points. Therefore, simply binding interfaces based on source interfaces or single topics cannot support the reuse requirement of "the same data having multiple usage scenarios," nor can it allow large models to retrieve the same data through different semantic entry points.
[0017] Therefore, this invention finds that currently, on the one hand, it is necessary to solve the problem of "too many interfaces and difficulty in direct exposure," and on the other hand, it is also necessary to solve the problem of "single data source path and insufficient semantic reuse in scenarios." If two separate mechanisms are designed separately, the system complexity will further increase. Moreover, this invention finds that simply expanding the number of interfaces cannot solve the problem of multi-semantic data reuse; simply tagging data may also reintroduce the problem of an excessively large candidate space after the number of tags increases. Therefore, this invention believes it is necessary to provide a unified hierarchical organizational structure and call control mechanism, so that interface paths and tag paths can work together under the same framework, which can both retain the interface-level structured access capability and support the joint call method of tag-level cross-source semantic retrieval capability.
[0018] Based on this, in order to at least partially solve one or more of the above-mentioned problems and other potential problems, this invention proposes a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models. In this method, for scenarios where large models in a user terminal invoke multi-source data tools, a hierarchical capability organization structure consisting of a topic layer, interface layer, tag layer, and multi-source data layer is constructed in a personal data management application. The problem text first flows into the topic layer to determine the target topic corresponding to the problem text. Then, based on the number of candidate interfaces and / or candidate tags corresponding to the target topic, the visible capability range of the large model in the interface layer and / or tag layer is controlled, thereby enabling… The large model allows for controlled selection, joint retrieval, and / or unified invocation between interface paths and tag paths. Furthermore, the large model doesn't need to initially face all interfaces and tags. Finally, in the multi-source data layer, the same data record can be hit through multiple interfaces and / or multiple usage scenario entry points, avoiding data omissions that could lead to biased answers from the large model. This solves the model invocation complexity problem caused by an excessive number of interfaces, as well as the omission problem caused by multiple usage scenarios for a single data point, improving the accuracy of the large model's answers. Additionally, it enhances the user experience of the large model on the user terminal, enabling personal data management and intelligent recommendations on the user terminal.
[0019] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models, as shown in an embodiment of the present invention. Figure 1 As shown, the multi-source data capability hierarchical discovery and joint invocation method for large models provided in this embodiment includes at least the following steps: Step S11: Obtain the current round of question text input by the user on the user terminal through the user interaction page of the personal data management application.
[0020] The multi-source data capability hierarchical discovery and joint invocation method for large models provided in this embodiment is deployed in a personal data management application on a user terminal. This personal data management application is used to manage and recommend user data on the user terminal. Furthermore, the user can only install and register the personal data management application on the user terminal after agreeing to its access to relevant user data. In other words, in this embodiment, all data obtained by the personal data management application is obtained with the user's consent, complying with relevant laws, regulations, and privacy requirements.
[0021] Users on the user terminal can enter question text or task description in natural language on the user interaction page of the personal data management application. Based on the input question text or task description, the personal data management application calls the big model to call and retrieve relevant data records of various client applications on the user terminal in order to answer the question text or task description and output the answer.
[0022] Specifically, users can also generate answers to multi-round questions in the personal data management application. For the current round of questions in the multi-round question, the personal data management application can obtain the text of the current round of questions input by the user on the user terminal through the user interaction page of the personal data management application.
[0023] Step S12: Expose each topic in the topic layer of the personal data management application to the large model, and match the semantics of each topic with the semantics of the current round of question text through the large model, and determine the target topic that matches the semantics of the current round of question text from each topic.
[0024] In this embodiment, after the personal data management application obtains the current round of question text, it can provide the current round of question text to the corresponding large model. Furthermore, the personal data management application can feed the current round of question text into the topic layer to expose various topics in the topic layer to the large model. After obtaining each topic, the large model can match the semantics of each topic with the semantics of the current round of question text, identifying one or more topics that semantically match the current round of question text and designating them as target topics.
[0025] Step S13: If the number of candidate interfaces that semantically match the target topic is greater than the number of interfaces, and the number of candidate tags that semantically match the target topic is greater than the number of tags, expose each candidate interface in the interface layer of the personal data management application to the large model, and expose each candidate tag in the tag layer of the personal data management application to the large model.
[0026] In this embodiment, each topic in the topic layer corresponds to one or more interfaces that match its semantics, and each topic also corresponds to one or more tags that match its semantics. After determining the target topic for processing, the personal data management application can identify interfaces that match the semantics of the target topic as candidate interfaces, and tags that match the semantics of the target topic as candidate tags. The semantics of the interfaces refer to data source semantics, and the semantics of the tags refer to usage scenario semantics.
[0027] This embodiment can control the visibility range of a large model at the interface and / or tag layers based on the number of candidate interfaces and / or candidate tags corresponding to the target topic. This visibility range characterizes whether the interface and / or tag layers expose all interfaces and / or tags, or only some, to the large model. Specifically, this embodiment pre-sets interface and tag number thresholds for determining the visibility range. The specific values of these thresholds can be freely set according to requirements, without any limitations.
[0028] This embodiment compares the number of candidate interfaces semantically matching the target topic with an interface number threshold; and compares the number of candidate tags semantically matching the target topic with a tag number threshold. If it is determined that the number of candidate interfaces semantically matching the target topic is greater than the interface number threshold, and the number of candidate tags semantically matching the target topic is greater than the tag number threshold, the data access mode can be determined to be a shared topic-driven joint discovery method. In this shared topic-driven joint discovery method, the visibility capabilities of both the tag layer and the interface layer are partially exposed: specifically, the personal data management application exposes each candidate interface in its interface layer to the large model, and exposes each candidate tag in its tag layer to the large model.
[0029] Step S14: The large model retrieves data records with at least one candidate label from the multi-source data layer of the personal data management application based on the candidate labels. The large model then calls each candidate interface to read data records that semantically match the current round of question text from the multi-source data layer of the personal data management application.
[0030] In this embodiment, after the large model obtains each candidate label, it can retrieve data records in the multi-source data layer of the personal data management application based on each candidate label to obtain data records with at least one candidate label. Specifically, if a data record in the multi-source data layer has at least one label that matches the candidate label, then that data record is determined to have at least one candidate label. Furthermore, after the large model obtains each candidate interface, it can call each candidate interface to read data records from the multi-source data layer of the personal data management application that semantically match the current round of question text.
[0031] Step S15: Using the large model, based on data records with at least one candidate label and data records that semantically match the current round question text, generate the answer text corresponding to the current round question text, and display the answer text corresponding to the current round question text through the user interaction page of the personal data management application.
[0032] In this embodiment, after the large model obtains data records with at least one candidate label and data records that semantically match the current round of question text, it can generate the answer text corresponding to the current round of question text based on the data records with at least one candidate label and the data records that semantically match the current round of question text. The answer text corresponding to the current round of question text is then displayed through the user interaction page of the personal data management application to complete the question and answer for the current round, thereby realizing personal data management and / or recommendation for the user terminal.
[0033] In this embodiment, for scenarios where large models in user terminals call multi-source data tools, a hierarchical capability organization structure of topic layer, interface layer, tag layer, and multi-source data layer is constructed in the personal data management application. The question text first flows into the topic layer to determine the target topic corresponding to the question text. Then, based on the number of candidate interfaces and / or candidate tags corresponding to the target topic, the visibility range of the large model in the interface layer and / or tag layer is controlled. This allows the large model to make controllable selections, joint searches, and / or unified calls between interface paths and tag paths, and the large model does not have to face all interfaces and all tags at the beginning. Finally, in the multi-source data layer, the same data record can be hit through multiple interfaces and / or multiple usage scenario entry points, avoiding data omissions that cause deviations in the large model's answer. This solves the model call complexity problem caused by too many interfaces and the omission problem caused by multiple usage scenarios for a single data point, improving the accuracy of the large model's answer and the user's experience of using the large model in the user terminal. This realizes personal data management and intelligent recommendation in the user terminal.
[0034] In one specific implementation, based on the above embodiments, the personal data management application maintains simultaneously for each topic in the topic layer: a subset of interfaces (APIs) related to that topic, and a subset of tags related to that topic. After receiving the question text, the large model can first determine the target topic corresponding to the question text, and then determine the data invocation mode based on the number of candidate interfaces and / or candidate tags under that target topic. In the case of a shared topic-driven joint discovery mode, the large model invokes an interface discovery tool to obtain each candidate interface corresponding to the target topic, and invokes a tag discovery tool to obtain each candidate tag corresponding to the target topic. Then, based on each candidate interface and each candidate tag, it selects to perform a search using only the API path, only the tag path, or a combination of the API and tag paths to return the search results and generate the answer text. That is, this shared topic-driven joint discovery mode simultaneously discovers candidate API subsets and candidate tag subsets under the same target topic, and then the large model selects the interface path, tag path, or a combination of both to complete the search and invocation.
[0035] like Figure 2 As shown, Figure 2 This is a timing diagram illustrating a shared-topic-driven joint discovery method according to an embodiment of the present invention. Figure 2 In this process, users can submit questions to the large model, which then requests available capabilities (visible capability range) from the MCP server of the personal data management application. The MCP server returns various topics and discovery tools from the topic layer to the large model. Based on the target topic, the large model requests candidate APIs from the API discovery module of the personal data management application, which returns a subset of candidate APIs. Similarly, based on the target topic, the large model requests candidate tags from the Tag discovery module of the personal data management application, which returns a subset of candidate tags. Based on each candidate API and each candidate tag, the large model selects an API path, a Tag path, or a combination of paths and sends a search request to the unified execution module. Finally, the unified execution module returns the search results to the large model, which then generates an answer based on the search results and outputs it to the user.
[0036] In a specific example of a topic-driven joint discovery approach, a user asks "What do I need to buy recently?". In the first round, the model selects a shopping topic based on the question's semantics. Under this shopping topic, the server simultaneously supports: returning a subset of shopping-related interfaces, such as order query interfaces, product information interfaces, and consumption statistics interfaces; and returning a subset of shopping-related tags, such as shopping intent tags, pending purchase tags, and repeat purchase tags. In the second round, the model selects a retrieval path based on the question's requirements: to cover multiple sources and implicit demand expressions, it can first use the tag path to retrieve chat logs, favorites, browsing history, etc., with shopping semantics; to obtain structured order or consumption statistics results, it can then use the interface path to call the order or statistics interfaces; or it can first call the interfaces to obtain structured results and then supplement cross-source semantic evidence with the tag path. For example, a chat log containing "Buy me a charger," although its physical source belongs to the social data table, is retrieved by the model using the tag path under the shopping question because it is attached with shopping-related tags and associated with the shopping topic, without first considering the social topic or chat log interfaces. In this way, the same data is no longer fixed to a single source interface path, but can be discovered and utilized by large models through different paths that share themes, tags and interfaces.
[0037] In conjunction with the above embodiments, in one implementation, the present invention also provides a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models. In addition to the steps described above, this method may further include steps S21 to S24: Step S21: Obtain the existing interface information and tag information in the personal data management application, and cluster them based on the semantics of the interface information and tag information to obtain multiple topics, so as to construct the topic layer of the personal data management application.
[0038] In this embodiment, a multi-layer capability organization structure is constructed on the relevant MCP tool calling system, and data access capabilities are uniformly abstracted into four layers, which include: theme layer, interface layer, tag layer and multi-source data layer.
[0039] At the topic layer, the personal data management application can obtain interface and tag information corresponding to the connected data sources. Interface information includes at least one or more of the following: interface name, interface description, input parameter structure, and corresponding data source or data table information. Tag information includes at least one or more of the following: tag name, number of sample records, or hit scale. The personal data management application can then cluster the obtained interface and tag information at the topic level to obtain multiple topics, and construct the topic layer of the personal data management application based on these multiple topics.
[0040] In this embodiment, the topic layer is used to represent higher-level task scenarios, business topics, or semantic capability domains. Each topic can be associated with one or more interface information or one or more tag information, enabling the subsequent large model to first select a topic and then discover the corresponding candidate interface subset or candidate tag subset under that topic.
[0041] In addition, in some other implementations, the topics in the topic layer can also be obtained through manual configuration, rule mapping, historical question clustering, etc.
[0042] Step S22: Obtain the data tables of each client application in the user terminal, add the subject information of each client application's data table at the data table granularity, and add interface information to the data table of each client application according to the external interface provided by each client application, so as to construct the interface layer of the personal data management application.
[0043] In this embodiment, for the interface layer, the personal data management application can obtain data tables from various client applications on the user terminal. Each client application's data table includes relevant data records for that client application; that is, a client application's data table includes multiple data records generated by that client application. The personal data management application can add topic information to each client application's data table at the data table level, and add interface information to each client application's data table based on the external interfaces provided by each client application, thereby constructing the interface layer of the personal data management application based on the interface information of each data table. In one embodiment, the interface information includes at least one or more of the following: interface name, interface description, input parameter structure, topic information, and corresponding data source or corresponding data table information.
[0044] Each interface in the interface layer corresponds to a unique interface (i.e., an external interface), and the large model can obtain the data table corresponding to the interface by calling the interface. That is, in this embodiment, the interface layer (API layer) represents a set of interfaces for specific data tables, specific data sources, or specific structured query capabilities.
[0045] In the interface layer, the interface information corresponds to multiple topic information through data tables. For example, data table 1 corresponds to interface information 2, and data table 5 corresponds to topic 9. Therefore, the topic information corresponding to interface information 2 includes topic 5 and topic 9. That is, there is a mapping relationship between interface information 2 and topic 5 and topic 9.
[0046] Step S23: Obtain each data record from the data tables of each client application in the user terminal, and add multiple tag information and corresponding multiple subject information to each data record at the data record granularity according to the multiple usage scenarios of each data record, so as to construct the tag layer of the personal data management application.
[0047] In this embodiment, for the tag layer, the personal data management application can obtain each data record from the data tables of various client applications on the user terminal. Then, using data records as the granularity, based on the various usage scenarios of each data record, it adds multiple tag information and corresponding multiple subject information to each data record, and constructs the tag layer of the personal data management application based on the tag information. In one embodiment, the tag information includes at least one or more of the following: tag name, tag description, subject information, optional number of sample records, hit scale, or statistical information. The tags can be manually labeled tags, rule-generated tags, model-generated tags, or tags mapped from an external knowledge system.
[0048] Large models can retrieve data from multiple data layers based on tags. In this embodiment, the tag layer represents fine-grained semantic tags for specific data records, representing usage scenarios.
[0049] In the tag layer, the tag information corresponds to the topic information through data records. For example, data record M corresponds to tag information e and topic 4 and topic 7. Therefore, the topic information corresponding to tag information e includes topic 4 and topic 7. That is, there is a mapping relationship between tag information e and topic 4 and topic 7.
[0050] Step S24: For each data record in the data table of each client application in the user terminal, establish a mapping relationship between topic, interface, and tag to construct a multi-source data layer of the personal data management application. The same topic has a mapping relationship with multiple interface information and a mapping relationship with multiple tag information.
[0051] In this embodiment, a mapping relationship between topic, interface, and tag can be established for each data record in the data tables of various client applications in the user terminal (such as establishing a mapping relationship between topic and interface, and establishing a mapping relationship between topic and tag). A multi-source data layer for the personal data management application is then constructed based on multiple data records. The same topic can have mapping relationships with multiple interface information, and the same topic can have mapping relationships with multiple tag information. The multi-source data layer in this embodiment represents a data record that can simultaneously have multiple tags and can be retrieved through different paths. An interface or tag information can also belong to one or more topics. In one embodiment, a data record is attached with one or more tags, allowing the same record to be retrieved from different usage scenario entry points.
[0052] It should be noted that this embodiment does not impose any restrictions on the execution order of steps S21 to S24.
[0053] In this embodiment, a multi-layered capability organization structure is established, consisting of a topic layer, an interface layer, a tag layer, and a multi-source data layer, rather than organizing capabilities solely according to data source interfaces. On the one hand, interface subsets and tag subsets are uniformly organized at the topic layer to achieve dual-channel discovery driven by shared topics. On the other hand, interface paths and tag paths are supported simultaneously within the same framework, thus balancing structured access and cross-scenario semantic retrieval. Furthermore, multiple tags are allowed to be attached to a single piece of data, enabling the same piece of data to be retrieved by multiple usage scenario entry points.
[0054] In conjunction with any of the above embodiments, the present invention also provides a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models. In addition to the steps described above, this method may further include the following steps S31 to S310: Step S31: If the number of candidate tags that semantically match the target topic is not greater than the tag number threshold, expose all tags in the tag layer of the personal data management application to the large model, but do not expose the interface layer of the personal data management application to the large model.
[0055] In this embodiment, if the number of candidate tags that semantically match the target topic is not greater than the tag number threshold, the data retrieval mode can be determined to be a direct tag retrieval method. In this direct tag retrieval method, the visibility range of the tag layer is full exposure: specifically, the personal data management application fully exposes all tags in the tag layer of the personal data management application to the large model, and the personal data management application does not expose the interface layer of the personal data management application to the large model.
[0056] Step S32: Match the semantics of all tags with the semantics of the current round question text using the large model, and determine the target tags that match the semantics of the current round question text from all tags.
[0057] In this embodiment, after the large model obtains all the tags in the tag layer, it can match the semantics of all the tags with the semantics of the current round of question text, and determine each target tag that matches the current round of question text from all the tags. It can be understood that candidate tags are tags that match the semantics of the target topic, and target tags are tags that match the semantics of the current round of question text.
[0058] Step S33: Using the large model, retrieve the multi-source data layer of the personal data management application based on the target tags to obtain data records with at least one target tag.
[0059] In this embodiment, the large model can retrieve data records in the multi-source data layer of the personal data management application based on various target tags to obtain data records with at least one target tag. Specifically, if any data record in the multi-source data layer has at least one tag that matches the target tag, then that data record is determined to be a data record with at least one target tag.
[0060] Step S34: Using the large model, generate the answer text corresponding to the current round of questions based on data records with at least one target label, and display the answer text corresponding to the current round of questions through the user interaction page of the personal data management application.
[0061] In this embodiment, after the large model obtains data records with at least one target label, it can generate the answer text corresponding to the current round of question text based on the data records with at least one target label, and display the answer text corresponding to the current round of question text through the user interaction page of the personal data management application, thus completing the question and answer for the current round and realizing the personal data management and / or recommendation of the user terminal.
[0062] In one embodiment, under the direct tag retrieval method, the server of the personal data management application can directly expose the full set of tags and a unified tag retrieval tool to the large model. The large model selects one or more target tags based on the semantics of the current round of questions and directly retrieves relevant data records in the multi-source data layer. This method is particularly suitable for scenarios where "the same data record has multiple use cases." Since data records can be attached with multiple tags, the large model is no longer forced to access data along the source interface path, but can directly retrieve data according to the semantics of the use case.
[0063] In other words, when the system has already attached tags to multi-source data records, the model can retrieve data directly by tags without going through the source interface path. For example, if a user asks "What do I need to buy recently?", the model can directly select shopping-related tags for retrieval. Even if the relevant information exists in a chat record, it can still be found, thus solving the problem of "multiple use cases for the same data" through tag paths.
[0064] Step S35: Obtain the user's response text to the current round of questions via the user interaction page of the personal data management application.
[0065] In this embodiment, after receiving the answer text corresponding to the current round of questions, the user on the user terminal can provide feedback on the answer text in the user interaction page of the personal data management application. This feedback includes positive and negative feedback, where positive feedback indicates user satisfaction with the answer text, and negative feedback indicates user dissatisfaction. Thus, the personal data management application can obtain user feedback on the answer text corresponding to the current round of questions through this user interaction page.
[0066] Step S36: If the user on the user terminal gives negative feedback on the answer text corresponding to the current round of question text, determine whether the number of each candidate interface that semantically matches the target topic is greater than the interface number threshold.
[0067] In this embodiment, if it is determined that the user's response to the answer text corresponding to the current round of questions is negative, the personal data management application can redetermine other data call patterns to obtain data; specifically, at this time, it can be determined whether the number of each candidate interface that semantically matches the target topic is greater than the interface number threshold.
[0068] Step S37: If the number of candidate interfaces that semantically match the target topic is not greater than the number of interfaces, expose all interfaces in the interface layer of the personal data management application to the large model, but do not expose the tag layer of the personal data management application to the large model.
[0069] In this embodiment, if the number of candidate interfaces that semantically match the target topic is not greater than the interface number threshold, the data call mode can be determined to be the direct interface access method. In this direct interface access method, the visibility of the interface layer is fully exposed: specifically, the personal data management application fully exposes all interfaces in the interface layer of the personal data management application to the large model, and the personal data management application does not expose the tag layer of the personal data management application to the large model.
[0070] Step S38: Match the semantics of all interfaces with the semantics of the current round question text using the large model, and determine the target interfaces that match the semantics of the current round question text from all interfaces.
[0071] In this embodiment, after the large model obtains all interfaces in the interface layer, it can match the semantics of each interface with the semantics of the current round of question text, and determine each target interface that matches the current round of question text from all interfaces. It can be understood that candidate interfaces are interfaces that match the semantics of the target topic, and target interfaces are interfaces that match the semantics of the current round of question text.
[0072] Step S39: Call each target interface through the large model to read the target data record that semantically matches the current round question text from the multi-source data layer of the personal data management application.
[0073] In this embodiment, after the large model obtains each target interface, it can call each target interface to read target data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. These target data records are the data records that semantically match the current round of question text, read through the target interfaces.
[0074] Step S310: Using the large model, based on the target data record that semantically matches the current round question text, generate the answer text corresponding to the current round question text again, and display the answer text corresponding to the current round question text again through the user interaction page of the personal data management application.
[0075] In this embodiment, after the large model obtains the target data record that semantically matches the current round of question text, it can generate the answer text corresponding to the current round of question text again based on the target data record that semantically matches the current round of question text, and then display the answer text corresponding to the current round of question text again through the user interaction page of the personal data management application.
[0076] In one embodiment, under the direct interface access method, the server of the personal data management application directly exposes a full set of API tools to the large model. The large model selects one or more target APIs from all APIs based on the current round of questions, constructs parameters, and then initiates the call. This method is suitable for scenarios with a small number of APIs, clear API meanings, and a relatively well-defined question scope. Its advantages are simple structure and direct calls; its disadvantage is that the complexity of model tool selection increases significantly as the number of APIs increases.
[0077] In other words, in scenarios with a limited number of interfaces, the server exposes multiple data interfaces directly to the large model via direct interface access. The model then selects the appropriate interface and constructs its parameters based on the user's question. For example, if a user asks what content they have recently watched on a video platform, the model directly selects the video history interface to query recent viewing records and generates the answer.
[0078] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models. This method, in addition to the steps described above, may further include the following steps S41 to S44: Step S41: If the number of candidate interfaces that semantically match the target topic is not greater than the number of interfaces, and the number of candidate tags that semantically match the target topic is greater than the number of tags, the interface layer of the personal data management application is not exposed to the large model, and only the candidate tags in the tag layer of the personal data management application are exposed to the large model.
[0079] In this embodiment, if the number of candidate interfaces semantically matching the target topic is no greater than the interface number threshold, and the number of candidate tags semantically matching the target topic is greater than the tag number threshold, the data retrieval mode can be determined to be a discovery-based tag retrieval method. In this discovery-based tag retrieval method, the visibility range of the tag layer is partial exposure: specifically, the personal data management application only exposes each candidate tag in the tag layer of the personal data management application to the large model, and the personal data management application does not expose the interface layer of the personal data management application to the large model.
[0080] Step S42: The semantics of each candidate label are matched with the semantics of the current round question text using the large model, and each target label that matches the semantics of the current round question text is determined from the candidate labels.
[0081] In this embodiment, after the large model obtains each candidate label in the label layer, it can match the semantics of each candidate label with the semantics of the current round of question text, and determine each target label that matches the current round of question text from each candidate label.
[0082] Step S43: Using the large model, retrieve the multi-source data layer of the personal data management application based on the target tags to obtain data records with at least one target tag.
[0083] In this embodiment, the large model can retrieve data records in the multi-source data layer of the personal data management application based on each target label, and obtain data records with at least one target label.
[0084] Step S44: Using the large model, generate the answer text corresponding to the current round of questions based on data records with at least one target label, and display the answer text corresponding to the current round of questions through the user interaction page of the personal data management application.
[0085] In this embodiment, after the large model obtains data records with at least one target label, it can generate the answer text corresponding to the current round of question text based on the data records with at least one target label, and display the answer text corresponding to the current round of question text through the user interaction page of the personal data management application, thus completing the question and answer for the current round and realizing the personal data management and / or recommendation of the user terminal.
[0086] In one embodiment, in the discovery-based tag retrieval method, the server of the personal data management application first exposes a set of topics and a tag discovery tool to the large model. The large model first selects a target topic, then invokes the tag discovery tool. The server returns a subset of tags related to the target topic (i.e., each candidate tag) and the number of selectable sample records. The large model then retrieves data records in the multi-source data layer using a unified tag retrieval tool based on the returned tag subset. This method addresses the problem of "the number of tags may also be large." By first narrowing the candidate tag set by topic before performing tag retrieval, the tag space faced by the model can be reduced, and invalid tag selections can be decreased.
[0087] In other words, when the total number of tags is large (e.g., exceeding a tag count threshold), a discovery-based tag retrieval method is used to first expose the topic and tag discovery tool. For example, if a user asks "What do I need to buy recently?", the model first selects the shopping topic, calls the tag discovery tool, and the server returns a subset of tags related to the shopping scenario, along with the number of sample records. The model then performs tag retrieval based on these tags, matching relevant records from multiple sources such as chat history, orders, and browsing history, thereby narrowing the tag candidate space through the topic layer.
[0088] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models. In addition to the steps described above, this method may further include steps S51 to S55: Step S51: If the number of candidate interfaces that semantically match the target topic is greater than the number of interfaces, and the number of candidate tags that semantically match the target topic is not greater than the number of tags, the tag layer of the personal data management application is not exposed to the large model, and only the candidate interfaces in the interface layer of the personal data management application are exposed to the large model.
[0089] In this embodiment, if the number of candidate interfaces semantically matching the target topic is greater than the interface number threshold, and the number of candidate tags semantically matching the target topic is not greater than the tag number threshold, the data access mode can be determined to be a discovery-based interface access method. In this discovery-based interface access method, the visibility of the interface layer is partially exposed: specifically, the personal data management application only exposes each candidate interface in its interface layer to the large model, and the personal data management application does not expose its tag layer to the large model.
[0090] Step S52: Match the semantics of each candidate interface with the semantics of the current round question text using the large model, and determine each target interface that matches the semantics of the current round question text from the candidate interfaces.
[0091] In this embodiment, after the large model obtains each candidate interface in the interface layer, it can match the semantics of each candidate interface with the semantics of the current round of question text, and determine each target interface that matches the current round of question text from each candidate interface.
[0092] Step S53: Call each target interface through the large model to read the target data record that semantically matches the current round of question text from the multi-source data layer of the personal data management application.
[0093] In this embodiment, after the large model obtains each target interface, it can call each target interface to read the target data record that semantically matches the current round of question text from the multi-source data layer of the personal data management application.
[0094] Step S54: Using the large model, generate the answer text corresponding to the current round of questions based on the target data record that semantically matches the current round of questions.
[0095] In this embodiment, after the large model obtains the target data record that semantically matches the current round of question text, it can generate the answer text corresponding to the current round of question text based on the target data record that semantically matches the current round of question text, and display the answer text corresponding to the current round of question text through the user interaction page of the personal data management application.
[0096] In one embodiment, under the discovery-based interface access method, the server of the personal data management application does not directly expose all interfaces. Instead, it first exposes a set of topics and an interface discovery tool in the topic layer. The large model first determines the target topic corresponding to the current round of questions, then calls the discovery tool, and the server returns a subset of candidate interfaces under that target topic. The large model then selects a specific target interface from the local candidate set (the subset of candidate interfaces) and initiates the call. This method is mainly used to solve the problem of "too many interfaces and uneconomical full exposure". Its essence is to shrink the tool space by topic and separate the interface discovery phase from the interface call phase. Specifically, the interface discovery phase (first phase) first discovers interfaces according to the topic, and the interface call phase (second phase) calls the selected specific target interface to return user data.
[0097] In other words, when there are a large number of APIs, the server uses a discovery-based API access method, which does not directly expose all APIs, but instead exposes a set of topics first. For example, if a user asks "What types of videos do I mainly watch recently?", the model first selects video topics, then calls the API discovery tool. The server returns relevant candidate APIs from video website A, such as the recent viewing history API, category statistics API, and tag statistics API. The model selects the statistics API from the partial candidate set and then generates the answer, thus solving the problem of "too many APIs" through the topic layer.
[0098] Among them, the server corresponding to the personal data management application is used to store data and provide interfaces; the large model is on the user side, requesting data and providing services.
[0099] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models. In this method, step S14 may specifically include steps S61 to S62: Step S61: Match the semantics of each candidate interface with the semantics of the current round of question text using the large model, determine multiple target interfaces from the candidate interfaces, and call multiple target interfaces to filter out the data tables corresponding to each target interface from the multi-source data layer of the personal data management application.
[0100] In this embodiment, the personal data management application can use a large model to match the semantics of each candidate interface with the semantics of the current round of question text, thereby determining multiple target interfaces that match the current round of question text from among the candidate interfaces. After obtaining the target interfaces, these multiple target interfaces can be invoked to filter out the data tables corresponding to each of the multiple target interfaces from the multi-source data layer of the personal data management application.
[0101] Step S62: Using the large model, retrieve the data tables corresponding to each of the multiple target interfaces based on the candidate tags to obtain data records with at least one candidate tag.
[0102] In this embodiment, after obtaining the data tables corresponding to each of the multiple target interfaces, the large model can retrieve the data records in the data tables corresponding to each of the multiple target interfaces through each candidate label, and obtain the data records with at least one candidate label.
[0103] Thus, in this embodiment, the range of data sources (i.e., the data table) is first determined by the candidate interface path, and then the candidate tag path is used to perform further semantic filtering in the data table to obtain the filtered data records for generating the answer text of the large model.
[0104] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a method for hierarchical discovery and joint invocation of multi-source data capabilities for large models. In this method, step S14 may specifically include steps S71 to S72: Step S71: Following the method of supplementing with tags as the main interface, the semantics of each candidate tag are matched with the semantics of the current round of question text through the large model, and multiple target tags are determined from each candidate tag. Based on the multiple target tags, the multi-source data layer of the personal data management application is retrieved to obtain data records with at least one target tag.
[0105] In this embodiment, the personal data management application can first match the semantics of each candidate tag with the semantics of the current round of question text using a large model, supplementing the main interface with tags, to determine multiple target tags that match the semantics of the current round of question text. Then, based on the multiple target tags, the application retrieves data records in the multi-source data layer to obtain data records with at least one target tag.
[0106] Step S72: Based on the data record with at least one target label, call each candidate interface through the large model to read data records from the multi-source data layer of the personal data management application to supplement the data record with at least one candidate label.
[0107] In this embodiment, after obtaining data records with at least one target label, the personal data management application calls various candidate interfaces through the large model to read data records used to supplement the data records with at least one candidate label from the multi-source data layer of the personal data management application. Based on the data records with at least one target label and the supplementary data records, the large model generates the answer text. The data records used to supplement the data records with at least one candidate label can be other data records whose semantics match those of the data records with at least one candidate label.
[0108] In this embodiment, data records that semantically match the current round's question text can be quickly identified first through candidate label paths, and then supplemented with data records using candidate interface paths to generate the answer text for the large model. For example, data records with shopping semantics can be quickly found first through candidate label paths; then, order, product, or consumption statistics information can be supplemented through candidate interface paths.
[0109] In one embodiment, such as Figure 3 As shown, Figure 3 This is a schematic diagram illustrating a data retrieval mode according to an embodiment of the present invention. Wherein, Figure 3 The MCP client in the system can be used for personal data management applications on user terminals. It can determine the data access mode (i.e., the current mode) by calling tools or lists. There are a total of 5 data access modes: Mode 1 is the direct interface access method (corresponding to...). Figure 3 In the Direct API, the server directly exposes the entire set of API tools to the large model. The model selects one or more interfaces from all APIs based on the problem, constructs parameters, and then initiates the call; Mode 2 is a discovery-based interface access method (corresponding to...). Figure 3 In the Discovery_API section, the server does not directly expose all APIs; instead, it first exposes a set of topics and an API discovery tool. The model first determines the topic corresponding to the question, then calls the discovery tool. The server returns a subset of candidate APIs under that topic, and the model then selects a specific API from the partial candidate set and initiates the call. Mode 3 is a direct tag retrieval method (corresponding to...). Figure 3 In the Direct_Tag framework, the server directly exposes the entire tag set and a unified tag retrieval tool to the large model. The model selects one or more tags based on the semantics of the question and directly retrieves relevant data records; Mode 4 is a discovery-based tag retrieval method (corresponding to...). Figure 3In the Discovery_Tag framework, the server first exposes a set of topics and a tag discovery tool to the model. The model first selects a topic, then calls the tag discovery tool. The server returns a subset of tags related to that topic and the number of selectable sample records. The model then retrieves data records using a unified tag retrieval tool based on the returned tag subset. Mode 5 is a shared topic-driven joint discovery method (corresponding to...). Figure 3 The Discovery mode (as described in the example) maintains the following for each topic: a subset of APIs related to that topic and a subset of tags related to that topic. Upon receiving a question, the large model first determines the target topic. Then, under the same target topic, the model can: call the API discovery tool to obtain a subset of candidate APIs; and call the tag discovery tool to obtain a subset of candidate tags. In the second and subsequent rounds, it chooses to use only the API path, only the tag path, or a combination of both paths. The API path is suitable for methods requiring precise field and statistical queries using traditional SQL, while the tag path is suitable for cross-source retrieval. This fifth mode unifies the aforementioned four modes under a single framework, making API discovery and tag discovery no longer two separate systems, but rather forming a dual-channel discovery structure at the shared topic layer.
[0110] In one embodiment, after constructing a multi-layered capability organization structure and abstracting data access capabilities into four unified layers: topic layer, interface layer, tag layer, and multi-source data layer, two types of capability exposure mechanisms are provided on this layered structure: a direct exposure mechanism (the large model directly obtains the complete visible set of a certain capability layer) and a phased discovery mechanism (the large model first obtains the entry point of a higher-level capability, and then the server dynamically returns a local candidate subset according to that entry point). Simultaneously, this embodiment also provides two types of data arrival paths: 1. Interface access path: the model obtains data through API calls; 2. Tag retrieval path: the model obtains data through tag retrieval.
[0111] As can be seen, this invention is not simply adding a new discovery interface, nor is it merely adding tags on top of an interface. Instead, it establishes a multi-layered semantic capability structure of topic-interface-tag-data, enabling the server to control the visible capabilities of the large model at different stages (controlling the visible capability range of the model through two mechanisms: direct exposure and phased discovery), and allowing the same data to be hit through multiple use case entry points. For example, although a chat record is physically stored in a social data table and can only be accessed directly through the chat interface, in this invention it can carry shopping tags, to-do tags, and travel tags, and can be simultaneously associated with topics such as shopping, to-do lists, and travel. In this way, when a user asks "What do I need to buy recently?", the model can find the record either through the tag path under the shopping topic, or by first discovering relevant tags or relevant interfaces through a joint discovery method, and then performing a second round of retrieval, thereby avoiding being missed simply because its source belongs to "social data".
[0112] In addition, in one embodiment, the topics in the topic layer are divided into multiple levels. First, a first-level topic can be selected, then a second-level subtopic under the first-level topic can be selected, and then the corresponding interface or tag can be discovered based on the second-level subtopic, thereby forming a multi-stage convergence structure.
[0113] In another embodiment, the personal data management application receives a user's natural language question and provides it to the corresponding large model. The server of the personal data management application can determine the scope of exposure to the large model based on the current working method (e.g., full interface exposure, full tag exposure, or topic and discovery tool exposure). Then, the large model performs the first round of capability selection. If the first round adopts a direct approach, the large model directly selects a tag or interface; if the first round adopts a discovery approach, the large model first selects a topic. For the server, if the discovery approach is adopted, the server returns a subset of candidate interfaces or a subset of candidate tags based on the topic; in the preferred combined approach, two types of candidate sets can be returned respectively. Next, the model performs a second round of retrieval or invocation decision-making, where the model can select a specific interface for invocation, select a specific tag for retrieval, or select a combination path of interface and tag for filtering from a local candidate set; then the server executes the specific interface invocation, tag retrieval, or a combination of retrieval followed by invocation and returns the results. The model generates a natural language answer based on the data results returned by the server and can continue to add the next round of discovery, retrieval, or invocation if necessary.
[0114] As can be seen, this invention only controls the visible capabilities of the model in the first stage, while retaining the model's autonomous selection ability in the local candidate space in the second stage. Therefore, this invention balances server-side control over the exposure surface, flexible decision-making by the large model regarding the final tool path, and semantic reuse of the same data across multiple scenarios. Furthermore, this invention introduces a local candidate set between the discovery and execution stages, reducing the initial interface or label scale faced by the large model.
[0115] In summary, this invention addresses two main problems in current large-scale model data access scenarios: "too many interfaces leading to difficulty in tool selection" and "single data points being easily missed due to multiple use cases." It proposes a layered exposure and phased discovery method for data capabilities based on a multi-layered capability organization structure of "topic-interface-tag-multi-source data." This method uniformly supports direct interface access, discovery-based interface access, direct tag retrieval, discovery-based tag retrieval, and shared topic-driven joint discovery within the same framework. This allows the server to control the scope of capabilities visible to the model at different stages, enabling the same data point to be accessed through multiple use case entry points, while also considering structured interface capabilities, semantic tagging capabilities, server-side governance capabilities, and model autonomous decision-making capabilities. Therefore, this invention not only solves the problem of excessive interface numbers but also addresses the issue of missed detections due to a single data source path.
[0116] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0117] Based on the same inventive concept, one embodiment of the present invention provides a system for hierarchical discovery and joint invocation of multi-source data capabilities for large models. (Reference) Figure 4 , Figure 4 This is a structural block diagram of a multi-source data capability hierarchical discovery and joint invocation system for large models, provided by an embodiment of the present invention. Figure 4 As shown, the system is deployed in a user terminal as a personal data management application, including: The text acquisition module is used to acquire the current round of question text input by the user on the user terminal through the user interaction page of the personal data management application; The first matching module is used to expose each topic in the topic layer of the personal data management application to the large model, and then match the semantics of each topic with the semantics of the current round of question text through the large model, and determine the target topic that matches the semantics of the current round of question text from each topic. The data exposure module is used to expose each candidate interface in the interface layer of the personal data management application to the large model, and to expose each candidate tag in the tag layer of the personal data management application to the large model when the number of each candidate interface that semantically matches the target topic is greater than the number of interfaces and the number of each candidate tag that semantically matches the target topic is greater than the number of tags. The record reading module is used to retrieve data records with at least one candidate label from the multi-source data layer of the personal data management application based on the candidate labels using the large model, and to call the candidate interfaces through the large model to read data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. The text generation module is used to generate the answer text corresponding to the current round of questions by using the large model, based on data records with at least one candidate label and data records that semantically match the current round of questions, and to display the answer text corresponding to the current round of questions through the user interaction page of the personal data management application.
[0118] Optionally, the system further includes: The first construction module is used to obtain the existing interface information and tag information in the personal data management application, and to perform clustering based on the semantics of the interface information and the tag information to obtain multiple topics, so as to construct the topic layer of the personal data management application. The second construction module is used to obtain the data tables of each client application in the user terminal, add the subject information of each client application's data table at the granularity of the data table, and add interface information to the data table of each client application according to the external interface provided by each client application, so as to construct the interface layer of the personal data management application; the data table of a client application includes multiple data records generated by the client application; The third construction module is used to obtain each data record in the data table of each client application in the user terminal, and add multiple tag information and corresponding multiple subject information to the data record at the granularity of the data record according to the multiple usage scenarios of each data record, so as to construct the tag layer of the personal data management application. The fourth construction module is used to establish a mapping relationship between topic, interface, and tag for each data record in the data table of each client application in the user terminal, so as to construct the multi-source data layer of the personal data management application. There is a mapping relationship between the same topic and multiple interface information, and there is a mapping relationship between the same topic and multiple tag information.
[0119] Optionally, the system further includes: The first exposure module is used to expose all tags in the tag layer of the personal data management application to the large model, provided that the number of candidate tags that semantically match the target topic is not greater than the tag number threshold, and not to expose the interface layer of the personal data management application to the large model. The tag matching module is used to match the semantics of all tags with the semantics of the current round of question text using the large model, and to determine each target tag that matches the semantics of the current round of question text from all tags; The data retrieval module is used to retrieve data records with at least one target label from the multi-source data layer of the personal data management application based on the target labels using the large model. The answer generation module is used to generate the answer text corresponding to the current round of questions based on the data records with at least one target label using the large model, and to display the answer text corresponding to the current round of questions through the user interaction page of the personal data management application; The feedback acquisition module is used to acquire the user's response text to the current round of question text through the user interaction page of the personal data management application; The numerical judgment module is used to determine whether the number of candidate interfaces semantically matching the target topic is greater than an interface number threshold when the user's feedback on the answer text corresponding to the current round of questions on the user terminal is negative. The second exposure module is used to expose all interfaces in the interface layer of the personal data management application to the large model, provided that the number of candidate interfaces that semantically match the target topic is not greater than the number of interfaces, but does not expose the tag layer of the personal data management application to the large model. The interface matching module is used to match the semantics of all interfaces with the semantics of the current round question text using the large model, and to determine each target interface that matches the semantics of the current round question text from all interfaces. The interface calling module is used to call various target interfaces through the large model in order to read target data records that semantically match the current round of question text from the multi-source data layer of the personal data management application; The text display module is used to generate the answer text corresponding to the current round of questions again based on the target data record that semantically matches the current round of questions, using the large model, and then display the answer text corresponding to the current round of questions again through the user interaction page of the personal data management application.
[0120] Optionally, the system further includes: The third exposure module is used to expose only the candidate tags in the tag layer of the personal data management application to the large model when the number of each candidate interface that semantically matches the target topic is not greater than the number of interfaces and the number of each candidate tag that semantically matches the target topic is greater than the number of tags. The first matching module is used to match the semantics of each candidate label with the semantics of the current round question text using the large model, and to determine each target label that matches the semantics of the current round question text from the candidate labels. The first retrieval module is used to retrieve data records with at least one target label from the multi-source data layer of the personal data management application based on the target labels using the large model. The first generation module is used to generate the answer text corresponding to the current round of questions based on data records with at least one target label using the large model, and to display the answer text corresponding to the current round of questions through the user interaction page of the personal data management application.
[0121] Optionally, the system further includes: The fourth exposure module is used to expose only the candidate interfaces in the interface layer of the personal data management application to the large model when the number of candidate interfaces that semantically match the target topic is greater than the number of interfaces and the number of candidate tags that semantically match the target topic is not greater than the number of tags. The second matching module is used to match the semantics of each candidate interface with the semantics of the current round question text using the large model, and to determine each target interface that matches the semantics of the current round question text from the candidate interfaces. The first calling module is used to call various target interfaces through the large model in order to read target data records that semantically match the current round of question text from the multi-source data layer of the personal data management application; The second generation module is used to generate the answer text corresponding to the current round of questions by using the large model and target data records that semantically match the current round of questions.
[0122] Optionally, the record reading module includes: The matching and invocation module is used to match the semantics of each candidate interface with the semantics of the current round of question text through the large model, determine multiple target interfaces from the candidate interfaces, and invoke multiple target interfaces to filter out the data tables corresponding to each target interface from the multi-source data layer of the personal data management application. The second retrieval module is used to retrieve data records with at least one candidate label from the data tables corresponding to the multiple target interfaces based on the candidate labels using the large model.
[0123] Optionally, the record reading module includes: The matching and retrieval module is used to match the semantics of each candidate tag with the semantics of the current round of question text through the large model, supplemented by the tag as the main interface, to determine multiple target tags from the candidate tags, and to retrieve the multi-source data layer of the personal data management application based on the multiple target tags to obtain data records with at least one target tag. The record supplementation module is used to call various candidate interfaces through the large model based on data records with at least one target label, so as to read data records from the multi-source data layer of the personal data management application to supplement the data records with at least one candidate label.
[0124] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0125] The multi-source data capability hierarchical discovery and joint invocation system for large models in this embodiment of the invention can be a device, or a component, integrated circuit, or chip in a terminal. The system can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This embodiment of the invention does not impose specific limitations.
[0126] The multi-source data capability hierarchical discovery and joint invocation system for large models in this embodiment of the invention can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this embodiment of the invention does not impose specific limitations.
[0127] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, such as... Figure 5 As shown, Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory, a processor, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps in the multi-source data capability hierarchical discovery and joint invocation method for large models described in any of the above embodiments of the present invention.
[0128] It should be noted that the electronic devices in the embodiments of the present invention include the mobile electronic devices and non-mobile electronic devices described above.
[0129] Based on the same inventive concept, another embodiment of the present invention provides a readable storage medium storing a program or instructions. When executed by a processor, the program or instructions implement the steps in the multi-source data capability hierarchical discovery and joint invocation method for large models as described in any of the above embodiments of the present invention. The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0130] As the system implementation is basically similar to the method implementation, it is described in a relatively simple way. For relevant details, please refer to the description of the method implementation.
[0131] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0132] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0133] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. A large model-oriented multi-source data capability hierarchical discovery and joint calling method, characterized in that, A personal data management application deployed on a user terminal, the method comprising: The user terminal's input text for the current round of questions is obtained through the user interaction page of the personal data management application; Each topic in the topic layer of the personal data management application is exposed to the large model. The large model then matches the semantics of each topic with the semantics of the current round of question text, and identifies the target topic that semantically matches the current round of question text from each topic. If the number of candidate interfaces that semantically match the target topic is greater than the number of interfaces, and the number of candidate tags that semantically match the target topic is greater than the number of tags, then each candidate interface in the interface layer of the personal data management application is exposed to the large model, and each candidate tag in the tag layer of the personal data management application is exposed to the large model. The large model retrieves data records with at least one candidate label from the multi-source data layer of the personal data management application based on the candidate labels. The large model also calls various candidate interfaces to read data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. Using the large model, based on data records with at least one candidate label and data records that semantically match the current round of question text, the answer text corresponding to the current round of question text is generated, and the answer text corresponding to the current round of question text is displayed on the user interaction page of the personal data management application.
2. The large model-oriented multi-source data capability layered discovery and joint invocation method according to claim 1, characterized in that, The method further includes: Obtain existing interface information and tag information from the personal data management application, and cluster based on the semantics of the interface information and tag information to obtain multiple topics, thereby constructing the topic layer of the personal data management application; The system retrieves data tables from various client applications on the user terminal, adds topic information to the data tables of each client application at the data table level, and adds interface information to the data tables of each client application based on the external interfaces provided by each client application, in order to construct the interface layer of the personal data management application; a data table of a client application includes multiple data records generated by the client application. Each data record in the data tables of various client applications in the user terminal is obtained, and based on the data record as the granularity, multiple tag information and corresponding multiple subject information are added to the data record according to the multiple usage scenarios of each data record, so as to construct the tag layer of the personal data management application; For each data record in the data tables of various client applications in the user terminal, a mapping relationship is established between topic, interface, and tag to construct a multi-source data layer for the personal data management application. The same topic has a mapping relationship with multiple interface information and a mapping relationship with multiple tag information.
3. The method for hierarchical discovery and joint invocation of multi-source data capabilities for large models according to claim 1, characterized in that, The method further includes: If the number of candidate tags that semantically match the target topic is not greater than the tag number threshold, all tags in the tag layer of the personal data management application are fully exposed to the large model, but the interface layer of the personal data management application is not exposed to the large model. The large model is used to match the semantics of all tags with the semantics of the current round of question text, and to determine each target tag that matches the semantics of the current round of question text from all tags. The large model retrieves data from the multi-source data layer of the personal data management application based on the target labels to obtain data records with at least one target label. Using the large model, based on data records with at least one target label, the answer text corresponding to the current round of questions is generated, and the answer text corresponding to the current round of questions is displayed on the user interaction page of the personal data management application; The user terminal obtains the user's response text to the current round of questions through the user interaction page of the personal data management application; If the user on the user terminal provides negative feedback on the answer text corresponding to the current round of questions, determine whether the number of candidate interfaces that semantically match the target topic is greater than the interface number threshold: If the number of candidate interfaces that semantically match the target topic is not greater than the number of interfaces, all interfaces in the interface layer of the personal data management application are fully exposed to the large model, but the tag layer of the personal data management application is not exposed to the large model. The large model is used to match the semantics of all interfaces with the semantics of the current round of question text, and to determine each target interface that matches the semantics of the current round of question text from all interfaces. The large model calls various target interfaces to read target data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. Using the large model, based on the target data record that semantically matches the current round of question text, the answer text corresponding to the current round of question text is generated again, and the answer text corresponding to the current round of question text is displayed again through the user interaction page of the personal data management application.
4. The method for hierarchical discovery and joint invocation of multi-source data capabilities for large models according to claim 1, characterized in that, The method further includes: If the number of candidate interfaces that semantically match the target topic is not greater than the number of interfaces, and the number of candidate tags that semantically match the target topic is greater than the number of tags, the interface layer of the personal data management application will not be exposed to the large model, and only the candidate tags in the tag layer of the personal data management application will be exposed to the large model. The large model is used to match the semantics of each candidate label with the semantics of the current round of question text, and to determine each target label that matches the semantics of the current round of question text from the candidate labels. The large model retrieves data from the multi-source data layer of the personal data management application based on the target labels to obtain data records with at least one target label. Using the large model, based on data records with at least one target label, the answer text corresponding to the current round of questions is generated, and the answer text corresponding to the current round of questions is displayed on the user interaction page of the personal data management application.
5. The method for hierarchical discovery and joint invocation of multi-source data capabilities for large models according to claim 1, characterized in that, The method further includes: If the number of candidate interfaces that semantically match the target topic is greater than the number of interfaces, and the number of candidate tags that semantically match the target topic is not greater than the number of tags, the tag layer of the personal data management application will not be exposed to the large model, and only the candidate interfaces in the interface layer of the personal data management application will be exposed to the large model. The semantics of each candidate interface are matched with the semantics of the current round of question text using the large model, and each target interface that matches the semantics of the current round of question text is determined from the candidate interfaces. The large model calls various target interfaces to read target data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. Using the large model, the answer text corresponding to the current round of questions is generated based on the target data records that semantically match the current round of questions.
6. The method for hierarchical discovery and joint invocation of multi-source data capabilities for large models according to any one of claims 1 to 5, characterized in that, The large model retrieves data records from the multi-source data layer of the personal data management application based on the candidate tags to obtain data records with at least one candidate tag. Furthermore, the large model calls various candidate interfaces to read data records from the multi-source data layer of the personal data management application that semantically match the current round of question text, including: The semantics of each candidate interface are matched with the semantics of the current round of question text using the large model. Multiple target interfaces are determined from the candidate interfaces, and multiple target interfaces are called to filter out the data tables corresponding to each target interface from the multi-source data layer of the personal data management application. The large model retrieves data records with at least one candidate label from the data tables corresponding to the multiple target interfaces based on the candidate labels.
7. The method for hierarchical discovery and joint invocation of multi-source data capabilities for large models according to any one of claims 1 to 5, characterized in that, The large model retrieves data records from the multi-source data layer of the personal data management application based on the candidate tags to obtain data records with at least one candidate tag. Furthermore, the large model calls various candidate interfaces to read data records from the multi-source data layer of the personal data management application that semantically match the current round of question text, including: By supplementing the main interface with tags, the semantics of each candidate tag are matched with the semantics of the current round of question text through the large model, and multiple target tags are determined from each candidate tag. Based on the multiple target tags, the multi-source data layer of the personal data management application is retrieved to obtain data records with at least one target tag. Based on data records with at least one target label, the large model calls various candidate interfaces to read data records from the multi-source data layer of the personal data management application to supplement the data records with at least one candidate label.
8. A hierarchical discovery and joint invocation system for multi-source data capabilities in large-scale models, characterized in that, The system includes: The text acquisition module is used to acquire the current round of question text input by the user on the user terminal through the user interaction page of the personal data management application; The first matching module is used to expose each topic in the topic layer of the personal data management application to the large model, and then match the semantics of each topic with the semantics of the current round of question text through the large model, and determine the target topic that matches the semantics of the current round of question text from each topic. The data exposure module is used to expose each candidate interface in the interface layer of the personal data management application to the large model, and to expose each candidate tag in the tag layer of the personal data management application to the large model when the number of each candidate interface that semantically matches the target topic is greater than the number of interfaces and the number of each candidate tag that semantically matches the target topic is greater than the number of tags. The record reading module is used to retrieve data records with at least one candidate label from the multi-source data layer of the personal data management application based on the candidate labels using the large model, and to call the candidate interfaces through the large model to read data records that semantically match the current round of question text from the multi-source data layer of the personal data management application. The text generation module is used to generate the answer text corresponding to the current round of questions by using the large model, based on data records with at least one candidate label and data records that semantically match the current round of questions, and to display the answer text corresponding to the current round of questions through the user interaction page of the personal data management application.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method for hierarchical discovery and joint invocation of multi-source data capabilities for large models as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method for hierarchical discovery and joint invocation of multi-source data capabilities for large models as described in any one of claims 1 to 7.