Enterprise data retrieval method and system based on large model and medium
By configuring database priority and vectorized data in the enterprise data retrieval system, and combining large models to analyze user needs, the problem of insufficient database priority and user needs analysis capabilities in the existing technology is solved, and efficient and accurate enterprise data retrieval and result generation are achieved.
Patent Information
- Application Number
- CN202510550147.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art fails to effectively consider the priority differences between different databases in enterprise data retrieval, and lacks an active guidance mechanism for user needs, resulting in unreasonable weight allocation of search results and semantic ambiguity affecting accuracy.
By creating an enterprise database and configuring the database, searching for constraint prompt words and preset search priority, the data is processed vectorized, user demand data is obtained through the user interaction interface, searching for constraint prompt words and determining database priority, and data analysis and output generation are performed in combination with the big model.
Dynamic priority configuration is realized, retrieval efficiency and accuracy are improved, user needs analysis capabilities are enhanced, data representation space is unified, and result generation quality is improved.
Smart Images

Figure CN120067141A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of enterprise data retrieval, and particularly to an enterprise data retrieval method, system and medium based on a large model. Background Art
[0002] In recent years, with the improvement of enterprise informatization, the scale of enterprise data has shown an explosive growth, and the data form has become increasingly diverse, covering structured data (such as database forms) and unstructured data (such as documents, logs, emails, etc.). How to efficiently and accurately extract effective information from massive enterprise data has become an important challenge for enterprise intelligent management and decision-making. Traditional data retrieval technologies mostly rely on keyword matching or rule engines, such as search engines based on inverted indexes or conditional queries with preset rules. However, such methods have significant limitations: on the one hand, keyword matching is difficult to capture the deep semantics of user needs, especially in dealing with complex queries or polysemous words, which is prone to misdetection or missed detection; on the other hand, the rule engine depends on manually predefined logic, with poor flexibility and difficulty in adapting to dynamic business scenarios and personalized needs.
[0003] In the prior art, some solutions attempt to introduce natural language processing (NLP) technology to improve the retrieval semantic understanding ability. For example, the user query is vectorized through a pre-trained language model and matched with the database content for similarity. However, such methods still have the following problems: 1) In the scenario of joint retrieval of multiple enterprise databases, the priority differences of different databases are not considered, resulting in unreasonable weight allocation of retrieval results; 2) There is a lack of an active guidance mechanism for user needs. When the user's needs are expressed vaguely, the retrieval accuracy is easily affected by semantic ambiguity. Summary of the Invention
[0004] This application provides an enterprise data retrieval method, system and medium based on a large model to solve the problems that the prior solutions do not consider the priority differences of different databases and lack an active guidance mechanism for user needs.
[0005] In a first aspect, this application provides an enterprise data retrieval method based on a large model, and the method includes: Create a number of enterprise databases, and configure the corresponding relationships among the enterprise databases, retrieval constraint prompt words, and preset retrieval priorities; Perform vectorization processing on the data in the enterprise databases to convert the text-form data into a numerical vector representation; Through the user interface, obtain user demand data, then extract the retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database; Calculate the vector representation with the highest similarity to the retrieval constraint prompt words with the adjustment weights corresponding to the preset retrieval priorities; Perform data analysis on the retrieval constraint prompt words to determine the corresponding guiding data; among them, the guiding data at least includes: role positioning prompt words, ability positioning prompt words, task objectives, and output requirements. Input the guiding data and vector representation into the large model to obtain output data.
[0006] In one implementation manner of the present application, after creating several enterprise databases and configuring the preset retrieval priorities of the enterprise databases, the method further includes: Modify the data in the enterprise database through the backend interface, and / or modify the corresponding relationship between the enterprise database, retrieval constraint prompt words, and preset retrieval priorities.
[0007] In one implementation manner of the present application, perform vectorization processing on the data in the enterprise database to convert the data in text form into a numerical vector representation, specifically including: Obtain the first slice information through the preset slice acquisition interface. Based on the first slice information, split the data in the enterprise database to obtain several split segments. Obtain the segmentation granularity, and use the segmentation granularity to further split the split segments to obtain several knowledge fields. Convert the knowledge fields in text form into a numerical vector representation.
[0008] In one implementation manner of the present application, through the user interaction interface, obtain the user demand data, then extract the retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database, specifically including: Obtain the user demand data through the user interaction interface. Input the user demand data into the large model, extract the key features in the user demand data, and then determine whether the key features belong to the preset trigger conditions. When it belongs to the preset trigger conditions, extract the retrieval constraint prompt words from the user demand data. Based on the corresponding relationship between the enterprise database, retrieval constraint prompt words, and preset retrieval priorities, determine the preset retrieval priorities corresponding to each enterprise database.
[0009] In one implementation manner of the present application, calculate the vector representation with the highest similarity to the retrieval constraint prompt words with the adjustment weights corresponding to the preset retrieval priorities, specifically including: Calculate the initial similarity between the retrieval constraint prompt words and the vector representation through the language similarity algorithm. Multiply the adjustment weights by the initial similarity to obtain the final similarity, and then obtain the vector representation with the highest similarity.
[0010] In one implementation of the present application, before analyzing the data of the retrieval constraint prompt words to determine the corresponding guiding data, the method further includes: Configuring the correspondence between the retrieval constraint prompt words and the guiding data.
[0011] In a second aspect, the present application provides an enterprise data retrieval system based on a large model. The system includes: A RAG (Retrieval-augmented Generation) module for creating a number of enterprise databases, configuring the correspondence between the enterprise databases, retrieval constraint prompt words, and preset retrieval priorities; performing vectorization processing on the data in the enterprise databases to convert the text-form data into a numerical vector representation; A user interaction module for obtaining user requirement data through a user interaction interface; A large model module for extracting retrieval constraint prompt words, and then determining the preset retrieval priorities corresponding to each enterprise database; calculating the vector representation with the highest similarity to the retrieval constraint prompt words with the adjustment weights corresponding to the preset retrieval priorities; analyzing the data of the retrieval constraint prompt words to determine the corresponding guiding data; where the guiding data at least includes: role positioning guiding words, ability positioning guiding words, task objectives, and output requirements; inputting the guiding data and the vector representation into the large model to obtain output data.
[0012] In one implementation of the present application, the RAG module includes a revision unit, for modifying the data in the enterprise database and / or modifying the correspondence between the enterprise database, retrieval constraint prompt words, and preset retrieval priorities through a backend interface.
[0013] In one implementation of the present application, the large model module includes a determination unit, for extracting key features from the user requirement data, and then determining whether the key features belong to a preset trigger condition; When belonging to the preset trigger condition, extracting the retrieval constraint prompt words from the user requirement data; Based on the correspondence between the enterprise database, retrieval constraint prompt words, and preset retrieval priorities, determining the preset retrieval priorities corresponding to each enterprise database.
[0014] In a third aspect, the present application provides a non-volatile computer storage medium, on which computer instructions are stored, and when the computer instructions are executed, they implement a method for retrieving enterprise data based on a large model as described in any one of the above.
[0015] From the above technical solutions, it can be seen that the present application has the following advantages: 1. Dynamic priority configuration enhances retrieval efficiency and accuracy: By presetting retrieval priorities, it ensures that high-value databases (such as core business systems) have a higher weight in retrieval, avoiding interference from low-priority databases (such as log backups) on key results. Retrieving high-weight databases first reduces unnecessary computational overhead and improves the system response speed. The combination of similarity calculation and priority weights makes the retrieval results more in line with the actual business needs. For example, financial data is presented first instead of non-critical documents.
[0016] 2. Active guidance mechanism enhances the ability to analyze user needs: Existing solutions lack active guidance for user needs and are prone to retrieval deviations due to ambiguous requirements. By role positioning (such as "financial analyst") and ability positioning (such as "data visualization"), it clarifies the user's identity and skill background, reducing semantic ambiguity. The introduction of task goals (such as "generate quarterly reports") and output requirements (such as "Excel format") transforms ambiguous requirements into executable retrieval conditions. Generating differentiated guidance strategies for different user roles (such as sales, R & D) improves the adaptability of the retrieval experience.
[0017] 3. Cross-modal vectorization processing unifies the data representation space: Existing solutions have insufficient capabilities in processing unstructured data and are difficult to unify the semantic space. Vectorization representation breaks through the limitations of text forms and supports semantic matching of cross-modal data (such as documents, tables, logs). Numerical vectors can be directly input into large models for reasoning. For example, relevant content in different databases can be associated through vector similarity calculation. Support for combined multi-condition queries (such as "find product defect records in customer complaints") improves retrieval flexibility.
[0018] 4. Collaborative optimization with large models improves the quality of result generation: Existing solutions rely on a single retrieval model and are difficult to handle complex tasks. Large models combine task goals (such as "risk assessment report") in the guiding data and vector retrieval results to generate structured and highly readable outputs (such as reports, charts). Large models can correct potential biases in retrieval results (such as similarity calculation errors) and supplement missing information through context reasoning. Support for generating diverse output forms (such as text summaries, visual charts) according to user needs improves the practicality of the results. Description of the Drawings
[0019] To more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 This is a flowchart of a method for retrieving enterprise data based on a large model provided by an embodiment of the present application.
[0021] Figure 2 This is a schematic diagram of the internal structure of a system for retrieving enterprise data based on a large model provided by an embodiment of the present application. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] Those skilled in the art should understand that the embodiments described below are only preferred embodiments of the present disclosure, which does not mean that the present disclosure can only be implemented through these preferred embodiments. These preferred embodiments are only used to explain the technical principles of the present disclosure and are not used to limit the protection scope of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the preferred embodiments provided by the present disclosure without creative efforts shall still fall within the protection scope of the present disclosure.
[0024] It should also be noted that the term "including", "comprising" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.
[0025] This application builds a knowledge retrieval Agent based on the RAG architecture of a large model on the open-source Dify platform. Its basic idea is to use RAG technology to build a knowledge database, improve the retrieval accuracy by designing a text vectorization processing method adapted to knowledge retrieval, and at the same time configure roles, tasks and retrieval capabilities for the large model according to the actual query requirements.
[0026] Next, the technical solutions proposed in the embodiments of this application will be described in detail with reference to the accompanying drawings.
[0027] The embodiment provides a method for retrieving enterprise data based on a large model. As Figure 1 shown, the method provided by the embodiment of this application mainly includes the following steps: Step 110: Create a number of enterprise databases and configure the correspondence relationships among the enterprise databases, retrieval constraint prompt words, and preset retrieval priorities.
[0028] In some embodiments, after creating a number of enterprise databases and configuring the preset retrieval priorities of the enterprise databases, the method further includes: Modify the data in the enterprise database through the backend interface and / or modify the correspondence relationships among the enterprise database, retrieval constraint prompt words, and preset retrieval priorities.
[0029] This step supports users to create and dynamically update enterprise databases as needed. There can be multiple enterprise databases, and the retrieval priorities of different databases can be configured and adjusted by users in the interaction interface or backend interface of the Agent.
[0030] Those skilled in the art can understand that this step enables enterprises to create dedicated databases according to different business modules (such as customer management, supply chain, and finance), and control the retrieval weights through preset priorities. For example, in an e-commerce scenario, the "order database" is set with the highest priority to ensure the timeliness of transaction data retrieval; while the "user portrait database" has a lower priority and is used to assist the recommendation algorithm.
[0031] In a dynamic adjustment scenario, if a promotional activity is temporarily launched, the priority of the "inventory database" can be raised to the top through the backend interface to give priority to processing inventory query requests and avoid over-selling problems.
[0032] By binding the retrieval constraint prompt words (such as "region = North China", "product category = electronic products") to the database, the retrieval range can be accurately narrowed. For example, in a customer service system, when a user's question involves the "return and exchange policy", the Agent automatically associates with the "after-sales service database" and preferentially calls its data, improving the response speed.
[0033] Support parallel retrieval of databases with different priorities. For example, high-priority databases (such as core transaction data) use real-time response, and low-priority databases (such as historical logs) use asynchronous processing to balance the system load.
[0034] Step 120: Perform vectorization processing on the data in the enterprise database to convert the data in text form into a numerical vector representation.
[0035] It should be noted that by converting the data in text form into a vector representation in a high-dimensional space, it is convenient for subsequent fast retrieval and efficient matching.
[0036] As an example, this step can be specifically: Obtain the first slice information through a preset slice acquisition interface; Based on the first slice information, the data in the enterprise database is sliced to obtain several sliced segments; Obtain the slicing granularity, and use the slicing granularity to further slice the sliced segments to obtain several knowledge fields; Convert the knowledge fields in text form into numerical vector representations.
[0037] Those skilled in the art can understand that by obtaining the slicing rules through a preset interface (such as by paragraph, semantic block, or fixed character length), a long text is sliced into logically complete segments.
[0038] Example: In a medical knowledge base, a clinical guideline document is sliced by "chapter title" into independent segments such as "indications", "medication specifications", "contraindications", etc., ensuring that each segment carries single-topic information.
[0039] Dynamically define the slicing granularity (such as sentence level, phrase level) based on business requirements to refine the extraction of knowledge fields.
[0040] Example: After a financial contract text is first sliced by "clause item", the keyword fields (such as "interest rate = 5%", "liquidated damages clause = Article 3.2") are extracted by secondary slicing, which is convenient for subsequent precise matching.
[0041] Through hierarchical slicing, redundant content (such as irrelevant descriptions) is removed to focus on the core knowledge units.
[0042] The finer the granularity, the more accurate the semantic representation after vectorization, reducing the risk of "false recall" (such as confusing "Apple Inc." with "fruit apple").
[0043] Use a pre-trained model (such as BERT, Sentence-BERT) to convert the text into dense vectors to capture deep semantic features.
[0044] Example: In an e-commerce scenario, after vectorizing the product description "waterproof Bluetooth headset", the vector similarity with "sports sweatproof wireless earbuds" reaches 0.92, achieving semantic matching across keywords.
[0045] Traditional keyword retrieval cannot recognize synonyms or context differences (such as "Java" referring to a programming language or coffee beans), while vector similarity can distinguish semantics.
[0046] Case: In a customer service system, a user's question "how to reset the password" can be associated with multiple documents such as "account security - password modification process" and "solutions for login anomalies", even if the word "reset" is not included.
[0047] After vectorization, the retrieval time can be reduced from minutes to milliseconds through approximate nearest neighbor search (ANN) algorithms (such as FAISS and HNSW), which is especially suitable for scenarios with data volumes in the hundreds of millions.
[0048] Step 130: Obtain user demand data through the user interaction interface, then extract retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database.
[0049] This step can be specifically as follows: Obtain user demand data through the user interaction interface; Input the user demand data into the large model, extract the key features in the user demand data, and then determine whether the key features belong to the preset trigger conditions; When it belongs to the preset trigger conditions, extract the retrieval constraint prompt words from the user demand data; Based on the corresponding relationship between the enterprise database, the retrieval constraint prompt words, and the preset retrieval priorities, determine the preset retrieval priorities corresponding to each enterprise database.
[0050] Those skilled in the art can understand that this step obtains the original demand (such as "I want to find customers with sales over 1 million in North China in the last three months") through the interaction interface (such as a chat window or a form). Use large language models (such as GPT and ERNIE) to parse the demand text and extract key features (such as the time range "in the last three months", the region "North China", and the business indicator "sales > 1 million"). Compare the key features with the preset rules (such as "region is mandatory" and "amount threshold triggers the risk control library") to screen out the effective constraint words.
[0051] Case: The user demand "need to handle customer complaints" may imply different scenarios (such as "product quality problems" or "logistics delays"). Extract the "complaint type" feature through the large model and associate it with the "after-sales problem library" (priority 8) or the "logistics tracking library" (priority 6).
[0052] Case: When the user asks "which suppliers meet the ESG standards", the large model recognizes "ESG" as the environmental, social, and governance standard, extracts the constraint word "ESG rating ≥ A" and gives priority to retrieving the "supplier compliance library" (priority 9).
[0053] This step can pre-define the association relationship between the enterprise database and the constraint words (such as "sales > 1 million" is bound to the "high-value customer library" with a priority of 10). When multiple constraint words are triggered, the priority is accumulated according to the weight (such as "North China region + sales > 1 million" increases the priority of the "high-value customer library" from 10 to 12).
[0054] In this step, business rules can be preset (such as "amount > threshold", "contains sensitive words"), and automatic matching can be performed through regular expressions or classification models. When the trigger condition is met, the default priority is temporarily overwritten (for example, the daily priority is 5, and it rises to 12 after the risk control condition is triggered).
[0055] Step 140: Calculate the vector representation with the highest similarity to the retrieval constraint prompt words using the adjustment weight corresponding to the preset retrieval priority; perform data analysis on the retrieval constraint prompt words to determine the corresponding guiding data; input the guiding data and the vector representation into the large model to obtain the output data.
[0056] Among them, the guiding data at least includes: role positioning guiding words, ability positioning guiding words, task objectives, and output requirements; In some embodiments, calculating the vector representation with the highest similarity to the retrieval constraint prompt words using the adjustment weight corresponding to the preset retrieval priority specifically includes: Calculate the initial similarity between the retrieval constraint prompt words and the vector representation through a language similarity algorithm; Multiply the adjustment weight by the initial similarity to obtain the final similarity, and then obtain the vector representation with the highest similarity.
[0057] Before performing data analysis on the retrieval constraint prompt words to determine the corresponding guiding data, the method further includes: Configure the corresponding relationship between the retrieval constraint prompt words and the guiding data.
[0058] Those skilled in the art can understand that this step can use a language similarity algorithm (such as cosine similarity, BM25) to calculate the semantic matching degree between the retrieval constraint prompt words and the vector representation. Multiply the adjustment weight corresponding to the preset retrieval priority (for example, when the priority of the enterprise database is 10, the weight = 1.2) by the initial similarity to dynamically correct the matching result.
[0059] Case: In the financial risk control scenario, if the priority of the "high-risk transaction rule library" is 15 (weight = 1.5), even if its initial similarity is 0.7, it is increased to 1.05 after weighting, and the risk control interception logic is triggered preferentially.
[0060] Case: Retrieve the "customer demand library" (weight = 1.0) and the "compliance clause library" (weight = 1.3) simultaneously to ensure that the generated content meets both customer demands and regulatory requirements.
[0061] In this step, the identity of the generation subject is defined (such as "professional financial advisor", "medical expert") to constrain the language style and knowledge boundary. Clear structured requirements such as the generation format (such as JSON, table), content length, and prohibited words are defined.
[0062] Case: When configuring "Role = Legal Counsel", the generated text automatically quotes relevant legal articles (such as Article XXX of the Civil Code) to avoid insufficient professionalism caused by general descriptions.
[0063] Case: By restricting the generated content based on the "Output Requirements" to only the retrieved vector data, the probability of fictional information is reduced (such as descriptions of drug side effects without retrieval).
[0064] In this step, the weight coefficients of different priority databases can be predefined (such as Core Business Database Weight = 1.5, Auxiliary Database Weight = 0.8). Establish the association rules between retrieval constraint words and guiding data (such as "Complaint Handling" is mapped to "Role = Customer Service Specialist", "Task Goal = Resolve within 24 hours").
[0065] Case: When the user enters "Recommend financial products", according to the priority weight (1.2) of the "Product Library" and the guiding word "Risk Level = Stable", a plan that matches the user's risk preference is generated.
[0066] Based on the above description, the present application discloses an enterprise data retrieval method based on a large model. By presetting the retrieval priority, it ensures that high-value databases (such as the core business system) have a higher weight in the retrieval, avoiding interference from low-priority databases (such as log backups) on key results. Retrieving high-weight databases first reduces unnecessary computational overhead and improves the system response speed. The combination of similarity calculation and priority weight makes the retrieval results more in line with the actual business needs, for example, giving priority to displaying financial data rather than non-critical documents.
[0067] Existing solutions lack active guidance for user needs and are prone to retrieval deviations due to vague requirements. By role positioning (such as "Financial Analyst") and ability positioning (such as "Data Visualization"), the user's identity and skill background are clarified, reducing semantic ambiguity. The introduction of task goals (such as "Generate quarterly reports") and output requirements (such as "Excel format") transforms vague requirements into executable retrieval conditions. Generating differentiated guidance strategies for different user roles (such as sales, R & D) improves the adaptability of the retrieval experience.
[0068] Existing solutions have insufficient capabilities in processing unstructured data and it is difficult to unify the semantic space. Vector representation breaks through the text form limitation and supports semantic matching of cross-modal data (such as documents, tables, logs). Numerical vectors can be directly input into the large model for reasoning, for example, by calculating vector similarity to associate relevant content in different databases. Support for multi-condition combination queries (such as "Find product defect records in customer complaints") improves the retrieval flexibility.
[0069] Existing solutions rely on a single retrieval model and are difficult to handle complex tasks. The large model combines the task objectives in the guiding data (such as "risk assessment report") and the vector retrieval results to generate structured and highly readable outputs (such as reports, charts). The large model can correct potential biases in the retrieval results (such as similarity calculation errors) and supplement missing information through context reasoning. It supports generating diverse output forms according to user needs (such as text summaries, visual charts), enhancing the practicality of the results.
[0070] In addition, this application Figure 2 provides an enterprise data retrieval system based on a large model according to an embodiment of this application. As Figure 2 shown, the system provided by the embodiment of this application mainly includes: The RAG module 210 is used to create several enterprise databases and configure the corresponding relationships among the enterprise databases, retrieval constraint prompt words, and preset retrieval priorities; perform vectorization processing on the data in the enterprise databases to convert the data in text form into numerical vector representations.
[0071] The RAG module 210 includes a revision unit, which is used to modify the data in the enterprise databases through the backend interface and / or modify the corresponding relationships among the enterprise databases, retrieval constraint prompt words, and preset retrieval priorities.
[0072] The user interaction module 220 is used to obtain user requirement data through the user interaction interface; The large model module 230 is used to extract retrieval constraint prompt words, and then determine the preset retrieval priorities corresponding to each enterprise database; calculate the vector representation with the highest similarity to the retrieval constraint prompt words with the adjustment weights corresponding to the preset retrieval priorities; perform data analysis on the retrieval constraint prompt words to determine the corresponding guiding data; where the guiding data at least includes: role positioning guide words, ability positioning guide words, task objectives, and output requirements; input the guiding data and the vector representation into the large model to obtain output data.
[0073] The large model module 230 includes a determination unit, which is used to extract the key features in the user requirement data, and then determine whether the key features belong to the preset trigger conditions; When belonging to the preset trigger conditions, extract the retrieval constraint prompt words from the user requirement data; Based on the corresponding relationships among the enterprise databases, retrieval constraint prompt words, and preset retrieval priorities, determine the preset retrieval priorities corresponding to each enterprise database.
[0074] In addition, the embodiments of the present application also provide a non-volatile computer storage medium, on which executable instructions are stored, and when the executable instructions are executed, a method for enterprise data retrieval based on a large model as described above is implemented.
[0075] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for enterprise data retrieval based on a large model, characterized in that: The method comprises: Create several enterprise databases, and configure the corresponding relationships among the enterprise databases, search constraint prompt words, and preset search priorities; Vectorize the data in the enterprise database and convert the text data into numerical vector representation; Through the user interaction interface, user demand data is obtained, and then search constraint prompt words are extracted, and then the preset search priority corresponding to each enterprise database is determined; Using the adjustment weight corresponding to the preset search priority, calculate the vector representation with the highest similarity to the search constraint prompt word; Performing data analysis on the search constraint prompt words to determine corresponding guide data; wherein the guide data at least includes: role positioning guide words, ability positioning guide words, task objectives and output requirements; The bootstrap data and vector representation are input into the large model to obtain the output data.
2. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: After creating a number of enterprise databases and configuring the preset search priorities of the enterprise databases, the method further includes: Modify the data in the enterprise database through the back-end interface, and / or modify the correspondence between the enterprise database, the search constraint prompt words and the preset search priority.
3. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: Vectorize the data in the enterprise database and convert the text data into numerical vector representation, including: Obtain the first slice information through the preset slice acquisition interface; Based on the first slicing information, the data in the enterprise database is segmented to obtain a number of segmented fragments; Obtain segmentation granularity, and use the segmentation granularity to segment the segmented fragments again to obtain several knowledge fields; Convert textual knowledge fields into numerical vector representations.
4. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: Through the user interaction interface, user demand data is obtained, and then the search constraint prompt words are extracted, and then the preset search priority corresponding to each enterprise database is determined, including: Obtain user demand data through the user interaction interface; Input user demand data into the big model, extract key features from the user demand data, and then determine whether the key features belong to the preset trigger conditions; When the preset trigger condition is met, the search constraint prompt words are extracted from the user demand data; Based on the correspondence between the enterprise databases, the search constraint prompt words and the preset search priorities, the preset search priorities corresponding to the various enterprise databases are determined.
5. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: The vector representation with the highest similarity to the search constraint prompt word is calculated using the adjustment weight corresponding to the preset search priority, including: Calculate the initial similarity between the retrieval constraint prompt word and the vector representation through the language similarity algorithm; By adjusting the weight and multiplying it with the initial similarity, the final similarity is obtained, and then the vector representation with the highest similarity is obtained.
6. The enterprise data retrieval method based on a large model according to claim 1 is characterized in that: Before performing data analysis on the search constraint prompt words to determine corresponding guide data, the method further includes: Configure the correspondence between search constraint hint words and guide data.
7. An enterprise data retrieval system based on a large model, characterized in that: The system comprises: RAG module is used to create several enterprise databases and configure the corresponding relationship between enterprise databases, search constraint prompt words and preset search priorities; vectorize the data in the enterprise database and convert the text data into numerical vector representation; A user interaction module is used to obtain user demand data through a user interaction interface; The large model module is used to extract search constraint prompt words, and then determine the preset search priority corresponding to each enterprise database; calculate the vector representation with the highest similarity to the search constraint prompt word with the adjustment weight corresponding to the preset search priority; perform data analysis on the search constraint prompt word to determine the corresponding guide data; wherein the guide data at least includes: role positioning guide words, ability positioning guide words, task objectives and output requirements; input the guide data and the vector representation into the large model to obtain output data.
8. The enterprise data retrieval system based on a large model according to claim 7, characterized in that: The RAG module includes revision units, Used to modify data in the enterprise database through the back-end interface, and / or modify the correspondence between the enterprise database, search constraint prompt words and preset search priorities.
9. The enterprise data retrieval system based on a large model according to claim 7, characterized in that: The large model module includes a determination unit, Used to extract key features from user demand data and then determine whether the key features belong to the preset trigger conditions; When the preset trigger condition is met, the search constraint prompt words are extracted from the user demand data; Based on the correspondence between the enterprise databases, the search constraint prompt words and the preset search priorities, the preset search priorities corresponding to the various enterprise databases are determined.
10. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, the method for enterprise data retrieval based on a large model as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Large model knowledge question and answer optimization method based on semantic enhanced knowledge graph
CN118939779A
AI large model enterprise knowledge base system based on cue words
CN119106143A
Industrial automatic product selection method, device and equipment and storage medium
CN119338005A
Information retrieval method and device and medium
CN119577126A
Information retrieval method and corporate information retrieval system
RU2019100812A
Cited By
Intelligent supervision system based on enterprise digitization
CN120822924A