Industrial economic query word segmentation retrieval method, system and equipment based on context enhancement, large model fine tuning and vector library and medium
By introducing context enhancement, large model fine-tuning and vector library into word segmentation technology, the problem of low matching between word segmentation results and search results in the existing technology is solved, and efficient and accurate word segmentation search for industrial economic query is achieved, which is suitable for complex industrial economic data analysis scenarios.
Patent Information
- Application Number
- CN202411838183.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing word segmentation technology has failed to effectively combine context enhancement and vectorized search mechanisms, resulting in a low matching between word segmentation results and search results, affecting query efficiency and accuracy, especially in the field of industrial economy, professional data support is limited.
The industrial economic query word segmentation search method based on context enhancement, large model fine-tuning and vector library is adopted. By preprocessing user query data, fine-tuning the big model, generating word segmentation models and keywords, and using vector library to search similarity, optimize SQL query statements and database query results.
It realizes automated, efficient and accurate database query, reduces the need for manual intervention, adapts to diversified industrial economic query scenarios, improves the accuracy and efficiency of query, and supports dynamic adaptive optimization.
Smart Images

Figure CN119988404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer technology and artificial intelligence technology, and in particular to an industrial economic query word segmentation retrieval method, system, device and medium based on context enhancement, large model fine-tuning and vector library. Background Art
[0002] Existing word segmentation methods include: 1. Dictionary-based word segmentation method: The text is segmented according to the longest match principle through a pre-defined dictionary. This method is simple to implement and is suitable for high-frequency word segmentation in specific fields.
[0003] 2. Statistical word segmentation method: A model is constructed by counting the word frequency, co-occurrence relationship and other information in the text to achieve word segmentation in a way that maximizes probability. Typical technologies include word segmentation algorithms based on hidden Markov models (HMM) and conditional random fields (CRF).
[0004] 3. Word segmentation method based on deep learning: Relying on deep neural networks, such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs), the word segmentation results are dynamically generated in combination with context information. This type of technology learns language rules through large-scale training corpus and has strong generalization ability.
[0005] 4. Word segmentation method based on pre-trained large models: With the rise of the Transformer architecture, large models such as BERT and GPT are used for word segmentation tasks. These models capture semantic relationships through contextual information and significantly improve the accuracy of word segmentation.
[0006] The existing word segmentation methods have the following main shortcomings: 1. Lack of dictionary dependence and scalability: The dictionary-based word segmentation method has weak ability to segment new words, domain-specific words, and misspelled words. It is necessary to frequently update the dictionary to meet the needs of different fields, and the maintenance cost is high.
[0007] 2. Insufficient use of contextual information: Statistical word segmentation methods usually rely only on local information and have difficulty accurately identifying words or phrases that are highly context-dependent. For example, the segmentation of "industrial economy" may have different meanings in different contexts.
[0008] 3. High demand for computing resources: Although word segmentation technology based on deep learning and large models has significant effects, its model training and reasoning process requires a lot of computing resources, and its performance is insufficient in real-time query or high-concurrency scenarios.
[0009] 4. Limited support for professional field data: Existing word segmentation models are mostly trained on general corpus, and lack support for data in specific fields (such as industrial economy). The word segmentation results often lack professionalism and accuracy.
[0010] 5. Lack of query optimization: In specific application scenarios such as industrial economic queries, the existing word segmentation technology fails to combine context enhancement and vectorized retrieval mechanisms, resulting in a low match between the word segmentation results and the retrieval results, affecting query efficiency and accuracy. Summary of the invention
[0011] Aiming at the problem that the existing word segmentation technology fails to combine context enhancement and vectorized retrieval mechanism, resulting in low matching degree between word segmentation results and retrieval results, thus affecting query efficiency and accuracy, the present invention proposes an industrial economic query word segmentation retrieval method, system, device and medium based on context enhancement, large model fine-tuning and vector library; the method first pre-processes the acquired original user inquiry data to obtain a user inquiry data set; secondly, fine-tunes the large model according to the user inquiry data set to obtain a word segmentation model and user inquiry keywords; then, inputs the acquired user questions into the word segmentation model to decompose and obtain multiple new keywords; matches preset SQL query statements according to the keywords to generate SQL query statements related to the new keywords; finally, performs query operations on the database according to the SQL query statements to obtain query results; and generates vectors according to the user questions, query results and keywords, and calls the vector library for similarity retrieval to generate answers; automatic, efficient and accurate database query is realized, and the method has the characteristics of low cost, easy implementation and dynamic adaptive optimization.
[0012] The specific implementation contents of the present invention are as follows: A method for industrial economic query word segmentation retrieval based on context enhancement, large model fine-tuning and vector library, specifically comprising the following steps: Step S1: preprocessing the acquired original user query data to obtain a user query data set; Step S2: fine-tune the large model according to the user query data set to obtain the word segmentation model and user query keywords; Step S3: input the acquired user questions into the word segmentation model to decompose and obtain multiple new keywords; Step S4: Generate a SQL query statement related to the new keyword according to the preset SQL query statement matching the keyword; Step S5: query the database according to the SQL query statement to obtain the query result; Step S6: Generate vectors based on user questions, query results, and keywords, and call the vector library for similarity search to generate answers.
[0013] Step S7: Collect the user's satisfaction with the query result. When the satisfaction exceeds a preset threshold, add the new question and the multiple new keywords to the user query data set to obtain a new user query data set; use the new user query data set to fine-tune the fine-tuned large model.
[0014] In order to better implement the present invention, further, the step S1 specifically includes the following steps: Step S11: obtaining and preprocessing original user query data to obtain a preliminary user query data set; Step S12: calling the large model to complete the context information of the preliminary user query data set to obtain a completed user query data set.
[0015] In order to better implement the present invention, further, the step S2 specifically includes the following steps: Step S21: converting the original user query data into a format matching the model according to the constructed preprocessing function; Step S22: selecting a target data volume, and preprocessing the converted user inquiry data according to the constructed preprocessing function to obtain user inquiry training data; Step S23: Initialize the data organizer and set the evaluation index function; Step S24: calling the Lora algorithm to fine-tune the large model according to the user query training data to obtain a word segmentation model; Step S25: split the user's question into keywords to be searched according to the word segmentation model to obtain the user's query keywords.
[0016] In order to better implement the present invention, further, the specific operation of step S21 is: converting the original user query data into json format according to the constructed preprocessing function.
[0017] In order to better implement the present invention, further, the user query data set includes a plurality of questions related to industrial economy and their corresponding keyword groups.
[0018] In order to better implement the present invention, further, the large model adopts a Qwen model with a scale of more than 6B and trained with more than billions of corpora.
[0019] Based on the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library, in order to better realize the present invention, further, an industrial economic query word segmentation retrieval system based on context enhancement, large model fine-tuning and vector library is proposed, which is used to execute the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library; including a data cleaning module, a model fine-tuning module, a keyword analysis module, an SQL matching module, and a vector query module; The data cleaning module is used to pre-process the acquired original user query data to obtain a user query data set; The model fine-tuning module is used to fine-tune the large model according to the user query data set to obtain a word segmentation model and user query keywords; The keyword analysis module inputs the acquired user questions into the word segmentation model to decompose and obtain multiple new keywords; The SQL matching module is used to match the preset SQL query statement according to the keyword and generate the SQL query statement related to the new keyword; The vector query module is used to query the database according to the SQL query statement to obtain the query result; generate vectors according to the user question, query result and keywords, and call the vector library to perform similarity search to generate answers.
[0020] Based on the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library, in order to better realize the present invention, further, an electronic device is proposed, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library is implemented.
[0021] Based on the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library, in order to better realize the present invention, further, a computer-readable storage medium is proposed, on which computer instructions are stored; when the computer instructions are executed on the above-mentioned electronic device, the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library is implemented.
[0022] The present invention has the following beneficial effects: (1) The present invention significantly reduces the need for manual intervention through context enhancement and automated multimodal analysis functions; it can intelligently adapt to diverse query scenarios in industrial economics, such as policy analysis, market forecasting, and industry dynamics tracking, thereby saving a lot of time and human resources.
[0023] (2) The present invention has excellent generalization ability through fine-tuning of professional corpus in the field of industrial economics, and can adapt to various query scenarios such as enterprise demand analysis and market competitiveness assessment. At the same time, multimodal fine-tuning technology supports input forms such as text, charts and voice, enabling the system to perform well in applications such as industrial report interpretation and data visualization analysis.
[0024] (3) The present invention supports implicit user feedback and retrieval effect analysis based on industrial economic queries, and realizes dynamic adjustment of model parameters through the vector library. For example, when a user queries "investment trends in new energy industries", the system can gradually optimize the retrieval results, improve accuracy and relevance, and meet the user's real-time needs for changes in industry data.
[0025] (4) The intelligent recommendation and real-time prompt functions of the present invention lower the technical threshold. Even users with no database experience can obtain accurate query results through simple operations and easily complete complex analysis tasks such as "industry growth forecast" or "regional economic comparison".
[0026] (5) The present invention utilizes the historical data matching capability of the vector library to rapidly generate preliminary query results; through automated learning and feedback mechanisms, it is able to dynamically identify and correct potential errors that may affect industry analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A schematic block diagram of the structure of an industrial economic query word segmentation retrieval system based on context enhancement, large model fine-tuning and vector library provided in an embodiment of the present invention.
[0028] Figure 2 A schematic flowchart of an industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be understood that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, and therefore should not be regarded as limiting the scope of protection. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technical personnel in this field without making creative work are within the scope of protection of the present invention.
[0030] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "disposed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0031] Embodiment 1: This embodiment proposes an industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library, which specifically includes the following steps: Step S1: pre-process the acquired original user query data to obtain a user query data set.
[0032] The step S1 specifically includes the following steps: Step S11: obtaining and preprocessing original user query data to obtain a preliminary user query data set; The user query data set includes a plurality of questions related to industrial economics and their corresponding keyword groups.
[0033] Step S12: calling the large model to complete the context information of the preliminary user query data set to obtain a completed user query data set.
[0034] The large model uses the Qwen model, which is over 6B in size and has been trained with over billions of corpora.
[0035] Step S2: Fine-tune the large model according to the user query data set to obtain the word segmentation model and user query keywords.
[0036] The step S2 specifically includes the following steps: Step S21: converting the original user query data into a format matching the model according to the constructed preprocessing function; The specific operation of step S21 is: converting the original user query data into json format according to the constructed preprocessing function.
[0037] Step S22: selecting a target data volume, and preprocessing the converted user inquiry data according to the constructed preprocessing function to obtain user inquiry training data; Step S23: Initialize the data organizer and set the evaluation index function; Step S24: calling the Lora algorithm to fine-tune the large model according to the user query training data to obtain a word segmentation model; Step S25: split the user's question into keywords to be searched according to the word segmentation model to obtain the user's query keywords.
[0038] Step S3: Input the acquired user question into the word segmentation model to decompose and obtain multiple new keywords.
[0039] Step S4: Generate a SQL query statement related to the new keyword according to the preset SQL query statement matched with the keyword.
[0040] Step S5: perform a query operation on the database according to the SQL query statement to obtain the query result.
[0041] Step S6: Generate vectors based on user questions, query results, and keywords, and call the vector library for similarity search to generate answers.
[0042] Step S7: Collect the user's satisfaction with the query result. When the satisfaction exceeds a preset threshold, add the new question and the multiple new keywords to the user query data set to obtain a new user query data set; use the new user query data set to fine-tune the fine-tuned large model.
[0043] Working principle: This embodiment first uses the user query data set to fine-tune the large model that has completed the corpus training, so that the large model can learn and master the process of context understanding and multimodal fusion, and obtain a fine-tuned large model that can obtain keyword groups from user questions, and then uses the fine-tuned large model to perform semantic analysis and context completion on the new questions input by the user to obtain multiple new keywords. Finally, according to the multiple new keywords, the database is queried and the query results are returned to the user. In this way, not only an automated, efficient and accurate database query solution is provided, but also low-cost, easy to implement and dynamically adaptive optimization are provided, providing a new solution for database query, which is convenient for practical application and promotion. In addition, a vector library is introduced to further improve the response speed and efficiency of the system. By vectorizing the keyword groups generated by the large model and storing them in the vector library, the system can use the vectors in the vector library to perform fast similarity matching after receiving the new questions input by the user, and find the answer closest to the semantics of the historical questions without retraining the large model or re-inference to generate answers. With the assistance of the vector library, the large model can quickly adjust and optimize the final answer based on similar historical query results, thereby achieving more efficient data.
[0044] Embodiment 2: This embodiment is based on the above embodiment 1. Figure 2 As shown, a specific embodiment is described in detail.
[0045] Step S1: Data cleaning.
[0046] The original user query data is preprocessed to remove redundant and irrelevant information and obtain a preliminary user query data set.
[0047] Use Tongyi Qianwen Large Model 2.5-32b to filter out which conversations are useful and which are useless by combining prompt words with thought chains, and then use Tongyi Qianwen Large Model 2.5-32b to optimize the question and answer format and optimize the user's questions.
[0048] Data enhancement: Based on the original data cleaning, the contextual data enhancement function is added. Not only does it clean up redundant information, but it also automatically identifies and completes the contextual information related to the user's query. For example, when a user asks the question "What is the situation of enterprises in Shanghai?" the system uses a large model combined with the context to automatically associate and supplement the background data related to the question, such as "enterprise size" and "industry distribution", to better support subsequent keyword analysis and SQL generation.
[0049] Step S2: Fine-tune the large model.
[0050] a. Multimodal large model selection: Select a basic large model that can process multimodal information such as text, images, and audio, such as the latest multimodal transformer architecture. Through the joint training of Lora with different types of data, the system's understanding and retrieval capabilities of complex, multi-dimensional data are improved. b. Use the preliminary user query data set to fine-tune the large model.
[0051] Fine-tuning steps: 1. Data preprocessing: Define the preprocessing function (preprocess_function_train), which is responsible for converting the raw conversation data into a format acceptable to the model. Tokenize each conversation instance and pad and truncate it appropriately.
[0052] 2. Prepare training data: Select the amount of data you need, which could be the entire dataset or just a subset.
[0053] Use the preprocessing function to process the raw data to obtain data suitable for training.
[0054] 3. Set training parameters and tools: Initialize the data collator, which is responsible for processing data during training.
[0055] Define evaluation metric functions such as Rouge and BLEU scores.
[0056] 4. Model fine-tuning: Use transformers to fine-tune Lora on the preprocessed data.
[0057] Perform model evaluation and prediction as needed.
[0058] 5. Evaluation and prediction: Once the model is fine-tuned, it can be evaluated and predicted on the validation and test datasets to check the performance of the model.
[0059] Record and save evaluation metrics.
[0060] 6. Save the results: Decode the prediction output of the model and save it to a file.
[0061] b. Split user questions into keywords that need to be searched, for example, (What is the number of artificial intelligence companies in Shanghai? - Shanghai, artificial intelligence, number of companies), so that the big model can learn and master the process of text segmentation (dividing the questions raised by users into multiple keywords) to obtain the key information in the user's inquiry.
[0062] The fine-tuning operation involved in this embodiment adopts the LoRA large model fine-tuning method, the Adapter large model fine-tuning method, the Prefix-tuning large model fine-tuning method, the P-tuning large model fine-tuning method or the Prompt-tuning large model fine-tuning method.
[0063] Step S3: Keyword analysis.
[0064] a. Receive user input of questions.
[0065] b. Use the fine-tuned large model to perform text analysis on the user’s question and break it down into multiple keywords.
[0066] Semantic completion and analysis: a. Contextual semantic completion: By using a large model to automatically complete the context of the user's question, the system can better analyze the user's potential intentions. For example, if the user asks "AI companies in Shanghai", the system automatically completes "policy support", "market share" and other potential key information related to the question. b. Deep semantic network analysis: Using semantic network technology, combined with knowledge graphs, the user's questions are semantically decomposed and associated, so as to extract more comprehensive key information.
[0067] Step S4: SQL matching.
[0068] a. Match the preset SQL query statements based on the keywords extracted from the user's question.
[0069] b. Generate specific SQL query statements related to keywords.
[0070] Step S5: Data retrieval.
[0071] a. Use the generated SQL query statement to query the database.
[0072] b. Get the query results and return them to the user.
[0073] Model learning: a. Collect user feedback and retrieval effect data. b. Use this data to further train and learn the large model. c. Continuously update and optimize the model to improve the ability and accuracy of data retrieval.
[0074] Step S6: fast vector query.
[0075] a. For each question entered by the user, use the keywords to generate the corresponding vector, and use the vector library to perform similarity search.
[0076] b. Through vector query, quickly locate keywords and related answers similar to historical questions.
[0077] c. Use the search results in the vector library to help the large model optimize the answer, thereby achieving efficient data retrieval and accurate answer generation.
[0078] Step S7: Real-time recommendation.
[0079] a. Real-time question recommendation: When a user enters a query, the system recommends related questions or answers in real time based on existing query records and keywords, helping users narrow the query scope and speed up the query.
[0080] b. Intelligent prompts and completion: During the user input process, the system can automatically provide keyword completion and intelligent prompts to help users construct clearer query content.
[0081] Specific implementation example of fine-tuning model: 1. First, manually write a batch of raw data: for example: What is the development status of AI companies in Shanghai? ——[“Beijing”, “AI company”, “development status”], how many AI companies will there be in Beijing in 2024? ——[“Beijing”, “AI company”, “2024”]; 2. Convert the above question “How is the development of AI companies in Shanghai?” into a voice use case; 3. Convert the raw data into json format for fine-tuning; 4. Use the Lora algorithm to fine-tune Qwen-2.5-9B to obtain a new word segmentation model.
[0082] 5. Use the Lora algorithm to perform multi-modal fine-tuning on Qwen2-VL-7B-Instruct.
[0083] 1. User question: "What is the development status of AI companies in Beijing?"; 2. Context enhancement and word segmentation: The system parses the question into keywords: "Beijing", "AI company", "development status", and further associates it with related words such as "policy" and "technological innovation" through context completion (using Qwen-2.5-32B to complete in combination with user context).
[0084] 3. Vector library search: The segmented keywords "Beijing" and "AI company" are converted into vector representations and matched with historical questions and keywords in the vector library for similarity.
[0085] Use the rearrangement model to sort the results and return them; The vector library is used to quickly find the historical query results closest to the current question. If similar query records exist, the cached query results are directly returned, thus avoiding the need to reconstruct the SQL statement and execute the query.
[0086] 4. If the user uses voice, Qwen2-VL-7B-Instruct will be used to segment the user's question and repeat the above steps 5. Convert the segmented content into the corresponding SQL query statement: SELECT company_name, establishment_date, funding_rounds, total_funding FROM ai_companies WHERE location = 'Beijing' ORDER BY total_funding DESC; 6. Find the result and return it to the user.
[0087] 7. Vector library update: The current user question and its word segmentation results, SQL query statements, and returned results are stored in the vector library for quick matching in subsequent queries to optimize response speed.
[0088] 8. The system automatically records query logs and periodically iterates itself based on model feedback. By analyzing the logs, the system can identify new user needs or query trends and take them into consideration when the model is updated next time, ensuring that the system always maintains optimal performance.
[0089] 9. Use Qwen-2.5-32B regularly to clean up user questions.
[0090] 10. Use the model to extract records where users are dissatisfied.
[0091] 11. Manually complete unsatisfactory questions.
[0092] 12. Use the Lora algorithm to further fine-tune Qwen-2.5-9B.
[0093] This embodiment reduces the need for manual configuration. Data queries in the traditional industrial economy usually require professionals to perform complex manual configuration and optimization of the database structure. This technology significantly reduces the need for manual intervention through context enhancement and automated multimodal analysis functions. The system can intelligently adapt to diverse query scenarios in the industrial economy, such as policy analysis, market forecasting, and industry dynamics tracking, thereby saving a lot of time and human resources.
[0094] This embodiment has strong generalization ability. The large model of this embodiment has excellent generalization ability through fine-tuning of professional corpus in the field of industrial economy, and can adapt to various query scenarios such as enterprise demand analysis and market competitiveness evaluation. At the same time, multimodal fine-tuning technology supports input forms such as text, charts and voice, enabling the system to perform well in applications such as industrial report interpretation and data visualization analysis.
[0095] This embodiment has the characteristics of continuous learning and optimization. The system supports implicit user feedback and retrieval effect analysis based on industrial economic queries, and realizes dynamic adjustment of model parameters through the vector library. For example, when a user queries "investment trends in new energy industries", the system can gradually optimize the retrieval results, improve accuracy and relevance, and meet the user's real-time needs for industry data changes.
[0096] This embodiment lowers the entry threshold. Industrial economic data analysis often requires professional skills, and the intelligent recommendation and real-time prompt functions of this system lower the technical threshold. Even users without database experience can obtain accurate query results through simple operations and easily complete complex analysis tasks such as "industry growth forecast" or "regional economic comparison".
[0097] This embodiment improves the query effect. The system combines the semantic analysis of the big model in the field of industrial economy and the similarity retrieval capability of the vector library to quickly return the analysis results related to the query. For example, if the user enters "2023 intelligent manufacturing policy impact analysis", the system can accurately match relevant industry reports, policy interpretations and expert comments, improving the query efficiency and accuracy.
[0098] This embodiment achieves rapid iteration. In the face of rapidly changing market environments and industry dynamics, the system can quickly respond through the vector library without retraining large models. For example, the newly issued "green financial policy" data can be quickly integrated into the system, providing relevant query support in real time, helping enterprises to respond to policy changes in a timely manner and optimize decision-making efficiency.
[0099] This embodiment enhances data utilization. This embodiment introduces multimodal training and context enhancement technology to fully utilize text such as reports, images such as statistical charts, and voice such as expert interpretation in industrial economic data. For example, the system can associate text in economic data reports with visual charts to provide users with more in-depth industry trend analysis.
[0100] This embodiment improves user satisfaction. The system supports personalized sorting and intelligent feedback, and can adjust search results according to the preferences of industry analysts or corporate users. For example, for users who are interested in the "smart home appliance market", the system will prioritize displaying relevant data and in-depth reports, thereby improving the relevance and satisfaction of each query.
[0101] This embodiment reduces storage costs. The vector library optimizes storage requirements through efficient compression technology, significantly reducing the storage and processing costs of industrial economic data. Compared with the traditional item-by-item comparison method, this technology improves the retrieval speed while saving the operation and maintenance resources of the industry database.
[0102] This embodiment has flexible scalability, and the system design supports dynamic expansion. It can easily integrate new data modules such as "new energy industry" or "regional economic forecast", helping users to quickly respond to changing market demands and policy environments.
[0103] This embodiment reduces the cold start problem. For new industries or data sources (such as startup industry analysis), the system uses the historical data matching capabilities of the vector library to quickly generate preliminary query results. For example, data analysis related to emerging fields such as the "metaverse economy" can be quickly deployed and provide high-quality support.
[0104] This embodiment improves the model security. Through automated learning and feedback mechanisms, the system can dynamically identify and correct potential errors that may affect the results of industry analysis. For example, when key terms in policy interpretation are misunderstood, the system will automatically optimize the processing logic to ensure the reliability of the analysis results and reduce the risk of decision-making errors.
[0105] The other parts of this embodiment are the same as those of the above-mentioned embodiment 1, and thus will not be described in detail.
[0106] Embodiment 3: Based on any one of the above-mentioned embodiments 1-2, this embodiment proposes an industrial economic query word segmentation retrieval system based on context enhancement, large model fine-tuning and vector library, which is used to execute the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library; it includes a data cleaning module, a model fine-tuning module, a keyword analysis module, an SQL matching module, and a vector query module; The data cleaning module is used to pre-process the acquired original user query data to obtain a user query data set; The model fine-tuning module is used to fine-tune the large model according to the user query data set to obtain a word segmentation model and user query keywords; The keyword analysis module inputs the acquired user questions into the word segmentation model to decompose and obtain multiple new keywords; The SQL matching module is used to match the preset SQL query statement according to the keyword and generate the SQL query statement related to the new keyword; The vector query module is used to query the database according to the SQL query statement to obtain the query result; generate vectors according to the user question, query result and keywords, and call the vector library to perform similarity search to generate answers.
[0107] like Figure 1 As shown, the industrial economic query word segmentation retrieval system based on context enhancement, large model fine-tuning and vector library provided in this embodiment includes a data set acquisition module, a model fine-tuning module, a question receiving module, a model application module and a query operation module; The data set acquisition module is used to acquire a user query data set for fine-tuning the large model, wherein the user query data set includes a plurality of user questions and a plurality of user question keyword groups corresponding to the plurality of user questions one by one, and the user question keyword groups include a plurality of keywords manually extracted from the corresponding user questions in advance; The model fine-tuning module is communicatively connected to the data set acquisition module and is used to use the user query data set to perform a fine-tuning operation on the large model that has completed the corpus training, so as to allow the large model to learn and master the process of text segmentation, and obtain a fine-tuned large model that can obtain keyword groups from user questions; The question receiving module is used to receive new questions input by users; The model application module is respectively connected to the model fine-tuning module and the question receiving module for performing text segmentation processing on the new question using the fine-tuned large model to obtain a plurality of new keywords; The query operation module is communicatively connected to the model application module, and is used to perform a query operation on the database according to the multiple new keywords, and return the query results to the user.
[0108] It also includes a satisfaction collection module and a data set addition module; The satisfaction collection module is communicatively connected to the query operation module and is used to collect the user's satisfaction with the query result; The data set adding module is respectively connected to the satisfaction collection module and the model application module for adding the new question and the multiple new keywords to the user inquiry data set to obtain a new user inquiry data set when the satisfaction exceeds a preset threshold; The model fine-tuning module is also communicatively connected to the data set adding module, and is used to use the new user query data set to perform fine-tuning operations on the fine-tuned large model.
[0109] This embodiment also proposes an electronic device, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library is implemented.
[0110] This embodiment also proposes a computer-readable storage medium, on which computer instructions are stored; when the computer instructions are executed on the above-mentioned electronic device, the above-mentioned industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library is implemented.
[0111] The other parts of this embodiment are the same as any one of the above-mentioned embodiments 1-2, so they will not be repeated here.
[0112] The processor involved in the embodiments of the present application may be a chip. For example, it may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0113] The memory involved in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0114] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0115] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0116] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0117] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0118] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one device or distributed on multiple devices. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0119] In addition, each functional module in each embodiment of the present application may be integrated into one device, or each module may exist physically separately, or two or more modules may be integrated into one device.
[0120] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or may include one or more servers, data centers and other data storage devices that can be integrated with the medium. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).
[0121] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for industrial economic query word segmentation retrieval based on context enhancement, large model fine-tuning and vector library, characterized in that: The specific steps include: Step S1: preprocessing the acquired original user query data to obtain a user query data set; Step S2: fine-tune the large model according to the user query data set to obtain the word segmentation model and user query keywords; Step S3: input the acquired user questions into the word segmentation model to decompose and obtain multiple new keywords; Step S4: Generate a SQL query statement related to the new keyword according to the preset SQL query statement matching the keyword; Step S5: query the database according to the SQL query statement to obtain the query result; Step S6: Generate vectors based on user questions, query results, and keywords, and call the vector library for similarity search to generate answers.
2. According to claim 1, the industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library is characterized in that: The step S1 specifically includes the following steps: Step S11: obtaining and preprocessing original user query data to obtain a preliminary user query data set; Step S12: calling the large model to complete the context information of the preliminary user query data set to obtain a completed user query data set.
3. The method for industrial economic query word segmentation retrieval based on context enhancement, large model fine-tuning and vector library according to claim 2 is characterized in that: The step S2 specifically includes the following steps: Step S21: converting the original user query data into a format matching the model according to the constructed preprocessing function; Step S22: selecting a target data volume, and preprocessing the converted user inquiry data according to the constructed preprocessing function to obtain user inquiry training data; Step S23: Initialize the data organizer and set the evaluation index function; Step S24: calling the Lora algorithm to fine-tune the large model according to the user query training data to obtain a word segmentation model; Step S25: split the user's question into keywords to be searched according to the word segmentation model to obtain the user's query keywords.
4. The method for industrial economic query word segmentation retrieval based on context enhancement, large model fine-tuning and vector library according to claim 3 is characterized in that: The specific operation of step S21 is: converting the original user query data into json format according to the constructed preprocessing function.
5. The industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library according to claim 1 is characterized in that: The industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library also includes: Step S7: Collect the user's satisfaction with the query result. When the satisfaction exceeds a preset threshold, add the new question and the multiple new keywords to the user query data set to obtain a new user query data set; use the new user query data set to fine-tune the fine-tuned large model.
6. The method for industrial economic query word segmentation retrieval based on context enhancement, large model fine-tuning and vector library according to claim 1 is characterized in that: The user query data set includes a plurality of questions related to industrial economics and their corresponding keyword groups.
7. The industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library according to claim 1 is characterized in that: The large model uses the Qwen model, which is over 6B in size and has been trained with over billions of corpora.
8. An industrial economic query word segmentation retrieval system based on context enhancement, large model fine-tuning and vector library, used to execute the industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library as claimed in claim 1; characterized in that: Including data cleaning module, model fine-tuning module, keyword analysis module, SQL matching module, and vector query module; The data cleaning module is used to pre-process the acquired original user query data to obtain a user query data set; The model fine-tuning module is used to fine-tune the large model according to the user query data set to obtain a word segmentation model and user query keywords; The keyword analysis module inputs the acquired user questions into the word segmentation model to decompose and obtain multiple new keywords; The SQL matching module is used to match the preset SQL query statement according to the keyword and generate the SQL query statement related to the new keyword; The vector query module is used to query the database according to the SQL query statement to obtain the query result; It generates vectors based on user questions, query results, and keywords, and calls the vector library for similarity retrieval to generate answers.
9. An electronic device, characterized in that: It includes a memory and a processor; the memory stores a computer program; when the computer program is executed on the processor, it implements the industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions; when the computer instructions are executed on the electronic device as described in claim 8, the industrial economic query word segmentation retrieval method based on context enhancement, large model fine-tuning and vector library as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Word segmentation retrieval method, device and equipment based on large model fine tuning and storage medium
CN117555992A
Class case retrieval system and method based on retrieval enhancement generation technology
CN118260391A
Vector database retrieval method and device based on large model, terminal and medium
CN118312594A
Retrieval enhancement method combining keyword extraction and semantic analysis
CN118535682A
Continuous question answering method and device based on large language model and electronic equipment
CN118964585A
Cited By
Knowledge-driven multi-mode large model lung cancer postoperative rehabilitation guidance method and system
CN120674011A
Huangmuo paddling inheriting and innovating method based on large model
CN120930743A