Method for optimizing networking retrieval capability of large model
By optimizing word segmentation algorithms and parameter settings, combining advanced network protocols and distributed architectures, intelligent cache is introduced, which solves the network delay, privacy leakage, concurrent processing and cache problems of network search systems, improves retrieval efficiency and security, and provides results that are more in line with user needs.
Patent Information
- Application Number
- CN202510141588.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-07-04
AI Technical Summary
The existing networked retrieval systems have problems such as high network latency, risk of privacy leakage, limited concurrent request processing capabilities, unintelligent search logic and lack of effective caching mechanisms, resulting in low retrieval efficiency and poor user experience.
By optimizing word segmentation algorithm and parameter settings, combining professional field knowledge, dynamically adjusting word segmentation strategies, introducing advanced network protocols and caching mechanisms, adopting distributed microservice architecture, optimizing retrieval algorithms using TF-IDF and BERT hybrid strategies, and introducing an intelligent cache system.
It improves the search speed and system response sensitivity, enhances information security management, improves concurrent request processing capabilities, improves the relevance of search results, reduces database access pressure, and ensures the stability and efficiency of the system.
Smart Images

Figure CN120256700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and particularly to an optimization method for the internet retrieval ability of large models. Background Art
[0002] With the rapid development of information technology, the Internet has become one of the main channels for people to obtain information. In the face of a vast amount of information, how to retrieve the required content efficiently and accurately has become an urgent problem to be solved. Traditional search engines usually provide search results based on keyword matching. However, this method often fails to well understand the user's query intention, resulting in a deviation between the returned results and the user's true needs. In addition, with the progress of natural language processing and machine learning technologies, large models (such as deep learning models) have begun to be applied in the field of information retrieval in order to optimize the search quality through more advanced understanding capabilities.
[0003] Nevertheless, the existing search technologies still face a series of challenges. First, the high network latency problem in the public network environment seriously affects the user's search experience. The time from submitting a query to receiving a response is relatively long, reducing the search efficiency. Second, in terms of data protection, when processing content containing personal or enterprise sensitive information, if there is a lack of effective security measures, there may be a risk of privacy leakage. Moreover, the system's processing ability for concurrent requests is limited. Once there is a surge in user requests, the background server may experience performance bottlenecks or even crashes due to excessive occupation of computing resources, which directly affects the stability and availability of the service. In addition, the current retrieval logic and rules are not yet intelligent enough to accurately capture the user's query intention, and sometimes it will feedback non-desired results, thus affecting the user experience. Finally, the lack of an effective caching mechanism is also a major defect. Each search needs to directly interact with the background database to obtain the latest data, which not only increases the burden on the storage system but also further slows down the overall operation speed.
[0004] To solve the above problems, the present invention proposes a new method for optimizing the large model's online retrieval ability. This method aims to improve search efficiency, accuracy, and enhance the security and stability of the system by improving the retrieval process and introducing more intelligent algorithms and technical means. Specifically, this method includes but is not limited to: deeply parsing the received user query and invoking the search engine to obtain preliminary retrieval results; screening and analyzing these results to extract key text content; further processing the text using preprocessing techniques such as word segmentation and part-of-speech tagging, and using statistical methods such as TF-IDF to evaluate the importance of words; on this basis, selecting the optimal results through relevance score ranking and integrating their summaries to form the final summary for users' reference. At the same time, the present invention also pays special attention to dynamically adjusting the word segmentation strategy to adapt to the characteristics of different types of text, and proposes several innovative parameter setting rules and conflict resolution methods to better serve diverse application scenarios. Through these comprehensive measures, the present invention strives to create a new generation of search solution that is both fast and accurate while effectively protecting user privacy. Summary of the Invention
[0005] In view of the problems existing in the above-mentioned existing methods for optimizing the large model's online retrieval ability, the present invention is proposed.
[0006] Therefore, the problems to be solved by the present invention are as follows: First, due to the high network latency of the public network, the retrieval request takes a long time during transmission, resulting in a slow retrieval response speed perceived by the user side; second, due to the imperfect protection measures of the system for personal and enterprise data, there may be a security risk of privacy leakage when processing content containing sensitive information; third, the number of concurrent requests of the system is limited. When the user request volume surges beyond the threshold, the background server is difficult to quickly respond to a large number of simultaneous online requests, resulting in high occupancy of computing resources and even the risk of collapse; fourth, the retrieval logic and rules need to be further intelligently upgraded, and there are currently problems such as inability to accurately identify intentions and feedback of non-desired results; fifth, there is no effective caching mechanism to support repeated requests, and each search needs to directly interact with the background data source to obtain the latest information, which increases the burden on the storage system and reduces the overall operation rate. These five aspects are all thorny problems that need to be seriously considered and solved when optimizing information retrieval using large pre-trained models in a networked environment.
[0007] To solve the above technical problems, the present invention provides the following technical solution: A method for optimizing the large model's online retrieval ability, comprising the following steps:
[0008] Receive the user query statement, call the search engine to retrieve the results and list them;
[0009] Perform a preliminary screening on the search result list, obtain the HTML source code through web links, and analyze the HTML source code to extract the text of the web page body;
[0010] Preprocess the extracted text content, extract the keyword list, and perform keyword matching in the preprocessed text content to count the number of keyword occurrences;
[0011] Calculate the relevance score of each search result, select several search results with the highest scores, perform text summarization on the selected results and integrate them to generate a summary of the query statement;
[0012] Output the summary as the final search result and present it to the user.
[0013] As a preferred solution of the optimization method for the large model's online retrieval ability described in the present invention, the preprocessing includes word segmentation, part-of-speech tagging, and calculation of the TFIDF value of each word;
[0014] The process of the word segmentation includes:
[0015] Download and configure the required dependency libraries and versions for the model environment;
[0016] Set the main parameters of the word segmentation algorithm;
[0017] Define the segmentation strategy.
[0018] As a preferred solution of the optimization method for the large model's online retrieval ability described in the present invention, defining the segmentation strategy includes:
[0019] Set the main parameters of the word segmentation algorithm;
[0020] Define the specific dictionary matching logic and conflict resolution method, and select the priority according to the word frequency ranking when multiple segmentation methods are applicable at the same time;
[0021] Dynamically set parameters to handle different types of texts. If it is recognized that the currently processed text is a subjective comment, reduce the empirical coefficient value.
[0022] As a preferred solution of the optimization method for the large model's online retrieval ability described in the present invention, the specific details of setting the main parameters of the word segmentation algorithm include:
[0023] Including the memory management for allocating task processor resources and setting steps such as dynamic adjustment of the sliding window:
[0024] The rule for setting the longest word search threshold is as follows;
[0025] In the initial state, estimate the most likely length based on the statistical information of the text obtained in the previous steps;
[0026] When adjusting the threshold, consider the text characteristics and add an empirical value η for calculation;
[0027] If the detected text content is relatively complex or has a large number of special symbols, adjust the empirical factor to make it more flexible and adaptive. For a document with a higher complexity, the adjustment ratio η of the search threshold should be increased. The new calculation formula becomes Threshold = original length γδ + η correction, where γ represents the length multiplier and δ is the text difficulty adjustment factor.
[0028] As a preferred solution of the optimization method for the large model's online retrieval ability described in the present invention, the correction factors to be added as needed during the dynamic configuration adjustment process include the following detailed steps:
[0029] Measure the average correction factor for each word type in the current text;
[0030] When encountering a change in text category, calculate the estimated approximate value of the correction factor based on the existing approximate categories and use this value to calculate the exact value required for subsequent operations;
[0031] Use the formula original value adjustment coefficient (ξ) to determine the new correction factor F;
[0032] Implement special processing for data where the correction factor exceeds the set interval: when the calculated new word correction factor is greater than the threshold λ, the correction value of this entry deserves extra attention and an extra weight is assigned to it.
[0033] As a preferred solution of the optimization method for the large model's online retrieval ability described in the present invention, the process of judging the importance of the final words includes the following specific operation procedures to accurately determine their actual contribution degrees:
[0034] Detect whether it exceeds the pre-established importance benchmark. If it is determined that the importance degree of a specific word exceeds the established important level, that is, whether C > C base standard C magnification after multiplying the importance degree C of the word by the upper limit adjustment multiple, then mark it as this highest level P_HIGH;
[0035] Assign the maximum priority status flag to each word that exceeds the importance threshold for quick access;
[0036] If the word does not meet the requirements of the highest level, mark it according to the usual method. When necessary, the magnification factor can be reduced until it equals the basic multiple 1 to prevent overemphasis on any entry;
[0037] Special measures need to be taken for some words that have been assigned the top priority flag. Place a mark at the corresponding index entry of these entries for efficient tracking.
[0038] As a preferred solution of the optimization method for the large model's online retrieval ability described in the present invention, when the word segmentation process detects an overlap conflict in the dictionary entries:
[0039] The algorithm will give priority to words with higher IDF scores and TF that meet certain range requirements as the preferred word boundary cutting points;
[0040] Determine whether to use the specified term to participate in the final word segmentation, the condition is that the document appearance frequency of the term should exceed the pre-set standard threshold value X;
[0041] If the system automatically parses and finds that the document being operated is of the news report style, the accuracy factor of the term boundary division is enhanced. The formula is defined as: "segmentation accuracy = original accuracy level + δX news type bonus", where δX is a constant value dedicated to news documents;
[0042] And in the content that determines that the sentence is a statement of personal opinion, the empirical value η of the parameter ξ used to adjust the cutting acuity is dynamically lowered.
[0043] As a preferred solution of the method for optimizing the online retrieval capability of the large model described in the present invention, if the average length of the words measured is shorter than the initially set limit, it is necessary to:
[0044] Reduce the sliding window size to adapt it to vocabulary capture operations at a finer scale;
[0045] If the proportion of special punctuation marks and heterogeneous characters exceeds the limit of 20% of the total, the sliding window size L' needs to be adjusted to adapt to the current discourse environment: "Modified window width L' = standard width * (1 + ξ)", where ξ reflects the influence coefficient of character structure diversity on the scanning unit length;
[0046] The exploration range of word length is fine-tuned based on the η coefficient value, "adjusting the word detection depth D = default search depth limit + (η * D base growth amount)" to better fit the changes in the performance of actual contextual demand characteristics;
[0047] If the complexity of the text increases, the contribution ratio of the constant coefficient η to the total score will be appropriately increased to the second component, that is, η ratio = η original basis + ω, to ensure that the mechanism has stronger adaptive adjustment performance.
[0048] As a preferred solution of the method for optimizing the large model online retrieval capability of the present invention, in the process of re-evaluating the importance level of novel morphemes, the following rules should be considered to update their weight attribute scores:
[0049] After identifying and analyzing the unknown style of the sentence category, refer to the document categories of similar styles previously stored in the knowledge base, and use this as a basis to estimate the preliminary influencing factor coefficient μ;
[0050] If the proportion of proper nouns in the recognized text breaks through the key threshold of one-tenth of the total number of words, corresponding adjustment operations need to be performed on the initial μ factor: "The new factor ν = μ initial factor * 1.1 + α revised variable" to reflect the specific influence of the characteristics of the new type of text on the final score and the trend of the result information;
[0051] If the main body of the text content is a personal name, a specific constant is selected as an additional item to supplement "the specified adjustment factor α = α base amount + Δα revised increment = 0.17" to ensure that the highlighting effect of exclusive information such as personal names in the calculation is fully reflected and the display form is consistent;
[0052] For technical reports or materials with professional terms, "when this type of text style is found, set the adjustment factor Δα to a relatively high positive value" to strengthen the measurement of the importance of technical terms and reflect the improvement of the reliability and accuracy of the evaluation work and the improvement of safeguard measures.
[0053] As a preferred solution of the optimization method for the large model network retrieval ability described in the present invention, in the process of determining the importance of vocabulary:
[0054] By comparing the weight score value of the word with the preset important index threshold, it is determined that when the word importance score I > the predefined important level I0, it is regarded as the "highest-priority key element item A" and is specially marked;
[0055] Assign the top-level flag T to the most concerned word element to ensure that such keywords enjoy the highest degree of attention treatment conditions during the search process;
[0056] At the same time, general vocabulary that does not meet the high-value evaluation criteria still needs to be assigned the appropriate grade code B in accordance with the existing consistent procedures to classify, store, organize, and facilitate retrieval, query, and access for use.
[0057] The beneficial effects of the present invention are as follows: Through a series of meticulous Chinese word segmentation processes, parameter settings, and optimizations, it aims to comprehensively improve the efficiency and accuracy of large-scale networked retrieval systems, thereby solving the technical challenges caused by multiple factors. First, when dealing with Chinese word segmentation, an advanced language model is adopted and combined with professional domain knowledge to ensure the accuracy and consistency of vocabulary segmentation, improving the language understanding and query parsing capabilities of the retrieval system; Second, in the initialization stage, each processing step is carefully defined, from model-dependent configurations to the customization of word segmentation strategies, especially increasing the importance of domain words, making the system more suitable for professional text retrieval; Then, for different scenarios, such as changes in text content or documents in specific fields, the key parameters are dynamically adjusted, including the adaptive adjustment of empirical coefficients and thresholds, to optimize the retrieval speed; In the dynamic adjustment configuration, it is also specifically mentioned that different types of correction factors should be increased or decreased based on actual needs, enabling the model to make more flexible decisions according to different conditions; And in order to more accurately evaluate the importance of each vocabulary and reasonably allocate resources, a detailed weight calculation and importance marking mechanism is proposed.
[0058] This method effectively solves several major problems:
[0059] 1. Reduce the slow retrieval response caused by network latency, and improve the retrieval speed and system response sensitivity through optimized algorithms.
[0060] 2. Strengthen the protection of users' sensitive information, integrate multi-level protection mechanisms in word segmentation, retrieval and other links, and improve the quality of information security management.
[0061] 3. Improve the system's concurrent request processing ability and load balancing management strategy, effectively reducing the burden on the server.
[0062] 4. Improve the core retrieval algorithm, making it more intelligent and personalized, ensuring the relevance of the results, and providing results that better meet the needs of users.
[0063] 5. Finally, an intelligent caching mechanism is introduced to relieve the pressure caused by frequent database access, ensuring that the service response is more rapid and efficient. All the above measures comprehensively enhance the overall function and stability of the large model networked retrieval system. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 It is a flowchart of an optimization method for the large model networked retrieval ability. Detailed Implementation Modes
[0066] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific implementation modes of the present invention in conjunction with the accompanying drawings of the specification.
[0067] In the following description, numerous specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0068] Secondly, as used herein, an "embodiment" or "embodiments" refer to specific features, structures, or characteristics that can be included in at least one implementation of the present invention. The phrase "in an embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separate from or mutually exclusive of other embodiments.
[0069] Embodiment 1
[0070] Referring to Figure 1 , a method for optimizing the large model network retrieval ability of the present invention is described, aiming to improve the efficiency and accuracy of the network retrieval system while ensuring data security, including the following steps:
[0071] Receive a user query statement, call the search engine to retrieve results and list them;
[0072] Perform a preliminary screening on the retrieved result list, obtain the HTML source code through the web link, and analyze the HTML source code to extract the text of the web page body;
[0073] Preprocess the extracted text content, extract a keyword list, and perform keyword matching in the preprocessed text content to count the number of keyword occurrences;
[0074] Calculate the relevance score of each retrieved result, select several retrieved results with the highest scores, perform text summarization on the selected results and integrate them to generate a summary of the query statement;
[0075] Output the summary as the final retrieved result and present it to the user;
[0076] The process of performing Chinese word segmentation based on a pre-trained language model includes:
[0077] Initialize the word segmentation model environment and define the strategy for vocabulary segmentation, which includes the following specific steps:
[0078] Download and configure the required dependency libraries and versions for the model environment;
[0079] Set the main parameters of the word segmentation algorithm, such as window size and maximum word length limit;
[0080] Define the segmentation strategy, including the use of the dictionary, the word segmentation method, and the handling of polysemous words;
[0081] If the text processing involves domain vocabulary, the weight of professional words should be increased when setting the main parameters of the word segmentation algorithm. If the current word belongs to the professional vocabulary of a specific domain, its weight W = W_original(1 + α), where α is the professional weighting correction value, ranging from 0 to 1, and the specific value depends on the matching degree between the text and the domain;
[0082] Initializing the word segmentation model environment and defining the steps of the vocabulary segmentation strategy include:
[0083] In the specific operation of defining the segmentation strategy, it is further refined as:
[0084] Set the main parameters of the word segmentation algorithm, including the sliding window size, the longest word search threshold, and the memory configuration;
[0085] Define the specific dictionary matching logic and conflict resolution method, and select the priority according to the word frequency sorting when multiple segmentation methods are applicable;
[0086] Dynamically set parameters to handle different types of texts. If the currently processed text is a subjective comment, reduce the empirical coefficient value β;
[0087] In the process of dynamically adjusting the configuration parameters, if the text to be processed has a clear professional background, use the following adjustment mechanism to enhance the model performance. If the currently processed text is classified as mainly covering technical terms, the correction factor θ value of technical terms should be appropriately increased. The formula for the adjusted factor is: θ_adjusted = θβ, and the value of the correction factor should conform to the actual situation and experimental data support;
[0088] The specific details of setting the main parameters of the word segmentation algorithm include:
[0089] Including the memory management of allocating task processor resources and setting the dynamic adjustment of the sliding window, etc.:
[0090] The rule for setting the longest word search threshold is as follows;
[0091] In the initial state, estimate the most likely length based on the statistical information of the text obtained in the previous steps;
[0092] When adjusting the threshold, consider adding the empirical value η according to the text characteristics for calculation;
[0093] If the detected text content is relatively complex or contains a large number of special symbols, adjust the experience factor to make it more flexible and adaptive. For documents with higher complexity, increase the adjustment ratio η of the search threshold. The new calculation formula becomes Threshold = Original length γδ + η Correction, where γ represents the length multiplier and δ is the text difficulty adjustment factor;
[0094] The detailed steps for adding correction factors as needed during the dynamic configuration adjustment process are as follows:
[0095] Measure the average correction factor for each word type in the current text;
[0096] When encountering a change in text category, calculate the estimated approximate value of the correction factor based on existing approximate categories and use this value to calculate the exact value required for subsequent operations;
[0097] Use the formula Original value adjustment coefficient (ξ) to determine the new correction factor F;
[0098] Implement special processing for data where the correction factor exceeds the set range: When the calculated new word correction factor is greater than the threshold λ, the correction value of this entry deserves extra attention and an extra weight is assigned to it;
[0099] The process of judging the importance of the final words includes the following specific operation procedures to accurately determine their actual contribution:
[0100] Detect whether it exceeds the pre - established importance benchmark. If it is determined that the importance of a specific word exceeds the established important level, that is, whether C > C base standard C magnification after multiplying the importance degree C of the word by the upper - limit adjustment multiple, then mark it as this highest level P_HIGH;
[0101] Assign the maximum priority status flag to each word that exceeds the importance threshold for quick access;
[0102] If the word does not meet the requirements of the highest level, mark it according to the usual method. When necessary, the magnification factor can be reduced until it equals the basic multiple 1 to prevent over - emphasizing any entry;
[0103] Special measures need to be taken for some words that have been assigned the top - level priority flag. Place marks at the corresponding index entries of these entries for efficient tracking;
[0104] Referring to the attached drawings, an optimization method for the large - model network retrieval ability of the present invention is described. This method improves the system efficiency and accuracy by improving the way of language model processing and optimizing network data interaction, and solves the problems of network delay, data security risks, high - concurrency processing challenges, poor retrieval effects, and frequent database access. The following details the core steps in this method and the technical means adopted in solving the above challenges:
[0105] The first step is to improve the network latency problem. First, advanced protocol stack optimization and network accelerators are introduced at the transmission layer between the server and the user. This step shortens the time overhead of connection establishment and closure by selecting a better network transmission protocol, such as using HTTP / 2 or HTTP / 3 instead of the traditional HTTP, and leveraging functions such as header compression; and enables content compression to reduce the bandwidth burden, thereby effectively reducing the problem of slow retrieval speed caused by network latency.
[0106] Specifically, in an embodiment of implementing this optimization method, assuming that the client sends a query request containing a large-scale data stream, the above mechanism can establish a server connection with the target resource for effective data interaction in a very short time, greatly enhancing the service real-time performance and the user experience perception. In addition, the CDN (Content Delivery Network) and edge caching technologies are also combined to store resources close to the user's terminal node, alleviating the pressure on the data center while ensuring the access rate.
[0107] Secondly, the present invention proposes a data protection strategy to avoid the risk of theft of important privacy data caused by illegal intrusion. For this purpose, a strict identity authentication process is established to distinguish between legitimate logged-in personnel and ordinary visitor accounts; data encryption measures are adopted to ensure the security of the transmission channel information; a fine-grained authorization management scheme is implemented, allowing only specific people within a specific group to access sensitive files, and even if there is an accidental loss, the scope of harm can be limited. In addition, in order to ensure the integrity and confidentiality of sensitive information during storage and transmission and prevent it from being tampered with, a transparent file encryption function is introduced at the application layer, and irreversible functions such as HMAC and SHA series are used to construct hash key verification integrity. When the system detects an unexpected behavior, it will automatically lock the account and trigger an internal audit, further enhancing the overall system's anti-leakage effect.
[0108] The third step focuses on dealing with the risk of increased service pressure during peak periods caused by the limited computing capacity of a single server. For this purpose, the design concept of a distributed microservices platform architecture is introduced - by horizontally expanding to improve the performance throughput of the entire cluster, rather than blindly trying to increase the level of single hardware configuration (because the latter is prone to hitting the hardware wall). The microservices architecture allows the business to be broken down into independent small services, and each component can be independently developed, deployed, and maintained. When a problem occurs in a certain part, it will not cause a chain reaction to affect the work of other modules. The load balancing request scheduling is achieved through the API gateway, thereby effectively dispersing the instantaneous request traffic pressure brought by hot spots.
[0109] In practice, for example, this solution was adopted when a large number of online orders exploded during the Double 11 Shopping Festival. The platform can quickly increase or decrease the number of nodes to adapt to the changes, greatly alleviating the intensive I / O read and write pressure of the backend database storage while maintaining service quality. Moreover, the container orchestration framework K8s based on the microservice model can also implement an automatic scaling mechanism. The system will automatically adjust the scale of the working machine according to the actual business traffic situation to achieve the dual goals of saving costs and responding to emergencies;
[0110] The next step is to upgrade and iterate the search engine algorithm to improve matching accuracy and timeliness of result presentation. The BM25 text retrieval technology widely used in early versions was transformed and integrated with machine understanding technology. Specialized data in downstream fields was added on the basis of pre-training to fine-tune. The recall rate was optimized by introducing a hybrid strategy of TFIDF and BERT while retaining sensitivity to capturing important phrases in documents. When processing long text information, the divide-and-conquer idea was used to load and pre-process in batches to overcome the shortcoming that the original version was only applicable to short sentence searches. It showed superior performance in extracting core topics and significantly reduced the probability of noise fragments. On the basis of achieving accurate retrieval, it continued to explore deep connections to provide auxiliary information, creating new possibilities for the development of knowledge-driven businesses.
[0111] For example, in an application scenario, if an enterprise wants to efficiently find clauses related to customer query intent from a huge amount of historical document warehouses, the algorithm improvement can help quickly filter out irrelevant noise and lock high-probability target locations, greatly improving the daily office efficiency of legal workers. In addition, combined with modules such as knowledge graph and natural question-answer generation, it can realize the automated workflow mode of fuzzy intent analysis, similar fragment search, key sentence summary extraction, answer construction, and reply construction, presenting users with more comprehensive and intuitive answers instead of simple independent records;
[0112] The last measure is to explore and practice the establishment of a complete cache system framework to reduce the direct read pressure of the central storage unit to improve the read and write performance. In this area, we first adopt a hierarchical storage architecture for commonly used access records (for example, temporarily storing hot data in SSD instead of HHD) to achieve faster read and write response time and reduce the impact of head seek delay, and combine Redis database technology to temporarily cache the result sets of the most recent high-frequency requests in memory for easy reuse. When the client initiates the same query, it can quickly give feedback without having to pass the request back to the background for redundant verification calculations. Secondly, in order to overcome the problem of simple memory space size limitations, the concept of secondary cache technology is introduced: for some objects that are too large to be completely placed in the high-speed buffer, SSHD hybrid hard disks are used as transitional temporary storage pools to store them. When the user issues the next round of retrieval, relevant information can be provided efficiently and accurately.
[0113] An optimization method for the large model's online retrieval ability of the present invention includes: through a series of meticulous Chinese word segmentation processes, parameter settings and optimizations, aiming to comprehensively improve the efficiency and accuracy of the large-scale online retrieval system, and further solve the technical challenges caused by multiple factors. First, when dealing with Chinese word segmentation, an advanced language model is adopted and combined with professional domain knowledge to ensure the accuracy and consistency of vocabulary segmentation, improving the language understanding and query parsing ability of the retrieval system; second, in the initialization stage, each processing step is carefully defined, from model dependency configuration to tokenization strategy customization, especially increasing the importance of domain words, making the system more suitable for professional text retrieval; then, for different scenarios, such as text content changes or documents in specific fields, by dynamically adjusting key parameters, including the adaptive adjustment of empirical coefficients and thresholds, to optimize the retrieval speed; in the dynamic adjustment configuration, it is also particularly mentioned to increase or decrease different types of correction factors based on actual needs, enabling the model to make more flexible decisions according to different conditions; and in order to more accurately evaluate the importance of each vocabulary and reasonably allocate resources, a detailed weight calculation and importance marking mechanism is proposed.
[0114] This method effectively solves several main problems:
[0115] 1. Reduce the slow retrieval response caused by network latency, and improve the retrieval speed and system response sensitivity through optimizing the algorithm.
[0116] 2. Strengthen the protection of user sensitive information, integrate multi-level protection mechanisms in the word segmentation, retrieval and other links, and improve the quality of information security management.
[0117] 3. Improve the system's concurrent request processing ability and load balancing management strategy, effectively reducing the burden on the server.
[0118] 4. Improve the core retrieval algorithm to make it more intelligent and personalized, ensure the relevance of the results, and provide results more in line with user needs.
[0119] 5. Finally, introduce an intelligent caching mechanism to relieve the pressure caused by frequent database access, ensuring that the service response is more rapid and efficient. All the above measures comprehensively enhance the overall function and stability of the large model online retrieval system.
[0120] Such a design not only greatly improves the effect of the search engine and the user experience, but also strengthens the security and efficient operation of the entire architecture.
[0121] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An optimization method for the network retrieval ability of large models, characterized in that It includes the following steps: Receive the user's query statement, call the search engine to retrieve results and list them; Conduct a preliminary screening on the retrieved result list, obtain the HTML source code through web links, and analyze the HTML source code to extract the text of the web page body; Preprocess the extracted text content, extract the keyword list, and perform keyword matching in the preprocessed text content to count the number of keyword occurrences; Calculate the relevance score of each retrieved result, select several retrieved results with the highest scores, perform text summarization on the selected results and integrate them to generate a summary of the query statement; Output the summary and present it to the user as the final retrieval result.
2. The optimization method for the large model's online retrieval ability according to claim 1, wherein The preprocessing includes word segmentation, part-of-speech tagging, and calculating the TFIDF value of each word; The process of word segmentation includes: Download and configure the required dependency libraries and versions for the model environment; Set the main parameters of the word segmentation algorithm; Define the segmentation strategy.
3. The optimization method for the large model's online retrieval ability according to claim 2, characterized in that The definition of the segmentation strategy includes: Set the main parameters of the word segmentation algorithm; Define the specific dictionary matching logic and conflict resolution method, and select the priority according to the word frequency sorting when multiple segmentation methods are applicable at the same time; Dynamically set parameters to handle different types of texts. If it is recognized that the current processed text is a subjective comment, reduce the empirical coefficient value.
4. The optimization method for the large model's online retrieval ability according to claim 3, wherein The specific details of setting the main parameters of the word segmentation algorithm include: Including the memory management of allocating task processor resources and setting steps such as dynamic adjustment of the sliding window: The rule for setting the longest word search threshold is as follows; Estimate the most likely length based on the statistical information of the text obtained in the previous steps in the initial state; When adjusting the threshold, consider the text characteristics and add the empirical value η for calculation; If it is detected that the text content is relatively complex or has a large number of special symbols, adjust the empirical factor, and increase the adjustment ratio of the search threshold for documents with higher complexity.
5. The optimization method for the large model's online retrieval ability as described in claim 4, characterized in that, The correction factors that need to be added during the dynamic adjustment configuration process include the following detailed steps: Measure the average correction factor of each lexical type in the current text; When encountering a change in text category, calculate the estimated approximate value of the correction factor based on the existing approximate categories and use this value to calculate the exact value required for subsequent operations; Use the formula original value adjustment coefficient to determine the new correction factor.
6. The optimization method for the large model's online retrieval ability according to claim 5, wherein The process of determining the importance of the final vocabulary includes the following specific operation processes to accurately determine its actual contribution: Detect whether it exceeds the pre-set importance benchmark. If it is determined that the importance of a specific vocabulary exceeds the established important level; Assign the maximum priority status flag to each word that exceeds the importance threshold for quick access; If the vocabulary does not meet the requirements of the highest level, mark it in the usual way; For some of the words that have been assigned the top priority flag, place a mark at the corresponding index entry.
7. The optimization method for the large model's online retrieval ability as described in claim 3, characterized in that, When the word segmentation process detects an overlap conflict in the dictionary entries: The algorithm will give priority to considering the vocabulary with a higher IDF score and the TF meeting a certain range requirement as the preferred word boundary cutting point; When the document frequency of the entry exceeds the pre-set standard threshold X value, use the specified entry to participate in the final word segmentation; If the system automatically analyzes and determines that the operated document belongs to the news report style, the accuracy factor for enhancing the boundary division of entries is increased, and the formula is defined as: "Segmentation accuracy = Original accuracy level + δX News type bonus", where δX is a constant value specific to news manuscripts.
8. The optimization method for the large model's online retrieval ability according to claim 3, characterized in that If the average length of the measured words is relatively short compared to the initially set limit, then it is necessary to: Reduce the size of the sliding window; If the proportion of special punctuation marks and heterogeneous characters exceeds the 20% threshold value of the total amount, it is necessary to adjust the sliding window size L' to adapt to the current discourse environment: "The modified window width L' = standard width (1 + ξ)", where ξ reflects the influence coefficient of the character structure diversity on the scanning unit length; In the case of an increase in text complexity, appropriately increase the proportion of the contribution of the constant coefficient η to the total score to two parts, i.e., η proportion = η original basis + ω.
9. The optimization method for the large model's online retrieval ability according to claim 3, wherein In the process of re-evaluating the importance level of novel morphemes, the following rules should be considered for updating their weight attribute scores: After identifying and analyzing the statement category of the unknown style, refer to the document categories of similar styles previously stored in the knowledge base, and based on this, estimate the preliminary influence factor coefficient μ; If the proportion of the number of proper nouns in the recognized text breaks through the key boundary point of one-tenth of the total number of words, then a corresponding adjustment operation needs to be performed on the preliminary factor μ: "New factor ν = μ preliminary factor 1.1 + α revised variable"; If the main body of the text content is a personal name, a specific constant is selected as an additional item to supplement "Specified adjustment factor α = α base amount + Δα correction increment = 0.17"; For technical reports or materials with professional terms, "When this type of document style is found, set the adjustment factor Δα to a relatively high positive value".
10. The optimization method for the large model's online retrieval ability as described in claim 9, characterized in that, In the process of determining the importance of vocabulary: By comparing the weight score value of a word with a preset important index threshold, it is determined that "when the important score I of a word > the predefined important level ", it is regarded as an item of "the most priority key element label item A" and is specially marked; Assign the top-level flag T to the most concerned word elements to ensure that such keywords receive the highest level of attention guarantee during the search process; At the same time, general vocabulary that does not meet the high-value evaluation criteria still needs to be classified and stored according to the existing consistent procedures and regulations by continuing to assign the appropriate level code B.