A Method for Constructing a Large-Scale Writing Model that Integrates Knowledge Injection and Retrieval Enhancement
By integrating knowledge injection and retrieval enhancement methods, the writing model is optimized in real time, which solves the problem of insufficient stability of the writing model, improves the model's adaptability in high-frequency and low-frequency domains, and enhances the logical continuity and accuracy of content generation.
Patent Information
- Application Number
- CN202511432707.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-09
AI Technical Summary
The existing writing model suffers from insufficient stability due to the excessive number of key components, resulting in repeated data transport and inadequate system availability.
By integrating knowledge injection and retrieval enhancement methods, high-frequency and low-frequency knowledge data are injected respectively. The retrieval enhancement mechanism is used to optimize query semantics and adjust the update frequency of regional knowledge cache, the forced retention ratio of low-frequency domains in query text, and the contextual dynamic expansion coefficient of low-frequency semantics in real time to improve model stability.
By dynamically adjusting the knowledge cache update frequency and the forced retention ratio, the retrieval latency and content logic breakage rate are reduced, the injection coverage of low-frequency knowledge and semantic capture capabilities are improved, and the stability of the integrated construction of the writing big model is enhanced.
Smart Images

Figure CN120911623B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model building technology, and in particular to a method for building a large writing model that integrates knowledge injection and retrieval enhancement. Background Technology
[0002] In the current era of rapid development in artificial intelligence technology and the emergence of large language models as a core component of digital infrastructure, intelligent writing technology has become a crucial technology urgently needed in scenarios such as medical reports, financial compliance documents, and legal documents. Its core value lies in its ability to deeply integrate domain knowledge with dynamic information, balancing content professionalism with fluent output, thus addressing issues such as low efficiency and lagging knowledge updates in professional scenarios. This drives the intelligent upgrade of professional content production and has significant practical and industrial value. However, existing large-scale writing models still have obvious limitations, with insufficient semantic adaptability between query text and writing requirements, making it difficult to meet the demands for "precise and professional" writing results.
[0003] Chinese Patent Publication No. CN120562570A discloses a retrieval enhancement system and method for generating a large language model. The method, based on a large language model, includes: receiving a query transmitted by a host-side network interface card (NIC) device via a deployed first orchestrator; wherein the query is received by the host from a user client targeting the large language model via a question-and-answer service module, and then sent by the host to the host-side NIC device; both the question-and-answer service module and the large language model are deployed on the host; scheduling the query to at least one of multiple computing nodes in a big data cluster via a deployed second orchestrator, so that the computing node converts the query into a query vector using a deployed embedded model; receiving the query vector sent by the computing node; accessing the knowledge database of the big data cluster based on the query vector using the second orchestrator, obtaining retrieval knowledge corresponding to the query vector from the knowledge database, and sending the retrieval knowledge to the host-side NIC device, so that the host-side NIC device, through the first orchestrator, sends the retrieval knowledge to the host, and the host generates an answer based on the retrieval knowledge using the large language model. Therefore, it can be seen that the aforementioned retrieval enhancement generation large language model system and method have problems such as insufficient stability in the fusion construction of the writing large model due to the excessive number of key components leading to repeated data transportation and insufficient system availability. Summary of the Invention
[0004] To address this, the present invention provides a method for constructing a large-scale writing model that integrates knowledge injection and retrieval enhancement, thereby overcoming the problem in the prior art where the large-scale writing model suffers from insufficient stability due to the excessive number of key components leading to repeated data transport and insufficient system availability.
[0005] To achieve the above objectives, this invention provides a method for constructing a large-scale writing model that integrates knowledge injection and retrieval enhancement, comprising:
[0006] The collected knowledge data and text information are respectively injected into the initial model for training to obtain a knowledge-enhanced writing model, wherein the knowledge data includes high-frequency knowledge and low-frequency knowledge;
[0007] The query semantics are retrieved through a retrieval enhancement mechanism to obtain the query text. The query semantics and the query text are respectively input into the knowledge-enhanced writing model to obtain the enhanced text. The knowledge-enhanced writing model is then optimized in real time based on the enhanced text to obtain the large writing model.
[0008] The query semantics include high-frequency semantics and low-frequency semantics;
[0009] Obtain the average response time of several searches, and determine whether the stability of the fusion construction of the writing big model meets the requirements based on the average response time of the several searches;
[0010] If the stability of the fusion construction of the writing big model does not meet the requirements, then determine whether it is necessary to increase the dynamic weight of the regional knowledge cache update frequency.
[0011] If there is no need to increase the dynamic weight of the regional knowledge cache update frequency, then obtain the injection efficiency of low-frequency knowledge to determine whether the processing capability of low-frequency semantics meets the requirements.
[0012] If the processing capability of the low-frequency semantics does not meet the requirements, then determine whether it is necessary to increase the mandatory retention ratio of low-frequency domains in the query text;
[0013] If it is not necessary to increase the mandatory retention ratio of low-frequency domains in the query text, then the context dynamic expansion coefficient of low-frequency semantics is determined based on the matching accuracy of low-frequency semantics.
[0014] Furthermore, based on the average response time of the aforementioned retrievals, the stability of the fusion construction of the writing big model is determined to meet the requirements, including:
[0015] The average response time of several searches is compared with the preset first response time;
[0016] If the average response time of the aforementioned retrievals is less than or equal to the preset first response time, then the stability of the fusion construction of the writing big model is determined to meet the requirements.
[0017] If the average response time of the aforementioned retrievals exceeds the preset first response time, then it is determined that the stability of the fusion construction of the writing big model does not meet the requirements.
[0018] Furthermore, determine whether the dynamic weight of the regional knowledge cache update frequency needs to be increased, including:
[0019] The average response time of the several searches is compared with the preset first response time and the preset second response time, respectively;
[0020] If the average response time of the several retrievals is greater than the preset first response time and less than or equal to the preset second response time, then it is determined that there is no need to increase the dynamic weight of the regional knowledge cache update frequency.
[0021] If the average response time of the aforementioned retrievals exceeds the preset second response time, then it is determined that the dynamic weight of the regional knowledge cache update frequency needs to be increased.
[0022] Furthermore, the increase in the dynamic weight of the regional knowledge cache update frequency is determined by the difference between the average response time of several retrievals and the preset second response time.
[0023] Furthermore, determining whether the low-frequency semantic processing capability meets the requirements based on the injection efficiency of the low-frequency knowledge includes:
[0024] The injection efficiency of the low-frequency knowledge is compared with a preset second efficiency;
[0025] If the injection efficiency of the low-frequency knowledge is greater than the preset second efficiency, then it is determined that the processing capability of the low-frequency semantics meets the requirements, and it is determined whether the dynamic weight of the regional knowledge cache update frequency meets the requirements.
[0026] If the injection efficiency of the low-frequency knowledge is less than or equal to the preset second efficiency, then it is determined that the processing capability of low-frequency semantics does not meet the requirements.
[0027] Further, determine whether it is necessary to increase the mandatory retention ratio of low-frequency domains in the query text, including:
[0028] The injection efficiency of the low-frequency knowledge is compared with the preset first efficiency and the preset second efficiency, respectively.
[0029] If the injection efficiency of the low-frequency knowledge is greater than the preset first efficiency and less than or equal to the preset second efficiency, then it is determined that the forced retention ratio of the low-frequency domain in the query text needs to be increased, and the forced retention ratio of the low-frequency domain in the query text needs to be increased.
[0030] If the injection efficiency of the low-frequency knowledge is less than or equal to the preset first efficiency, then it is determined that there is no need to increase the forced retention ratio of low-frequency domains in the query text.
[0031] Furthermore, the increase in the mandatory retention ratio of low-frequency domains in the query text is determined by the difference between the preset first efficiency and the injection efficiency of the low-frequency knowledge.
[0032] Furthermore, the contextual dynamic expansion coefficient of low-frequency semantics is determined based on the matching accuracy of low-frequency semantics, including:
[0033] The matching accuracy of the low-frequency semantics is compared with the preset accuracy.
[0034] If the matching accuracy of the low-frequency semantics is greater than or equal to the preset accuracy, then the retrieval validity of the low-frequency semantics is determined to meet the requirements, and it is not necessary to increase the context dynamic expansion coefficient of the low-frequency semantics. It is also determined whether the forced retention ratio of the low-frequency domain meets the requirements.
[0035] If the matching accuracy of the low-frequency semantics is less than the preset accuracy, it is determined that the retrieval effectiveness of the low-frequency semantics does not meet the requirements, and it is necessary to increase the context dynamic expansion coefficient of the low-frequency semantics.
[0036] Furthermore, the matching accuracy of the low-frequency semantics is the ratio of the number of accurate low-frequency semantic matches to the total number of matches.
[0037] Furthermore, the increase in the context dynamic expansion coefficient of the low-frequency semantics is determined by the difference between the preset accuracy and the matching accuracy of the low-frequency semantics.
[0038] Compared with existing technologies, the beneficial effects of this invention are as follows: The method of this invention dynamically adjusts the update frequency of regional knowledge cache based on the average response time of several retrievals. Since retrieval nodes deployed across regions experience increased response time for retrieval requests due to network bandwidth limitations, the logical breakage rate of generated content increases with increased retrieval latency in real-time writing. By increasing the dynamic weight of the update frequency of regional knowledge cache, cross-regional dependencies can be reduced by improving cache hit rate, thereby fundamentally reducing retrieval latency and content logical breakage rate. Furthermore, the method adjusts the forced retention ratio of low-frequency domains in the query text based on the injection efficiency of low-frequency knowledge. Because the training set is overly concentrated in high-frequency domains, the model's writing ability in low-frequency domains will significantly decrease, leading to... The model exhibits over-adaptation to high-frequency domains and under-adaptation to low-frequency domains. By increasing the mandatory retention ratio of low-frequency domains in the query text, it is possible to ensure that low-frequency samples are exposed in each round of training, thereby improving the injection coverage of low-frequency knowledge. The contextual dynamic expansion coefficient of low-frequency semantics is adjusted based on the matching accuracy of low-frequency semantics. Since the semantics of low-frequency concepts are highly dependent on the context, and the static semantic representation of the retrieval model is difficult to capture this dynamic association, it leads to incorrect association with irrelevant domains. By increasing the contextual dynamic expansion coefficient of low-frequency semantics, the window can be expanded to obtain more comprehensive semantic constraints, thereby capturing more complete contextual clues, reducing ambiguity, and thus reducing the probability of being incorrectly associated with traditional databases, improving the stability of the fusion construction of the writing big model.
[0039] Furthermore, the method of the present invention adjusts the dynamic weight of the regional knowledge cache update frequency by setting a preset first response time and a preset second response time. Since the retrieval request response time of cross-regional retrieval nodes is limited by network bandwidth, the logical breakage rate of generated content increases when the retrieval latency increases in real-time writing. By increasing the dynamic weight of the regional knowledge cache update frequency, the cross-regional dependency can be reduced by improving the cache hit rate, thereby reducing the retrieval latency and content logical breakage rate from the root, and further improving the stability of the fusion construction of the writing big model.
[0040] Furthermore, the method of the present invention adjusts the forced retention ratio of low-frequency domains in the query text by setting a preset first efficiency and a preset second efficiency. Since the training set is overly concentrated in the high-frequency domain, the model's writing ability in the low-frequency domain will decrease significantly, resulting in overfitting of the model to the high-frequency domain and underfitting of the low-frequency domain. By increasing the forced retention ratio of low-frequency domains in the query text, it can be ensured that low-frequency samples are exposed in each round of training, improve the injection coverage of low-frequency knowledge, and further improve the stability of the fusion construction of the writing big model.
[0041] Furthermore, the method described in this invention adjusts the contextual dynamic expansion coefficient of low-frequency semantics by setting a preset accuracy rate. Since the semantics of low-frequency concepts are highly dependent on the context, and the static semantic representation of the retrieval model is difficult to capture this dynamic association, it leads to incorrect association with irrelevant domains. By increasing the contextual dynamic expansion coefficient of low-frequency semantics, the window can be expanded to obtain more comprehensive semantic constraints, thereby capturing more complete contextual clues, reducing ambiguity, and thus reducing the probability of being incorrectly associated with traditional databases, further improving the stability of the fusion construction of the writing big model. Attached Figure Description
[0042] Figure 1 This is an overall flowchart of the writing big model construction method that integrates knowledge injection and retrieval enhancement according to an embodiment of the present invention;
[0043] Figure 2 The following is a flowchart illustrating the dynamic weighting process for determining whether to increase the update frequency of the regional knowledge cache in the writing big model construction method that integrates knowledge injection and retrieval enhancement in this embodiment of the invention.
[0044] Figure 3 The flowchart below illustrates the process of determining whether to increase the mandatory retention ratio of low-frequency domains in the query text in the writing big model construction method that integrates knowledge injection and retrieval enhancement, as described in this embodiment of the invention.
[0045] Figure 4 This is a flowchart illustrating the process of determining the contextual dynamic expansion coefficients of low-frequency semantics in the writing big model construction method that integrates knowledge injection and retrieval enhancement, as described in an embodiment of the present invention. Detailed Implementation
[0046] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0047] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0048] Please see Figure 1 , Figure 2 , Figure 3 as well as Figure 4The diagrams shown are, respectively, the overall flowchart of the writing model construction method integrating knowledge injection and retrieval enhancement according to an embodiment of the present invention, the logical flowchart of the dynamic weight process for determining whether to increase the update frequency of regional knowledge cache, the logical flowchart of the process for determining whether to increase the forced retention ratio of low-frequency domains in the query text, and the logical flowchart of the process for determining the context dynamic expansion coefficient of low-frequency semantics. The present invention provides a writing model construction method integrating knowledge injection and retrieval enhancement, comprising:
[0049] Step S1: The collected knowledge data and text information are injected into the initial model for training to obtain a knowledge-enhanced writing model. The knowledge data includes high-frequency knowledge and low-frequency knowledge.
[0050] Step S2: The query semantics are retrieved through a retrieval enhancement mechanism to obtain the query text. The query semantics and the query text are respectively input into the knowledge-enhanced writing model to obtain the enhanced text. The knowledge-enhanced writing model is then optimized in real time based on the enhanced text to obtain the large writing model. The query semantics include high-frequency semantics and low-frequency semantics.
[0051] Step S3: Obtain the average response time of several searches, and determine whether the stability of the fusion construction of the writing big model meets the requirements based on the average response time of the several searches.
[0052] Step S4: If the stability of the fusion construction of the writing big model does not meet the requirements, determine whether it is necessary to increase the dynamic weight of the regional knowledge cache update frequency.
[0053] Step S5: If it is not necessary to increase the dynamic weight of the regional knowledge cache update frequency, then obtain the injection efficiency of low-frequency knowledge to determine whether the processing capability of low-frequency semantics meets the requirements.
[0054] Step S6: If the processing capability of the low-frequency semantics does not meet the requirements, determine whether it is necessary to increase the proportion of mandatory retention of low-frequency domains in the query text.
[0055] Step S7: If it is not necessary to increase the proportion of low-frequency domains in the query text, then determine the context dynamic expansion coefficient of low-frequency semantics based on the matching accuracy of low-frequency semantics.
[0056] Specifically, knowledge data includes literature, textbooks, and encyclopedic content.
[0057] Specifically, the initial model is a general-purpose, large-scale pre-trained language model with basic language understanding and generation capabilities, built on the Transformer architecture.
[0058] Specifically, low-frequency knowledge refers to knowledge that occurs less than 0.1% of the total knowledge volume.
[0059] Specifically, high-frequency knowledge refers to knowledge that appears more than or equal to 20% of the total knowledge content.
[0060] Specifically, textual information includes news briefs, web articles, and journal articles.
[0061] Specifically, knowledge-enhanced writing models include Wenxin Yiyan AI, OmniThink model, and GPT-4.
[0062] Specifically, the query semantics include the structural composition of computers, commonly used tools for circuit repair, and commonly used frameworks for software development.
[0063] Specifically, the query text includes details of computer structure and parameters, basic multimeter knowledge, and the classification and characteristics of software development frameworks.
[0064] Specifically, high-frequency semantics are those semantics that appear more than 70% of the total number of query semantics.
[0065] Specifically, low-frequency semantics are semantics that appear less than 10% of the total number of query semantics.
[0066] Specifically, the retrieval enhancement mechanism involves converting the query semantics into a semantic vector, then performing a similarity search in a vectorized database of text information to recall the most relevant knowledge fragments, and finally combining these fragments with the original query semantics to create enhanced hints.
[0067] Specifically, enhanced text refers to text content generated by the knowledge-enhanced writing model based on enhanced prompts and model generation.
[0068] Specifically, the writing models include Tongyi 1000 Questions, Grok-3, and AIscolan AI.
[0069] Specifically, the process of real-time optimization of the knowledge-enhanced writing model using enhanced text involves selecting high-quality generated results through automated metrics, constructing them as triple training data with the corresponding query semantics and enhanced text, and then updating the parameters of the large writing model through supervised fine-tuning.
[0070] Specifically, the dynamic weight of the regional knowledge cache update frequency is the control coefficient for dynamically adjusting the priority of updating the knowledge data cache within the writing big model.
[0071] Specifically, the mandatory retention ratio of low-frequency domains in the query text is the ratio of the number of low-frequency domain texts that are forcibly retained by the retrieval enhancement mechanism to the total amount of query text.
[0072] Specifically, the contextual dynamic expansion coefficient of low-frequency semantics represents the degree to which the context of low-frequency semantics is expanded when the knowledge-enhanced model processes text.
[0073] Specifically, accurate matching of low-frequency semantics with enhanced text means that the enhanced text generated based on low-frequency semantics has semantic consistency with the standard interpretation corresponding to the semantics.
[0074] In implementation, the method of this invention adjusts the dynamic weight of the regional knowledge cache update frequency based on the average response time of several retrievals. Because cross-regional retrieval nodes experience increased retrieval request response time due to network bandwidth limitations, the logical breakage rate of generated content increases with increased retrieval latency in real-time writing. By increasing the dynamic weight of the regional knowledge cache update frequency, cross-regional dependencies can be reduced by improving cache hit rate, thereby fundamentally reducing retrieval latency and content logical breakage rate. The forced retention ratio of low-frequency domains in the query text is adjusted based on the injection efficiency of low-frequency knowledge. Since the training set is overly concentrated in high-frequency domains, the model's writing ability in low-frequency domains will significantly decrease, leading to a decline in the model's ability to handle high-frequency domains. Overfitting and underfitting low-frequency domains can be addressed by increasing the mandatory retention ratio of low-frequency domains in the query text. This ensures that low-frequency samples are exposed in each training round, improving the injection coverage of low-frequency knowledge. The contextual dynamic expansion coefficient of low-frequency semantics is adjusted based on the matching accuracy of low-frequency semantics. Since the semantics of low-frequency concepts are highly dependent on the context, and the static semantic representation of the retrieval model is difficult to capture this dynamic association, it leads to incorrect association with irrelevant domains. By increasing the contextual dynamic expansion coefficient of low-frequency semantics, the window can be expanded to obtain more comprehensive semantic constraints, thereby capturing more complete contextual clues, reducing ambiguity, and thus reducing the probability of being incorrectly associated with traditional databases, improving the stability of the fusion construction of the writing big model.
[0075] Specifically, the stability of the fusion construction of the writing model is determined based on the average response time of the aforementioned retrievals, including:
[0076] The average response time of several searches is compared with the preset first response time;
[0077] If the average response time of the aforementioned retrievals is less than or equal to the preset first response time, then the stability of the fusion construction of the writing big model is determined to meet the requirements.
[0078] If the average response time of the aforementioned retrievals exceeds the preset first response time, then it is determined that the stability of the fusion construction of the writing big model does not meet the requirements.
[0079] One possible reason for the inadequate stability of the fusion construction of the writing model is that the processing capability for low-frequency semantics is insufficient, or the dynamic weight of the regional knowledge cache update frequency is inadequate. The next step is to determine which specific reason it is, which is essentially the process of deciding whether to increase the dynamic weight of the regional knowledge cache update frequency.
[0080] Specifically, determining whether a dynamic weight needs to be increased for the frequency of regional knowledge cache updates includes:
[0081] The average response time of the several searches is compared with the preset first response time and the preset second response time, respectively;
[0082] If the average response time of the several retrievals is greater than the preset first response time and less than or equal to the preset second response time, then it is determined that there is no need to increase the dynamic weight of the regional knowledge cache update frequency.
[0083] If the average response time of the aforementioned retrievals exceeds the preset second response time, then it is determined that the dynamic weight of the regional knowledge cache update frequency needs to be increased.
[0084] Specifically, when the average response time of several searches exceeds the preset second response time, it is determined that the reason for the unsatisfactory stability of the writing model's fusion construction is that the dynamic weight of the regional knowledge cache update frequency does not meet the requirements. Therefore, it is necessary to increase the dynamic weight of the regional knowledge cache update frequency. When the average response time of several searches exceeds the preset first response time but is less than or equal to the preset second response time, it can be preliminarily determined that the processing capability of low-frequency semantics does not meet the requirements. Next, it is necessary to make a final determination based on the injection efficiency of low-frequency knowledge to determine whether the processing capability of low-frequency semantics meets the requirements, that is, to determine whether the reason for the unsatisfactory stability of the writing model's fusion construction is the unsatisfactory processing capability of low-frequency semantics.
[0085] It is understandable that the preset first response time is shorter than the preset second response time, and the three intervals divided by the preset first and second response times correspond to three different scenarios:
[0086] The first interval is when the average response time of several searches is less than or equal to the preset first response time. The corresponding situation is that the stability of the fusion construction of the writing model meets the requirements, and no adjustment is needed.
[0087] The second interval is the average response time of several retrievals, which is greater than the preset first response time and less than or equal to the preset second response time. The corresponding situation is: because the training set is overly concentrated in the high-frequency domain, the model's writing ability in the low-frequency domain will be significantly reduced, resulting in the model's overfitting to the high-frequency domain and underfitting to the low-frequency domain. At this time, it is necessary to further judge whether the processing ability of low-frequency semantics meets the requirements.
[0088] The third interval is when the average response time of several searches is greater than the preset second response time. The corresponding situation is: due to network bandwidth limitations, the response time of search requests increases for cross-regional search nodes. In real-time writing, when the search latency increases, the logical breakage rate of the generated content increases. At this time, it is necessary to adjust the dynamic weight of the regional knowledge cache update frequency.
[0089] Understandably, in the process of retrieving relevant knowledge from external knowledge bases, using preset first and second response durations to characterize the response efficiency and stability of the retrieval process is essentially a hierarchical control logic based on retrieval timeliness and knowledge completeness. This avoids the one-sidedness of evaluating system performance using a single response threshold, while also achieving a dynamic linkage between qualitative judgment of response efficiency and resource scheduling strategies, ultimately adapting to the dual requirements of response speed and system reliability in writing scenarios. The core role of the first response duration is to provide a quantitative basis for determining the stability of the integrated construction of the writing model; the second response duration is the direct criterion for determining whether the system response is effective. The preset first and second response durations can be set according to actual working conditions. The setting of the preset first and second response durations aims to ensure the stability and practicality of the integrated construction of the writing model. Optionally, the preset first and second response durations are determined through a limited number of experiments by evaluating the effect of different response durations on the integrated construction of the writing model. The determined preset first and second response durations should satisfy the condition that they are neither too small nor cause excessive interference to the integrated construction process of the writing model. For example, the preset first response time is generally selected in the range of [180ms, 220ms], and the preset second response time is generally selected in the range of [280ms, 320ms].
[0090] Preferably, the first response duration is 200ms in a preferred embodiment, and the second response duration is 300ms in a preferred embodiment.
[0091] Specifically, the average response time of a number of searches is the ratio of the total response time of a number of semantic search model searches to the total number of searches.
[0092] Specifically, the increase in the dynamic weight of the regional knowledge cache update frequency is determined by the difference between the average response time of several retrievals and the preset second response time.
[0093] Specifically, when the difference between the average response time of several searches and the preset second response time is within 20ms, the dynamic weight of the regional knowledge cache update frequency increases to 1.2 times the original value. When the difference between the average response time of several searches and the preset second response time exceeds 20ms, in addition to increasing to 1.2 times the original value, the dynamic weight of the regional knowledge cache update frequency increases by 0.1 for every 5ms exceeding the original value. For example, when the difference between the average response time of several searches and the preset second response time is 30ms, the current dynamic weight of the regional knowledge cache update frequency is 2.0, and the increased dynamic weight is 2.0×1.2+0.1×2=2.6.
[0094] In practice, the method of this invention adjusts the dynamic weight of the regional knowledge cache update frequency by setting a preset first response time and a preset second response time. Due to network bandwidth limitations, the response time of retrieval requests increases for cross-regional retrieval nodes. In real-time writing, when the retrieval latency increases, the logical breakage rate of the generated content rises. By increasing the dynamic weight of the regional knowledge cache update frequency, the cross-regional dependency can be reduced by improving the cache hit rate, thereby reducing the retrieval latency and content logical breakage rate from the root, and further improving the stability of the fusion construction of the writing big model.
[0095] Specifically, determining whether the processing capability of low-frequency semantics meets the requirements based on the injection efficiency of the low-frequency knowledge includes:
[0096] The injection efficiency of the low-frequency knowledge is compared with a preset second efficiency;
[0097] If the injection efficiency of the low-frequency knowledge is greater than the preset second efficiency, then it is determined that the processing capability of the low-frequency semantics meets the requirements, and it is determined whether the dynamic weight of the regional knowledge cache update frequency meets the requirements.
[0098] If the injection efficiency of the low-frequency knowledge is less than or equal to the preset second efficiency, then it is determined that the processing capability of low-frequency semantics does not meet the requirements.
[0099] Among them, when the injection efficiency of low-frequency knowledge is greater than the preset second efficiency, it is determined that the processing capability of low-frequency semantics meets the requirements. However, the stability of the fusion construction of the writing big model has been determined to be unsatisfactory. Therefore, it is necessary to further determine whether the dynamic weight of the regional knowledge cache update frequency meets the requirements.
[0100] In implementation, the dynamic weight of the actual regional knowledge cache update frequency is compared with the predetermined dynamic weight threshold to determine whether the dynamic weight of the regional knowledge cache update frequency meets the requirements. If the dynamic weight of the actual regional knowledge cache update frequency is less than the predetermined dynamic weight threshold, the dynamic weight of the regional knowledge cache update frequency is determined to be unacceptable. The predetermined dynamic weight threshold is the average value of the dynamic weight of the regional knowledge cache update frequency monitored in the previous three months of the historical period.
[0101] If the dynamic weight of the regional knowledge cache update frequency does not meet the requirements, then increase the dynamic weight of the regional knowledge cache update frequency; if the dynamic weight of the regional knowledge cache update frequency meets the requirements, then re-collect the average response time of several retrievals and re-determine whether the stability of the fusion construction of the writing big model meets the requirements.
[0102] When the injection efficiency of low-frequency knowledge is less than or equal to the preset second efficiency, it can be determined that the reason for the failure of the fusion construction stability of the writing model to meet the requirements is that the processing capability of low-frequency semantics is not up to standard. The reasons for the failure to meet the requirements of low-frequency semantic processing capability may be that the mandatory retention ratio of low-frequency domains in the query text is not up to standard, or that the retrieval effectiveness of low-frequency semantics is not up to standard. The next step is to determine which specific reason it is, which is also the process of whether to increase the mandatory retention ratio of low-frequency domains in the query text.
[0103] Specifically, determining whether it is necessary to increase the mandatory retention ratio of low-frequency domains in the query text includes:
[0104] The injection efficiency of the low-frequency knowledge is compared with the preset first efficiency and the preset second efficiency, respectively.
[0105] If the injection efficiency of the low-frequency knowledge is greater than the preset first efficiency and less than or equal to the preset second efficiency, then it is determined that the forced retention ratio of the low-frequency domain in the query text needs to be increased, and the forced retention ratio of the low-frequency domain in the query text needs to be increased.
[0106] If the injection efficiency of the low-frequency knowledge is less than or equal to the preset first efficiency, then it is determined that there is no need to increase the forced retention ratio of low-frequency domains in the query text.
[0107] Specifically, when the injection efficiency of low-frequency knowledge is greater than the preset first efficiency but less than or equal to the preset second efficiency, it is determined that the reason why the processing capability of low-frequency semantics does not meet the requirements is that the mandatory retention ratio of low-frequency domains in the query text does not meet the requirements. Therefore, it is necessary to increase the mandatory retention ratio of low-frequency domains in the query text. When the injection efficiency of low-frequency knowledge is less than or equal to the preset first efficiency, it can be preliminarily determined that the retrieval validity of low-frequency semantics does not meet the requirements. Next, it is necessary to make a final determination on whether the retrieval validity of low-frequency semantics meets the requirements based on the matching accuracy of low-frequency semantics, that is, to determine whether the reason why the processing capability of low-frequency semantics does not meet the requirements is that the retrieval validity of low-frequency semantics does not meet the requirements.
[0108] It is understandable that the preset first efficiency is less than the preset second efficiency, and the three intervals divided by the preset first efficiency and the preset second efficiency correspond to three different situations:
[0109] The first interval is when the injection efficiency of low-frequency knowledge is less than or equal to the preset first efficiency. The corresponding situation is: because the semantics of low-frequency concepts are highly dependent on the context, and the static semantic representation of the retrieval model is difficult to capture this dynamic association, it leads to incorrect association with irrelevant domains. At this time, it is necessary to further judge whether the retrieval effectiveness of low-frequency semantics meets the requirements.
[0110] The second interval is where the injection efficiency of low-frequency knowledge is greater than the first preset efficiency and less than or equal to the second preset efficiency. The corresponding situation is: because the training set is overly concentrated in the high-frequency domain, the model's writing ability in the low-frequency domain will decrease significantly, resulting in the model's overfitting to the high-frequency domain and underfitting to the low-frequency domain. At this time, it is necessary to adjust the forced retention ratio of the low-frequency domain in the query text.
[0111] The third interval is where the injection efficiency of low-frequency knowledge is greater than the preset second efficiency. The corresponding situation is that the processing capability of low-frequency semantics meets the requirements. At this time, it is necessary to further determine whether the dynamic weight of the regional knowledge cache update frequency meets the requirements.
[0112] Understandably, introducing preset first efficiency and preset second efficiency to characterize the quality of knowledge coverage and the economy of knowledge injection during the retrieval process from external knowledge bases is essentially a hierarchical control logic based on the value density of low-frequency knowledge and system resource consumption. This avoids the one-sidedness of evaluating low-frequency knowledge processing with a single efficiency indicator, while also achieving a dynamic linkage between qualitative judgment of the necessity of low-frequency knowledge and injection strategies. Ultimately, it adapts to the dual needs of the writing model for professional scenario knowledge coverage and lightweight system operation. The core role of preset first efficiency is to provide a quantitative basis for the value necessity of low-frequency knowledge, ensuring that the injected low-frequency knowledge truly supports the professionalism of the content. Preset second efficiency is a direct criterion for judging the rationality of resource consumption in the low-frequency knowledge injection process, avoiding system redundancy caused by excessive injection of low-frequency knowledge. Preset first efficiency and preset second efficiency can be set according to actual working conditions. The setting of preset first efficiency and preset second efficiency aims to ensure the stability and practicality of the integrated construction of the writing model. Optionally, the preset first efficiency and preset second efficiency are determined through a limited number of experiments by evaluating the effect of different injection efficiencies on the fusion construction of the writing big model. The determined preset first efficiency and preset second efficiency should satisfy the condition that they are neither too small nor cause excessive interference to the fusion construction process of the writing big model. For example, the preset first efficiency is generally selected in the range of [90ms, 110ms], and the preset second efficiency is generally selected in the range of [140ms, 160ms].
[0113] Preferably, the preferred embodiment of the preset first efficiency is 100ms, and the preferred embodiment of the preset second efficiency is 150ms.
[0114] Specifically, the low-frequency knowledge injection efficiency is the speed at which the initial model effectively integrates and applies low-frequency knowledge into the writing process.
[0115] Specifically, the increase in the mandatory retention rate of low-frequency domains in the query text is determined by the difference between the injection efficiency of low-frequency knowledge and the preset first efficiency.
[0116] Specifically, when the difference between the injection efficiency of low-frequency knowledge and the preset first efficiency is within 10ms, the mandatory retention ratio of low-frequency domains in the query text increases to 1.1 times the original value. When the difference between the injection efficiency of low-frequency knowledge and the preset first efficiency exceeds 10ms, in addition to increasing to 1.1 times the original value, the mandatory retention ratio of low-frequency domains increases by 1% for every 5ms exceeding the original value. For example, when the difference between the injection efficiency of low-frequency knowledge and the preset first efficiency is 20ms, the mandatory retention ratio of low-frequency domains in the current query text is 8%. After the increase, the mandatory retention ratio of low-frequency domains in the query text is 8×1.1+1×2=10.8%.
[0117] In practice, the method of this invention adjusts the forced retention ratio of low-frequency domains in the query text by setting a preset first efficiency and a preset second efficiency. Since the training set is overly concentrated in the high-frequency domain, the model's writing ability in the low-frequency domain will decrease significantly, resulting in overfitting of the model to the high-frequency domain and underfitting of the low-frequency domain. By increasing the forced retention ratio of low-frequency domains in the query text, it can be ensured that low-frequency samples are exposed in each round of training, improve the injection coverage of low-frequency knowledge, and further improve the stability of the fusion construction of the writing big model.
[0118] Specifically, the contextual dynamic expansion coefficient of low-frequency semantics is determined based on the matching accuracy of low-frequency semantics, including:
[0119] The matching accuracy of the low-frequency semantics is compared with the preset accuracy.
[0120] If the matching accuracy of the low-frequency semantics is greater than or equal to the preset accuracy, then the retrieval validity of the low-frequency semantics is determined to meet the requirements, and it is not necessary to increase the context dynamic expansion coefficient of the low-frequency semantics. It is also determined whether the forced retention ratio of the low-frequency domain meets the requirements.
[0121] If the matching accuracy of the low-frequency semantics is less than the preset accuracy, it is determined that the retrieval effectiveness of the low-frequency semantics does not meet the requirements, and it is necessary to increase the context dynamic expansion coefficient of the low-frequency semantics.
[0122] Specifically, when the matching accuracy of low-frequency semantics is greater than or equal to the preset accuracy, it is determined that the retrieval validity of low-frequency semantics meets the requirements. However, if it has been previously determined that the processing capability of low-frequency semantics does not meet the requirements, then it is necessary to further determine whether the mandatory retention ratio of low-frequency domains in the query text meets the requirements.
[0123] In implementation, the mandatory retention ratio of low-frequency domains in the actual query text is compared with the predetermined mandatory retention ratio threshold to determine whether the mandatory retention ratio of low-frequency domains in the query text meets the requirements. If the mandatory retention ratio of low-frequency domains in the actual query text is less than the predetermined mandatory retention ratio threshold, it is determined that the mandatory retention ratio of low-frequency domains in the query text does not meet the requirements. The predetermined mandatory retention ratio threshold is the average value of the mandatory retention ratio of low-frequency domains in the query text monitored in the previous three months of the historical period.
[0124] If the mandatory retention ratio of low-frequency domains in the actual query text does not meet the requirements, then the mandatory retention ratio of low-frequency domains in the actual query text will be increased; if the mandatory retention ratio of low-frequency domains in the actual query text meets the requirements, then the injection efficiency of low-frequency knowledge will be re-collected, and the processing capability of low-frequency semantics will be re-evaluated to see if it meets the requirements.
[0125] When the matching accuracy of low-frequency semantics is less than the preset accuracy, it can be determined that the reason why the processing capability of low-frequency semantics does not meet the requirements is that the retrieval effectiveness of low-frequency semantics does not meet the requirements. Therefore, it is necessary to increase the context dynamic expansion coefficient of low-frequency semantics.
[0126] It is understandable that the two intervals for the preset accuracy rate division correspond to two different scenarios:
[0127] The first interval is where the matching accuracy of low-frequency semantics is greater than or equal to the preset accuracy. The corresponding situation is: the retrieval validity of low-frequency semantics meets the requirements. At this time, it is necessary to further determine whether the forced retention ratio of the low-frequency domain meets the requirements.
[0128] The second interval is when the matching accuracy of low-frequency semantics is less than the preset accuracy. The corresponding situation is that the semantics of low-frequency concepts are highly dependent on the context, and the static semantic representation of the retrieval model is difficult to capture this dynamic association, resulting in incorrect association with irrelevant domains. In this case, it is necessary to adjust the context dynamic expansion coefficient of low-frequency semantics.
[0129] Understandably, using a preset accuracy rate to characterize the completeness and precision of knowledge retrieval is essentially a hierarchical evaluation logic based on the characteristics of knowledge distribution and application scenarios. This design avoids the limitations of a single indicator in evaluating the system's retrieval capabilities, while achieving coordinated control over the breadth of knowledge coverage and the depth of semantic understanding, ultimately adapting to the needs of comprehensive knowledge discovery and relevance of content generation in writing scenarios. The core function of the preset accuracy rate is to provide a quantitative basis for judging the effectiveness of low-frequency semantic retrieval. The preset accuracy rate can be set according to actual working conditions. The setting of the preset accuracy rate aims to ensure the stability and practicality of the fusion construction of the writing big model. Optionally, the preset accuracy rate is determined through a limited number of experiments by evaluating the matching accuracy of different low-frequency semantics on the fusion construction effect of the writing big model. The determined preset accuracy rate should be neither too low nor too high, and should not cause excessive interference to the fusion construction process of the writing big model. For example, the preset accuracy rate is generally selected in the range of [85%, 95%].
[0130] Preferably, the preset accuracy rate is 90% in this preferred embodiment.
[0131] Specifically, the increase in the context dynamic expansion coefficient of low-frequency semantics is determined by the difference between the preset accuracy and the matching accuracy of low-frequency semantics.
[0132] Specifically, when the difference between the preset accuracy and the matching accuracy of low-frequency semantics is within 5%, the context dynamic expansion coefficient of low-frequency semantics increases to 1.1 times the original value. When the difference between the preset accuracy and the matching accuracy of low-frequency semantics exceeds 5%, the context dynamic expansion coefficient of low-frequency semantics increases by 0.05 for every 1% increase beyond the original 1.1 times. For example, when the difference between the preset accuracy and the matching accuracy of low-frequency semantics is 7%, the current context dynamic expansion coefficient of low-frequency semantics is 0.4. After the increase, the context dynamic expansion coefficient of low-frequency semantics is 0.5×1.1+0.05×2=0.65.
[0133] In practice, the method described in this invention adjusts the contextual dynamic expansion coefficient of low-frequency semantics by setting a preset accuracy rate. Since the semantics of low-frequency concepts are highly dependent on the context, and the static semantic representation of the retrieval model is difficult to capture this dynamic association, it leads to incorrect association with irrelevant domains. By increasing the contextual dynamic expansion coefficient of low-frequency semantics, the window can be expanded to obtain more comprehensive semantic constraints, thereby capturing more complete contextual clues, reducing ambiguity, and thus reducing the probability of being incorrectly associated with traditional databases, further improving the stability of the fusion construction of the writing big model.
[0134] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for building a writing large model with enhanced knowledge injection and retrieval, characterized in that, The method comprises the following steps: injecting the collected knowledge data and text information into an initial model for training to obtain a knowledge-enhanced writing model, wherein the knowledge data comprises high-frequency knowledge and low-frequency knowledge; retrieving query semantics through a retrieval enhancement mechanism to obtain query text, inputting the query semantics and the query text into the knowledge-enhanced writing model to obtain enhanced text, and optimizing the knowledge-enhanced writing model in real time according to the enhanced text to obtain a writing large model; wherein the query semantics comprises high-frequency semantics and low-frequency semantics; obtaining the average response time of a plurality of retrievals, and determining whether the fusion construction stability of the writing large model meets the requirements based on the average response time of the plurality of retrievals; if the fusion construction stability of the writing large model does not meet the requirements, determining whether the dynamic weight of the regional knowledge cache update frequency needs to be increased; if the dynamic weight of the regional knowledge cache update frequency does not need to be increased, obtaining the injection efficiency of the low-frequency knowledge to determine whether the processing capacity of the low-frequency semantics meets the requirements; if the processing capacity of the low-frequency semantics does not meet the requirements, determining whether the forced retention proportion of the low-frequency field in the query text needs to be increased; if the forced retention proportion of the low-frequency field in the query text does not need to be increased, determining the context dynamic expansion coefficient of the low-frequency semantics based on the matching accuracy of the low-frequency semantics.
2. The method of claim 1, wherein the method further comprises: determining whether the fusion construction stability of the writing large model meets the requirements based on the average response time of the plurality of retrievals, comprising: comparing the average response time of a plurality of retrievals with a preset first response time; if the average response time of the plurality of retrievals is less than or equal to the preset first response time, it is determined that the fusion construction stability of the writing large model meets the requirements; if the average response time of the plurality of retrievals is greater than the preset first response time, it is determined that the fusion construction stability of the writing large model does not meet the requirements.
3. The method of claim 2, wherein the method further comprises: determining whether the dynamic weight of the regional knowledge cache update frequency needs to be increased, comprising: comparing the average response time of the plurality of retrievals with the preset first response time and the preset second response time respectively; if the average response time of the plurality of retrievals is greater than the preset first response time and less than or equal to the preset second response time, it is determined that the dynamic weight of the regional knowledge cache update frequency does not need to be increased; if the average response time of the plurality of retrievals is greater than the preset second response time, it is determined that the dynamic weight of the regional knowledge cache update frequency needs to be increased.
4. The method of claim 3, wherein the method further comprises: The increase range of the dynamic weight of the regional knowledge cache update frequency is determined by the difference between the average response time of the plurality of retrievals and the preset second response time.
5. The method of claim 4, wherein the method further comprises: determining whether the processing capacity of the low-frequency semantics meets the requirements based on the injection efficiency of the low-frequency knowledge, comprising: comparing the injection efficiency of the low-frequency knowledge with a preset second efficiency; if the injection efficiency of the low-frequency knowledge is greater than the preset second efficiency, it is determined that the processing capacity of the low-frequency semantics meets the requirements, and it is determined whether the dynamic weight of the regional knowledge cache update frequency meets the requirements; If the injection efficiency of the low-frequency knowledge is less than or equal to the preset second efficiency, it is determined that the processing capability of the low-frequency semantic does not meet the requirement.
6. The method of claim 5, wherein the method further comprises: It is determined whether the forced retention proportion of the low-frequency field in the query text needs to be increased, including: The injection efficiency of the low-frequency knowledge is compared with the preset first efficiency and the preset second efficiency respectively. If the injection efficiency of the low-frequency knowledge is greater than the preset first efficiency and less than or equal to the preset second efficiency, it is determined that the forced retention proportion of the low-frequency field in the query text needs to be increased, and the forced retention proportion of the low-frequency field in the query text is increased. If the injection efficiency of the low-frequency knowledge is less than or equal to the preset first efficiency, it is determined that the forced retention proportion of the low-frequency field in the query text does not need to be increased.
7. The method of claim 6, wherein the method further comprises: The increase range of the forced retention proportion of the low-frequency field in the query text is determined by the difference between the injection efficiency of the low-frequency knowledge and the preset first efficiency.
8. The method of claim 7, wherein the method further comprises: The context dynamic expansion coefficient of the low-frequency semantic is determined based on the matching accuracy of the low-frequency semantic, including: The matching accuracy of the low-frequency semantic is compared with the preset accuracy. If the matching accuracy of the low-frequency semantic is greater than or equal to the preset accuracy, it is determined that the retrieval effectiveness of the low-frequency semantic meets the requirement, the context dynamic expansion coefficient of the low-frequency semantic does not need to be increased, and it is determined whether the forced retention proportion of the low-frequency field meets the requirement. If the matching accuracy of the low-frequency semantic is less than the preset accuracy, it is determined that the retrieval effectiveness of the low-frequency semantic does not meet the requirement, the context dynamic expansion coefficient of the low-frequency semantic needs to be increased, and the context dynamic expansion coefficient of the low-frequency semantic is increased.
9. The method of claim 8, wherein the method further comprises: The matching accuracy of the low-frequency semantic is the ratio of the number of accurate matches between the low-frequency semantic and the enhanced text to the total number of matches.
10. The method of claim 9, wherein the method further comprises: The increase range of the context dynamic expansion coefficient of the low-frequency semantic is determined by the difference between the preset accuracy and the matching accuracy of the low-frequency semantic.
Citation Information
Patent Citations
Retrieval enhancement generation large language model system and method
CN120562570A
Retrieval enhanced semantic instruction response method based on knowledge fusion
CN120179801A