A Large Model Deployment Method and System Based on a Multi-Level Caching Mechanism
Through the multi-level caching mechanism, the intelligent customer service system captures semantic features and builds dynamic correlation maps in real time, recognizes cross-session context dependencies, and generates personalized responses, solving the problems of slow response and incoherent response of existing systems, and improving the processing capabilities and user experience of intelligent customer service.
Patent Information
- Application Number
- CN202510245377.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing intelligent customer service system lacks an effective caching mechanism, resulting in long response times, incoherent responses and lack of personalization, making it difficult to deal with complex and cross-session context dependencies.
The multi-level caching mechanism is adopted, including the first-level dynamic semantic cache, the second-level context-related cache and the third-level intent decision-making cache. Through semantic identifiers, implicit semantic trajectory vectors and knowledge topology networks, user requests are captured and processed in real time to generate coherent and personalized responses.
It significantly improves the response speed and accuracy of the intelligent customer service system, can handle complex dialogue situations, provide coherent and personalized answers, and improve user experience.
Smart Images

Figure CN119739809B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of large model deployment, and in particular, to a large model deployment method and system based on a multi-level caching mechanism. Background Art
[0002] In the intelligent AI customer service Q&A scenario, users expect to obtain instant, accurate, and personalized answers. This requires the system to not only quickly respond to users' inquiries but also understand complex semantic information and context associations to provide coherent answers. In addition, as the complexity of the conversation increases, the system needs to handle cross-session context dependencies and dynamically adjust the response strategy according to different interaction patterns. Therefore, an intelligent AI customer service system needs to have efficient processing capabilities, accurate semantic understanding, and a flexible multi-modal response generation mechanism.
[0003] Currently, many intelligent customer service systems adopt methods based on large models to improve the quality and accuracy of Q&A. These methods usually include a single-level caching mechanism for storing common question-and-answer pairs and some simple context processing logics. To improve the response speed, some systems also directly use pre-trained models for matching or retrieval-based answers. However, this method mainly focuses on the construction and maintenance of a static knowledge base and has limited capabilities in capturing and processing dynamic semantic features.
[0004] Although existing intelligent customer service solutions can meet basic Q&A needs to a certain extent, they still have several significant deficiencies: due to the lack of an effective caching mechanism, especially for dynamic semantic features, existing systems often cannot quickly process repeated or similar query requests, resulting in long response times; most systems are difficult to effectively identify and utilize cross-session context information, resulting in incoherent or inaccurate answers in continuous conversation scenarios, affecting the user experience; traditional methods usually generate answers based on fixed templates or simple keyword matching, lacking the ability to deeply understand the true intentions of users and are difficult to provide highly personalized and rich response content. Summary of the Invention
[0005] The embodiments of the present application provide a large model deployment method and system based on a multi-level caching mechanism to solve the problem of the lack of an effective caching mechanism in the prior art.
[0006] In a first aspect, the embodiments of the present application provide a large model deployment method based on a multi-level caching mechanism, including the following steps:
[0007] Build the first-level dynamic semantic cache, capture the semantic feature encoding of the online session in real time to generate a unique identifier, statistically analyze the high-frequency semantic fragments based on the dynamic access weight. When the overlap degree between the semantic identifier of the new request and the cached fragment meets the preset condition, aggregate the discrete semantic fragments to reconstruct the complete response logic and give priority to feedback;
[0008] For requests that miss the cache, start the second-level context-related cache, extract the implicit semantic trajectory vector of the session, construct a dynamic association graph by combining the historical interaction path, identify the potential context dependencies across sessions. When continuous semantic nodes that form a logical link with the current request are detected in the association graph, trigger the progressive answer splicing mechanism;
[0009] When the second-level cache is not covered, activate the third-level intent decision cache. Classify the scenario and mark the status of the request through a multi-level parsing structure, generate a structured query instruction based on the knowledge topology network, dynamically adjust the weight of the knowledge retrieval path, and extract multi-modal response elements to generate a combined response;
[0010] Synchronously execute the collaborative optimization of multi-level caches, dynamically adjust the semantic coverage range and matching priority of the cache according to the response results of the large model and user feedback, and realize the linkage between the adaptive expansion of the cache capacity and the elimination logic based on the session traffic characteristics;
[0011] Among them, the semantic identifier of the first-level cache is orthogonally complementary to the semantic trajectory vector space of the second-level cache, the association graph nodes of the second-level cache and the knowledge topology network of the third-level cache are mapped through the semantic bridging layer, and the cache update conditions at all levels are dynamically negatively correlated with the session flow complexity.
[0012] Optionally, the step of starting the second-level context-related cache for requests that miss the cache, extracting the implicit semantic trajectory vector of the session, constructing a dynamic association graph by combining the historical interaction path, identifying the potential context dependencies across sessions, and triggering the progressive answer splicing mechanism when continuous semantic nodes that form a logical link with the current request are detected in the association graph includes:
[0013] Extract the context encoding features of the continuous interaction sequence in the current session to generate an implicit semantic trajectory vector with time series dependence;
[0014] Align the implicit semantic trajectory vector with the interaction paths of the same user identifier or similar user groups in the historical session library in space-time to construct a dynamic association graph containing semantic nodes and relationship edges, where the graph nodes store the semantic context fragments across sessions, and the relationship edges record the logical jump probability and timeliness weight between nodes;
[0015] Perform two-way semantic propagation calculation on the dynamic association graph, activate adjacent nodes on the association path based on the semantic trajectory vector of the current request, and identify multi-hop semantic chains that form logical coherence with the current session through a path confidence screening algorithm;
[0016] When it is detected that there are at least two or more associated nodes with temporal continuity and consistency in the multi-hop semantic chain, trigger the reconstruction engine for cross-session answer fragments, and dynamically sort and connect discrete answer elements according to the logical order and weight ratio between nodes to generate a progressive splicing response that is coherent with the current context. Figure 1 When it is detected that there are at least two or more associated nodes with temporal continuity and consistency in the multi-hop semantic chain, trigger the reconstruction engine for cross-session answer fragments, and dynamically sort and connect discrete answer elements according to the logical order and weight ratio between nodes to generate a progressive splicing response that is coherent with the current context.
[0017] Optionally, when the second-level cache is not covered, activate the third-level intention decision cache, classify the scenario and mark the status of the request through a multi-level parsing structure, generate a structured query instruction based on the knowledge topology network, dynamically adjust the weight of the knowledge retrieval path, and extract multi-modal response elements to generate a combined response, including:
[0018] Capture the intention expression pattern and scenario feature parameters in the current interaction sequence through the session state awareness module, and generate a multi-dimensional classification vector containing domain labels and state variables;
[0019] Map the multi-dimensional classification vector to the pre-constructed knowledge topology network, and generate a structured query instruction with path constraints according to the domain association strength of each sub-graph in the network. The query instruction includes the retrieval condition priority and the cross-domain fusion rule;
[0020] Real-time monitor the focus shift and feedback signals in the user interaction behavior, dynamically calculate the credibility attenuation factor of each node in the knowledge retrieval path, and adjust the scope and depth of cross-domain retrieval based on the path weight re-distribution algorithm;
[0021] Parallelly extract text, image, and operation instruction multi-modal data units from heterogeneous data sources according to the updated retrieval path, filter out conflicting elements through the semantic consistency verification module, and embed the multi-modal units into the dynamically generated dialogue framework according to the preset response assembly template to form a combined response.
[0022] Optionally, synchronously execute multi-level cache collaborative optimization, dynamically adjust the cache semantic coverage range and matching priority according to the large model response result and user feedback, and realize the logical linkage of cache capacity adaptive expansion and elimination based on the session traffic characteristics, including:
[0023] Construct a semantic traceability link, reversely decompose the final response generated by the large model into a contribution degree distribution map of multi-level cache hit results, and calculate the semantic alignment error of each level of cache according to the dependency relationship between the map nodes;
[0024] Collect user satisfaction scores and follow-up behavior data when the collection session ends, establish a cache effect evaluation matrix, and obtain the semantic coverage blind spots and matching rule deviations that need to be corrected for each level of cache through matrix eigenvalue decomposition;
[0025] Analyze the spatio-temporal distribution characteristics of session traffic in real time, construct a cache load prediction model based on a sliding window, and dynamically calculate the capacity pressure coefficient and hot data migration trend of each level of cache;
[0026] According to the coupling relationship between the semantic alignment error and the capacity pressure coefficient, synchronously adjust the semantic matching threshold and storage partition strategy of each level of cache, and trigger the priority rearrangement and redundant data elimination of cross-layer cache based on the hot migration trend.
[0027] Optionally, monitor the focus transfer and feedback signals in the user interaction behavior in real time, dynamically calculate the credibility decay factor of each node in the knowledge retrieval path, and adjust the scope and depth of cross-domain retrieval based on the path weight reallocation algorithm, including:
[0028] Construct an attention transfer matrix based on the user interaction behavior sequence, extract the current interaction focus and historical focus sequence through a sliding window mechanism, calculate the focus transfer probability and generate a focus transfer vector;
[0029] Define the credibility decay factor in the knowledge topology network, calculate the node weight correction value in combination with the user feedback signal, and update the node weight;
[0030] Based on the updated node weights, use the path weight reallocation algorithm to calculate the priority and depth limit of the cross-domain retrieval path;
[0031] Generate a multi-modal retrieval instruction according to the path weight coefficient and depth limit, and drive the parallel extraction module to extract data units according to the weight ratio and input them into the semantic consistency verification module.
[0032] Optionally, the synchronously adjusting the semantic matching threshold and storage partition strategy of each level of cache according to the coupling relationship between the semantic alignment error and the capacity pressure coefficient, and triggering the priority rearrangement and redundant data elimination of cross-layer cache includes:
[0033] Construct a cross-layer association model of the semantic alignment error matrix and the capacity pressure coefficient matrix, extract the collaborative optimization parameters of multi-level cache through a feature fusion algorithm, and generate a dynamic weight allocation map;
[0034] Based on the dynamic weight allocation map, use a path optimization algorithm to calculate the adjustment amount of the semantic matching threshold of each level of cache, and generate a storage partition strategy update instruction;
[0035] Construct a priority requeueing queue according to the hot spot migration trend, and generate a cross-layer cache priority mapping table by combining the spatio-temporal distribution characteristics of session traffic;
[0036] Input the semantic matching threshold adjustment amount, the storage partition policy update instruction, and the cross-layer cache priority mapping table into the cache policy execution engine, synchronously trigger threshold update, partition reconstruction, and redundant data elimination operations, and feedback the adjustment result to the semantic traceability link for calibration.
[0037] Optionally, the spatio-temporal distribution characteristics of the real-time analysis session traffic are used to construct a cache load prediction model based on a sliding window, and the capacity pressure coefficients and hot data migration trends of each level of cache are dynamically calculated, including:
[0038] Divide the session traffic time series based on the sliding window mechanism, extract the interaction frequency, semantic complexity, and response delay parameters within the window, and generate a spatio-temporal feature vector;
[0039] Input the spatio-temporal feature vector into a hierarchical prediction model, extract short-term traffic fluctuation features through a temporal convolutional network, and capture the cross-window traffic evolution law in combination with a long short-term memory network, and output the capacity pressure coefficients of each level of cache;
[0040] Construct a hot data recognition engine based on the capacity pressure coefficient, screen high-frequency accessed semantic fragments through a dynamic threshold algorithm, and generate a hot spot migration path map in combination with the semantic topological relationship;
[0041] Input the capacity pressure coefficient and the hot spot migration path map into the cache expansion decision module, trigger the dynamic reorganization of the storage partition and the redundant data elimination mechanism, and synchronously update the load status mark of the semantic traceability link.
[0042] In a second aspect, an embodiment of the present application provides a large model deployment system based on a multi-level cache mechanism, including:
[0043] A capture module for establishing a first-level dynamic semantic cache, real-time capturing the semantic feature encoding of an online session to generate a unique identifier, statistically counting high-frequency semantic fragments based on dynamic access weights, and when the overlap degree between the semantic identifier of a new request and the cache fragment meets a preset condition, aggregating discrete semantic fragments to reconstruct a complete response logic and giving priority to feedback;
[0044] An identification module for starting a second-level context-related cache for an unhit request, extracting the implicit semantic trajectory vector of the session, constructing a dynamic association map in combination with the historical interaction path, identifying potential context dependencies across sessions, and triggering a progressive answer splicing mechanism when detecting continuous semantic nodes in the association map that form a logical link with the current request;
[0045] An extraction module, configured to activate a third-level intent decision cache when the second-level cache is not covered, classify the scenarios and mark the status of requests through a multi-level parsing structure, generate structured query instructions based on a knowledge topology network, dynamically adjust the weights of knowledge retrieval paths, and extract multi-modal response elements to generate a combined response;
[0046] An adjustment module, configured to synchronously execute multi-level cache collaborative optimization, dynamically adjust the cache semantic coverage range and matching priority according to the large model response result and user feedback, and implement cache capacity adaptive expansion and elimination logic linkage based on session traffic characteristics.
[0047] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a large model deployment method based on a multi-level cache mechanism as described in the first aspect.
[0048] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it implements a large model deployment method based on a multi-level cache mechanism as described in the first aspect.
[0049] In the embodiment of the present application, a first-level dynamic semantic cache is established, the semantic feature encoding of the online session is captured in real time to generate a unique identifier, and high-frequency semantic fragments are statistically analyzed based on the dynamic access weight. When the overlap degree between the semantic identifier of the new request and the cache fragment meets the preset condition, the discrete semantic fragments are aggregated to reconstruct the complete response logic and feedback is given preferentially; for requests that miss, a second-level context-related cache is started, the implicit semantic trajectory vector of the session is extracted, a dynamic association graph is constructed in combination with the historical interaction path, and potential context dependencies across sessions are identified. When it is detected that there are continuous semantic nodes in the association graph that form a logical link with the current request, a progressive answer splicing mechanism is triggered; when the second-level cache is not covered, a third-level intent decision cache is activated, the scenarios of the request are classified and the status is marked through a multi-level parsing structure, structured query instructions are generated based on the knowledge topology network, the weights of the knowledge retrieval paths are dynamically adjusted, and multi-modal response elements are extracted to generate a combined response;
[0050] The technical solution of the present application has the following beneficial effects:
[0051] Synchronously execute multi-level cache collaborative optimization, dynamically adjust the cache semantic coverage range and matching priority according to the large model response results and user feedback, and realize the linkage between cache capacity adaptive expansion and elimination logic based on session traffic characteristics; among them, the semantic identifier of the first-level cache is orthogonally complementary to the semantic trajectory vector space of the second-level cache, the associated graph nodes of the second-level cache and the knowledge topology network of the third-level cache are mapped through the semantic bridging layer, and the cache update conditions at all levels are dynamically negatively correlated with the session flow complexity.
[0052] By establishing a first-level dynamic semantic cache, it can capture the semantic features of online sessions in real time and generate unique identifiers. When the overlap degree between the semantic identifier of a new request and the cache fragment meets the conditions, it can quickly aggregate discrete semantic fragments to reconstruct the complete response logic and achieve fast feedback, thus significantly improving the response speed of the question-and-answer system; for requests that miss the first-level cache, start the second-level context-associated cache. By extracting the implicit semantic trajectory vector of the session and constructing a dynamic associated graph, identify the potential context dependencies across sessions and trigger the progressive answer splicing mechanism. This helps the system better understand and process complex and continuous dialogue scenarios and provide more coherent and accurate answers; the third-level intention decision cache classifies the scenarios and marks the states of requests through a multi-level parsing structure, and generates structured query instructions based on the knowledge topology network, which can more accurately capture the actual needs of users, extract multi-modal response elements to generate combined answers, and further improve the intention parsing ability of the question-and-answer system and the richness of answers; synchronously execute multi-level cache collaborative optimization, dynamically adjust the cache strategy according to the large model response results and user feedback, which can not only effectively manage the cache capacity, avoid unnecessary storage consumption, but also realize the linkage between adaptive expansion and elimination logic based on session traffic characteristics, ensure the stable operation of the system under high concurrency, and improve the user experience at the same time; the mutual complementation and dynamic adjustment mechanism between caches at all levels enables the system to flexibly handle various complex user interaction modes, especially for those tasks that require multiple dialogue rounds to complete, and can provide a more smooth and natural communication experience.
[0053] Furthermore, this solution includes extracting the implicit semantic trajectory vector of the current session and aligning it spatiotemporally with the historical interaction path to construct a dynamic associated graph. Identify the multi-hop semantic chain that forms logical coherence with the current request through two-way semantic propagation calculation. When it is detected that there are at least two temporally consecutive and Figure 1When reaching relevant associated nodes, the reconstruction engine for cross - session answer fragments is triggered to sort and connect discrete answer elements according to logical order and weight ratio, generating a progressive spliced response. By adopting the method described in the core solution, the ability of the intelligent AI customer service system to handle complex dialogue scenarios can be significantly enhanced. Through the effective utilization of current session and historical interaction data, the system can more accurately identify and understand cross - session context dependencies, providing more coherent and natural answers. This method not only improves the relevance and accuracy of answers but also enhances the user experience, enabling the system to maintain logical consistency in continuous conversations, effectively solving the challenges faced by traditional methods in dealing with complex and continuous conversations, and achieving a more intelligent and personalized customer service experience.
[0054] Furthermore, when the second - level cache does not cover a request, the process of activating the third - level intent decision cache is triggered. This process involves classifying the user request into scenarios and marking its status through a multi - level parsing structure, and generating structured query instructions based on the constructed knowledge topology network. At the same time, the weights of knowledge retrieval paths are dynamically adjusted to optimize search efficiency, and multi - modal response elements such as text and images are extracted from heterogeneous data sources. After semantic consistency verification, these elements are integrated into a combined response according to a preset template. This significantly improves the ability of the intelligent AI customer service system to handle complex and diverse user requests. By accurately classifying the user request into scenarios and marking its status, and using the knowledge topology network to optimize the retrieval path, this method can not only quickly locate the most relevant answers but also effectively integrate various types of response elements to provide rich and personalized answers. This method greatly enhances the accuracy of intent parsing and the diversity of response content of the system, enabling the intelligent customer service to not only understand the direct needs of users but also capture and respond to their potential intentions, thus providing a more comprehensive and satisfactory customer service experience.
[0055] These aspects or other aspects of this application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of this application or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 The flowchart of a large - model deployment method based on a multi - level cache mechanism provided by this application is shown;
[0058] Figure 2Shows a schematic structural diagram of a large model deployment system based on a multi - level caching mechanism provided by the present application;
[0059] Figure 3 Shows a schematic structural diagram of a computing device provided by the present application. Detailed implementation manners
[0060] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application.
[0061] In some processes described in the specification and claims of the present application and the above - mentioned drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0062] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.
[0063] Figure 1 The flowchart of a large model deployment method based on a multi - level caching mechanism provided for the embodiments of the present application is as Figure 1 shown. The method includes:
[0064] 101. Establish a first - level dynamic semantic cache, capture the semantic feature encoding of the online session in real - time to generate a unique identifier, statistically analyze the high - frequency semantic fragments based on the dynamic access weight. When the overlap degree between the semantic identifier of the new request and the cached fragments meets the preset condition, aggregate the discrete semantic fragments to reconstruct the complete response logic and give priority to feedback;
[0065] In this step, the first - level dynamic semantic cache includes the semantic feature encoding of the online session captured in real - time and the unique identifier. These data are used to quickly match the new request with the stored high - frequency semantic fragments to achieve an efficient response to the user's query.
[0066] In the embodiments of the present application, the system first establishes a dynamic semantic cache, generates semantic feature encodings by analyzing the user input in the intelligent AI customer service Q&A scenario in real time, and creates unique identifiers for each session segment. The high-frequency semantic segments statistically obtained based on the access frequency are preferentially stored. When a new user request arrives, the system calculates the overlap degree between its semantic identifier and the segments in the cache. If the preset conditions are met, relevant semantic segments are aggregated to reconstruct the complete response logic and quickly feedback to the user.
[0067] Suppose a user asks about the battery life of a certain mobile phone on the intelligent customer service platform. The system captures the semantic features of this query in real time and finds that it highly overlaps with a high-frequency semantic segment in the cache. Then, relevant information is directly extracted from the cache, combined into a complete answer and immediately replied to the user.
[0068] 102. For requests that miss the first-level cache, activate the second-level context correlation cache, extract the implicit semantic trajectory vector of the session, construct a dynamic correlation graph in combination with the historical interaction path, identify potential context dependencies across sessions, and trigger the progressive answer splicing mechanism when continuous semantic nodes that form a logical link with the current request are detected in the correlation graph;
[0069] In this step, the second-level context correlation cache includes the implicit semantic trajectory vector of the session and the dynamic correlation graph constructed by the historical interaction path. These elements help identify potential context dependencies across sessions, thereby supporting more complex dialogue management.
[0070] In the embodiments of the present application, in the intelligent AI customer service Q&A scenario, for requests that cannot find answers in the first-level cache, the system activates the second-level cache mechanism, extracts the implicit semantic trajectory vector of the current session, and constructs a dynamic correlation graph in combination with the historical interaction path. By identifying the continuous semantic nodes related to the current request in the graph, the system triggers the progressive answer splicing mechanism to provide a more coherent answer.
[0071] Suppose the user mentions the mobile phone model problem consulted before in a session. The system analyzes the correlation between the current session and the historical interaction path, identifies the relevant context, and provides more personalized service recommendations accordingly, such as recommending mobile phone models suitable for the user and their features.
[0072] 103. When the second-level cache is not covered, activate the third-level intention decision cache, classify the scenario and mark the status of the request through a multi-level parsing structure, generate a structured query instruction based on the knowledge topology network, dynamically adjust the weight of the knowledge retrieval path, and extract multi-modal response elements to generate a combined response;
[0073] In this step, the third-level intent decision cache involves multi-level parsing structures, knowledge topology networks, and structured query instructions. These tools work together to parse the scene classification and state tags of user requests in order to accurately extract multimodal response elements and generate combined responses.
[0074] In the embodiment of the present application, in the intelligent AI customer service question and answer scenario, when the second-level cache cannot cover the request, the system activates the third-level cache, uses a multi-level parsing structure to analyze the request in detail, and uses the knowledge topology network to generate structured query instructions. Adjust the retrieval path weight according to the actual needs of the user, extract multimodal data units from different data sources, and integrate them into the final answer after consistency verification.
[0075] Suppose a user asks a multi-faceted question on the intelligent customer service platform, such as asking about the battery life, camera quality and price of a certain mobile phone. The system deeply analyzes the user's intention, uses the knowledge topology network to locate relevant information, and integrates data in various forms such as text and pictures to provide a comprehensive and accurate answer, including battery life test reports, camera samples and the latest market price information.
[0076] 104. Synchronously perform multi-level cache collaborative optimization, dynamically adjust cache semantic coverage and matching priority according to large model response results and user feedback, and realize cache capacity adaptive expansion and elimination logic linkage based on session traffic characteristics;
[0077] Among them, the semantic identifiers of the first-level cache are orthogonally complementary to the semantic trajectory vector space of the second-level cache, the association graph nodes of the second-level cache and the knowledge topology network of the third-level cache are mapped through the semantic bridging layer, and the update conditions of the caches at all levels are dynamically negatively correlated with the complexity of the session flow.
[0078] In this step, the multi-level cache collaborative optimization mechanism covers functions such as cache semantic coverage, matching priority adjustment, adaptive expansion, and elimination logic linkage. These functions ensure the flexibility and efficiency of the system, enabling it to self-adjust according to actual usage.
[0079] In the embodiment of the present application, in the intelligent AI customer service Q&A scenario, the system synchronously performs multi-level cache collaborative optimization, continuously adjusts the cache strategy based on the large model response results and user feedback, and monitors the session traffic characteristics to achieve automatic capacity adjustment. In addition, the semantic bridge layer connects the caches at all levels to ensure effective collaboration between them.
[0080] Suppose that with the arrival of the new product release season, the number of consultations on the intelligent customer service platform surges, especially for information inquiries about the new mobile phone. The system automatically expands the cache capacity and optimizes the matching rules to ensure that high-priority issues (such as new product features, release time, etc.) are promptly responded to, while maintaining the efficient operation of the overall system, reducing waiting time and improving user satisfaction.
[0081] In summary, steps 101 to 104 cover a full range of intelligent customer service solutions from rapid response to complex context processing to personalized service provision, aiming to provide a flexible and efficient Q&A platform to meet the diverse consultation needs of users in the intelligent AI customer service Q&A scenario.
[0082] To further improve the processing ability for unhit requests, in some embodiments, in step 102, a second-level context association cache is initiated for unhit requests, extracting session implicit semantic trajectory vectors, constructing a dynamic association graph by combining historical interaction paths, identifying potential context dependencies across sessions, and when detecting continuous semantic nodes in the association graph that form a logical link with the current request, triggering a progressive answer splicing mechanism, including:
[0083] Extracting the context encoding features of the continuous interaction sequence in the current session to generate an implicit semantic trajectory vector with temporal dependence; aligning the implicit semantic trajectory vector with the interaction paths of the same user identifier or similar user groups in the historical session library in space and time to construct a dynamic association graph containing semantic nodes and relationship edges, where the graph nodes store cross-session semantic context fragments, and the relationship edges record the logical jump probability and timeliness weight between nodes; performing two-way semantic propagation calculation on the dynamic association graph, activating adjacent nodes on the association path based on the semantic trajectory vector of the current request, and identifying multi-hop semantic chains that form logical coherence with the current session through a path confidence screening algorithm; when detecting that there are at least two or more associated nodes with time continuity and consistency in the multi-hop semantic chain, triggering a cross-session answer fragment reconstruction engine, dynamically sorting and connecting discrete answer elements according to the logical order and weight ratio between nodes to generate a progressive splicing response that is coherent with the current context. Figure 1 When detecting that there are at least two or more associated nodes with time continuity and consistency in the multi-hop semantic chain, trigger a cross-session answer fragment reconstruction engine, dynamically sort and connect discrete answer elements according to the logical order and weight ratio between nodes to generate a progressive splicing response that is coherent with the current context.
[0084] In this embodiment, the second-level context association cache involves data structures such as implicit semantic trajectory vectors, dynamic association graphs, and multi-hop semantic chains. These data are used to capture and understand the user's interaction history, identify potential context dependencies, and generate coherent answers accordingly.
[0085] In the embodiments of the present application, first, context encoding features in the current session are extracted to generate an implicit semantic trajectory vector; second, this vector is spatio-temporally aligned with the historical interaction path to construct a dynamic association graph; third, two-way semantic propagation calculation is performed to identify multi-hop semantic chains related to the current request; finally, when eligible associated nodes are detected, the answer fragment reconstruction engine is triggered to generate a progressive spliced answer according to the logical order and weight ratio;
[0086] The following is a specific example:
[0087] Suppose a user mentions the previously asked router configuration problem during the consultation on how to set up a home network. The system first extracts relevant context features of the current session and generates an implicit semantic trajectory vector; second, aligns this vector with the interaction path of the same user in the historical session library to construct a dynamic association graph containing multiple semantic nodes and their relationship edges; third, performs two-way semantic propagation calculation in the graph to find multi-hop semantic chains related to the current query; finally, discovers that there are more than two associated nodes with continuity and consistency, triggers the answer fragment reconstruction engine, combines relevant information according to the logical order and weight ratio, and provides a coherent and detailed answer to help the user successfully complete the setup of the home network; through the above steps, not only the consistency and accuracy of the conversation are improved, but also the user experience is enhanced, enabling the intelligent customer service to better understand and respond to complex requirements.
[0088] To further improve the processing ability for complex user requests, in some embodiments, when the second-level cache is not covered in step 103, the third-level intent decision cache is activated, the request is classified by scenario and marked by status through a multi-level parsing structure, a structured query instruction is generated based on the knowledge topology network, the weight of the knowledge retrieval path is dynamically adjusted, and multi-modal response elements are extracted to generate a combined answer, including:
[0089] The intent expression pattern and scenario feature parameters in the current interaction sequence are captured by the session state awareness module to generate a multi-dimensional classification vector containing domain labels and status variables; the multi-dimensional classification vector is mapped to a pre-constructed knowledge topology network, and a structured query instruction with path constraints is generated according to the domain association strength of each sub-graph in the network, and the query instruction includes the retrieval condition priority and cross-domain fusion rules; the focus shift and feedback signals in the user interaction behavior are monitored in real time, the credibility attenuation factor of each node in the knowledge retrieval path is dynamically calculated, and the scope and depth of cross-domain retrieval are adjusted based on the path weight reallocation algorithm; according to the updated retrieval path, text, image, and operation instruction multi-modal data units are parallelly extracted from heterogeneous data sources, and after filtering conflict elements through the semantic consistency verification module, the multi-modal units are embedded into the dynamically generated dialogue framework according to the preset answer assembly template to form a combined answer.
[0090] In this embodiment, the third-level intent decision cache involves data structures such as multi-dimensional classification vectors, knowledge topology networks, structured query instructions, and multi-modal data units. These data are used to accurately capture the user's intent and scenario requirements, and generate personalized response content accordingly.
[0091] In the embodiment of the present application, first, the session state awareness module analyzes the current interaction sequence, identifies the intent expression pattern and scenario features, and generates a multi-dimensional classification vector; secondly, maps this vector into the knowledge topology network, and generates structured query instructions according to the domain association strength; thirdly, real-time monitors the user's interaction behavior and feedback, and adjusts the node weights in the knowledge retrieval path; finally, extracts multi-modal data from different data sources according to the optimized retrieval path, integrates it into the dialogue framework after consistency verification, and provides a comprehensive answer.
[0092] The following is a specific example:
[0093] Suppose a user asks about how to solve the problem that their smartwatch cannot be synchronized with the mobile phone. First, the system analyzes the user's specific problem through the session state awareness module, determines their intent and scenario, and generates a multi-dimensional classification vector; secondly, maps this vector into the knowledge topology network, identifies the relevant technical fields, and generates structured query instructions with priorities and cross-domain fusion rules; thirdly, the system monitors the user's interaction behavior in real time, such as whether the user's attention has shifted to a specific solution or new questions have arisen, and dynamically adjusts the node weights in the retrieval path; finally, based on the updated retrieval path, the system extracts relevant information from various data sources such as text guides, operation videos, and troubleshooting steps, integrates it into a detailed answer according to a preset template after semantic consistency verification, and helps the user gradually solve the synchronization problem between the watch and the mobile phone; through the above steps, not only the understanding and response efficiency of complex problems are improved, but also the provided solutions are ensured to be comprehensive and accurate, greatly improving the user experience.
[0094] In order to further improve the accurate capture and response efficiency of the user's interaction behavior, in some embodiments, the focus shift and feedback signals in the real-time monitoring of the user's interaction behavior are used to dynamically calculate the credibility decay factor of each node in the knowledge retrieval path, and adjust the scope and depth of cross-domain retrieval based on the path weight reallocation algorithm, including:
[0095] Construct an attention transfer matrix based on the user interaction behavior sequence, extract the current interaction focus and historical focus sequence through the sliding window mechanism, calculate the focus transfer probability and generate a focus transfer vector; define a credibility attenuation factor in the knowledge topology network, calculate the node weight correction value in combination with the user feedback signal, and update the node weight; based on the updated node weight, use the path weight reallocation algorithm to calculate the priority and depth limit of the cross-domain retrieval path; generate a multimodal retrieval instruction according to the path weight coefficient and depth limit, and drive the parallel extraction module to extract data units according to the weight ratio and input them into the semantic consistency verification module.
[0096] In this embodiment, concepts such as an attention transfer matrix, a focus transfer vector, a credibility attenuation factor, and a path weight reallocation algorithm are involved. These data structures are used to dynamically analyze the user's interaction pattern and interest changes, optimize the knowledge retrieval path, and ensure that the provided information is both relevant and accurate.
[0097] In the embodiment of the present application, first, an attention transfer matrix is established based on the user's behavior sequence, the sliding window technology is used to identify the current and historical interaction foci, and the focus transfer probability is calculated to form a focus transfer vector; second, a credibility attenuation factor is set in the knowledge topology network, and the node weights are adjusted in combination with the user feedback; third, according to the updated node weights, the path weight reallocation algorithm is used to determine the priority and depth of the cross-domain retrieval path; finally, a multimodal retrieval instruction is generated based on the path weight and depth limit, guiding the parallel extraction module to obtain data according to the weight ratio, and integrating it into the response content after semantic consistency verification;
[0098] The following is a specific example: Suppose a user is querying how to solve the problem of frequent computer restarts. The system first constructs an attention transfer matrix based on the user's interaction behavior sequence, identifies whether the user is mainly concerned about hardware failures or software problems through the sliding window mechanism, calculates the focus transfer probability and generates a focus transfer vector; second, defines a credibility attenuation factor in the knowledge topology network, and adjusts the weights of each node in combination with the user's feedback on different solutions; third, based on the updated node weights, the system uses the path weight reallocation algorithm to determine the best retrieval path and its depth from hardware diagnosis to operating system inspection; finally, the system generates a multimodal retrieval instruction according to the optimized path weight, driving the parallel extraction module to extract relevant information from text guides, video tutorials, and forum discussions, and integrating it into a comprehensive and targeted answer after semantic consistency verification to help the user effectively solve the problem; through the above steps, not only the accuracy of understanding the user's needs is improved, but also the system's ability to provide personalized and efficient solutions is enhanced, greatly improving the user experience.
[0099] To further improve the flexibility and efficiency of the cache system, in some embodiments, in step 104, the synchronous execution of multi-level cache collaborative optimization is performed, the cache semantic coverage range and matching priority are dynamically adjusted according to the large model response result and user feedback, and the cache capacity adaptive expansion and elimination logic linkage are realized based on the session traffic characteristics, including:
[0100] Construct a semantic traceability link, decompose the final response generated by the large model reversely into a contribution degree distribution map of multi-level cache hit results, and calculate the semantic alignment error of each level of cache according to the dependency relationship between the map nodes; collect the user satisfaction score and follow-up behavior data at the end of the session, establish a cache effect evaluation matrix, and obtain the semantic coverage blind area and matching rule deviation that need to be corrected for each level of cache through matrix eigenvalue decomposition; analyze the spatio-temporal distribution characteristics of the session traffic in real time, construct a cache load prediction model based on a sliding window, and dynamically calculate the capacity pressure coefficient and hot data migration trend of each level of cache; according to the coupling relationship between the semantic alignment error and the capacity pressure coefficient, synchronously adjust the semantic matching threshold and storage partition strategy of each level of cache, and trigger the priority rearrangement and redundant data elimination of cross-layer caches based on the hot migration trend.
[0101] In this embodiment, concepts such as semantic traceability link, cache effect evaluation matrix, and cache load prediction model are involved. These tools are used to accurately analyze the effectiveness of the system response, user satisfaction, and the performance of the cache system, so as to dynamically adjust the cache policy to improve the overall service quality.
[0102] In the embodiment of the present application, first, a semantic traceability link is constructed to determine the contribution degree of multi-level caches by analyzing the response of the large model and calculate the semantic alignment error; second, the user satisfaction score and follow-up behavior data are collected, a cache effect evaluation matrix is established, and the semantic coverage blind area and matching rule deviation that need to be improved are identified; third, the sliding window mechanism is used to monitor the session traffic characteristics in real time, predict the cache load and calculate the pressure coefficient of each level of cache and the hot data migration trend; finally, according to the above analysis results, the semantic matching threshold and storage partition strategy of the cache are synchronously adjusted, and the cache priority is rearranged and redundant data is eliminated according to the hot migration trend;
[0103] The following is a specific example:
[0104] Suppose in an intelligent customer service system of an online education platform, a teacher frequently asks questions about the course management function. First, the system analyzes the answers to these questions through a large model of semantic traceability link analysis to determine the contribution of each level of cache, and calculates the semantic alignment error; second, the system collects the satisfaction score of the teacher at the end of the session and whether there is a behavior of further asking questions, establishes a cache effect evaluation matrix, and finds that the semantic coverage of some specific queries is insufficient; third, the system uses the sliding window technology to analyze the temporal and spatial distribution characteristics of the session traffic, predicts the cache load and determines which data are the current hotspots and which can be eliminated; finally, based on the relationship between the semantic alignment error and the capacity pressure coefficient, the system adjusts the semantic matching threshold and storage strategy of the cache, and rearranges the cache priority to ensure that the most relevant course management information can be quickly retrieved, and at the same time eliminates some outdated or infrequently used data; through the above steps, not only the response speed and accuracy of the cache system are improved, but also the resource allocation is optimized, so that the system can more effectively handle the query requirements under high concurrency, and significantly improves the user experience.
[0105] To further improve the adaptability and optimization performance of the cache system, in some embodiments, according to the coupling relationship between the semantic alignment error and the capacity pressure coefficient, synchronously adjusting the semantic matching threshold and storage partition strategy of each level of cache, and triggering the priority rearrangement and redundant data elimination of cross-layer cache based on the hotspot migration trend, includes:
[0106] Construct a cross-layer association model of the semantic alignment error matrix and the capacity pressure coefficient matrix, extract the collaborative optimization parameters of multi-level cache through the feature fusion algorithm, and generate a dynamic weight distribution map; based on the dynamic weight distribution map, use the path optimization algorithm to calculate the adjustment amount of the semantic matching threshold of each level of cache, and generate a storage partition strategy update instruction; construct a priority rearrangement queue according to the hotspot migration trend, and generate a cross-layer cache priority mapping table in combination with the temporal and spatial distribution characteristics of the session traffic; input the semantic matching threshold adjustment amount, the storage partition strategy update instruction and the cross-layer cache priority mapping table into the cache policy execution engine, synchronously trigger the threshold update, partition reconstruction and redundant data elimination operations, and feedback the adjustment results to the semantic traceability link for calibration.
[0107] In this embodiment, concepts such as semantic alignment error matrix, capacity pressure coefficient matrix, dynamic weight distribution map, and cross-layer cache priority mapping table are involved. These tools are used to accurately evaluate the system performance, and accordingly dynamically adjust the cache policy to ensure efficient data management and quick response to user requests.
[0108] In the embodiments of the present application, first, a cross-layer association model of a semantic alignment error matrix and a capacity pressure coefficient matrix is constructed, and a collaborative optimization parameter of a multi-level cache is extracted by using a feature fusion algorithm to form a dynamic weight distribution map; second, based on this map, a path optimization algorithm is used to calculate the adjustment amount of the semantic matching threshold of each level of cache and generate a new storage partition strategy; third, a priority re-queue is established according to the hot spot migration trend, and a cross-layer cache priority mapping table is generated by combining the spatio-temporal distribution characteristics of session traffic; finally, all adjustment information is input into a cache policy execution engine to synchronously perform threshold update, partition reconstruction, and redundant data elimination, and the final adjustment result is fed back to the semantic traceability link for calibration;
[0109] The following is a specific example:
[0110] Suppose in an intelligent customer service system of an e-commerce platform, in the face of a large number of inquiries about the product return policy, the system first constructs a cross-layer association model of a semantic alignment error matrix and a capacity pressure coefficient matrix, analyzes the collaborative optimization requirements of the current cache system through a feature fusion algorithm, and generates a dynamic weight distribution map; second, based on this map, the system calculates the semantic matching threshold that needs to be adjusted for each level of cache and formulates a new storage partition strategy; third, the system observes that inquiries about the return policy of a specific product have become hot spots, so a priority re-queue is established, and a cross-layer cache priority mapping table is generated by combining the temporal and spatial distribution characteristics of session traffic; finally, the system inputs all the above adjustment information into the cache policy execution engine, synchronously performs threshold update, partition reconstruction, and redundant data elimination operations, and feeds back these adjustment results to the semantic traceability link for effect verification; through the above steps, not only the flexibility and response speed of the cache system are improved, but also the efficiency of data management is ensured, significantly improving the user experience, especially performing well in handling high-concurrency inquiries.
[0111] In order to further improve the adaptability of the cache system to dynamic traffic and the optimization performance, in some embodiments, the spatio-temporal distribution characteristics of the session traffic are analyzed in real time, and a cache load prediction model based on a sliding window is constructed to dynamically calculate the capacity pressure coefficient and the hot spot data migration trend of each level of cache, including:
[0112] Divide the session traffic time series based on the sliding window mechanism, extract the interaction frequency, semantic complexity, and response delay parameters within the window, and generate a spatio-temporal feature vector; input the spatio-temporal feature vector into a hierarchical prediction model, extract short-term traffic fluctuation features through a temporal convolutional network, capture the cross-window traffic evolution law in combination with a long short-term memory network, and output the capacity pressure coefficients of each level of the cache; construct a hot data recognition engine based on the capacity pressure coefficients, screen high-frequency accessed semantic segments through a dynamic threshold algorithm, and generate a hot migration path map in combination with the semantic topological relationship; input the capacity pressure coefficients and the hot migration path map into the cache expansion decision module, trigger the dynamic reorganization of the storage partition and the elimination of redundant data, and synchronously update the load status mark of the semantic traceability link.
[0113] In this embodiment, concepts such as a sliding window mechanism, a spatio-temporal feature vector, a hierarchical prediction model (including a temporal convolutional network and a long short-term memory network), a hot data recognition engine, and a hot migration path map are involved. These tools are used to monitor and analyze the dynamic changes of session traffic in real time to optimize the capacity management and data storage strategy of the cache system.
[0114] In the embodiment of the present application, first, use the sliding window mechanism to process the time series of session traffic, extract parameters such as interaction frequency, semantic complexity, and response delay from it to form a spatio-temporal feature vector; second, input these feature vectors into a hierarchical prediction model, use a temporal convolutional network to capture short-term traffic fluctuation features, and identify long-term traffic patterns through a long short-term memory network, so as to calculate the capacity pressure coefficients of each level of the cache; third, construct a hot data recognition engine based on the capacity pressure coefficients, use a dynamic threshold algorithm to screen out high-frequency accessed semantic segments, and generate a hot migration path map in combination with the semantic topological relationship; finally, input the capacity pressure coefficients and the hot migration path map into the cache expansion decision module, trigger the dynamic adjustment of the storage partition and the elimination of redundant data, and synchronously update the load status mark in the semantic traceability link;
[0115] The following is a specific example:
[0116] Suppose in an intelligent customer service system of an online travel platform, in the face of the situation where users frequently query information about popular destinations, the system first uses a sliding window mechanism to analyze the time series of session traffic, extracts the interaction frequency, semantic complexity, and response latency within each window, and generates spatio-temporal feature vectors; secondly, these feature vectors are input into a hierarchical prediction model, which identifies short-term traffic fluctuation features through a temporal convolutional network, and then the long short-term memory network captures the traffic evolution law across windows to calculate the capacity pressure coefficients of each level of cache; thirdly, the system constructs a hot data recognition engine based on these coefficients, uses a dynamic threshold algorithm to filter out high-frequency access semantic fragments about specific popular destinations, and generates a hot migration path map; finally, the system inputs the capacity pressure coefficients and the hot migration path map into the cache expansion decision module, triggers the dynamic reorganization of the storage partition and the elimination operation of redundant data, and synchronously updates the load status markers in the semantic traceability link; through the above steps, not only the response speed and resource utilization rate of the cache system are significantly improved, but also the efficiency of data management is ensured. Especially in dealing with high-concurrency queries, the multi-level cache mechanism shows its advantages, greatly improving the user experience.
[0117] Figure 2 The embodiment of the present application provides a structural schematic diagram of a system based on large model deployment, as Figure 2 shown, the device includes:
[0118] A capture module 21, configured to establish a first-level dynamic semantic cache, capture the semantic feature encoding of the online session in real time to generate a unique identifier, count high-frequency semantic fragments based on dynamic access weights, and when the overlap degree between the semantic identifier of the new request and the cache fragment meets the preset condition, aggregate discrete semantic fragments to reconstruct the complete response logic and give priority to feedback;
[0119] An identification module 22, configured to start a second-level context-related cache for requests that miss, extract the implicit semantic trajectory vector of the session, construct a dynamic association map in combination with the historical interaction path, identify potential context dependencies across sessions, and when it is detected that there are continuous semantic nodes in the association map that form a logical link with the current request, trigger a progressive answer splicing mechanism;
[0120] An extraction module 23, configured to activate a third-level intention decision cache when the second-level cache is not covered, classify the scenario and mark the status of the request through a multi-level parsing structure, generate a structured query instruction based on the knowledge topology network, dynamically adjust the weight of the knowledge retrieval path, and extract multi-modal response elements to generate a combined response;
[0121] An adjustment module 24, configured to synchronously execute the collaborative optimization of the multi-level cache, dynamically adjust the semantic coverage range and matching priority of the cache according to the large model response result and user feedback, and realize the linkage between the cache capacity adaptive expansion and the elimination logic based on the session traffic characteristics.
[0122] Figure 2 The described large model deployment system based on a multi-level cache mechanism can execute Figure 1 The large model deployment method based on the multi-level cache mechanism described in the illustrated embodiment, its implementation principle and technical effects will not be elaborated. For the large model deployment device based on the multi-level cache mechanism in the above embodiment, the specific ways for each module and unit to execute operations have been described in detail in the embodiment related to this method, and will not be elaborated here.
[0123] In a possible design, Figure 2 A large model deployment system of the illustrated embodiment can be implemented as a computing device, such as Figure 3 As shown, the computing device can include a storage component 31 and a processing component 32;
[0124] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.
[0125] The processing component 32 is used for the Figure 1 Large model deployment method based on the multi-level cache mechanism of the illustrated embodiment.
[0126] Among them, the processing component 32 can include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0127] The storage component 31 is configured to store various types of data to support operations on the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0128] Of course, the computing device may also necessarily include other components, such as input / output interfaces, display components, communication components, etc.
[0129] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module can be an output device, an input device, etc.
[0130] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.
[0131] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device can refer to a cloud server. The above processing component, storage component, etc. can be basic server resources leased or purchased from a cloud computing platform.
[0132] An embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 A large model deployment method based on a multi-level cache mechanism shown in the embodiment.
[0133] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for deploying large models based on a multi-level caching mechanism, characterized in that, It includes the following steps: Establish a first-level dynamic semantic cache, capture the semantic feature encoding of the online session in real time to generate a unique identifier, statistically analyze high-frequency semantic fragments based on dynamic access weights. When the overlap degree between the semantic identifier of the new request and the high-frequency semantic fragments meets the preset condition, aggregate the high-frequency semantic fragments to reconstruct the complete response logic and give priority to feedback; For requests that miss the first-level cache, start the second-level context association cache, extract the implicit semantic trajectory vector of the session, construct a dynamic association graph by combining the historical interaction path, identify potential context dependencies across sessions. When continuous semantic nodes that form a logical link with the current request are detected in the association graph, trigger the progressive answer splicing mechanism; When the second-level cache is not covered, activate the third-level intent decision cache, classify the scenario and mark the status of the request through a multi-level parsing structure, generate a structured query instruction based on the knowledge topology network, dynamically adjust the weight of the knowledge retrieval path, and extract multi-modal response elements to generate a combined response; Synchronously execute the collaborative optimization of multi-level caches, dynamically adjust the semantic coverage range and matching priority of the cache according to the response results of the large model and user feedback, and realize the linkage between the adaptive expansion of the cache capacity and the elimination logic based on the session traffic characteristics; Among them, the semantic identifier of the first-level cache is orthogonally complementary to the semantic trajectory vector space of the second-level cache, the association graph nodes of the second-level cache are mapped to the knowledge topology network of the third-level cache through the semantic bridging layer, and the cache update conditions of each level are dynamically negatively correlated with the session flow complexity.
2. The method according to claim 1, wherein The step of starting the second-level context association cache for requests that miss the first-level cache, extracting the implicit semantic trajectory vector of the session, constructing a dynamic association graph by combining the historical interaction path, and identifying potential context dependencies across sessions. When continuous semantic nodes that form a logical link with the current request are detected in the association graph, triggering the progressive answer splicing mechanism includes: Extract the context encoding features of the continuous interaction sequence in the current session to generate an implicit semantic trajectory vector with time-dependent relationships; Align the implicit semantic trajectory vector with the interaction paths of the same user identifier or similar user groups in the historical session library in space-time to construct a dynamic association graph containing semantic nodes and relationship edges, where the graph nodes store cross-session semantic context fragments, and the relationship edges record the logical jump probability and timeliness weight between the nodes; Perform two-way semantic propagation calculation on the dynamic association graph, activate adjacent nodes on the association path based on the semantic trajectory vector of the current request, and identify multi-hop semantic chains that form logical coherence with the current session through the path confidence screening algorithm; When it is detected that there are at least two or more associated nodes with time continuity and intention consistency in the multi-hop semantic chain, trigger the reconstruction engine of cross-session answer fragments, dynamically sort and connect the discrete answer elements according to the logical order and weight ratio between the nodes, and generate a progressive splicing response that is coherent with the current context.
3. The method according to claim 1, characterized in that, When the second-level cache is not covered, activate the third-level intent decision cache, classify the scenarios and mark the status of the request through a multi-level parsing structure, generate a structured query instruction based on the knowledge topology network, dynamically adjust the weight of the knowledge retrieval path, and extract multi-modal response elements to generate a combined response, including: Capture the intent expression pattern and scenario feature parameters in the current interaction sequence through the session state awareness module, and generate a multi-dimensional classification vector including domain labels and state variables; Map the multi-dimensional classification vector to a pre-constructed knowledge topology network, and generate a structured query instruction with path constraints according to the domain association strength of each sub-graph in the network. The query instruction includes the retrieval condition priority and the cross-domain fusion rule; Real-time monitor the focus transfer and feedback signals in the user interaction behavior, dynamically calculate the credibility attenuation factor of each node in the knowledge retrieval path, and adjust the scope and depth of cross-domain retrieval based on the path weight reallocation algorithm; Parallelly extract text, image, and operation instruction multi-modal data units from heterogeneous data sources according to the updated retrieval path. After filtering conflict elements through the semantic consistency verification module, embed the multi-modal units into the dynamically generated dialogue framework according to the preset response assembly template to form a combined response.
4. The method according to claim 1, characterized in that The synchronous execution of multi-level cache collaborative optimization dynamically adjusts the cache semantic coverage range and matching priority according to the large model response result and user feedback, and realizes the linkage of cache capacity adaptive expansion and elimination logic based on the session traffic characteristics, including: Construct a semantic traceability link, reverse decompose the final response generated by the large model into a contribution degree distribution map of multi-level cache hit results, and calculate the semantic alignment error of each level of cache according to the dependency relationship between the map nodes; Collect the user satisfaction score and follow-up behavior data at the end of the session, establish a cache effect evaluation matrix, and obtain the semantic coverage blind area and matching rule deviation that need to be corrected for each level of cache through matrix eigenvalue decomposition; Real-time analyze the spatio-temporal distribution characteristics of session traffic, construct a cache load prediction model based on a sliding window, and dynamically calculate the capacity pressure coefficient and hot data migration trend of each level of cache; According to the coupling relationship between the semantic alignment error and the capacity pressure coefficient, synchronously adjust the semantic matching threshold and storage partition strategy of each level of cache, and trigger the priority rearrangement and redundant data elimination of cross-layer cache based on the hot migration trend.
5. The method according to claim 3, characterized in that, The real-time monitoring of the focus transfer and feedback signals in the user interaction behavior, dynamically calculating the credibility attenuation factor of each node in the knowledge retrieval path, and adjusting the scope and depth of cross-domain retrieval based on the path weight reallocation algorithm, including: Construct an attention transfer matrix based on the user interaction behavior sequence, extract the current interaction focus and historical focus sequences through the sliding window mechanism, calculate the focus transfer probability and generate a focus transfer vector; Define the credibility attenuation factor in the knowledge topology network, calculate the node weight correction value in combination with the user feedback signal, and update the node weight; Based on the updated node weights, use the path weight reallocation algorithm to calculate the priority and depth limit of the cross-domain retrieval path; Generate multimodal retrieval instructions based on path weight coefficients and depth limits, and drive the parallel extraction module to extract data units according to the weight ratio and input them into the semantic consistency verification module.
6. The method according to claim 4, characterized in that, Synchronously adjust the semantic matching thresholds and storage partition strategies of each level of cache according to the coupling relationship between the semantic alignment error and the capacity pressure coefficient, and trigger the priority rearrangement and redundant data elimination of cross-layer caches based on the hot migration trend, including: Construct a cross-layer association model of the semantic alignment error matrix and the capacity pressure coefficient matrix, extract the collaborative optimization parameters of multiple levels of cache through a feature fusion algorithm, and generate a dynamic weight allocation map; Based on the dynamic weight allocation map, use a path optimization algorithm to calculate the adjustment amount of the semantic matching threshold of each level of cache, and generate a storage partition strategy update instruction; Construct a priority requeueing queue according to the hot migration trend, and generate a cross-layer cache priority mapping table in combination with the spatio-temporal distribution characteristics of session traffic; Input the semantic matching threshold adjustment amount, the storage partition strategy update instruction, and the cross-layer cache priority mapping table into the cache policy execution engine, synchronously trigger threshold update, partition reconstruction, and redundant data elimination operations, and feedback the adjustment results to the semantic traceability link for calibration.
7. The method according to claim 4, characterized in that Real-time analyze the spatio-temporal distribution characteristics of session traffic, construct a cache load prediction model based on a sliding window, and dynamically calculate the capacity pressure coefficients and hot data migration trends of each level of cache, including: Divide the session traffic time series based on the sliding window mechanism, extract the interaction frequency, semantic complexity, and response delay parameters within the window, and generate a spatio-temporal feature vector; Input the spatio-temporal feature vector into a hierarchical prediction model, extract short-term traffic fluctuation characteristics through a temporal convolutional network, and capture the cross-window traffic evolution law in combination with a long short-term memory network, and output the capacity pressure coefficients of each level of cache; Construct a hot data recognition engine based on the capacity pressure coefficient, screen high-frequency accessed semantic segments through a dynamic threshold algorithm, and generate a hot migration path map in combination with the semantic topological relationship; Input the capacity pressure coefficient and the hot migration path map into the cache expansion decision module, trigger the dynamic reorganization of the storage partition and the redundant data elimination mechanism, and synchronously update the load status mark of the semantic traceability link.
8. A large model deployment system based on a multi-level caching mechanism, characterized in that, Including: A capture module for establishing a first-level dynamic semantic cache, real-time capturing the semantic feature encoding of an online session to generate a unique identifier, statistically counting high-frequency semantic segments based on dynamic access weights, and when the overlap degree between the semantic identifier of a new request and the high-frequency semantic segments meets a preset condition, aggregating the high-frequency semantic segments to reconstruct a complete response logic and giving priority feedback; An identification module for starting a second-level context association cache for an unhit request, extracting the implicit semantic trajectory vector of the session, constructing a dynamic association map in combination with the historical interaction path, identifying potential context dependencies across sessions, and triggering a progressive answer splicing mechanism when continuous semantic nodes that form a logical link with the current request are detected in the association map; An extraction module, which is used to activate the third-level intent decision cache when the second-level cache is not covered, classify the scenarios and mark the status of requests through a multi-level parsing structure, generate structured query instructions based on the knowledge topology network, dynamically adjust the weights of the knowledge retrieval paths, and extract multi-modal response elements to generate a combined response; An adjustment module, which is used to synchronously execute multi-level cache collaborative optimization, dynamically adjust the cache semantic coverage and matching priority according to the large model response results and user feedback, and realize the linkage between cache capacity adaptive expansion and elimination logic based on the session traffic characteristics.
9. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a large model deployment method based on a multi-level cache mechanism as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a computer, it implements a large model deployment method based on a multi-level cache mechanism as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent question and answer interaction system based on semantic analysis engine
CN117056479A
Retrieval enhanced intelligent question answering method and system based on large model
CN118503392A