Inference acceleration optimization method and system applied to intelligent dialogue large model

By jointly deconstructing the dialogue sequence to be reasoned and the configuration information of the reasoning environment in the intelligent dialogue big model, a reasoning node dependency graph and a resource requirement list are generated. This optimizes the reasoning link of the intelligent dialogue big model, solves the problems of slow reasoning speed and unreasonable resource allocation in the existing intelligent dialogue big model, and realizes the rapid and accurate response of the intelligent dialogue system.

CN120725158BActive Publication Date: 2025-11-28XINGFAN XINGQI (CHENGDU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511211400.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-28
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing methods for accelerating reasoning in large-scale intelligent dialogue models fail to effectively combine hardware resources and algorithm optimization, resulting in slow reasoning speeds that cannot meet the demands of real-time interaction. Furthermore, they neglect the logical relationships and dependencies between semantic processing steps during the reasoning process, leading to unreasonable resource allocation and an inability to fully leverage the advantages of hardware performance.

Method used

By acquiring the dialogue sequence to be inferred and the configuration information of the inference environment from the intelligent dialogue big model, joint process deconstruction is performed to generate an inference node dependency graph and a resource elasticity requirement list. Based on this information, inference link optimization is performed to determine the parallel scheduling rules and resource pre-allocation strategies for inference nodes, and the inference operation process is adjusted to generate an accelerated dialogue response sequence.

Benefits of technology

It enables the rational scheduling of computing nodes and allocation of resources based on the dependencies and resource requirements of inference nodes, thereby improving inference efficiency, shortening inference time, providing a smoother and more real-time intelligent dialogue experience, and enhancing the overall performance and user satisfaction of the intelligent dialogue system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725158B_ABST
    Figure CN120725158B_ABST
Patent Text Reader

Abstract

The application provides a reasoning acceleration optimization method and system applied to an intelligent dialogue large model, and belongs to the technical field of large models. First, a dialogue sequence to be reasoned and reasoning environment configuration information are acquired, wherein the dialogue sequence to be reasoned comprises real-time input text and a historical interactive sentence chain of a user, and the reasoning environment configuration information covers operation node load states and cache resource occupation information. Next, joint process deconstruction processing is performed on the two, to obtain a reasoning node dependency graph and a resource elasticity demand list. Then, reasoning link optimization processing is performed based on the above results, to generate a reasoning acceleration execution scheme, which comprises reasoning node parallel scheduling rules and resource pre-allocation strategies. The reasoning acceleration execution scheme is used to regulate a reasoning operation process, to generate an accelerated processed dialogue response sequence. Finally, the accelerated processed dialogue response sequence is pushed to a user interactive terminal to complete intelligent dialogue output, so as to effectively improve the reasoning speed of the intelligent dialogue large model and optimize the dialogue interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large models, in particular to a reasoning acceleration optimization method and system applied to intelligent dialogue large models. BACKGROUND

[0002] In the field of intelligent dialogue systems, with the continuous development of large model technology, intelligent dialogue large models provide users with a more natural and smooth dialogue experience due to their strong semantic understanding and generation capabilities. However, intelligent dialogue large models usually have a large number of parameters and complex computational structures, which consume a large amount of computational resources during dialogue reasoning, resulting in slow reasoning speed and difficulty in meeting the needs of real-time interaction.

[0003] Currently, existing intelligent dialogue large model reasoning acceleration methods mainly focus on single aspect optimization. Some methods only focus on hardware resource optimization, increasing the number of operation nodes or improving the performance of individual operation nodes to improve reasoning speed, but the above methods often ignore the logical relationship and dependency relationship between each semantic processing step in the reasoning process, which may lead to unreasonable resource allocation and failure to fully utilize the performance advantages of hardware. Some other methods focus on algorithm-level optimization, such as pruning, quantization, and other operations on the model to reduce the computational load and storage requirements of the model, but these methods may sacrifice the accuracy and semantic understanding capabilities of the model to some extent, affecting the quality of the dialogue. In addition, existing methods do not fully consider the dynamic changes of the reasoning environment, such as real-time fluctuations in operation node load state and changes in cache resource occupancy, and cannot dynamically adjust the reasoning strategy according to the actual environment, resulting in low reasoning efficiency. SUMMARY

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present application, a reasoning acceleration optimization method applied to intelligent dialogue large models is provided, which comprises:

[0005] Obtaining a dialogue sequence to be reasoned and reasoning environment configuration information of an intelligent dialogue large model, the dialogue sequence to be reasoned containing real-time input text and historical interactive sentence chain of the user, the reasoning environment configuration information containing operation node load state and cache resource occupancy information, the operation node being a hardware processing unit for performing reasoning calculation;

[0006] Joint process deconstruction processing of the dialogue sequence to be reasoned and the reasoning environment configuration information to obtain a reasoning node dependency graph and a resource elasticity demand list, the reasoning node dependency graph containing reasoning node association hierarchy and reasoning order constraint, the reasoning node being a logical unit corresponding to a semantic processing step in the reasoning process, the resource elasticity demand list containing operation resource dynamic threshold and cache resource allocation range;

[0007] perform inference link optimization processing based on the inference node dependency graph and the resource elasticity requirement list to obtain an inference acceleration execution scheme, the inference acceleration execution scheme including an inference node parallel scheduling rule and a resource pre-allocation strategy;

[0008] According to the inference acceleration execution scheme, the inference operation process of the intelligent dialogue large model is regulated to generate an accelerated processed dialogue response sequence.

[0009] The accelerated processed dialogue response sequence is pushed to a user interaction terminal to complete intelligent dialogue output.

[0010] In another aspect, the present application also provides an inference acceleration optimization system applied to an intelligent dialogue large model, which includes a processor and a machine readable storage medium, the machine readable storage medium is connected with the processor, the machine readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine readable storage medium to realize the above method.

[0011] Based on the above aspects, by obtaining the dialogue sequence to be inferred and inference environment configuration information of the intelligent dialogue large model, the dialogue sequence to be inferred and the inference environment configuration information are subjected to joint process deconstruction processing to obtain an inference node dependency graph and a resource elasticity requirement list, the correlation level and inference order constraint between each semantic processing step in the inference process are depicted, and the dynamic demand range of the operation resource and the cache resource is also determined. The inference link optimization processing is performed based on the inference node dependency graph and the resource elasticity requirement list to determine the inference node parallel scheduling rule and the resource pre-allocation strategy. According to the dependency relationship and resource requirement of the inference node, the operation node can be reasonably scheduled and the resource can be allocated, the performance advantage of the hardware can be fully utilized, the inference efficiency is improved, the inference operation process of the intelligent dialogue large model is regulated according to the inference acceleration execution scheme to generate an accelerated processed dialogue response sequence, the inference time is effectively shortened, the dialogue response is quickly and accurately realized, and finally the accelerated processed dialogue response sequence is pushed to the user interaction terminal to provide the user with a more smooth and real-time intelligent dialogue experience, and the overall performance and user satisfaction of the intelligent dialogue system are improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is an execution flow schematic diagram of the inference acceleration optimization method applied to the intelligent dialogue large model provided by the embodiments of the present application.

[0013] Figure 2 is a schematic diagram of exemplary hardware and software components of the inference acceleration optimization system applied to the intelligent dialogue large model provided by the embodiments of the present application. DETAILED DESCRIPTION

[0014] The application will be described in detail below with reference to the accompanying drawings, Figure 1 is a flowchart of the inference acceleration optimization method applied to the intelligent dialogue large model provided by an embodiment of the application. The inference acceleration optimization method applied to the intelligent dialogue large model will be described in detail below.

[0015] Step S110: Obtain the inference dialogue sequence and inference environment configuration information of the intelligent dialogue large model, wherein the inference dialogue sequence comprises real-time input text and a historical interactive sentence chain of the user, and the inference environment configuration information comprises an operation node load state and cache resource occupation information, and the operation node is a hardware processing unit for performing inference calculation.

[0016] In an office scenario, the user may be handling a large project involving multiple sub-projects, cross-department collaboration, and a tight time node. At this time, the real-time text input by the user through the intelligent office assistant often contains multi-dimensional requirements, which are interrelated and each has its own focus, and the logical structure is complex. For example, the user may input a text containing project progress checking, resource conflict coordination, cross-department communication arrangement, and risk plan adjustment, and there is a certain causal relationship and time sequence between these contents. At the same time, the historical interaction of the user with the intelligent dialogue large model may involve the initial planning of the large project, the division of responsibilities of each sub-project, the allocation of resources in the early stage, and some problems and solution records that have occurred, etc. The above sentence content jointly constitutes a historical interactive sentence chain.

[0017] To obtain the inference dialogue sequence, it is necessary to receive the above-mentioned complex text input by the user at the interactive terminal in real time, and to retrieve all historical dialogue records related to the current user from the database storing historical interaction data. For the inference environment configuration information, the hardware processing unit for performing inference calculation needs to be monitored to collect the current running state of each operation node, such as the number of tasks being processed, the resource occupation of each task, etc., to form the operation node load state; at the same time, the storage situation and access frequency of data in different level caches are queried and integrated into cache resource occupation information.

[0018] Step S111: Receive the dialogue input data stream transmitted by the user interactive terminal, extract the real-time input text of the user therefrom in chronological order, and mark the input sequence of the text segments.

[0019] In this embodiment, the user may input the above complex office content in multiple times, and each input text segment carries part of the information, and there is a logical connection between the front and rear segments. For example, the user first inputs "For the overall promotion of project B, it is necessary to check the current progress of each sub-project, especially the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2", and after a period of time, inputs "In addition, the equipment procurement of sub-project B1 may be delayed due to supplier reasons, which will affect the installation and commissioning of sub-project B3, and it is necessary to coordinate the technical department and the procurement department to synchronously promote the solution", and then may supplement "At the same time, the participants of the cross-department coordination meeting on next Wednesday need to be determined in advance to ensure that the key persons of the marketing department, the finance department and the technical department can attend to discuss the resource reallocation and risk response strategy synchronously".

[0020] The above text segments are transmitted in the form of data stream, and each text segment has a corresponding timestamp. After receiving the data stream, the text segments input by the user in real time are extracted according to the chronological order of the timestamps, and a sequential marker is added to each semantic segment, such as "segment one", "segment two", "segment three", etc., so as to ensure that the input order of the text segments is not confused.

[0021] Step S112: accessing the historical interaction database of the intelligent dialogue large model, retrieving the historical interaction records associated with the current user identifier, and arranging the historical interaction records into a historical interaction sentence chain according to the interaction time, wherein the historical interaction sentence chain includes the user historical input content and the model historical response content.

[0022] The database storing the historical interaction data is accessed through the unique identifier of the user, such as the user account. All interaction records related to the user identifier are filtered out in the database, which covers all the content of the user's past communication with the intelligent dialogue large model about project B. For example, it may include the input content when the user asked "the budget upper limit of equipment procurement of sub-project B1" last week, and the response "the budget upper limit of equipment procurement of sub-project B1 is set according to the proportion of the total budget of project B" given by the intelligent dialogue large model at that time; it also includes the input when the user previously proposed "personnel training plan of sub-project B2" and the specific training plan suggestion provided by the intelligent dialogue large model.

[0023] The above filtered historical interaction records are arranged in chronological order from early to late, forming a continuous historical interaction sentence chain, which presents the communication history of the user and the model about project B.

[0024] Step S113: merging the user real-time input text and the historical interaction sentence chain, adding sequence separator and semantic association marker, and generating a dialogue sequence to be reasoned.

[0025] The user real-time input text segment with sequential marks is merged with the historical interactive sentence chain sorted by time. In order to clearly distinguish the historical content and the current input content, a sequence separator is added, such as using "[end of history]" to mark the end of the historical interactive sentence chain, and using "[current start]" to mark the start of the user real-time input text.

[0026] At the same time, semantic association marks are added according to the semantic association between the text contents. For example, the historical interactive sentence chain mentions "the supplier selection standard of sub-project B1", and the user real-time input text involves "the equipment procurement of sub-project B1 may be delayed due to the supplier reason", so the mark "[association: sub-project B1 supplier related]" is added; the historical interactive sentence chain discusses "the organization process of cross-department meeting", and the current input mentions "the cross-department coordination meeting next Wednesday", so the mark "[association: cross-department meeting arrangement related]" is added. Through the above processing, a dialogue sequence to be reasoned is generated, which contains complete context and clear semantic association.

[0027] Step S114: Collect the current running state of each operation node through the operation node monitoring program, and integrate it into the operation node load state. The operation node is a hardware processing unit for executing reasoning calculation, and the running state includes the number of active processes and the length of task waiting queue.

[0028] The monitoring program running on each operation node continuously collects the running data of the node. For the number of active processes, the monitoring program scans all the processes currently in the running state of the node and counts the total number; for the length of task waiting queue, the task list waiting to be executed in the node can be viewed to count the number of tasks.

[0029] The collected number of active processes and the length of task waiting queue of each operation node are summarized. For example, the number of active processes of operation node A is a certain number, and the length of task waiting queue is a certain number; the number of active processes of operation node B is another number, and the length of task waiting queue is another number, etc. After integrating the above information, the operation node load state is formed, reflecting the current busy degree of each node.

[0030] Step S115: Call the cache resource management interface to query the usage of each level cache, and integrate it into the cache resource occupation information, which includes the used storage space proportion and data access heat.

[0031] The usage data of different levels of caches such as the first-level cache and the second-level cache is queried by calling the cache resource management interface. For the used storage space ratio, the proportion of the capacity of the stored data in each level of cache to the total capacity of the level of cache is calculated; for the data access heat, the number of times that different data in the cache is accessed in the recent period of time is counted, and the more the number of times, the higher the access heat of the data.

[0032] The queried information is integrated, such as the used storage space ratio of the first-level cache being a certain ratio, the data with higher access heat including the basic information data of project B, and the used storage space ratio of the second-level cache being another ratio, the data with higher access heat including the progress record data of each sub-project, to form cache resource occupation information.

[0033] Step S116: Establishing a time association mark of the to-be-reasoned dialogue sequence and the reasoning environment configuration information to form an association data set.

[0034] A time stamp of the acquisition time is added to the to-be-reasoned dialogue sequence and the reasoning environment configuration information respectively to ensure the consistency of the accuracy of the time stamp. Then, the to-be-reasoned dialogue sequence and the reasoning environment configuration information acquired at the same time point are associated using a common association identifier.

[0035] For example, the time stamp of the to-be-reasoned dialogue sequence is a specific time, and the time stamp of the reasoning environment configuration information is also the time, then the two are bound using “association identifier-the time” to form an association data set. In this way, it can be ensured that the dialogue sequence and the environment configuration information used in the subsequent processing process are at the same time point, avoiding the problem of information mismatch caused by time difference.

[0036] Step S120: Joint process deconstruction processing is performed on the to-be-reasoned dialogue sequence and the reasoning environment configuration information to obtain a reasoning node dependency graph and a resource elasticity demand list, the reasoning node dependency graph contains reasoning node association levels and reasoning sequence constraints, the reasoning node is a logical unit corresponding to a semantic processing step in a reasoning process, and the resource elasticity demand list contains an operation resource dynamic threshold and a cache resource allocation range.

[0037] After obtaining the association data set, the to-be-reasoned dialogue sequence and the reasoning environment configuration information in the association data set need to be jointly processed and deconstructed. For the to-be-reasoned dialogue sequence, the semantic structure is analyzed in depth, and complex user requirements are decomposed into multiple mutually associated semantic processing steps, each step corresponding to a reasoning node. At the same time, in combination with the reasoning environment configuration information, the dependency relationship between each reasoning node and the demand range of each node in terms of operation resources and cache resources are determined.

[0038] For example, the user inputs in real time the text "check the current progress of each sub-project", "coordinate the solution to the delay problem of sub-project B1", "determine the participants of the cross-department coordination meeting", and the like, which correspond to different reasoning nodes. Through analysis, it is known that "coordinate the solution to the delay problem of sub-project B1" needs to be based on the result of "check the current progress of each sub-project", and "determine the participants of the cross-department coordination meeting" may have information interaction with the former two, thereby forming the association level and sequence constraint between the reasoning nodes. In combination with the operation node load state and cache resource occupation information, the dynamic threshold of the operation resource required by each reasoning node and the cache resource allocation range are determined, and then the reasoning node dependency graph and the resource elasticity demand list are obtained.

[0039] Step S121: performing semantic level splitting on the to-be-reasoned dialogue sequence to identify semantic blocks therein, the semantic blocks including core semantic blocks and associated semantic blocks, the core semantic blocks corresponding to user core demands, and the associated semantic blocks being supplementary information for auxiliary understanding.

[0040] The to-be-reasoned dialogue sequence is subjected to detailed semantic analysis and is split into multiple semantic blocks. The core semantic blocks are the parts directly embodying the user core demands, and the associated semantic blocks are supplementary explanations of the core semantic blocks, which are helpful for more accurately understanding the user intent.

[0041] In the complex text input by the user about project B, "check the current progress of each sub-project", "coordinate the solution to the delay problem of sub-project B1", "determine the participants of the cross-department coordination meeting", and the like belong to the core semantic blocks, because these are the main demands explicitly proposed by the user. "Especially the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2" is a supplement to "check the current progress of each sub-project", "may be delayed due to supplier reasons, which will affect the installation and commissioning of sub-project B3" is a further explanation of "delay problem of sub-project B1", and "ensure that the key persons of the marketing department, the finance department, and the technology department can all attend to discuss resource reallocation and risk response strategies synchronously" is a supplementary requirement for "determine the participants of the cross-department coordination meeting", all of which belong to the associated semantic blocks.

[0042] Step S1211: scanning the to-be-reasoned dialogue sequence by using a semantic boundary detection algorithm to identify semantic pause points and logical transition markers therein as semantic block splitting boundaries.

[0043] The semantic boundary detection algorithm is used to comprehensively scan the to-be-reasoned dialogue sequence. The semantic pause points usually occur at comma, period, semicolon, and the like, which are the natural pause positions of semantic expression. The logical transition markers include conjunctions such as "in addition", "at the same time", "however", "especially", and the like, which play a role in connecting different semantic contents and indicating logical relationships in the text.

[0044] In the user input text, the comma in "For the overall progress of project B, it is necessary to check the current progress of each sub-project, especially the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2" is a semantic pause point; "In addition, the equipment procurement of sub-project B1 may be delayed due to supplier reasons" is a logical transition marker, and the comma after it is also a semantic pause point; "At the same time, the participants of the cross-department coordination meeting on next Wednesday need to be determined in advance" is a logical transition marker, and the comma after it is a semantic pause point. These are identified as semantic block splitting boundaries.

[0045] Step S1212: The to-be-reasoned dialogue sequence is segmented into multiple semantic segments according to the semantic block splitting boundaries, and each semantic segment contains complete semantic information.

[0046] According to the identified semantic block splitting boundaries, the to-be-reasoned dialogue sequence is segmented into multiple semantic segments. Each semantic segment can independently express a relatively complete semantic.

[0047] For example, according to the above boundaries, the user input text is segmented into semantic segments such as "For the overall progress of project B, it is necessary to check the current progress of each sub-project", "the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2", "the equipment procurement of sub-project B1 may be delayed due to supplier reasons, which will affect the installation and commissioning of sub-project B3, and it is necessary to coordinate the technical department and the procurement department to synchronously promote the solution", "the participants of the cross-department coordination meeting on next Wednesday need to be determined in advance", "the key persons of the marketing department, the finance department and the technical department can attend to discuss the resource reallocation and risk response strategy synchronously", etc. Each semantic segment contains complete semantic information.

[0048] Step S1213: Determine the intent weight of each semantic segment based on the occurrence frequency and semantic importance of the core verbs and nouns in the semantic segment, and mark the semantic segments with an intent weight greater than a preset weight threshold as core semantic blocks, and mark the rest as associated semantic blocks.

[0049] Each semantic segment is subjected to lexical analysis to extract the core verbs and nouns therein, the occurrence frequency of which in the segment is counted, and the intent weight is determined in combination with their semantic importance in the entire to-be-reasoned dialogue sequence. The closer the association between the core verbs and nouns and the user's core demand, the higher the semantic importance, and the greater the intent weight.

[0050] The core verb in "For the overall progress of Project B, it is necessary to check the current progress of each sub-project" is "check", and the core nouns are "Project B", "sub-project" and "progress". These words directly point to one of the core needs of the user, and although the frequency is not high, the semantic importance is very high. The calculated intent weight is greater than the preset weight threshold, so it is marked as a core semantic block.

[0051] The core nouns in "The equipment procurement link of sub-project B1 and the personnel training link of sub-project B2" are "sub-project B1", "equipment procurement link", "sub-project B2" and "personnel training link". They are specific descriptions of "the current progress of each sub-project", and the semantic importance is relatively low. The intent weight is less than the preset weight threshold, so it is marked as an associated semantic block.

[0052] Step S1214: Calculate the semantic distance between the associated semantic block and the core semantic block based on the number of co-occurring words and the grammatical dependency relationship, and bind the associated semantic block with the corresponding core semantic block if the semantic distance is less than the preset distance threshold to form a semantic block group.

[0053] Calculate the number of co-occurring words between the associated semantic block and the core semantic block, that is, the number of words that appear in both semantic blocks. At the same time, analyze the grammatical dependency relationship between them, such as whether the associated semantic block is a modifier of a noun in the core semantic block. According to these factors, calculate the semantic distance. The smaller the semantic distance, the closer the association between the two semantic blocks.

[0054] "Sub-project B1 equipment procurement link and sub-project B2 personnel training link" and "For the overall progress of Project B, it is necessary to check the current progress of each sub-project" have the co-occurring word "sub-project", and the former is a specific enumeration of the latter "each sub-project", which has a grammatical modification relationship. Therefore, the semantic distance between them is small, less than the preset distance threshold, so they are bound together to form a semantic block group.

[0055] Step S1215: Add a hierarchical label to each semantic block group, where the core semantic block is level one, the directly associated associated semantic block is level two, and the indirectly associated is level three.

[0056] In the formed semantic block group, the core semantic block is at the highest level and is marked as level one. The associated semantic block directly associated with the core semantic block is at the second level and is marked as level two. If there are other associated semantic blocks associated with the level two associated semantic block, they are marked as level three.

[0057] For example, in the above semantic block group, “For the overall progress of project B, it is necessary to check the current progress of each sub-project” is marked as a first-level core semantic block; “the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2” are marked as second-level associated semantic blocks. If there is “specific brand requirements for sub-project B1 equipment procurement” related to the second-level associated semantic block, it will be marked as a third-level.

[0058] Step S1216: output the semantic hierarchy splitting result containing the core semantic block, the associated semantic block, the semantic block group, and the hierarchical marking.

[0059] The core semantic block, the associated semantic block, the semantic block group to which they belong, and the corresponding hierarchical marking obtained by processing are arranged to form a semantic hierarchy splitting result and output. The output result can be as follows: semantic block group 1 contains first-level core semantic block “For the overall progress of project B, it is necessary to check the current progress of each sub-project” and second-level associated semantic block “the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2”; semantic block group 2 contains first-level core semantic block “The equipment procurement of sub-project B1 may be delayed due to supplier reasons, which will affect the installation and commissioning of sub-project B3, and it is necessary to coordinate the technical department and the procurement department to synchronously promote solutions”; semantic block group 3 contains first-level core semantic block “The cross-department coordination meeting on next week Wednesday needs to determine the attendees in advance” and second-level associated semantic block “The key persons of the marketing department, the finance department, and the technical department can attend to discuss the resource reallocation and risk response strategies synchronously” and the like.

[0060] Step S122: constructing a preliminary reasoning node based on the logical association between the semantic blocks, the preliminary reasoning node being an initial construction form of the reasoning node, each preliminary reasoning node corresponding to a specific semantic processing step, and being used to record the input semantic block and the output semantic block of the preliminary reasoning node.

[0061] The preliminary reasoning node is constructed according to the logical association between the semantic blocks. Each preliminary reasoning node corresponds to a specific semantic processing step, and needs to clearly define the input semantic block (i.e. the information to be processed by the step) and the output semantic block (i.e. the result obtained after the step is processed).

[0062] For example, for the semantic block group 1, a preliminary reasoning node 1 is constructed, the input semantic block of which is the first-level core semantic block “For the overall progress of project B, it is necessary to check the current progress of each sub-project” and the second-level associated semantic block “the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2”, and the output semantic block of which is the checking result of the current progress of each sub-project, including the specific progress of the equipment procurement link of sub-project B1 and the personnel training link of sub-project B2.

[0063] For semantic block group 2, a preliminary reasoning node 2 is constructed, with the input semantic block being the first-level core semantic block "The procurement of equipment for sub-project B1 may be delayed due to supplier reasons, which will affect the installation and commissioning of sub-project B3, and coordination between the technical and procurement departments is needed to synchronize the solution", and the output semantic block being the solution to the delay problem of sub-project B1 obtained after coordination between the technical and procurement departments.

[0064] For semantic block group 3, a preliminary reasoning node 3 is constructed, with the input semantic block being the first-level core semantic block "The cross-department coordination meeting next Wednesday needs to determine the attendees in advance" and the second-level associated semantic block "The key persons from the marketing, finance and technical departments can attend to discuss the resource reallocation and risk response strategies synchronously", and the output semantic block being the list of attendees of the cross-department coordination meeting, including the names and positions of the key persons from the marketing, finance and technical departments.

[0065] In addition, if there are other semantic block groups in the dialogue sequence to be reasoned, such as content related to budget adjustment of project B, corresponding preliminary reasoning nodes will also be constructed for them. For example, semantic block group 4 includes the first-level core semantic block "According to the progress and resource usage of each sub-project, the budget of project B needs to be adjusted" and the second-level associated semantic block "The procurement budget of sub-project B1 and the training budget of sub-project B2 need to be adjusted", and a preliminary reasoning node 4 is constructed, with the input semantic block being the core semantic block and the associated semantic block, and the output semantic block being the budget adjustment plan of project B, which clearly specifies the increase and decrease of the budget of each sub-project and the adjustment basis.

[0066] The input semantic block and output semantic block of each preliminary reasoning node are recorded in detail to form a node information table, which contains node identification, input semantic block identification, output semantic block identification, etc., to facilitate subsequent analysis of the dependency relationship between nodes.

[0067] Step S123: Analyze the dependency relationship between the preliminary reasoning nodes. If the input semantic block of one preliminary reasoning node is the output semantic block of another preliminary reasoning node, it is marked as having a direct dependency relationship, forming a reasoning node association hierarchy.

[0068] All constructed preliminary reasoning nodes are analyzed one by one to see if the input semantic block of each node comes from the output semantic block of other nodes. If such a situation exists, it means that there is a direct dependency relationship between the two nodes, i.e. the output of the former is the basis for the processing of the latter.

[0069] For example, the input semantic block of the preliminary reasoning node 2 mentions that “the equipment procurement of sub-project B1 may be delayed due to supplier reasons”, which needs to be based on the “specific progress of the equipment procurement link of sub-project B1” output by the preliminary reasoning node 1. Therefore, the input semantic block of the preliminary reasoning node 2 contains part of the output semantic block of the preliminary reasoning node 1, indicating that the preliminary reasoning node 2 has a direct dependency relationship with the preliminary reasoning node 1, the preliminary reasoning node 1 is the predecessor of the preliminary reasoning node 2, and the preliminary reasoning node 2 is the successor of the preliminary reasoning node 1.

[0070] For another example, if there is a preliminary reasoning node 5, whose input semantic block is “solution to the delay problem of sub-project B1” output by the preliminary reasoning node 2 and “list of participants in the cross-department coordination meeting” output by the preliminary reasoning node 3, it is marked that the preliminary reasoning node 5 has a direct dependency relationship with the preliminary reasoning node 2 and the preliminary reasoning node 3, the preliminary reasoning node 2 and the preliminary reasoning node 3 are the predecessors of the preliminary reasoning node 5, and the preliminary reasoning node 5 is the successor of them.

[0071] According to these direct dependency relationships, the preliminary reasoning nodes are divided into different association levels. The preliminary reasoning nodes without predecessors are in the first level, the nodes with the first level nodes as predecessors are in the second level, the nodes with the second level nodes as predecessors are in the third level, and so on, forming the association levels of the reasoning nodes.

[0072] Step S1231: Extract the input semantic block identifier and the output semantic block identifier of each preliminary reasoning node, and establish a target mapping table.

[0073] A unique identifier is assigned to each input semantic block and output semantic block of each preliminary reasoning node, which is composed of node number and semantic block type to ensure uniqueness. For example, the input semantic block of the preliminary reasoning node 1 includes a first-level core semantic block and a second-level association semantic block, which are identified as “Node1-In-Core1” and “Node1-In-Assoc1” respectively, and the output semantic block is identified as “Node1-Out-Result1”; the input semantic block of the preliminary reasoning node 2 is identified as “Node2-In-Core2”, and the output semantic block is identified as “Node2-Out-Result2”; the input semantic block of the preliminary reasoning node 3 is identified as “Node3-In-Core3” and “Node3-In-Assoc3”, and the output semantic block is identified as “Node3-Out-Result3”.

[0074] These identifiers are arranged in order according to the nodes to establish a target mapping table. Each row of the target mapping table corresponds to a preliminary reasoning node, including node number, input semantic block identifier list, output semantic block identifier, etc. columns, clearly showing the input and output semantic block situation of each node.

[0075] Step S1232: Traversing the target mapping table to find the pair of preliminary reasoning nodes whose output semantic block identifier is consistent with the input semantic block identifier of other preliminary reasoning nodes.

[0076] Traverse the target mapping table in the order of node number, and compare the output semantic block identifier of each preliminary reasoning node with the input semantic block identifier of all other preliminary reasoning nodes. For example, when traversing the preliminary reasoning node 1, check whether the output semantic block identifier "Node1-Out-Result1" of the preliminary reasoning node 1 appears in the input semantic block identifier list of other nodes.

[0077] It is found through comparison that the input semantic block identifier "Node2-In-Core2" of the preliminary reasoning node 2 contains information related to "Node1-Out-Result1", that is, the input semantic block of the preliminary reasoning node 2 depends on the output semantic block of the preliminary reasoning node 1, so it is determined that the preliminary reasoning node 1 and the preliminary reasoning node 2 constitute a pair of preliminary reasoning nodes.

[0078] Similarly, if the input semantic block identifier list of the preliminary reasoning node 5 contains "Node2-Out-Result2" and "Node3-Out-Result3", it is determined that the preliminary reasoning node 2 and the preliminary reasoning node 5, and the preliminary reasoning node 3 and the preliminary reasoning node 5 constitute a pair of preliminary reasoning nodes, respectively.

[0079] Step S1233: Mark the found pair of preliminary reasoning nodes as having a direct dependency relationship, the former as the predecessor preliminary reasoning node of the latter, and the latter as the successor preliminary reasoning node of the former.

[0080] For the found pair of preliminary reasoning nodes, add a dependency relationship mark in the target mapping table. For example, for the pair of nodes of the preliminary reasoning node 1 and the preliminary reasoning node 2, mark the preliminary reasoning node 1 as the predecessor preliminary reasoning node of the preliminary reasoning node 2, and the preliminary reasoning node 2 as the successor preliminary reasoning node of the preliminary reasoning node 1; for the pair of nodes of the preliminary reasoning node 2 and the preliminary reasoning node 5, mark the preliminary reasoning node 2 as the predecessor preliminary reasoning node of the preliminary reasoning node 5, and the preliminary reasoning node 5 as the successor preliminary reasoning node of the preliminary reasoning node 2; for the pair of nodes of the preliminary reasoning node 3 and the preliminary reasoning node 5, mark the preliminary reasoning node 3 as the predecessor preliminary reasoning node of the preliminary reasoning node 5, and the preliminary reasoning node 5 as the successor preliminary reasoning node of the preliminary reasoning node 3.

[0081] Step S1234: Count the number of predecessor preliminary reasoning nodes and the number of successor preliminary reasoning nodes of each preliminary reasoning node, and calculate the dependency degree of the preliminary reasoning node, which is the ratio of the number of predecessor preliminary reasoning nodes to the number of successor preliminary reasoning nodes.

[0082] The number of front nodes and the number of back nodes of each preliminary reasoning node are counted one by one. For example, the preliminary reasoning node 1 has no front node, i.e. the number of front nodes is 0, and the back node is the preliminary reasoning node 2, i.e. the number of back nodes is 1, so its dependency is the ratio of 0 to 1, i.e. 0; the front node of the preliminary reasoning node 2 is the preliminary reasoning node 1, the number of front nodes is 1, and the back node is the preliminary reasoning node 5, the number of back nodes is 1, and its dependency is the ratio of 1 to 1, i.e. 1; the preliminary reasoning node 3 has no front node, the number of front nodes is 0, and the back node is the preliminary reasoning node 5, the number of back nodes is 1, and the dependency is the ratio of 0 to 1, i.e. 0; the front node of the preliminary reasoning node 5 is the preliminary reasoning node 2 and the preliminary reasoning node 3, the number of front nodes is 2, and if there is no back node, the number of back nodes is 0, and since the denominator cannot be 0, the dependency can be marked as a special value such as "high" at this time, indicating that the node depends on multiple front nodes and has no subsequent node.

[0083] By calculating the dependency, the dependency of each node in the reasoning process can be intuitively known. The node with a dependency of 0 is usually the starting node of the reasoning process, and the node with a high dependency is in the middle or later stage of the reasoning process.

[0084] Step S1235: According to the dependency, the preliminary reasoning nodes are divided into basic preliminary reasoning nodes, intermediate preliminary reasoning nodes and terminal preliminary reasoning nodes. The basic preliminary reasoning node has no front preliminary reasoning node, and the terminal preliminary reasoning node has no back preliminary reasoning node.

[0085] According to the calculated dependency and the number of front and back nodes, the preliminary reasoning nodes are classified. The basic preliminary reasoning node refers to a node without a front node. These nodes can be independently started and do not need to depend on the output results of other nodes, such as the preliminary reasoning node 1 and the preliminary reasoning node 3, which have a front node number of 0 and belong to the basic preliminary reasoning node.

[0086] The intermediate preliminary reasoning node refers to a node that has both front nodes and back nodes. They play a role in the reasoning process, such as the preliminary reasoning node 2, which has a front node of the preliminary reasoning node 1 and a back node of the preliminary reasoning node 5, and belongs to the intermediate preliminary reasoning node.

[0087] The terminal preliminary reasoning node refers to a node without a back node. The output result of these nodes is one of the final achievements of the reasoning process, such as the preliminary reasoning node 5, which belongs to the terminal preliminary reasoning node if it has no back node.

[0088] Step S1236: Arrange the preliminary reasoning nodes in a hierarchical order according to their execution order. The base preliminary reasoning node is in the first layer, the post-node of the base preliminary reasoning node is in the second layer, and so on, forming a reasoning node association hierarchy.

[0089] According to the dependency relationship and execution order between the preliminary reasoning nodes, the nodes are arranged in a hierarchical order. The base preliminary reasoning node is arranged in the first layer as the starting point of the reasoning process; the post-node of the base preliminary reasoning node, i.e. the node directly dependent on the base node, is arranged in the second layer; the post-node of the second layer node is arranged in the third layer, and so on.

[0090] For example, preliminary reasoning node 1 and preliminary reasoning node 3 are base preliminary reasoning nodes, in the first layer; preliminary reasoning node 2 is the post-node of preliminary reasoning node 1, in the second layer; preliminary reasoning node 5 is the post-node of preliminary reasoning node 2 and preliminary reasoning node 3, in the third layer. The above hierarchical arrangement shows the execution order and association relationship of the reasoning nodes.

[0091] Step S1237: Add an intra-layer priority to each reasoning node associated with the preliminary reasoning node in the hierarchy, which is determined based on the processing complexity of the preliminary reasoning node.

[0092] The preliminary reasoning nodes in the same hierarchy are arranged according to their processing complexity. The processing complexity can be evaluated according to the number of node input semantic blocks, the length of semantic blocks, the logical complexity of the required processing, and other factors. The higher the processing complexity of the node, the higher the intra-layer priority, which will be executed preferentially.

[0093] For example, preliminary reasoning node 1 and preliminary reasoning node 3 in the first layer, if the semantic block to be processed by preliminary reasoning node 1 contains more sub-item information and the processing logic is more complex, a higher intra-layer priority is assigned to it, and the priority of preliminary reasoning node 3 is relatively low. When scheduling, the processing flow of preliminary reasoning node 1 can be started preferentially.

[0094] Step S1238: Construct a reasoning node association hierarchy graph, distinguish different levels with different colors, use arrows to represent direct dependency relationship, and mark the number of preliminary reasoning nodes.

[0095] According to the reasoning node association hierarchy and dependency relationship, a visual reasoning node association hierarchy graph is constructed. In the graph, different levels are represented by different colors, such as blue for the first layer, green for the second layer, yellow for the third layer, etc. Nodes with direct dependency relationship are connected by lines with arrows, and the arrows point from the pre-node to the post-node, clearly showing the dependency direction.

[0096] Next to each node, the number of its predecessor nodes and successor nodes are marked, such as the preliminary inference node 1 next to the mark "Predecessor: 0, Successor: 1", the preliminary inference node 2 next to the mark "Predecessor: 1, Successor: 1", and the preliminary inference node 5 next to the mark "Predecessor: 2, Successor: 0". Through the above inference node association hierarchy atlas, the structure and dependency of the entire inference node can be intuitively known.

[0097] Step S124: determining the execution order of the preliminary inference nodes according to the direct dependency relationship, adding an inference order constraint mark, the inference order constraint mark including mandatory predecessor preliminary inference nodes and optional predecessor preliminary inference nodes.

[0098] According to the direct dependency relationship between the preliminary inference nodes, the execution order of the nodes is clear. For nodes with direct dependency relationship, the predecessor nodes must be executed first, and then the successor nodes, which is a mandatory order constraint. At the same time, for some nodes, there may be optional predecessor nodes, that is, the output results of these predecessor nodes can be used as a reference, but they are not mandatory, and the successor nodes can be executed without the output of these predecessor nodes, but the processing result may not be accurate enough.

[0099] For example, the mandatory predecessor node of the preliminary inference node 2 is the preliminary inference node 1, which must be executed after the preliminary inference node 1 is completed, so the inference order constraint mark "mandatory predecessor: Node1" is added to the preliminary inference node 2. If the preliminary inference node 5 can refer to the output result of the preliminary inference node 4 in addition to the mandatory predecessor nodes of the preliminary inference node 2 and the preliminary inference node 3, but can also be executed without the output of the preliminary inference node 4, then the mark "mandatory predecessor: Node2, Node3; optional predecessor: Node4" is added to it.

[0100] These inference order constraint marks ensure that the inference nodes are executed in the correct order, ensuring the accuracy of the inference result.

[0101] Step S125: integrating the inference node association hierarchy and the inference order constraint to construct an inference node dependency atlas containing the positions of the preliminary inference nodes, connection lines and inference order constraint marks.

[0102] Integrating the inference node association hierarchy and the inference order constraint mark together forms a complete inference node dependency atlas. In the atlas, not only the hierarchical position of the nodes and the connection lines between the nodes (i.e. direct dependency relationship) are shown, but also the inference order constraint mark is marked next to each node.

[0103] For example, in the inference node dependency graph, the position of the preliminary inference node 2 is in the second layer, connected with the preliminary inference node 1 in the first layer through an arrow, and labeled with “mandatory prerequisite: Node1” beside it; the preliminary inference node 5 is in the third layer, connected with the preliminary inference node 2 in the second layer and the preliminary inference node 3 in the first layer through arrows respectively, and labeled with “mandatory prerequisite: Node2, Node3; optional prerequisite: Node4” beside it.

[0104] Step S126: Analyzing the operation node load state, combining the operation intensity evaluation of the preliminary inference node, determining the minimum value and the maximum value of the operation resource required by each preliminary inference node as the dynamic threshold of the operation resource.

[0105] The number of active processes and the length of the task waiting queue of each operation node in the operation node load state are analyzed to evaluate the current load of each node. At the same time, the operation intensity of each preliminary inference node is evaluated, which is related to factors such as the complexity of the semantic block processed by the node and the data volume. According to the load condition and the operation intensity, the demand range of each preliminary inference node on the operation resource, i.e. the minimum value and the maximum value, is determined.

[0106] For example, the current load of operation node A is low, and the current load of operation node B is high. The operation intensity of the preliminary inference node 1 is high, which requires more operation resources. Combining the load condition of the operation node, the minimum value of the operation resource of the preliminary inference node 1 is determined as the resource amount that can meet the basic processing demand, and the maximum value is the maximum resource amount that can be allocated without affecting the operation of other nodes.

[0107] Step S1261: Analyzing the number of active processes and the length of the task waiting queue in the operation node load state, calculating the current load rate of each operation node, which is the ratio of the number of active processes to the maximum process capacity.

[0108] For each operation node, the number of active processes and the maximum process capacity are extracted from the operation node load state. The maximum process capacity is the maximum number of processes that can be run simultaneously by the operation node. The current load rate is obtained by dividing the number of active processes by the maximum process capacity.

[0109] For example, the number of active processes of operation node A is a certain number, and the maximum process capacity is another number, and the current load rate is the ratio of the two; the number of active processes and the maximum process capacity of operation node B are different, and the calculated current load rate is also different. By calculating the current load rate, the busy degree of each operation node can be quantitatively evaluated.

[0110] Step S1262: Dividing the operation nodes into multiple load level operation nodes according to the current load rate.

[0111] According to the size range of the current load rate, the operation nodes are divided into different load levels. For example, the operation nodes with a current load rate lower than a certain proportion are low-load level operation nodes; the operation nodes with a current load rate between the certain proportion and another higher proportion are medium-load level operation nodes; and the operation nodes with a current load rate higher than the higher proportion are high-load level operation nodes.

[0112] The above division helps to match the operation nodes of appropriate load levels to the preliminary reasoning nodes according to the operation requirements of the preliminary reasoning nodes, so as to improve the resource utilization efficiency and the reasoning speed.

[0113] Step S1263: The operation intensity of each preliminary reasoning node is evaluated based on the number of semantic blocks processed by the preliminary reasoning node, the complexity of each semantic block (such as the amount of vocabulary, the complexity of logical relationship, etc.), and the historical processing time required for processing similar semantic blocks, and the operation nodes of a corresponding load level are matched to each preliminary reasoning node according to the operation intensity, to obtain a matching result.

[0114] When evaluating the operation intensity of the preliminary reasoning node, the number of semantic blocks processed by the preliminary reasoning node, the complexity of each semantic block (such as the amount of vocabulary, the complexity of logical relationship, etc.), and the historical processing time required for processing similar semantic blocks are comprehensively considered. The nodes with high operation intensity need to be matched with operation nodes of low load level to ensure that there are enough resources for processing; the nodes with low operation intensity can be matched with operation nodes of high load level.

[0115] For example, the preliminary reasoning node 1 processes a large number of semantic blocks with high complexity, and has a long historical processing time, so the operation intensity is high, and the operation node of low load level is matched; the preliminary reasoning node 3 processes relatively simple semantic blocks, and the operation intensity is low, so the operation node of medium load level is matched, and the corresponding matching result is obtained.

[0116] Step S1264: The number of operation nodes required by each preliminary reasoning node is determined as the operation resource reference value according to the matching result.

[0117] According to the load level of the operation node matched to each preliminary reasoning node in the matching result, and in combination with the processing capacity of the operation node of the level, the number of operation nodes required, i.e., the operation resource reference value, is determined. The processing capacity refers to the amount of tasks or data that can be processed per unit time.

[0118] For example, the operation node of low load level matched to the preliminary reasoning node 1 has strong processing capacity, and according to the operation intensity, a certain number of operation nodes of the level are determined to be required; the operation node of medium load level matched to the preliminary reasoning node 3 has medium processing capacity, and another number of operation nodes of the level are determined to be required, which are the respective operation resource reference values.

[0119] Step S1265: Based on the load fluctuation rule of the operation node, an operation resource maximum value is obtained by adding a fluctuation buffer to the operation resource reference value, and an operation resource minimum value is obtained by reducing unnecessary redundancy.

[0120] The load fluctuation of the operation node in the historical time period is analyzed, and the load fluctuation rule is summarized, such as the load change in different time periods in a day, the influence of different task types on the load, etc. According to the fluctuation rule, a fluctuation buffer is added to the operation resource reference value to cope with possible load peaks. This value is the operation resource maximum value. At the same time, unnecessary resource redundancy is reduced to obtain the operation resource minimum value that can meet the basic processing needs of the node.

[0121] For example, the operation node has large load fluctuation in the morning period. A certain fluctuation buffer is added to the operation resource reference value of the preliminary reasoning node 1 as the maximum value. The unnecessary redundant part of the reference value is removed to obtain the minimum value, so as to ensure that the resources are not wasted when the load is low.

[0122] Step S1266: Verify whether the operation resource minimum value can meet the minimum operation requirement of the preliminary reasoning node. If not, it is adjusted to the minimum requirement value.

[0123] The determined operation resource minimum value is compared with the minimum operation requirement of the preliminary reasoning node. The minimum operation requirement of the preliminary reasoning node refers to the minimum operation resource amount required for the node to complete the basic processing task. This requirement is determined based on the most basic processing requirement of the semantic block processed by the node, such as the resource required for processing the most core semantic information.

[0124] For example, the minimum operation requirement of the preliminary reasoning node 1 is the resource amount required for processing the most core sub-project progress check task in its input semantic block. The previously determined operation resource minimum value is compared with the requirement. If the minimum value is less than the minimum operation requirement, it means that the current minimum value cannot meet the basic processing requirement of the node, and it needs to be adjusted to the minimum operation requirement value. If the minimum value is greater than or equal to the minimum operation requirement, the minimum value remains unchanged.

[0125] Through the above verification, it is ensured that the operation resource minimum value of each preliminary reasoning node can meet its most basic operation requirement, avoiding the situation that the reasoning task cannot be completed due to insufficient resources.

[0126] Step S1267: The operation resource minimum value and the operation resource maximum value of each preliminary reasoning node are arranged as operation resource dynamic thresholds, and are recorded in the resource elasticity requirement list in the order of the preliminary reasoning nodes.

[0127] After the determination and verification of the minimum and maximum values of the operation resources of each preliminary reasoning node are completed, these values are sorted in the order of the numbers of the preliminary reasoning nodes. The dynamic threshold of the operation resources of each node is composed of its corresponding minimum value and maximum value, for example, the dynamic threshold of the operation resources of the preliminary reasoning node 1 is "minimum value 1-maximum value 1", the dynamic threshold of the operation resources of the preliminary reasoning node 2 is "minimum value 2-maximum value 2", and so on.

[0128] These dynamic thresholds are sequentially recorded in the resource elasticity demand list, and the identification of each dynamic threshold corresponding to the preliminary reasoning node is marked in the list, so that subsequent resource allocation can be quickly queried and applied.

[0129] Step S127: Analyzing the cache resource occupation information, combining the cache access demand of the preliminary reasoning node, determining the lower limit and upper limit of the cache space available to each preliminary reasoning node as the cache resource allocation range.

[0130] The used storage space proportion and data access heat of each level of cache in the cache resource occupation information are analyzed to know the current use of the cache resources. At the same time, the cache access demand of each preliminary reasoning node in the processing process is analyzed, including the data type, data volume and access frequency that need to be cached. According to these information, the cache space range available to each preliminary reasoning node, i.e. the lower limit and the upper limit, is determined.

[0131] For example, the used storage space proportion of the first-level cache is low, and the data access heat in it is high, mainly storing commonly used basic data. The preliminary reasoning node 1 needs to frequently access the basic information data of sub-projects in the processing process, and the access demand of the first-level cache is large, so the lower limit of the cache space determined for it is the minimum space that can store these basic information, and the upper limit is the maximum space that can be allocated without affecting the use of the first-level cache by other nodes.

[0132] Step S1271: Analyzing the data type and data volume that need to be cached in the processing process of the preliminary reasoning node, evaluating the cache access frequency and cache data retention time.

[0133] For each preliminary reasoning node, the data type that needs to be temporarily stored or frequently accessed in the process of processing the input semantic block and generating the output semantic block is analyzed, such as the progress data of sub-projects, personnel information data, historical interaction record data, etc. At the same time, the data volume of these data is estimated, the access frequency of these data, i.e. the number of accesses per unit time, and the time interval from storage to being cleared, i.e. the retention time of the data in the cache, are evaluated.

[0134] For example, the preliminary reasoning node 3 needs to access the job information and schedule data of key personnel of each department when determining the participants of the cross-department coordination meeting. The data volume of these data is medium, the access frequency is high, and the data needs to be retained for a long time after the meeting is determined for subsequent queries. Therefore, the cache access requirement of the preliminary reasoning node 3 has the characteristics of fixed data type, high access frequency, and long retention time.

[0135] Step S1272: Calculate the remaining available storage space of each level of cache according to the used storage space ratio in the cache resource occupation information.

[0136] The total storage space and the used storage space ratio of each level of cache are extracted from the cache resource occupation information. The remaining available storage space of each level of cache is calculated by multiplying the total storage space by (1-used storage space ratio).

[0137] For example, the total storage space of the first level of cache is a certain capacity, and the used storage space ratio is a certain percentage. The remaining available storage space is the total storage space multiplied by (1-the percentage). The total storage space and the used storage space ratio of the second level of cache are different, and the calculated remaining available storage space is also different.

[0138] Step S1273: Allocate the corresponding cache level to each preliminary reasoning node according to the remaining available storage space of each level of cache and the cache access requirement of the preliminary reasoning node.

[0139] According to the cache access requirement of the preliminary reasoning node, such as data access frequency, data retention time, data volume, and the characteristics of each level of cache, such as fast access speed but small capacity of the first level of cache and large capacity but relatively slow access speed of the second level of cache, an appropriate cache level is allocated to each node.

[0140] Nodes with high access frequency and the need for fast response are preferentially allocated to the first level of cache. Nodes with large data volume and relatively low access frequency can be allocated to the second level of cache. For example, the preliminary reasoning node 1 needs to frequently access the sub-project progress data and has high requirements on access speed, so it is allocated to the first level of cache. The preliminary reasoning node 4 processes budget adjustment data with large volume and relatively low access frequency, so it is allocated to the second level of cache.

[0141] Step S1274: Determine the lower and upper limits of the cache space of each preliminary reasoning node according to the remaining available storage space of the allocated cache level and the cache access requirement of the preliminary reasoning node.

[0142] After the cache level is allocated to the preliminary reasoning node, the lower limit and upper limit of the cache space are determined according to the remaining available storage space of the level and the data size in the cache access demand of the node. The lower limit is the minimum space that can meet the basic cache demand of the node, that is, the space required to store the most critical and most frequently accessed data; the upper limit is the maximum space that can be allocated to the node without affecting the storage of other data in the cache of the level.

[0143] For example, the preliminary reasoning node 1 allocated to the level 1 cache, the data size in the cache access demand is a certain size, and the remaining available storage space of the level 1 cache is a certain capacity. The lower limit of the cache space determined for it is a space slightly larger than the data size, and the upper limit is a space not exceeding a certain proportion of the remaining available storage space of the level 1 cache.

[0144] Step S1275: Verify whether the lower limit of the cache space can meet the minimum cache demand of the preliminary reasoning node. If not, adjust to the minimum cache demand value.

[0145] The determined lower limit of the cache space is compared with the minimum cache demand of the preliminary reasoning node. The minimum cache demand refers to the minimum cache space required for the node to normally process tasks, that is, the space to store the most core cache data.

[0146] If the lower limit of the cache space is less than the minimum cache demand, it means that the lower limit cannot meet the basic cache demand of the node, and needs to be adjusted to the minimum cache demand value. If it is greater than or equal to, it remains unchanged. For example, the minimum cache demand of the preliminary reasoning node 2 is the space to store the supplier information data. If the previously determined lower limit of the cache space is less than this space, it is adjusted to the size of this space.

[0147] Step S1276: The lower limit and upper limit of the cache space of each preliminary reasoning node are arranged as a cache resource allocation range and recorded in the resource elasticity demand list in the order of the preliminary reasoning nodes.

[0148] The lower limit and upper limit of the cache space of each preliminary reasoning node are arranged in the order of the node numbers to form a cache resource allocation range, such as the cache resource allocation range of the preliminary reasoning node 2 being "lower limit 2-upper limit 2", which is recorded in the resource elasticity demand list together with the operation resource dynamic threshold value. The resource demand information of each node corresponds to each other, facilitating the unified allocation and management of resources in the subsequent.

[0149] Step S128: The operation resource dynamic threshold value and the cache resource allocation range are arranged in the order of the preliminary reasoning nodes to generate the resource elasticity demand list, and the corresponding relationship of the preliminary reasoning nodes is established with the reasoning node dependency graph.

[0150] The operation resource dynamic threshold and the cache resource allocation range of each preliminary reasoning node are arranged in order of node number to form a complete resource elasticity demand list. The list contains node identification, operation resource dynamic threshold, cache resource allocation range, and the like.

[0151] Meanwhile, the index information of each preliminary reasoning node in the resource elasticity demand list, such as the row number in the list, is marked beside the node in the reasoning node dependency graph to establish a correspondence between the two. Through the above correspondence, the resource demand information of the node can be quickly found when viewing the reasoning node dependency graph.

[0152] Step S130: performing reasoning link optimization processing based on the reasoning node dependency graph and the resource elasticity demand list to obtain a reasoning acceleration execution scheme, wherein the reasoning acceleration execution scheme contains reasoning node parallel scheduling rules and resource pre-allocation strategies.

[0153] The reasoning link is optimized by combining the associated level, dependency relationship, and reasoning order constraint of the node in the reasoning node dependency graph, and the operation resource dynamic threshold and cache resource allocation range of each node in the resource elasticity demand list. The optimization goal is to maximize the reasoning speed and reduce the overall processing time under the premise of meeting the node dependency relationship and resource demand.

[0154] By identifying the node groups that can be executed in parallel, parallel scheduling rules are developed to determine the execution order and priority of these node groups. At the same time, according to the resource demand and scheduling rules, a resource pre-allocation strategy is developed to reserve the required operation resources and cache resources for the node groups in advance, thereby forming a reasoning acceleration execution scheme.

[0155] Step S131: analyzing the reasoning node associated level and reasoning order constraint in the reasoning node dependency graph, identifying the preliminary reasoning node groups without direct dependency relationship and belonging to the same associated level, and forming a parallel preliminary reasoning node group list.

[0156] The associated level and reasoning order constraint of each preliminary reasoning node are viewed by traversing the reasoning node dependency graph. For nodes belonging to the same associated level, it is checked whether there is a direct dependency relationship between them, i.e., whether there is an association of the front or rear node. If there is no direct dependency relationship between the nodes and they are in the same associated level, they are grouped into a group to form a preliminary reasoning node group that can be executed in parallel.

[0157] For example, the first layer of preliminary inference nodes 1 and 3, which are in the same associated level and have no direct dependency relationship between each other, i.e. the preliminary inference node 1 does not contain the preliminary inference node 3 in its pre-node and post-node, and vice versa, so they are classified into a parallel preliminary inference node group; if there are multiple nodes without direct dependency relationship in the second layer, they are also classified into corresponding parallel node groups. After sorting these node groups, a list of parallel preliminary inference node groups is formed.

[0158] Step S1311: traversing all preliminary inference nodes in the inference node dependency graph, extracting the pre-preliminary inference node list and the post-preliminary inference node list of each preliminary inference node, wherein the pre-preliminary inference node list contains mandatory pre-preliminary inference nodes and optional pre-preliminary inference nodes.

[0159] Each preliminary inference node in the inference node dependency graph is accessed one by one, and the pre-preliminary inference node list and the post-preliminary inference node list are extracted from the inference order constraint mark of the node. The pre-preliminary inference node list specifies the pre-nodes that the node must depend on and the pre-nodes that the node can optionally depend on, and the post-preliminary inference node list specifies the post-nodes that depend on the node.

[0160] For example, the pre-preliminary inference node list of the preliminary inference node 5 is “mandatory pre-preliminary inference nodes: Node2, Node3; optional pre-preliminary inference nodes: Node4”, and the post-preliminary inference node list is empty; the pre-preliminary inference node list of the preliminary inference node 2 is “mandatory pre-preliminary inference nodes: Node1”, and the post-preliminary inference node list is “Node5”. These lists are sorted and archived according to the node identifier.

[0161] Step S1312: for any two preliminary inference nodes, if their pre-preliminary inference node list and post-preliminary inference node list have no intersection, and there is no direct dependency relationship, and they belong to the same level in the inference node associated level, they are marked as parallel preliminary inference node pairs.

[0162] Any two preliminary inference nodes are selected, and their pre-preliminary inference node list and post-preliminary inference node list are compared to see if there are common nodes, i.e. if there is an intersection. If there is no intersection, it is further checked whether there is a direct dependency relationship between the two, i.e. whether one node is the pre-node or post-node of the other node. If there is neither an intersection nor a direct dependency relationship, and they belong to the same associated level, they are marked as parallel preliminary inference node pairs.

[0163] For example, the preliminary inference nodes 1 and 3 have no common nodes in their pre-node and post-node lists, and there is no direct dependency relationship, and they belong to the first layer, so they are marked as parallel preliminary inference node pairs.

[0164] Step S1313: Based on the expansion of the pair of parallel preliminary reasoning nodes, preliminary reasoning nodes that have no direct dependency relationship with each other and belong to the same associated level are grouped into the same group, forming a preliminary reasoning node group.

[0165] Based on the pair of parallel preliminary reasoning nodes, more nodes that meet the conditions are included in the group. For each node in the marked pair of parallel preliminary reasoning nodes, find other nodes that have no direct dependency relationship with them and belong to the same associated level. These nodes and the original node pair form a larger preliminary reasoning node group.

[0166] For example, based on the pair of parallel nodes consisting of preliminary reasoning node 1 and preliminary reasoning node 3, it is found that preliminary reasoning node 6 also has no direct dependency relationship with them and belongs to the first level, so preliminary reasoning node 6 is included to form a preliminary reasoning node group containing three nodes.

[0167] Step S1314: Check whether there is a mandatory preliminary reasoning node in each preliminary reasoning node group. If there is, split the preliminary reasoning node group to ensure that the reasoning order constraints of the preliminary reasoning nodes in the group are conflict-free.

[0168] For each preliminary reasoning node group, check whether the mandatory preliminary reasoning node of each node in the group is also in the group. If the mandatory preliminary reasoning node of a certain node is in the group, it means that there is a mandatory order constraint between the node and the preliminary reasoning node, and they cannot be executed in parallel. The node needs to be split out of the group to form a new node group.

[0169] For example, a preliminary reasoning node group contains node A, node B, and node C, where the mandatory preliminary reasoning node of node B is node A. There is a reasoning order constraint conflict in the group, and node B needs to be split out to form a group of node A and node C and a separate group of node B.

[0170] Step S1315: Match the split preliminary reasoning node group with the reasoning node associated level to verify whether all preliminary reasoning nodes in the preliminary reasoning node group belong to the same level. If it crosses levels, re-group according to the level to which it belongs.

[0171] Compare the split preliminary reasoning node group with the reasoning node associated level to check whether all nodes in the group belong to the same level. If there are nodes that cross levels, i.e., some nodes belong to the first level and some nodes belong to the second level, the nodes need to be re-grouped according to the level to which they belong to ensure that all nodes in each group belong to the same associated level.

[0172] For example, a preliminary inference node group contains node D belonging to the first level and node E belonging to the second level. This preliminary inference node group spans multiple levels, so node D needs to be assigned to the node group of the first level and node E needs to be assigned to the node group of the second level.

[0173] Step S1316: Assign a preliminary inference node group identifier to each qualified preliminary inference node group, and record the list of preliminary inference nodes and their associated levels within the preliminary inference node group.

[0174] Each preliminary inference node group that meets the requirements after inspection and splitting is assigned a unique identifier, such as "Group1" or "Group2". At the same time, the identifiers of all preliminary inference nodes contained in each group are recorded to form a list of preliminary inference nodes in the group, as well as the association level to which the group belongs, such as "Level 1" or "Level 2".

[0175] For example, the initial inference node list within Group 1 is "Node1, Node3, Node6", and its association level is "Level 1"; the list for Group 2 is "Node2", and its association level is "Level 2".

[0176] Step S1317: Calculate the total processing complexity of each preliminary inference node group, and divide the preliminary inference node groups whose differences between the total processing complexity are less than a set threshold into the same parallel batch.

[0177] The total processing complexity of each initial inference node group is calculated. This total processing complexity is the sum of the processing complexities of all nodes within the group. The processing complexity of a node is determined based on factors such as the complexity of its input semantic block and the amount of data. The total processing complexities are compared, and node groups with a difference less than a set threshold are grouped into the same parallel batch. This means that these groups can be executed in parallel within the same time period, preventing resource allocation imbalance due to excessive differences in processing complexity.

[0178] For example, if the total processing complexity of Group 1 is a certain value and the total processing complexity of Group 3 is another value, and the difference between the two is less than a set threshold, then they are divided into the same parallel batch.

[0179] Step S1318: The preliminary inference node group identifier, the list of preliminary inference nodes within the group, the associated level and the parallel batch number are sequentially associated and sorted according to the preliminary inference node group identifier to form a list of parallel preliminary inference node groups.

[0180] The identification of each preliminary reasoning node group, the list of nodes within the group, the associated hierarchy level, and the associated parallel batch number are associated in order, such as "Group1: Node1, Node3, Node6 - first level - batch 1", "Group3: Node7, Node8 - first level - batch 1", and the like. Then, the final list of parallel preliminary reasoning node groups is sorted in alphabetical or numerical order of the preliminary reasoning node group identification.

[0181] Step S132: Based on the processing complexity of each preliminary reasoning node in the parallel preliminary reasoning node group and the position in the reasoning node associated hierarchy, a parallel execution priority is set for each parallel preliminary reasoning node group.

[0182] For each parallel preliminary reasoning node group, the processing complexity of each preliminary reasoning node within the group, i.e. the sum of the processing complexity of the nodes within the group, and the position of the group in the reasoning node associated hierarchy are considered. The higher the level, the higher the priority may be. The higher the sum of the processing complexity and the higher the level of the group, the higher the parallel execution priority, which will be given priority in resource allocation and scheduling.

[0183] For example, Group1 at the first level has a high sum of processing complexity, and is set to have a "high" parallel execution priority; Group3 at the first level has a low sum of processing complexity, and has a "medium" priority; Group2 at the second level has a medium sum of processing complexity, and has a "medium" priority.

[0184] Step S133: In combination with the parallel execution priority of the parallel preliminary reasoning node group and the reasoning order constraint, the start time and the execution time range of each parallel preliminary reasoning node group are determined to form the scheduling time window of each node group.

[0185] According to the parallel execution priority of the parallel preliminary reasoning node group, the group with high priority is arranged to start first. At the same time, in combination with the reasoning order constraint of the nodes within the group, if there is a mandatory pre-node within the group and the pre-node belongs to other group, the start time of the group needs to be after the execution of the pre-node in the group.

[0186] According to the processing complexity, the historical processing time consumption, and the allocated resources of the nodes within the group, the execution time range of each group, i.e. the shortest execution time and the longest execution time, is estimated. The start time and the execution time range jointly constitute the scheduling time window of each node group.

[0187] For example, the parallel execution priority of Group 1 is "high", and there is no pre-constraint group, and the starting time is immediately after the inference process starts, and the execution time range is a certain period of time; the group of the mandatory pre-node of Group 2 is Group 1, so the starting time is after the execution of Group 1 is completed, and the execution time range is another period of time.

[0188] Step S134: Analyze the operation resource dynamic threshold and the cache resource allocation range in the resource elasticity demand list, and calculate the resource reservation quota required by each parallel preliminary inference node group according to the scheduling time window and processing demand of the parallel preliminary inference node group. The resource reservation quota does not exceed the upper limit of the operation resource dynamic threshold and the upper limit of the cache resource allocation range.

[0189] The operation resource dynamic threshold and the cache resource allocation range of each node in each parallel preliminary inference node group are extracted from the resource elasticity demand list, and the sum range of the operation resource dynamic threshold and the sum range of the cache resource allocation range in the group are obtained. In combination with the scheduling time window of the parallel preliminary inference node group, the resource occupation of other node groups in the time period is analyzed to avoid resource conflicts.

[0190] According to the processing demand of the nodes in the group, such as the data volume and the operation intensity, the resource reservation quota is determined within the sum range of the operation resource dynamic threshold and the sum range of the cache resource allocation range. For example, the sum of the upper limits of the operation resource dynamic thresholds of the nodes in Group 1 is a certain value, and the sum of the upper limits of the cache resource allocation ranges is another value. In combination with the processing demand, the calculated resource reservation quota needs to be lower than the two upper limit values to ensure that the resource reservation is reasonable and does not exceed the available range.

[0191] Step S135: According to the operation resource dynamic threshold and the cache resource allocation range, the critical state description of the resource usage is set. When the actual resource usage of the preliminary inference node reaches the critical state, a resource reallocation process is triggered, and a resource dynamic adjustment trigger condition is formed.

[0192] The upper limit and the lower limit of the operation resource dynamic threshold, and the upper limit and the lower limit of the cache resource allocation range are analyzed, and the critical state of the resource usage is set. For example, when the actual operation resource usage of the preliminary inference node reaches a certain proportion of the upper limit of the operation resource dynamic threshold, or the actual cache resource usage reaches a certain proportion of the upper limit of the cache resource allocation range, the critical state is reached.

[0193] When the critical state is reached, a resource reallocation process is triggered, such as reducing the resource allocation of other low-priority nodes, or calling backup resources, to form a resource dynamic adjustment trigger condition to ensure that the node group has sufficient resources during execution.

[0194] Step S136: According to the list of parallel preliminary reasoning node groups, the scheduling time window of each node group, the resource reservation quota, and the resource dynamic adjustment trigger condition, the reasoning node parallel scheduling rule is constructed to determine the parallel preliminary reasoning node group, the parallel execution priority, and the starting condition.

[0195] The information in each group in the list of parallel preliminary reasoning node groups is combined with the scheduling time window, the resource reservation quota, and the resource dynamic adjustment trigger condition to construct the reasoning node parallel scheduling rule. The rule specifies which node groups can be executed in parallel, their parallel execution priority order, and the conditions for starting execution, such as the completion of the execution of the preceding node group, the availability of resource reservation, etc.

[0196] For example, the rule specifies that Group1 and Group3 can be executed in parallel in Batch1, Group1 has higher priority than Group3, and the starting condition is that the resource reservation quota is met and no other high-priority node group occupies the resources; Group2 starts after the execution of Group1 is completed, and the starting condition is that the Group1 output result is generated and the resource reservation of itself is in place.

[0197] Step S137: Based on the starting conditions and sequence in the reasoning node parallel scheduling rule, and combined with the scheduling time window and resource reservation quota of the parallel preliminary reasoning node group, the resource pre-allocation strategy is constructed by locking the corresponding computing resources and cache space in advance according to the set advance time, so that the parallel preliminary reasoning node group can obtain the corresponding resource reservation quota when the scheduling time window starts.

[0198] According to the starting conditions and sequence of each node group in the reasoning node parallel scheduling rule, the resource pre-allocation time of each node group is determined. The required computing resources and cache space of the node group are locked in advance according to the set advance time, such as a certain period of time before the start of the scheduling time window, to ensure that the corresponding resource reservation quota can be obtained when the node group starts.

[0199] For example, the scheduling time window of Group1 starts at a certain time, and the set advance time is a period of time before that time. The resource reservation quota corresponding to the computing resources and cache space of Group1 is locked at the advance time point to prevent being occupied by other node groups.

[0200] Step S138: The reasoning node parallel scheduling rule and the resource pre-allocation strategy are integrated to generate the reasoning acceleration execution scheme containing the list of parallel preliminary reasoning node groups, the scheduling time window of each node group, the resource reservation quota, and the resource dynamic adjustment trigger condition.

[0201] The inference node parallel scheduling rule and the resource pre-allocation strategy are integrated together to form an inference acceleration execution scheme. The inference acceleration execution scheme completely contains the list of parallel preliminary inference node groups, the scheduling time window of each node group, the resource reservation quota, and the resource dynamic adjustment trigger condition, and the like, and provides specific guidance for the inference operation process regulation of the intelligent dialogue large model.

[0202] Step S140: The inference operation process of the intelligent dialogue large model is regulated according to the inference acceleration execution scheme, and an accelerated processing dialogue response sequence is generated.

[0203] According to the inference node parallel scheduling rule and the resource pre-allocation strategy in the inference acceleration execution scheme, the inference operation process of the intelligent dialogue large model is regulated. The node groups that can be parallelly executed are started simultaneously, and are sequentially promoted according to the priority and the scheduling time window, and the dynamic adjustment is triggered when the resource is insufficient, and finally the output results of the node groups are integrated to generate an accelerated processing dialogue response sequence.

[0204] Step S141: The inference acceleration execution scheme is parsed to extract the parallel preliminary inference node groups, the parallel execution priority and the starting condition in the inference node parallel scheduling rule, and the resource reservation parameters and the adjustment mechanism in the resource pre-allocation strategy.

[0205] The inference acceleration execution scheme is parsed to determine which preliminary inference node groups can be parallelly executed, how to sort the parallel execution priority of the groups, and what the starting condition is, and the resource reservation parameters in the resource pre-allocation strategy are extracted, such as the amount of reserved computing resources and the size of cache space, and the resource adjustment mechanism, such as the trigger condition and the adjustment method of adjustment.

[0206] For example, the parallel node groups are parsed as Group1 and Group3, the priority of Group1 is high, and the starting condition is that the resource locking is completed; the resource reservation parameters are the amount of computing resources and the size of cache space of Group1, and the adjustment mechanism is to reduce the resources of low-priority nodes when the resource usage reaches the critical state.

[0207] Step S142: The information of the parallel preliminary inference node groups is loaded to the inference process controller of the intelligent dialogue large model to configure the parallel execution logic of the preliminary inference nodes in the groups.

[0208] The information of the parallel preliminary inference node groups, such as the node identification in the groups and the logical relationship between the nodes, is loaded to the inference process controller of the intelligent dialogue large model. The parallel execution logic of the nodes in the groups, such as the simultaneous starting of the nodes and the synchronous summary of the results, is configured in the controller to ensure that the node groups can be executed in parallel.

[0209] For example, the information of Node1, Node3, Node6 in Group1 is loaded to the controller, and they are configured to start processing at the same time, and the intermediate results output by each of them are shared in real time so as to be integrated subsequently.

[0210] Step S143: setting the scheduling priority register of the inference flow controller according to the parallel execution priority.

[0211] The parallel execution priority of each group of parallel preliminary inference nodes is set to the scheduling priority register of the inference flow controller, and the register stores the group identifiers in order of priority. When scheduling, the controller schedules the groups with high priority first according to the order in the register.

[0212] For example, the scheduling priority register stores Group1, Group3, and Group2 in order, and the controller processes Group1 first, then Group3, and finally Group2 when scheduling.

[0213] Step S144: locking the computing resources and cache space in advance according to the advance time in the resource pre-allocation strategy through the resource management module.

[0214] The resource management module locks the required computing resources and cache space before the corresponding node group scheduling time window according to the advance time set in the resource pre-allocation strategy. During the locking process, these resources are marked as "reserved" to prevent them from being occupied by other tasks with lower priority.

[0215] For example, the resource management module locks the computing nodes and cache areas required by Group1 at the advance time point before Group1 is started, and marks them as "reserved".

[0216] Step S145: When the start condition of the preliminary inference node group is met, the inference flow controller applies to the resource management module to call the reserved resources to start the parallel inference operation of the preliminary inference node group.

[0217] The inference flow controller monitors in real time whether the start condition of each preliminary inference node group is met, such as whether the pre-node group is executed, whether the resources are reserved, and the like. When the start condition is met, the controller sends an application to the resource management module to call the reserved resources to start the parallel inference operation of the node group.

[0218] For example, when the start condition of Group1 is met, the inference flow controller applies to the resource management module to call the resources reserved for Group1, the resource management module releases the locked resources, and Group1 starts the parallel inference operation.

[0219] Step S146: During the inference operation process, the resource usage is monitored in real time, and when the monitored resource usage reaches the upper limit or lower limit of the operation resource dynamic threshold, the resource dynamic adjustment mechanism is triggered.

[0220] During the inference operation process of the preliminary inference node group, the resource management module monitors the usage of operation resources and cache resources in real time. When it is monitored that the resource usage reaches the upper limit or lower limit of the operation resource dynamic threshold, the resource dynamic adjustment mechanism is triggered, such as applying for additional resources when the upper limit is reached, and releasing part of the redundant resources when the lower limit is reached.

[0221] For example, during the execution of Group1, the operation resource usage reaches the upper limit of the operation resource dynamic threshold, triggering the resource dynamic adjustment mechanism, and the resource management module allocates additional operation resources for it from the standby resource pool.

[0222] Step S147: Collect the intermediate inference results output by each preliminary inference node group, integrate the results according to the correlation level of the inference node dependency graph, process the result dependency relationship between the preliminary inference nodes, and obtain the integrated intermediate results.

[0223] After each preliminary inference node group is executed, the intermediate inference results are output. These results are collected, and according to the correlation level of the inference node dependency graph, the results are integrated layer by layer from the first layer. For the results of nodes with dependency relationship, such as the result of the post-node depending on the result of the pre-node, the result of the pre-node is integrated as input into the result processing of the post-node, and the integrated intermediate result is obtained.

[0224] For example, the intermediate results of Group1 (each sub-project progress checking result) and the intermediate results of Group3 (other related processing results) are collected, the results of the first layer are integrated first, and then they are integrated as input into the results of Group2 (sub-project B1 delay solution), and finally the integrated intermediate results are obtained.

[0225] Step S148: Perform semantic coherence verification on the integrated intermediate results, arrange the intermediate results that pass the verification according to the dialogue sequence order, and generate the accelerated processing dialogue response sequence.

[0226] The integrated intermediate results are subjected to semantic coherence verification to check whether there are logical contradictions, inconsistent expressions, etc. between the results. For example, check whether the sub-project progress result and the delay solution are logically coherent, whether the list of participants and the conference discussion content match, etc.

[0227] The intermediate results that pass the verification are arranged according to the dialogue sequence order input by the user to form a dialogue response sequence that conforms to the user's reading habits and logical order, i.e. the accelerated processing dialogue response sequence.

[0228] Step S150: pushing the accelerated processed dialogue response sequence to the user interaction terminal to complete the intelligent dialogue output.

[0229] The generated accelerated processed dialogue response sequence is pushed to the user interaction terminal, such as a computer client, a mobile phone APP, etc. through a data transmission channel. After the user interaction terminal receives the sequence, it is displayed to the user in order to complete the intelligent dialogue output, so that the user can obtain the processing results of each item B in a timely manner.

[0230] Figure 2 A schematic diagram of exemplary hardware and software components of the inference acceleration optimization system 100 for the intelligent dialogue large model that can implement the idea of the present application is shown. For example, the processor 120 can be used in the inference acceleration optimization system 100 for the intelligent dialogue large model and used to execute the functions in the present application.

[0231] For example, the inference acceleration optimization system 100 for the intelligent dialogue large model can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, a ROM, or a RAM, or any combination thereof. Exemplarily, the inference acceleration optimization system 100 for the intelligent dialogue large model can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The methods of the present application can be implemented according to these program instructions. The inference acceleration optimization system 100 for the intelligent dialogue large model also includes an I / O interface 150 between the computer and other input / output devices.

[0232] In addition, the present application also provides a readable storage medium, wherein computer executable instructions are pre-set in the readable storage medium, and when a processor executes the computer executable instructions, the inference acceleration optimization method for the intelligent dialogue large model is implemented.

[0233] It should be noted that, in order to simplify the description of the present application and to help understand one or more embodiments of the present application, in the foregoing description of the embodiments of the present application, various features are sometimes combined into one embodiment, drawing or description thereof.

Claims

1. A method for accelerating inference in a large-scale intelligent dialogue model, characterized in that, The method includes: The system obtains the dialogue sequence to be inferred and the configuration information of the inference environment of the intelligent dialogue model. The dialogue sequence to be inferred includes the user's real-time input text and historical interaction statement chain. The configuration information of the inference environment includes the load status of the computing node and the cache resource usage information. The computing node is a hardware processing unit that performs inference calculations. The dialogue sequence to be reasoned and the configuration information of the reasoning environment are subjected to joint process deconstruction to obtain a reasoning node dependency graph and a resource elasticity requirement list. The reasoning node dependency graph includes the relationship hierarchy of reasoning nodes and the reasoning order constraints. The reasoning node is the logical unit corresponding to the semantic processing step in the reasoning process. The resource elasticity requirement list includes the dynamic threshold of computing resources and the allocation range of cache resources. Based on the inference node dependency graph and the resource elasticity requirement list, inference link optimization is performed to obtain an inference acceleration execution scheme, which includes inference node parallel scheduling rules and resource pre-allocation strategies. The reasoning operation process of the intelligent dialogue model is adjusted according to the reasoning acceleration execution scheme to generate an accelerated dialogue response sequence. The accelerated dialogue response sequence is pushed to the user interaction terminal to complete the intelligent dialogue output. The joint process deconstruction of the dialogue sequence to be reasoned and the configuration information of the reasoning environment yields a reasoning node dependency graph and a list of resource elasticity requirements, including: The dialogue sequence to be reasoned is semantically segmented to identify semantic blocks. The semantic blocks include core semantic blocks and related semantic blocks. The core semantic blocks correspond to the user's core needs, and the related semantic blocks are supplementary information to aid understanding. Preliminary inference nodes are constructed based on the logical relationships between the semantic blocks. The preliminary inference nodes are the initial construction form of the inference nodes. Each preliminary inference node corresponds to a semantic processing step, which is used to record the input semantic blocks and output semantic blocks of the preliminary inference nodes. Analyze the dependencies between the preliminary inference nodes. If the input semantic block of one preliminary inference node is the output semantic block of another preliminary inference node, it is marked as having a direct dependency relationship, forming an inference node association hierarchy. The execution order of the preliminary inference nodes is determined based on the direct dependencies, and inference order constraint markers are added. The inference order constraint markers include mandatory and optional preliminary inference nodes. The inference node association hierarchy is integrated with the inference order constraint to construct an inference node dependency graph that includes the initial inference node positions, connecting lines, and inference order constraint markers. The load status of the computing nodes is analyzed, and combined with the computing intensity assessment of the preliminary inference nodes, the minimum and maximum computing resources required for each preliminary inference node are determined as dynamic thresholds for computing resources. The cache resource usage information is analyzed, and combined with the cache access requirements of the initial inference nodes, the lower and upper limits of the cache space that each initial inference node can use are determined as the range for cache resource allocation. Arrange the dynamic thresholds of computing resources and the range of cache resource allocation according to the order of the initial inference nodes, generate a list of elastic resource requirements, and establish a correspondence between the initial inference nodes and the inference node dependency graph. The inference link optimization process, based on the inference node dependency graph and the resource elasticity requirement list, yields an inference acceleration execution scheme, including: Analyze the association hierarchy and inference order constraints of the inference node dependency graph, identify preliminary inference node groups that have no direct dependency relationship and belong to the same association level, and form a list of preliminary inference node groups that can be parallelized. Based on the processing complexity of each preliminary inference node in the parallelizable preliminary inference node group and its position in the inference node association hierarchy, a parallel execution priority is set for each parallelizable preliminary inference node group. Based on the parallel execution priority and inference order constraints of the parallelizable preliminary inference node groups, the start-up timing and execution duration range of each parallelizable preliminary inference node group are determined, forming the scheduling time window for each node group. The dynamic threshold of computing resources and the allocation range of cache resources in the resource elasticity demand list are analyzed. Based on the scheduling time window and processing requirements of the parallel preliminary inference node group, the resource reservation quota required for each parallel preliminary inference node group is calculated. The resource reservation quota shall not exceed the upper limit of the corresponding dynamic threshold of computing resources and the upper limit of the allocation range of cache resources. Based on the dynamic threshold of computing resources and the range of cache resource allocation, a critical state description of resource usage is set. When the actual resource usage of the initial inference node reaches the critical state, the resource reallocation process is triggered, forming a dynamic resource adjustment trigger condition. Based on the list of parallel preliminary inference node groups, the scheduling time window of each node group, the resource reservation quota and the resource dynamic adjustment triggering conditions, the parallel scheduling rules for inference nodes are constructed, and the parallel preliminary inference node groups, parallel execution priorities and start conditions are determined. Based on the start conditions and order in the parallel scheduling rules of the inference nodes, and combined with the scheduling time window and resource reservation quota of the parallel preliminary inference node group, a resource pre-allocation strategy is constructed. By locking the corresponding computing resources and cache space in advance according to the set advance time, the parallel preliminary inference node group can obtain the corresponding resource reservation quota when the scheduling time window starts. The parallel scheduling rules for inference nodes are integrated with the resource pre-allocation strategy to generate an inference acceleration execution scheme that includes a preliminary list of parallel inference node groups, a scheduling time window for each node group, a resource reservation quota, and a resource dynamic adjustment trigger condition.

2. The reasoning acceleration optimization method for large-scale intelligent dialogue models according to claim 1, characterized in that, The acquisition of the dialogue sequence to be reasoned and the inference environment configuration information of the intelligent dialogue large model includes: Receive the dialogue input data stream transmitted from the user interaction terminal, extract the real-time input text of the user in the order of timestamps, and mark the input order of the text segments; Access the historical interaction database of the intelligent dialogue model, retrieve the historical interaction records associated with the current user identifier, and organize them into a historical interaction statement chain according to the interaction time. The historical interaction statement chain includes the user's historical input content and the model's historical response content. Merge the real-time user input text with the historical interaction statement chain, add sequence delimiters and semantic association markers, and generate a dialogue sequence to be inferred; The current operating status of each computing node is collected by the computing node monitoring program and integrated into the computing node load status. The computing node is a hardware processing unit that performs inference calculations. The operating status includes the number of active processes and the length of the task waiting queue. Call the cache resource management interface to query the usage of each level of cache and integrate it into cache resource usage information, which includes the percentage of used storage space and data access frequency. Establish time-related markers between the dialogue sequence to be reasoned and the configuration information of the reasoning environment to form a set of related data.

3. The reasoning acceleration optimization method for large-scale intelligent dialogue models according to claim 1, characterized in that, The step of performing semantic hierarchical segmentation of the dialogue sequence to be reasoned and identifying semantic blocks therein includes: The semantic boundary detection algorithm is used to scan the dialogue sequence to be reasoned, and the semantic pause points and logical transition marks are identified as semantic block splitting boundaries. The dialogue sequence to be reasoned is divided into multiple semantic segments according to the semantic block splitting boundaries, and each semantic segment contains complete semantic information; The intention weight of each semantic segment is determined based on the frequency of occurrence and semantic importance of the core verbs and nouns in the semantic segments. Semantic segments with intention weights greater than a preset weight threshold are marked as core semantic blocks, and the rest are marked as related semantic blocks. The semantic distance between the associated semantic block and the core semantic block is calculated based on the number of co-occurring words and grammatical dependency relationships. The associated semantic blocks with a semantic distance less than a preset distance threshold are then bound to the corresponding core semantic blocks to form a semantic block group. Add hierarchical tags to each semantic block group, where the core semantic block is level one, directly related semantic blocks are level two, and indirectly related semantic blocks are level three; The output includes the semantic hierarchy splitting results, which include core semantic blocks, related semantic blocks, semantic block groups, and hierarchical tags.

4. The reasoning acceleration optimization method for large-scale intelligent dialogue models according to claim 1, characterized in that, The analysis of the dependencies between the preliminary inference nodes indicates that if the input semantic block of one preliminary inference node is the output semantic block of another preliminary inference node, a direct dependency relationship exists, forming an inference node association hierarchy, including: Extract the input semantic block identifier and output semantic block identifier of each preliminary inference node, and establish a target mapping table; Traverse the target mapping table to find pairs of preliminary inference nodes whose output semantic block identifiers match the input semantic block identifiers of other preliminary inference nodes; The found preliminary inference nodes are marked as having a direct dependency relationship, where the former is the preceding preliminary inference node of the latter, and the latter is the following preliminary inference node of the former. The number of preceding and subsequent preliminary inference nodes for each preliminary inference node is counted, and the dependency of the preliminary inference node is calculated. The dependency is the ratio of the number of preceding preliminary inference nodes to the number of subsequent preliminary inference nodes. Based on dependency, the initial inference nodes are divided into basic initial inference nodes, intermediate initial inference nodes and terminal initial inference nodes. Basic initial inference nodes have no preceding initial inference nodes, and terminal initial inference nodes have no following initial inference nodes. The nodes are arranged hierarchically according to the execution order of the initial inference nodes. The basic initial inference nodes are the first level, the subsequent initial inference nodes of the basic initial inference nodes are the second level, and so on, forming a hierarchical association of inference nodes. Add an intra-level priority to the initial inference nodes of the associated hierarchy for each inference node. The intra-level priority is determined based on the processing complexity of the initial inference nodes. Construct a hierarchical graph of inference node relationships, using different colors to distinguish different levels, arrows to represent direct dependencies, and marking the number of preceding and following preliminary inference nodes for each preliminary inference node.

5. The reasoning acceleration optimization method for large-scale intelligent dialogue models according to claim 1, characterized in that, The process of analyzing the load status of the computing nodes, combined with the computational intensity assessment of the initial inference nodes, determines the minimum and maximum values ​​of computing resources required for each initial inference node, serving as dynamic thresholds for computing resources, including: The number of active processes and the length of the task waiting queue in the load status of the computing nodes are analyzed, and the current load rate of each computing node is calculated. The current load rate is the ratio of the number of active processes to the maximum process capacity. The computing nodes are divided into multiple load levels based on the current load rate; The computational intensity of each preliminary inference node is evaluated based on the number of semantic blocks processed, complexity, and historical processing time. Then, a corresponding load level computational node is matched for each preliminary inference node according to the computational intensity to obtain the matching result. The number of computing nodes required for each initial inference node is determined based on the matching results, serving as a baseline value for computing resources. Based on the load fluctuation pattern of the computing nodes, the maximum computing resource value is obtained by adding a fluctuation buffer to the baseline computing resource value, and the minimum computing resource value is obtained by reducing unnecessary redundancy. Verify whether the minimum computing resource value can meet the minimum computing requirements of the initial inference node. If it does not meet the requirements, adjust it to the minimum requirement value. The minimum and maximum computing resource values ​​for each initial inference node are compiled into dynamic computing resource thresholds and recorded in the resource elasticity requirement list in the order of the initial inference nodes.

6. The reasoning acceleration optimization method applied to large-scale intelligent dialogue models according to claim 1, characterized in that, The analysis of the inference node dependency graph reveals the inference node association hierarchy and inference order constraints, identifying preliminary inference node groups with no direct dependencies but belonging to the same association hierarchy, forming a list of parallel preliminary inference node groups, including: Traverse all preliminary inference nodes in the inference node dependency graph, and extract the list of preceding preliminary inference nodes and the list of subsequent preliminary inference nodes for each preliminary inference node. The list of preceding preliminary inference nodes includes mandatory preceding preliminary inference nodes and optional preceding preliminary inference nodes. For any two preliminary inference nodes, if their preceding preliminary inference node lists and subsequent preliminary inference node lists have no intersection and no direct dependency relationship, and both belong to the same level in the inference node association hierarchy, then they are marked as a pair of preliminary inference nodes that can be parallelized. Based on the parallel preliminary inference node pairs, preliminary inference nodes that have no direct dependency on each other and belong to the same association level are grouped into the same group to form a preliminary inference node group. Check if there are any required preceding preliminary inference nodes in each preliminary inference node group. If so, split the preliminary inference node group to ensure that the inference order constraints of the preliminary inference nodes in the group do not conflict. The split preliminary inference node groups are matched with the inference node association levels to verify whether all preliminary inference nodes in the preliminary inference node group belong to the same level. If they are across levels, they are regrouped according to their respective levels. Assign a preliminary inference node group identifier to each qualified preliminary inference node group, and record the list of preliminary inference nodes and their associated levels within the preliminary inference node group. The total processing complexity of each initial inference node group is calculated, and the initial inference node groups whose differences between the total processing complexity are less than a set threshold are divided into the same parallel batch. The preliminary inference node group identifier, the list of preliminary inference nodes within the group, the associated hierarchy, and the parallel batch number are sequentially associated and sorted according to the preliminary inference node group identifier to form a list of parallel preliminary inference node groups.

7. The reasoning acceleration optimization method for large-scale intelligent dialogue models according to claim 1, characterized in that, The process of regulating the inference operation flow of the intelligent dialogue model according to the inference acceleration execution scheme to generate an accelerated dialogue response sequence includes: The inference acceleration execution scheme is analyzed, and the parallel preliminary inference node group, parallel execution priority and start conditions in the parallel scheduling rules of inference nodes, as well as the resource reservation parameters and adjustment mechanism in the resource pre-allocation strategy are extracted. Load the information of the parallel preliminary inference node group into the inference process controller of the intelligent dialogue large model, and configure the parallel execution logic of the preliminary inference nodes within the group. Set the scheduling priority register of the inference process controller according to the parallel execution priority; According to the advance time in the resource pre-allocation strategy, the computing resources and cache space are locked in advance through the resource management module; When the startup conditions of the initial inference node group are met, the inference process controller requests reserved resources from the resource management module to start the parallel inference operation of the initial inference node group. During the inference operation, the resource usage is monitored in real time. When the monitored resource usage reaches the upper or lower limit of the dynamic threshold of computing resources, the dynamic resource adjustment mechanism is triggered. Collect the intermediate inference results output by each preliminary inference node group, integrate the results according to the association level of the inference node dependency graph, process the result dependencies between the preliminary inference nodes, and obtain the integrated intermediate results. The integrated intermediate results are subjected to semantic coherence verification. The intermediate results that pass the verification are arranged in the order of the dialogue sequence to generate a dialogue response sequence after accelerated processing.

8. A reasoning acceleration and optimization system applied to large-scale intelligent dialogue models, characterized in that, The device includes a processor and a memory, the memory being connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the reasoning acceleration optimization method for large-scale intelligent dialogue models as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Dynamic tool selection and optimization system and method for large model external tool calling

    CN119166318A