AI demand retrieval platform based on large model and implementation method thereof

By building a reinforcement learning environment and DQN model, combining NLP units and container orchestration tools, the optimal task allocation strategy is generated, which solves the problems of low task allocation efficiency and poor resource management flexibility, and achieves efficient task scheduling and resource utilization.

CN120336284AInactive Publication Date: 2025-07-18JUNCHU TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510397877.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has low task allocation efficiency and poor resource management flexibility in complex demand scenarios, making it difficult to meet the needs of enterprise-level multimodal data processing.

Method used

The NLP unit analyzes user requirements to generate structured data, builds a reinforcement learning environment training DQN model, generates the optimal task allocation strategy, and optimizes resource quotas through containerized processing and resource management, combines container orchestration tools to perform task scheduling and deployment, and monitors load conditions in real time to optimize resource quotas.

Benefits of technology

It realizes efficient task allocation and resource management, improves system efficiency and resource utilization, reduces understanding errors and manual intervention, and optimizes the intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336284A_ABST
    Figure CN120336284A_ABST
Patent Text Reader

Abstract

The invention discloses an AI demand retrieval platform based on a large model and an implementation method thereof, and relates to the technical field of product research and development management, and the implementation method comprises the following steps: constructing a reinforcement learning environment based on a structured data format, carrying out interactive training on an initialized DQN model through the reinforcement learning environment, and generating an optimal task allocation strategy; based on the optimal task allocation strategy, containerization processing is performed on the tasks, resource quotas are generated through a resource management unit, and based on task containerization and the resource quotas, task scheduling and deployment are performed through a container arrangement tool, and a task deployment state is generated; and monitoring the load condition of the task deployment state in real time, triggering an elastic expansion mechanism to optimize the resource quota according to a preset threshold value of the load condition, and generating an optimal demand retrieval strategy. By constructing a reinforcement learning environment and combining DQN model training, an optimal task allocation strategy is generated, the scheduling intelligence level is improved, the system efficiency and the resource utilization rate are optimized, and the problems that a traditional method is low in efficiency and wastes resources are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of product R & D management, and particularly to an AI requirement retrieval platform based on a large model and its implementation method. Background Art

[0002] In recent years, artificial intelligence technology has promoted the development of requirement retrieval platforms based on large language models (LLMs). Traditional retrieval systems rely on keyword matching or simple semantic analysis and have deficiencies in processing complex requirements. To improve the accuracy and efficiency of text parsing, researchers have introduced deep learning models such as GPT and BERT based on the Transformer architecture. Combining multi-modal data processing technology, the new generation of platforms can support multiple data inputs such as text, images, and tables to meet enterprise-level requirements. However, in the face of increasingly complex scenarios, the existing technologies still have deficiencies in dynamic task allocation and resource optimization management.

[0003] Containerization and cloud computing provide strong support for requirement retrieval platforms. Through tools such as Kubernetes, tasks can be quickly deployed and elastically scaled to ensure stable operation under high concurrency. The application of reinforcement learning further improves the task scheduling and resource allocation efficiency. For example, the optimal policy is generated through the DQN model. However, the existing technologies still face problems such as low task allocation efficiency and poor resource management flexibility in complex requirement scenarios, and better solutions are urgently needed. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an AI requirement retrieval method based on a large model to solve the problems of low task allocation efficiency and insufficient resource management flexibility in the existing technologies.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides an AI requirement retrieval method based on a large model, which includes parsing user requirements through an NLP unit, extracting functional and non-functional requirement information and converting it into a structured data format; based on the structured data format, constructing a reinforcement learning environment, and interacting and training an initialized DQN model through the reinforcement learning environment to generate an optimal task allocation strategy; based on the optimal task allocation strategy, containerizing the tasks, and generating resource quotas through a resource management unit, and based on task containerization and resource quotas, performing task scheduling and deployment through a container orchestration tool to generate a task deployment status; monitoring the load condition of the task deployment status in real time, and triggering an elastic expansion mechanism to optimize the resource quota according to a preset threshold of the load condition to generate an optimal requirement retrieval strategy.

[0008] As a preferred solution of the large model-based AI requirement retrieval method of the present invention, wherein: the user requirement is parsed by the NLP unit, and the specific steps are as follows:

[0009] The user requirement is parsed through input processing, semantic analysis, entity recognition, relationship extraction, classification annotation, data generation, verification feedback, and an extensible interface to generate structured data.

[0010] As a preferred solution of the large model-based AI requirement retrieval method of the present invention, wherein: based on the structured data format, a reinforcement learning environment is constructed, and the specific steps are as follows:

[0011] The task information, analysis node status, and task dependency relationship of the structured data are parsed by an XML parser and converted into a task queue, a computing node resource table, and a task dependency graph;

[0012] Based on the task queue, the computing node resource table, and the task dependency graph, it is converted into a state space through data preprocessing and then into an action space through action design;

[0013] A data parsing method is used to extract the state space and the action space, generate quantization feature values reflecting the system state and action influence, and integrate them into a reward function by weighted summation;

[0014] By integrating the task queue, the computing node resource table, the task dependency graph, the state space, the action space, and the reward function, a reinforcement learning environment is constructed.

[0015] As a preferred solution of the large model-based AI requirement retrieval method of the present invention, wherein: the initialized DQN model is interactively trained through the reinforcement learning environment to generate an optimal task allocation strategy; the specific steps are as follows:

[0016] The weight parameters of the DQN model are assigned using the initialization method Xavier, the neural network structure is defined through the user requirement, the state space, and the action space, and hyperparameters are set to generate an initialized DQN model;

[0017] The initialized DQN model interacts with the reinforcement learning environment for training to generate interaction data and store it in the experience replay buffer;

[0018] Interaction data is randomly extracted from the experience replay buffer, and the Q-learning algorithm is used to update the parameters of the DQN model to generate an optimal task allocation strategy for the current state.

[0019] As a preferred solution of the AI requirement retrieval method based on large models according to the present invention, wherein: based on the optimal task allocation strategy, tasks are containerized, and resource quotas are generated through a resource management unit. The specific steps are as follows:

[0020] Based on the optimal task allocation strategy, define the running environment, dependencies, and resource requirements of each task, and generate a requirements list for the tasks;

[0021] Based on the requirements list of the tasks, query whether there is a corresponding container image in the image repository;

[0022] When it does not exist, then based on the requirements list of the tasks, create a Dockerfile to define the build rules of the container, and based on the build rules, use Docker tools to build a container image;

[0023] When it exists, verify whether the task can run normally in the container by starting the container, checking the task output, verifying the function, and monitoring resource usage;

[0024] When it can run normally, the container image has successfully achieved containerization of the task;

[0025] When it cannot run normally, gradually troubleshoot the problem by checking the logs, dependency integrity, port mapping, resource limits, data mounts, and base image compatibility until the containerization of the task is completed;

[0026] Based on the requirements list of the tasks, use the resource management unit to perform requirement analysis and resource allocation on the requirements list of the tasks to generate resource quotas.

[0027] As a preferred solution of the AI requirement retrieval method based on large models according to the present invention, wherein: based on task containerization and resource quotas, task scheduling and deployment are performed through a container orchestration tool to generate a task deployment status. The specific steps are as follows:

[0028] Based on task containerization, use the container orchestration tool to pull the container image, and generate target nodes through resource quotas;

[0029] Based on the target nodes, complete environment initialization by mounting data volumes, configuring network parameters, and loading dependency libraries, and create task instances;

[0030] Collect the running status information of the task instances in real time to generate a task deployment status.

[0031] As a preferred solution of the AI requirement retrieval method based on large models according to the present invention, wherein: the load condition of the task deployment status is monitored in real time, and according to the preset threshold of the load condition, an elastic scaling mechanism is triggered to optimize the resource quota to generate an optimal requirement retrieval strategy. The specific steps are as follows:

[0032] Define a preset threshold based on historical load trend prediction data;

[0033] Use the monitoring function built into the container orchestration tool to collect the load of the task deployment status, generate monitoring data, and compare the load of the monitoring data with the preset threshold;

[0034] When the load exceeds the upper limit threshold, trigger an expansion operation; otherwise, trigger a reduction operation;

[0035] Optimize the resource quota based on the elastic expansion mechanism and generate an optimal retrieval strategy.

[0036] In a second aspect, the present invention provides an AI demand retrieval platform based on a large model, including a demand parsing module, a strategy generation module, a containerization processing module, a task scheduling module, a load monitoring module, and an optimal demand retrieval strategy generation module; the demand parsing module is used to parse user demands through an NLP unit, extract functional and non-functional demand information and convert it into a structured data format; the strategy generation module is used to construct a reinforcement learning environment based on the structured data format, interactively train the initialized DQN model through the reinforcement learning environment, and generate an optimal task allocation strategy; the containerization processing module is used to containerize the task based on the optimal task allocation strategy and generate a resource quota through a resource management unit; the task scheduling module is used to perform task scheduling and deployment through a container orchestration tool based on task containerization and resource quota, and generate a task deployment status; the load monitoring module is used to monitor the load of the task deployment status in real time; the optimal demand retrieval strategy generation module is used to trigger the elastic expansion mechanism to optimize the resource quota according to the preset threshold of the load situation and generate an optimal demand retrieval strategy.

[0037] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, it implements any step of the AI demand retrieval method based on a large model as described in the first aspect of the present invention.

[0038] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, it implements any step of the AI demand retrieval method based on a large model as described in the first aspect of the present invention.

[0039] The beneficial effects of the present invention are as follows: By parsing the user's needs through the NLP unit and generating structured data, accurate parsing and standardized output are achieved, providing high-quality data support for subsequent task allocation and resource management, reducing understanding errors and manual intervention; By constructing a reinforcement learning environment and combining DQN model training, an optimal task allocation strategy is generated, improving the intelligent level of scheduling, optimizing system efficiency and resource utilization rate, and solving the problems of low efficiency and resource waste in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0041] Figure 1 It is a flowchart of the AI requirement retrieval method based on a large model in Embodiment 1.

[0042] Figure 2 It is a schematic diagram of the AI requirement retrieval platform based on a large model in Embodiment 1.

[0043] Figure 3 It is a flowchart of the containerization processing and resource quota generation of the AI requirement retrieval method based on a large model in Embodiment 1.

[0044] Figure 4 It is a flowchart of the task scheduling and load monitoring of the AI requirement retrieval method based on a large model in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification.

[0046] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those specifically described herein. Those skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0047] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.

[0048] Embodiment 1, refer to Figures 1 to 4, which is the first embodiment of the present invention. This embodiment provides an AI requirement retrieval method based on a large model, including the following steps:

[0049] S1. Parse the user requirements through the NLP unit, extract functional and non-functional requirement information, and convert it into a structured data format.

[0050] Parse the user requirements through input processing, semantic analysis, entity recognition, relationship extraction, classification annotation, data generation, verification feedback, and an extensible interface to generate structured data.

[0051] Furthermore, through steps such as input processing, semantic analysis, entity recognition, relationship extraction, classification annotation, data generation, verification feedback, and an extensible interface, the present invention comprehensively analyzes the user requirements, extracts functional and non-functional requirement information, and converts it into standardized structured data. This process not only ensures the accuracy and integrity of requirement analysis but also provides guarantee for the scalability of the system through a flexible interface design, laying a solid data foundation for subsequent task allocation and resource management.

[0052] S2. Based on the structured data format, construct a reinforcement learning environment, and interactively train the initialized DQN model through the reinforcement learning environment to generate an optimal task allocation strategy.

[0053] Parse the task information, analysis node status, and task dependency relationships of the structured data through an XML parser, and convert them into a task queue, a computing node resource table, and a task dependency graph;

[0054] The specific process is to load the structured data through an XML parser, and detailedly analyze the task information (including task attributes, priorities, resource requirements), the computing node status (such as resource distribution, load conditions), and the dependency relationships between tasks. Based on the parsing results, generate a task queue to clarify the execution order, construct a computing node resource table to track resource availability, and represent the task dependency relationships through a graph structure to form a task dependency graph. The whole process ensures clear task scheduling logic and reasonable resource allocation, and can visually display the task dependency relationships through a visualization tool, thus providing a reliable data foundation and decision support for the subsequent construction of the reinforcement learning environment.

[0055] Based on the task queue, the computing node resource table, and the task dependency graph, convert them into a state space through data preprocessing, and then convert the action space through action design;

[0056] The specific process is as follows. Based on the task queue, the computing node resource table, and the task dependency graph, through data preprocessing, the task attributes, node resource information, and task dependency relationships are transformed into a unified state space representation. At the same time, through action design, the action space for task allocation, resource adjustment, and scheduling operations is defined. Specifically, the state space integrates multi-dimensional features such as task priority, resource requirements, node resource distribution, and task dependency relationships, while the action space covers various operations such as task allocation, resource adjustment, and scheduling order adjustment. By constructing the state space and the action space, a clear input-output framework is provided for the subsequent training of the reinforcement learning environment, thus realizing the intelligent generation and optimization of the task scheduling strategy.

[0057] The data parsing method is used to extract the state space and the action space, generate quantization feature values reflecting the system state and the impact of actions, and integrate them into a reward function in the way of weighted summation. The expression is as follows:

[0058]

[0059] Among them, A * is the optimal task allocation strategy, Q is the action value, R is the resource utilization rate in the current state, R0 is the target resource utilization rate, C is the conflict penalty coefficient, and Conflict(D) is the conflict flag of the task dependency relationship.

[0060] The specific process is as follows. Through the data parsing method, feature extraction is performed on the state space and the action space to generate quantization feature values reflecting the system state and the impact of actions, such as task priority, resource utilization rate, task completion time, and energy consumption change. According to the optimization objectives of reinforcement learning (such as task efficiency, resource utilization rate, and energy consumption control), weights are assigned to each quantization feature value, and they are integrated into a reward function in the way of weighted summation. This process ensures that the reward function can comprehensively measure the impact of the system state and actions, guiding the intelligent agent to learn the optimal strategy in a dynamic environment and realizing the efficient optimization of task scheduling and resource management.

[0061] By integrating the task queue, the computing node resource table, the task dependency graph, the state space, the action space, and the reward function, a reinforcement learning environment is constructed.

[0062] The specific process is as follows. By integrating the task queue, the computing node resource table, and the task dependency graph, a state space containing task attributes, node resource information, and task dependency relationships is constructed; at the same time, an action space covering task allocation, resource adjustment, and scheduling order adjustment is designed, and a reward function is formulated based on indicators such as task completion time, resource utilization rate, and energy consumption. On this basis, the state space, the action space, and the reward function are integrated into a reinforcement learning environment, simulating the task scheduling process and dynamically updating the state and the reward, and using the reinforcement learning algorithm to train the intelligent agent to realize the generation and optimization of the optimal scheduling strategy.

[0063] Use the Xavier initialization method to assign values to the weight parameters of the DQN model. Define the neural network structure based on user requirements, the state space, and the action space, and set hyperparameters to generate an initialized DQN model.

[0064] Use the Xavier initialization method to assign values to the weight parameters of the DQN model, ensuring good gradient propagation characteristics in the initial stage of the network. Then, based on the characteristics of user requirements, the system state space, and the action space, design a suitable neural network architecture and set necessary hyperparameters (such as learning rate, discount factor, experience replay buffer size, etc.). At the same time, initialize the target network to stabilize the Q-value estimation. Finally, generate an initialized DQN model, which is ready for subsequent reinforcement learning training and task allocation strategy optimization. This process not only ensures the stability of the model in the initial stage of training but also lays a solid foundation for its subsequent effective learning.

[0065] Interact and train the initialized DQN model with the reinforcement learning environment to generate interaction data and store it in the experience replay buffer.

[0066] Furthermore, first initialize the environment including the state space, action space, and reward function, as well as the DQN model (main network and target network), and set the experience replay buffer. In each training step, use the ∈-greedy policy to select an action. After executing the action, obtain the new state and reward, and store the experience in the buffer. Subsequently, randomly sample a batch of data from the buffer, calculate the TD error, and update the weights of the main network. Synchronize the weights of the main network and the target network every certain number of steps to ensure the stability of training. In this way, a large amount of high-quality interaction data is generated, enabling the DQN model to gradually learn the optimal task scheduling strategy.

[0067] Randomly extract interaction data from the experience replay buffer and use the Q-learning algorithm to update the parameters of the DQN model to generate the optimal task allocation strategy for the current state. The expression is:

[0068]

[0069] where y t is the target Q-value at time step t, r t is the immediate reward obtained at time step t, γ is the discount factor, s t+1 is the next state, a' is all actions in state s t+1 , Q(s t+1 , a'; θ - ) is the Q-value estimation of the target network for the state-action pair (s t+1 , a′), and θ - are the target network parameters.

[0070]

[0071] Among them, L(θ) is the loss function of the main network parameter θ, and y i is the target Q value of the i-th sample, B is the batch size, and Q(s i , a i ; θ) is the Q value estimation of the main network for the state-action (s i , a i ), s i is the state of the i-th sample, a i is the action of the i-th sample, and θ is the main network parameter.

[0072] Target network update:

[0073] θ - ← τθ + (1 - τ)θ - ;

[0074] Among them, τ is a coefficient less than 1.

[0075] Furthermore, based on the interaction data collected from the reinforcement learning environment, the Q-learning algorithm is used to randomly sample from the experience replay buffer, calculate the target Q value, and update the parameters of the DQN model, thereby gradually optimizing the model performance. The specific steps include calculating the target Q value, minimizing the loss between the predicted Q value and the target Q value to update the main network parameters, and regularly synchronizing or updating the target network parameters. Finally, after the model converges, the optimal task allocation policy is generated according to the current state to ensure that the system can make the best task allocation decisions under different states. This process not only improves the learning efficiency and stability of the model but also realizes efficient resource management and task scheduling.

[0076] S3. Based on task containerization and resource quotas, task scheduling and deployment are performed through a container orchestration tool to generate a task deployment status.

[0077] Based on the optimal task allocation policy, define the running environment, dependencies, and resource requirements of each task to generate a task requirements list;

[0078] Furthermore, first evaluate the characteristics of each task, including the specific environmental configurations required for its operation, other tasks or data resources it depends on, and specific resource requirements (such as CPU, memory, and storage space). Then, based on this information, define the specific operating environment and dependencies of each task, and list all the required resources in detail. Finally, generate a comprehensive task requirements list that precisely describes all the conditions and resources required for each task to execute, facilitating efficient task scheduling and resource allocation. This process ensures that all tasks can be allocated and executed according to the optimal strategy while meeting their operating requirements. In summary, this process formulates an accurate task requirements list by exhaustively analyzing and defining the operating environment, dependencies, and resource requirements of tasks to support efficient resource management and task scheduling.

[0079] Based on the task requirements list, query the image repository to check if there is a corresponding container image.

[0080] Furthermore, first extract the specific requirements of each task for the operating environment and dependencies, and then query the image repository based on these requirements to check if there is a matching container image; if a compliant image is found, it can be directly used, otherwise, a new image needs to be created or adjusted to meet the task requirements. This process can be summarized as: retrieving a suitable container image in the image repository according to the task requirements list to ensure that the operating environment and dependencies of the task are met, thus supporting the smooth execution of the task.

[0081] When it exists, directly use it to containerize the task.

[0082] It should be noted that when there is a container image in the image repository that meets the task requirements list, directly use this image to containerize the task, including configuring environment variables, mounting necessary storage volumes, and setting network parameters, and then start the container to execute the task. The whole process can be summarized as: after confirming the matching image, immediately use this image to complete the containerization configuration and startup of the task to quickly respond to and execute the task.

[0083] When it does not exist, based on the task requirements list, create a Dockerfile to define the build rules of the container, and based on the build rules, use Docker tools to build the container image.

[0084] The specific process is as follows. When the required container image does not exist, first analyze the requirements list of the task to determine the application and its dependencies. Then write a Dockerfile to clarify the container building rules, including steps such as selecting a base image, installing dependencies, and configuring the environment. Next, use the docker build command to build the container image according to the Dockerfile, and use the docker run command to verify whether the image can run the application correctly, so as to ensure the creation of a containerized application environment that meets the requirements. This process ensures the portability and consistency of the application and simplifies the deployment process.

[0085] When it exists, verify whether the task can run properly in the container by starting the container, checking the task output, verifying the function, and monitoring resource usage.

[0086] If it can run properly, the container image has successfully achieved containerization of the task.

[0087] If it cannot run properly, gradually troubleshoot the problem by checking the log, dependency integrity, port mapping, resource limits, data mounting, and base image compatibility until the containerization of the task is completed.

[0088] The specific process is as follows. When the container image exists, first start the container using the docker run command, and check the task output and function verification to ensure that the application works as expected. Then use tools such as docker stats to monitor resource usage. If everything is normal, it means that the container image has successfully achieved containerization of the task. If abnormalities are found, issues such as logs, dependency integrity, port mapping, resource limits, data mounting, and base image compatibility need to be checked in sequence until all obstacles are resolved and the containerized deployment of the task is completed. This process ensures the stability and reliability of the containerized application.

[0089] Based on the requirements list of the task, use the resource management unit to perform requirement analysis and resource allocation on the requirements list of the task to generate resource quotas.

[0090] It should be noted that the resource management unit is composed of a resource scheduling function, a quota limit function, a monitoring function, an allocation and recycling function, and a configuration management function.

[0091] The specific process is as follows. Based on the task requirement list, the resource management unit first conducts a detailed analysis of various task requirements (such as CPU, memory, storage, network bandwidth, etc.), and identifies the specific environment configurations and dependencies required for task operation. Then, according to these analysis results and the current resource usage situation of the system, a reasonable resource allocation plan is formulated, and specific resource quotas are generated. This process ensures that each task can obtain the necessary resources to meet its operation requirements, while optimizing the overall resource utilization rate and avoiding resource waste or over-allocation. In summary, this process involves the precise analysis of task requirements and the matching of system resources, so as to generate the optimal resource quotas for task execution.

[0092] S4. Based on task containerization and resource quotas, perform task scheduling and deployment through a container orchestration tool to generate a task deployment status;

[0093] Based on task containerization, use a container orchestration tool to pull the container image, and generate a target node through resource quotas;

[0094] The specific process is as follows. Based on task containerization, first define resource quotas according to the task requirement list through a container orchestration tool, and configure detailed container running parameters; then automatically pull the required container image from the image repository, and at the same time apply the preset resource quotas to ensure the reasonable allocation of system resources; then intelligently select the optimal target node for task deployment according to the resource usage situation and task requirements of each node in the cluster; finally, continuously monitor the task running status and dynamically adjust the resource allocation or re-schedule the task according to the actual load, so as to achieve efficient and stable task execution and resource management. The whole process ensures that the task optimizes the overall resource utilization rate while meeting its specific requirements through an automated tool chain.

[0095] Based on the target node, complete the environment initialization by mounting data volumes, configuring network parameters, and loading dependent libraries, and create a task instance;

[0096] The specific process is as follows. First, ensure that the container can access the required input and output paths or shared resources by mounting necessary data volumes, then configure network parameters to implement the correct port mapping and service discovery mechanism to ensure that the application inside the container can communicate effectively with the external environment or other containers; then load all necessary dependent libraries according to the task requirement list, and set key environment variables such as database connection information and API keys, etc., to provide the correct running environment for the application; then perform a health check to verify whether the application has been correctly initialized and whether its services are ready; finally, after completing the above steps to ensure that all configurations are correct, create and start a task instance, that is, start the container on the target node to make the task execute according to the predetermined configuration. The whole process ensures that the task can run smoothly in an optimized environment that meets its specific requirements, and at the same time realizes efficient task deployment and management.

[0097] Collect the running status information of task instances in real time and generate the task deployment status.

[0098] Furthermore, the process of collecting the running status information of task instances in real time and generating the task deployment status first involves configuring monitoring tools (such as Prometheus, Grafana) to track key performance indicators, including CPU usage, memory occupancy, network traffic, etc.; then continuously collect the standard output / error logs of tasks and application-specific health checkpoint data through the integrated logging and monitoring system. Subsequently, perform real-time analysis on the collected data to identify abnormal behaviors or performance bottlenecks, and generate a detailed task deployment status report based on this, covering the running status of task instances, performance indicators, resource usage, and early warnings of potential problems.

[0099] S5. Monitor the load situation of the task deployment status in real time, trigger the elastic scaling mechanism to optimize the resource quota according to the preset threshold of the load situation, and generate the best demand retrieval strategy.

[0100] Define the preset threshold based on the historical load trend prediction data;

[0101] It should be noted that the process of defining the preset threshold based on the historical load trend prediction data involves starting from the collection and preparation of historical data, going through detailed analysis and trend prediction, then setting reasonable resource usage thresholds according to the prediction results combined with business requirements, and ensuring its effectiveness through optimization and verification. Finally, apply these thresholds to the actual environment and continuously monitor. This process not only improves the efficiency of resource management but also enhances the ability to handle potential problems.

[0102] Use the built-in monitoring function of the container orchestration tool to collect the load situation of the task deployment status, generate monitoring data, and compare the load situation of the monitoring data with the preset threshold;

[0103] The specific process is as follows. First, define clear monitoring objectives and rules to guide data collection. Then, collect and process the performance indicator data of task instances in real time, generate a visual monitoring report, and compare these data with the preset threshold. Once any indicator is found to be close to or exceed the threshold, immediately trigger an alarm and take corresponding measures. Finally, continuously optimize and adjust the threshold and monitoring strategy based on the monitoring results, ensuring continuous improvement and efficient management. This method not only improves observability but also enhances the ability to quickly respond to potential problems.

[0104] When the load exceeds the upper limit threshold, trigger the expansion operation; conversely, trigger the reduction operation;

[0105] Specifically, when the load exceeds the upper threshold, the container orchestration tool will automatically trigger an expansion operation, such as increasing service instances or upgrading resource configurations; conversely, when the load is below the lower threshold, a contraction operation will be triggered to release unnecessary resources. This process begins with setting reasonable upper and lower thresholds, and through real-time monitoring and data analysis, automated expansion and contraction decisions are made to ensure that peak loads can be handled while resources are utilized efficiently and waste is avoided. The entire process emphasizes the importance of automation and continuous optimization, enabling efficient and stable operation while meeting business requirements.

[0106] Based on the elastic expansion mechanism, optimize resource quotas and generate the best retrieval strategy.

[0107] Furthermore, the process of optimizing resource quotas and generating the best retrieval strategy based on the elastic expansion mechanism begins with real-time tracking of the system load through the monitoring function of the container orchestration tool, and setting reasonable expansion and contraction strategies to respond to load changes. On this basis, the resource usage patterns are analyzed in depth to optimize resource quotas and improve overall resource utilization. At the same time, the best retrieval strategy is generated by using large language models and historical data analysis, covering aspects such as demand priority ranking, resource allocation optimization, and time schedule adjustment, to ensure efficient response to complex product design and R & D requirements. The entire process emphasizes the importance of automation and continuous optimization, not only maintaining stable operation during peak loads but also achieving cost savings during low-load periods, enhancing flexibility and efficiency.

[0108] This embodiment also provides an AI demand retrieval platform based on a large model, including: a demand parsing module, a strategy generation module, a containerization processing module, a task scheduling module, a load monitoring module, and a best demand retrieval strategy generation module; the demand parsing module is used to parse user demands through the NLP unit, extract functional and non-functional demand information and convert it into a structured data format; the strategy generation module is used to build a reinforcement learning environment based on the structured data format, and interactively train the initialized DQN model through the reinforcement learning environment to generate an optimal task allocation strategy; the containerization processing module is used to containerize tasks based on the optimal task allocation strategy and generate resource quotas through the resource management unit; the task scheduling module is used to schedule and deploy tasks through the container orchestration tool based on task containerization and resource quotas, and generate task deployment status; the load monitoring module is used to monitor the load of the task deployment status in real time; the best demand retrieval strategy generation module is used to trigger the elastic expansion mechanism to optimize resource quotas according to the preset threshold of the load situation and generate the best demand retrieval strategy.

[0109] This embodiment also provides a computer device, which is applicable to the case of an AI requirement retrieval method based on a large model, and includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the AI requirement retrieval method based on a large model as proposed in the above embodiment.

[0110] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or may also be a button, a trackball, or a touchpad provided on the housing of the computer device, or may also be an external keyboard, a touchpad, or a mouse, etc.

[0111] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the AI requirement retrieval method based on a large model as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0112] In summary, the present invention achieves the following: the NLP unit analyzes the user's requirements and generates structured data, enabling precise analysis and standardized output, providing high-quality data support for subsequent task allocation and resource management, reducing understanding errors and manual intervention; by constructing a reinforcement learning environment and training in combination with the DQN model, an optimal task allocation strategy is generated, improving the level of scheduling intelligence, optimizing system efficiency and resource utilization rate, and solving the problems of low efficiency and resource waste in traditional methods.

[0113] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An AI requirement retrieval method based on a large model, characterized in that: including Parsing user requirements through the NLP unit, extracting functional and non-functional requirement information and converting it into a structured data format; Based on the structured data format, constructing a reinforcement learning environment, and interacting and training the initialized DQN model through the reinforcement learning environment to generate an optimal task allocation strategy; Based on the optimal task allocation strategy, containerizing the tasks, and generating resource quotas through the resource management unit; Based on task containerization and resource quotas, performing task scheduling and deployment through a container orchestration tool to generate a task deployment status; Real-time monitoring of the load situation of the task deployment status, and triggering an elastic scaling mechanism to optimize resource quotas according to the preset threshold of the load situation, generating an optimal demand retrieval strategy.

2. The AI requirement retrieval method based on a large model according to claim 1, wherein: The specific steps for parsing user requirements through the NLP unit are as follows: Parsing user requirements through input processing, semantic analysis, entity recognition, relationship extraction, classification annotation, data generation, verification feedback, and extensible interfaces to generate structured data.

3. The AI requirement retrieval method based on a large model according to claim 2, characterized in that: The specific steps for constructing a reinforcement learning environment based on the structured data format are as follows: Parsing the task information, analysis node status, and task dependencies of the structured data through an XML parser, and converting them into a task queue, a computing node resource table, and a task dependency graph; Based on the task queue, the computing node resource table, and the task dependency graph, converting them into a state space through data preprocessing, and then converting the action space through action design; Using a data parsing method to extract the state space and the action space, generating a quantitative feature value reflecting the influence of the state and the action, and integrating them into a reward function by weighted summation; Constructing a reinforcement learning environment by integrating the task queue, the computing node resource table, the task dependency graph, the state space, the action space, and the reward function.

4. The AI requirement retrieval method based on a large model according to claim 3, wherein: The specific steps for interacting and training the initialized DQN model through the reinforcement learning environment to generate an optimal task allocation strategy are as follows: Using the initialization method Xavier to assign values to the weight parameters of the DQN model, defining the neural network structure through user requirements, the state space, and the action space, and setting hyperparameters to generate an initialized DQN model; Interacting and training the initialized DQN model with the reinforcement learning environment to generate interaction data, and storing it in the experience replay buffer; Randomly extracting interaction data from the experience replay buffer, and using the Q-learning algorithm to update the parameters of the DQN model to generate an optimal task allocation strategy for the current state.

5. The AI requirement retrieval method based on a large model according to claim 4, wherein: The specific steps for containerizing the tasks based on the optimal task allocation strategy and generating resource quotas through the resource management unit are as follows: Based on the optimal task allocation strategy, defining the running environment, dependencies, and resource requirements of each task to generate a demand list for the tasks; Based on the demand list of the tasks, querying whether there is a corresponding container image in the image repository; When it does not exist, then based on the demand list of the tasks, creating a Dockerfile, defining the container building rules, and using the Docker tool to build a container image based on the building rules; When it exists, verifying whether the task can run normally in the container by starting the container, checking the task output, verifying the function, and monitoring resource usage. When it can run properly, the container image has successfully containerized the task; When it cannot run properly, troubleshoot the problem step by step by checking the logs, dependency integrity, port mapping, resource limits, data mounts, and base image compatibility until the task is containerized; Based on the requirements list of the task, use the resource management unit to perform requirement analysis and resource allocation on the requirements list of the task to generate resource quotas.

6. The AI requirement retrieval method based on a large model according to claim 5, characterized in that: Based on the task containerization and resource quotas, perform task scheduling and deployment through the container orchestration tool to generate the task deployment status. The specific steps are as follows: Based on the task containerization, use the container orchestration tool to pull the container image and generate the target node through the resource quota; Based on the target node, complete the environment initialization by mounting the data volume, configuring network parameters, and loading the dependency library, and create a task instance; Collect the running status information of the task instance in real time to generate the task deployment status.

7. The AI requirement retrieval method based on a large model according to claim 6, characterized in that: Monitor the load of the task deployment status in real time. According to the preset threshold of the load situation, trigger the elastic scaling mechanism to optimize the resource quota and generate the best demand retrieval strategy. The specific steps are as follows: Define the preset threshold based on the historical load trend prediction data; Use the built-in monitoring function of the container orchestration tool to collect the load of the task deployment status, generate monitoring data, and compare the load situation of the monitoring data with the preset threshold; When the load exceeds the upper threshold, trigger the expansion operation. Conversely, trigger the reduction operation to generate the best retrieval strategy.

8. An AI requirement retrieval platform based on a large model, based on the AI requirement retrieval method based on a large model according to any one of claims 1 to 7, characterized in that: Including a requirement analysis module, a strategy generation module, a containerization processing module, a task scheduling module, a load monitoring module, and a best demand retrieval strategy generation module; The requirement analysis module is used to parse the user requirements through the NLP unit, extract the functional and non-functional requirement information and convert it into a structured data format; The strategy generation module is used to build a reinforcement learning environment based on the structured data format, and perform interactive training on the initialized DQN model through the reinforcement learning environment to generate the optimal task allocation strategy; The containerization processing module is used to containerize the task based on the optimal task allocation strategy and generate resource quotas through the resource management unit; The task scheduling module is used to perform task scheduling and deployment through the container orchestration tool based on the task containerization and resource quotas to generate the task deployment status; The load monitoring module is used to monitor the load of the task deployment status in real time; The best demand retrieval strategy generation module is used to trigger the elastic scaling mechanism to optimize the resource quota according to the preset threshold of the load situation and generate the best demand retrieval strategy.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the AI demand retrieval method based on the large model according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the AI demand retrieval method based on the large model according to any one of claims 1 to 7.