Service demand identification and on-demand AI service method based on large language model
Through the service demand recognition method based on the large language model and the stitchable neural network architecture, the problems of difficulty in extracting intentions and delayed communication in users' natural language descriptions are solved, and personalized and flexible on-demand AI services are realized, which improves user experience and service quality.
Patent Information
- Application Number
- CN202510233253.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-27
AI Technical Summary
The implicit intentions implicitly in the user's natural language descriptions are difficult to accurately extract, resulting in AI services being unable to meet user needs, and the communication delay is large, making it impossible to provide personalized and flexible on-demand AI services.
The service demand recognition method based on the large language model is adopted, and the user's natural language is preprocessed, intention and entity recognition is used to determine the KPI information required by the user, and the AI service network architecture is used to optimize the AI service network architecture to shorten the communication link.
It realizes accurate extraction of implicit intentions in user's natural language description, reduces communication delay, provides users with personalized and flexible on-demand AI services, and improves user experience and service quality.
Smart Images

Figure CN120218050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of service demand recognition and on-demand service, and particularly to a service demand recognition and on-demand AI service method based on a large language model. Background Art
[0002] With the rapid development of artificial intelligence technology, users' demands for personalized and intelligent services are increasing day by day. In many application scenarios, how to accurately identify users' intentions and provide corresponding on-demand AI services has become the key to improving user experience and service quality. Users may come from different industries and backgrounds, and their demands for services are complex and variable, and are usually expressed in the form of natural language. Therefore, developing a technology that can accurately extract key information from users' natural language and provide customized AI services based on these key information is crucial for meeting users' demands. In addition, with the progress of big data and machine learning technologies, more intelligent systems can be built, which can not only understand users' demands, but also dynamically adjust service content according to users' real-time behaviors and preferences to provide the best user experience. This requires the system to have a high degree of adaptability and intelligent decision support to achieve optimal allocation of resources and personalized customization of services.
[0003] User instructions can directly express specific intention requirements and are the simplest and most direct interaction method. Currently, for user demand recognition, methods based on traditional neural networks are usually adopted. However, although the method based on traditional neural networks has the ability to understand semantics, it is difficult to process long sentences, and the efficiency of identifying user demands is low, and it is difficult to complete the task of identifying complex intentions. Eventually, the AI services provided based on incorrect intentions cannot meet users' demands, resulting in the inability to accurately extract the intentions hidden in users' natural language descriptions. For on-demand AI services, the current AI services provided for users mainly run on cloud servers, and there is a long communication link between users and the server, resulting in a large communication delay overhead. Currently, the work of providing on-demand AI services for users is still lacking, and it is impossible to dynamically configure tasks according to users' demands, resulting in the inability to provide personalized and flexible AI services for users on demand. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a service demand recognition and on-demand AI service method based on a large language model, and solve the problems of large communication delay overhead, inability to accurately extract the intentions hidden in users' natural language descriptions, and inability to provide personalized and flexible AI services for users on demand for on-demand AI services.
[0005] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0006] The first aspect of the present invention provides a service demand recognition and on-demand AI service method based on a large language model, including:
[0007] When receiving the AI service request data initiated by the user, determine the KPI information required by the user according to the AI service request data, the BERT model, the new service request, the KPI and entity values in the knowledge base;
[0008] According to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchical relationship and the resource constraints, use the stitchable neural network and the AI service network architecture to determine the deployed AI service network architecture;
[0009] Use the deployed AI service network architecture to execute the user's task and obtain the execution result corresponding to the user's task.
[0010] The second aspect of the present invention provides a service demand recognition and on-demand AI service device based on a large language model, including:
[0011] The first determination module is used to determine the KPI information required by the user according to the AI service request data, the BERT model, the new service request, the KPI and entity values in the knowledge base when receiving the AI service request data initiated by the user;
[0012] The second determination module is used to determine the deployed AI service network architecture according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchical relationship and the resource constraints, using the stitchable neural network and the AI service network architecture;
[0013] The execution module is used to execute the user's task using the deployed AI service network architecture and obtain the execution result corresponding to the user's task.
[0014] Compared with the prior art, when receiving the AI service request data initiated by the user, the service demand recognition and on-demand AI service method and device based on the large language model provided by the present invention determine the KPI information required by the user according to the AI service request data, the BERT model, the new service request, the KPI and entity values in the knowledge base; according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchy relationship and the resource constraints, using the stitchable neural network and the AI service network architecture, determine the deployed AI service network architecture; use the deployed AI service network architecture to execute the user's task and obtain the execution result corresponding to the user's task. In this way, the communication link between the user and the server can be shortened through the deployed AI service network architecture, thereby reducing the communication delay; the implicit intention in the user's natural language description can be accurately extracted through the BERT model; the resource allocation in the deployed AI service network architecture can be optimized according to the KPI information required by the user, and personalized and flexible AI services can be provided for the user on demand. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:
[0016] Figure 1 Schematically shows a schematic diagram of the AI service network architecture;
[0017] Figure 2 Schematically shows a flowchart of the service demand recognition and on-demand AI service method based on the large language model;
[0018] Figure 3 Schematically shows a flowchart of an embodiment of the service demand recognition and on-demand AI service method based on the large language model;
[0019] Figure 4 Schematically shows a structural diagram of the service demand recognition and on-demand AI service device based on the large language model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0021] It should be noted that: Unless otherwise specified, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those skilled in the art to which the present invention pertains.
[0022] The method in the embodiments of the present invention will be described in detail below.
[0023] Figure 1 A schematic diagram of the AI service network architecture is schematically shown. Refer to Figure 1 As shown, the service demand recognition and on-demand AI service method based on the large language model of the present invention is carried out based on the AI service network architecture. The AI service network architecture consists of two cooperating layers: the cognitive layer and the decision layer. The user side initiates AI service request data. The main responsibility of the cognitive layer is to parse the user intention according to the AI service request data, and accurately identify the user's situation and needs during the service process. The decision layer, based on the demand recognition result of the cognitive layer, performs resource perception and on-demand allocation, and controls the network to complete the task requested by the user. Finally, the execution result of the task that completes the user request is uploaded to the cloud. The AI service network architecture is a network architecture that provides on-demand services and has multiple network element nodes. The AI service request data includes computer vision-related tasks, recommendation-related tasks, natural language processing or speech recognition-related tasks, reinforcement learning-related tasks, etc. The AI service network architecture includes: AI models uploaded by users, untrained AI models, trained AI models, AI model reuse, and combined use of AI models.
[0024] Figure 2 A flowchart of the service demand recognition and on-demand AI service method based on the large language model in the embodiments of the present invention is schematically shown. Refer to Figure 2 As shown, the method may include:
[0025] S201. When receiving the AI service request data initiated by the user, determine the KPI information required by the user according to the AI service request data, the BERT model, the new service request, the KPIs and entity values in the knowledge base.
[0026] The new service request includes a computer vision task request, a natural language processing task request, a speech recognition task request, a reinforcement learning task request, and a recommendation system task request.
[0027] On the user side, when the user initiates an AI service request through a local device, the cognitive layer starts to work. When the cognitive layer receives the AI service request data initiated by the user, it uses large model technology to accurately identify the user's needs. The execution of the cognitive layer mainly includes the following three key steps: text preprocessing, intent and entity recognition, and KPI mapping and extraction. The following step A1 is the text preprocessing step, step A3 is the intent and entity recognition step, and steps A4 and A5 are the KPI mapping and extraction steps.
[0028] Specifically, when receiving the AI service request data initiated by the user, according to the AI service request data, BERT model, new service request, KPI and entity values in the knowledge base, determine the KPI information required by the user, including:
[0029] Step A1: Perform text preprocessing on the AI service request data to obtain preprocessed data.
[0030] Specifically, step A1 includes:
[0031] Step A11: Use a generative pre-training model to perform data augmentation on the AI service request data to obtain augmented data.
[0032] Collect a small amount of data from the user, and then use the Generative Pre-Train (GPT) model to perform data augmentation on the collected data, that is, the AI service request data, to obtain augmented data.
[0033] Step A12: Clean the augmented data to obtain cleaned data.
[0034] Among them, the cleaned data is data that does not contain punctuation marks.
[0035] Clean the augmented data and remove all irrelevant characters and punctuation marks that may interfere. For example, the user input "I need to upload 100MB of data within 5 seconds." After cleaning, it will be simplified to "I need to upload 100MB of data within 5 seconds".
[0036] Step A13: Use the BERT tokenizer to tokenize the cleaned data and standardize the tokenized data to obtain standardized data.
[0037] Using the BERT tokenizer, continuous text information can be converted into discrete data units that the model can process. Continuing with the example in step A12, the text input by the user will be decomposed into ["I", "need", "to", "complete", "the", "upload", "of", "100MB", "of", "data", "within", "5 seconds"], and "5 seconds" will be standardized to "5s", and "100MB" will be standardized to "100MB".
[0038] Step A14: Perform data annotation on the standardized data to obtain preprocessed data.
[0039] Perform data annotation on the standardized data. Data annotation includes intent and entity annotation. Among them, intent annotation is mainly to annotate the type of task that the user needs to complete. For example, ([“I”, “need”, "to", "complete", "the", "upload", "of", "100MB", "of", "data", "within", "5s"], “data upload”) annotates the user's intent as “data upload”; entity annotation is mainly to annotate the KPIs that need to be satisfied to complete this task contained in the user's natural language. For example, [(“I”,0),(“need”,0),("to",0),("complete",0),("the",0),("upload",0),("of",0),("100MB", “data volume”),("data",0),("within",0),("5s", “time limit”)] annotates that the “time limit” to complete this task is 5s and the “data volume” is 100MB.
[0040] Step A2: Use the preprocessed data to train the BERT model to obtain a trained BERT large language model.
[0041] In order to obtain a large language model that can recognize user needs, the present invention adopts a BERT model and uses the collected data, namely AI service request data, to fine-tune and train the BERT model to obtain a large model that can accurately recognize user intent. This step A1 is executed before the entire system is deployed, and after the system is deployed, it will no longer be necessary to train the BERT model.
[0042] Step A3: Input a new service request into the trained BERT large language model so that the trained BERT large language model outputs the user intent.
[0043] When the user initiates a new service request, input the new service request into the trained BERT large language model, and use the trained BERT model in step A2 to recognize the user's needs, so that the trained BERT large language model outputs the user intent. The type of task that the user needs to complete and the numerical indicators that need to be satisfied to complete this task can be obtained.
[0044] Step A4: Among the KPIs in the knowledge base, search for the target KPI that matches the user's intention, and through natural language processing technology, extract the entity value corresponding to the target KPI.
[0045] Continuing with the example in Step A14, among the KPIs in the knowledge base, search for the target KPI that matches the identified user intention of "data upload". The target KPI is a key metric for measuring task performance, such as "task completion time", "task accuracy", "communication time", "data size", etc. By linking the user intention with these target KPIs, the user's needs can be understood more accurately.
[0046] Step A5: Map the entity value to the corresponding target KPI to obtain the KPI information required by the user.
[0047] According to the mapping logic, map the entity value corresponding to the extracted target KPI to the corresponding target KPI. For example, the time limit of "5s" is mapped to the target KPI "upload time", and the requirement "upload time ≤ 5 seconds" is set; the data volume of "100MB" is mapped to the target KPI "data size", and it is defined as "data size = 100MB". In this way, the specific performance requirements, that is, the KPI information required by the user, can be extracted from the user's new service request.
[0048] The KPI information required by the user obtained through cognitive layer processing is transmitted to the decision-making layer.
[0049] S202: According to the KPI information required by the user, the AI model library, service response time, service availability, AI model hierarchy relationship, and resource constraints, use the stitchable neural network and the AI service network architecture to determine the deployed AI service network architecture.
[0050] Design a refined on-demand service plan according to the KPI information required by the user. In order to meet different users' preferences for service latency and accuracy, the accuracy and latency of the neural network for executing AI services need to be adjustable. Therefore, the stitchable neural network (Stitchable Neural Networks, SNNet) is used to provide AI services.
[0051] Specifically, according to the KPI information required by the user, the AI model library, service response time, service availability, AI model hierarchy relationship, and resource constraints, use the stitchable neural network and the AI service network architecture to determine the deployed AI service network architecture, including:
[0052] Step B1: According to the KPI information required by the user, the AI model library, service response time, service availability, AI model hierarchy relationship, and resource constraints, use the stitchable neural network to determine a new AI model.
[0053] Specifically, step B1 includes:
[0054] Step B11: Select two AI models from the AI model library according to the KPI information required by the user.
[0055] Specifically, step B11 includes:
[0056] Step B111: Determine the corresponding tasks according to the KPI information required by the user.
[0057] Among them, the tasks include image classification tasks and sentiment analysis tasks.
[0058] Step B112: In the AI model library, select two AI models corresponding to the tasks.
[0059] Among them, the sizes of the two AI models are different.
[0060] Step B12: Establish an optimization problem with the goal of maximizing user satisfaction and on-demand service rate, and with service response time, service availability, AI model hierarchical relationship, and resource constraints as constraints.
[0061] Specifically, in order to determine the optimal splitting position of the AI model and the deployment position of the AI model in the AI service network architecture, with the goal of maximizing user satisfaction and on-demand service rate, while considering constraints such as service response time, service availability, AI model hierarchical relationship, and resource constraints, the following optimization problem is proposed, and the expression of the optimization problem is:
[0062]
[0063] s.t.C1:t q ≤D q
[0064]
[0065] Among them, is user satisfaction, β is the stitching coefficient matrix, L is the position matrix of the AI model deployment layer, C1 is the constraint of service response time, C2 is the constraint of service availability, C3 is the constraint of the hierarchical relationship between the pre and post AI models in stitching, C4 is the constraint of single AI model deployment, C5 is the constraint of the stitching coefficient of the pre AI model of the AI model, C6 is the constraint of the stitching coefficient matrix of the stitching model, C7 is the constraint of the stitching coefficient matrix of single AI model deployment, C8 is the constraint of resource limitation, t q is the completion time of each new service request q, D q is the preset deadline delay of each new service request q, is the availability of each new service request q, A qThe preset availability requirement for each new service request q The deployment location of the pre-model for each new service request q The deployment location of the post-model for each new service request q The deployment location of the single AI model corresponding to the new service request q The stitching coefficient of the pre-AI model related to the new service request q The stitching coefficient of the post-AI model related to the new service request q The pre-stitching coefficient matrix of the stitching model related to the new service request q The post-stitching coefficient matrix related to the new service request q The stitching coefficient matrix for the deployment of the single AI model related to the new service request q The pre-model The post-model The single model, C i The resource capacity, N i The number of nodes, m p One of the two AI models, m r The other AI model of the two AI models, where i is the i-th layer network in the AI service network architecture
[0066] Taking the AI model stitching coefficient and deployment location as optimization variables, a mathematical model aiming to maximize user satisfaction and on-demand service rate is established, as shown in the expression of the above optimization problem. The expression of the optimization problem contains eight constraint conditions from C1 to C8. Among them, for constraint condition C1, for each new service request q, its completion time t q Does not exceed the preset deadline D q , which ensures the timeliness of the service and meets the user's delay requirements. For constraint condition C2, for each new service request q, its availability Is not lower than the preset availability requirement A q , which ensures the reliability of the service and meets the user's minimum requirement for service availability. For constraint condition C3, for each service request q, the deployment location of the pre-model Is less than or equal to the deployment location of the post-model This ensures the correct logical order of AI model stitching and the correctness of the data flow. Constraint conditions C5 to C7 ensure the integrity and accuracy of the AI model stitching process. For constraint condition C8, for each layer network in the AI service network architecture, the deployment of the pre-model The post-model And the single model The total resources required cannot exceed the resource capacity C of this layer iMultiplied by the number of nodes N i , which ensures that the resource utilization of each layer does not exceed its carrying capacity.
[0067] The optimization objective is defined as follows:
[0068] In this optimization problem, the composition of the cost mainly includes the cost consumption of computing resources, the resource consumption required to maintain the AI model, and the cost consumption of communication resources. These costs reflect the costs of resource consumption in different network layers and the costs of inter-layer communication. The specific cost model E can be expressed as:
[0069]
[0070] Among them, E i represents the resource consumption cost of the i-th layer in the network, and E trans represents the communication cost, μ i is the weight factor associated with the i-th layer, C trans represents the communication resource consumption, and μ trans represents the weight factor of the inter-layer communication layer. is the pre-model, is the post-model, is a single model.
[0071] Define the successful service set F = {f1, f2,..., f q}, where f q is the completion status indicator for each new service request q, and is specifically defined as:
[0072]
[0073] The on-demand service rate θ can be calculated by the following formula:
[0074]
[0075] Among them, represents the total number of service requests, and the revenue R obtained from successfully completing the service is defined as the following formula:
[0076]
[0077] Among them, η q is the revenue weight corresponding to successfully completing the q-th new service request.
[0078] In summary, the total revenue I of the system can be determined by the difference between the revenue and the cost, that is, as shown in the following formula:
[0079] I = R - E;
[0080] In this optimization problem, the user satisfaction is also defined u q′ represents the satisfaction degree of user q′ with respect to the services provided by the system, which is mainly obtained based on the total revenue I of the system. represents a function with the total revenue of the system as the independent variable.
[0081] Among them, the optimization variables are as follows:
[0082] AI model stitching coefficient: The AI model stitching coefficient determines the stitching together of different AI models to adapt to specific user requests and service requirements.
[0083] AI model deployment location: Determines the specific deployment location of the AI model in the network architecture, which may involve edge computing nodes or cloud computing centers.
[0084] The constraint conditions are as follows:
[0085] Service response time constraint: The completion time of each service request does not exceed the user's latency requirement.
[0086] Service availability constraint: The availability of each service request meets the user's minimum requirements.
[0087] AI model hierarchical relationship constraint: During the stitching process, the level of the pre - model is less than or equal to the level of the post - model to ensure the correctness of the data flow and logical order.
[0088] Single AI model deployment constraint: Ensures that the deployment of a single AI model meets specific performance and resource requirements.
[0089] Stitching coefficient constraint: The stitching coefficient meets specific ranges and conditions to ensure the performance of the stitched model.
[0090] Resource limit constraint: The total resource consumption of each layer cannot exceed the resource capacity of that layer.
[0091] Step B13: Use the dung beetle algorithm to solve the optimization problem to obtain the splitting positions and deployment positions of the two AI models.
[0092] Due to the high non - convexity of the optimization problem and the complexity of the solution space, traditional optimization algorithms may be difficult to find the global optimal solution. Therefore, the dung beetle algorithm (Dung Beetle Optimizer Algorithm, DBO) is used as a heuristic search algorithm to simulate the foraging behavior of the dung beetle colony to find the global optimal solution or an approximate optimal solution in the complex search space.
[0093] Step B14: According to the stitched - together neural network and the splitting positions, split and splice the two AI models to obtain a new AI model.
[0094] Generally speaking, large AI models have higher accuracy but higher inference latency; while small AI models have lower inference latency but lower accuracy. By using a stitchable neural network to split and then reassemble two AI models of different sizes, a new AI model is formed, and the accuracy and latency of this new AI model are between those of the two original AI models. By controlling the splitting positions of the large AI model and the small AI model, dynamic adjustment of the accuracy and inference latency of the AI model can be achieved.
[0095] Step B2: Determine the deployed AI service network architecture according to the new AI model and the AI service network architecture.
[0096] Specifically, Step B2 includes:
[0097] Deploy the new AI model into the AI service network architecture using the deployment location to obtain the deployed AI service network architecture.
[0098] S203: Execute the user's task using the deployed AI service network architecture to obtain the execution result corresponding to the user task.
[0099] The present invention aims to achieve accurate recognition of user intentions through advanced natural language processing technologies and intelligent analysis methods, and provide highly personalized on-demand AI services for users according to the accurate recognition results. In this way, the response speed, flexibility, and user satisfaction of the services can be significantly improved, and at the same time, higher operational efficiency and market competitiveness can be brought to service providers.
[0100] Based on the above Figure 1 implementation method, it can be seen that when the embodiment of the present invention receives the AI service request data initiated by the user, it determines the KPI information required by the user according to the AI service request data, the BERT model, the new service request, the KPI and entity values in the knowledge base; according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchy relationship, and the resource constraints, it determines the deployed AI service network architecture by using the stitchable neural network and the AI service network architecture; and executes the user's task using the deployed AI service network architecture to obtain the execution result corresponding to the user task. In this way, the communication link between the user and the server can be shortened through the deployed AI service network architecture, thereby reducing the communication latency; the implicit intention in the user's natural language description can be accurately extracted through the BERT model; and the resource allocation in the deployed AI service network architecture can be optimized according to the KPI information required by the user to provide personalized and flexible AI services for the user on demand.
[0101] As an alternative embodiment of the present invention, two embodiments are given below, namely a computer vision task embodiment and a natural language processing task embodiment, to illustrate in detail the method for identifying service requirements based on a large language model and providing on-demand AI services according to the present invention.
[0102] Embodiment of computer vision task:
[0103] Background for solving technical problems: This computer vision task embodiment solves the problem of how to efficiently and accurately identify the content in an image uploaded by a user and provide feedback to the user. Through the AI service network architecture deployed according to the present invention, computing resources can be dynamically scheduled and the most suitable AI model can be called, reducing communication latency and providing personalized computer vision services according to user needs.
[0104] Implementation process:
[0105] (1) Cognitive layer:
[0106] Text preprocessing: The user uploads an image through the client and simultaneously inputs a request such as "Please identify all the people in the picture and classify them". The system first preprocesses the text, cleans irrelevant characters, standardizes the language, and inputs the processed text into the BERT model.
[0107] Intention and entity recognition: The BERT model, after training, can recognize that the main intention of the user is "image classification" and extract the entity as "person".
[0108] KPI mapping and extraction: By retrieving the knowledge base, the system associates "image classification" with relevant KPIs (such as processing time, recognition accuracy) and sets target performance indicators, such as "recognition accuracy ≥ 95%", "processing time ≤ 3 seconds".
[0109] (2) Decision-making layer:
[0110] Task parsing and model selection: Based on the information extracted by the cognitive layer, the decision-making layer determines that this is an image classification task that requires high-precision processing and selects the ResNet-50 model, which has a high accuracy rate in image classification tasks. The decision-making layer also allocates sufficient computing resources to meet the processing time requirement.
[0111] Fine-grained service scheduling: During the actual scheduling process, the system detects that the current network load is high. To ensure a processing time of 3 seconds, the system decides to reduce the batch size of the ResNet model and simultaneously process the image data in parallel on multiple computing nodes to reduce latency.
[0112] Model Execution and Result Return: The system calls the selected ResNet-50 model in the intelligent layer and uses this ResNet-50 model to identify and classify the people in the image. After parallel computing processing, the system completed the identification of all people within 2.8 seconds and returned the execution results to the user in the form of image annotation and text description.
[0113] Result Analysis and Feedback Optimization: The system analyzes the results of this task and finds that the accuracy reaches 96%, meeting the user's KPI requirements. The system records the configuration parameters of this task for reference and optimization in subsequent similar tasks.
[0114] Example of Natural Language Processing Task:
[0115] This example of natural language processing task solves the problem of how to efficiently and accurately parse text content and perform corresponding natural language processing tasks after the user inputs through natural language. Through the AI service network architecture deployed by the present invention, it can dynamically schedule computing resources and call the most suitable AI model, reduce communication latency, and provide personalized natural language processing services according to user needs.
[0116] Implementation Process:
[0117] (1) Cognitive Layer:
[0118] Text Preprocessing: The user inputs a text request through the client, such as "Please analyze the sentiment tendency of this text and give a detailed description". The system first preprocesses the text, removes redundant characters and punctuation marks, and standardizes the language expression. The processed text will be tokenized and input into the pre-trained BERT model.
[0119] Intention and Entity Recognition: Through training, the BERT model can recognize that the main intention of the user is "sentiment analysis" and extract relevant entities (such as sentiment tendency words, sentiment intensity, etc.) from the text. For example, the system recognizes sentiment words such as "happy" and "satisfied" in the text.
[0120] KPI Mapping and Extraction: The system retrieves the knowledge base, associates the "sentiment analysis" task with relevant KPIs (such as response time, analysis accuracy), and sets the target performance indicators. For example, "analysis accuracy ≥ 90%", "response time ≤ 2 seconds". The system prepares for subsequent task processing according to these indicators.
[0121] (2) Decision-making Layer:
[0122] Task parsing and model selection: Based on the information extracted by the cognitive layer, the decision layer determines that this is a sentiment analysis task that requires efficient processing, and selects the RoBERTa model, which has high accuracy and fast response time in sentiment analysis tasks. The decision layer also allocates appropriate computing resources to meet the response time requirements.
[0123] Refined service scheduling: During the actual scheduling process, the system detects that the current network resources are relatively abundant. In order to ensure that the results are returned within 2 seconds, the system chooses to use the more accurate RoBERTa model and processes text data in parallel on multiple computing nodes to further reduce the response time.
[0124] Model execution and result return: The system calls the selected RoBERTa model in the intelligent layer, which performs sentiment analysis on the text provided by the user. After parallel computing processing, the system completes the sentiment analysis task in 1.8 seconds and returns the analysis results (such as sentiment tendency and specific description) to the user.
[0125] Result analysis and feedback optimization: The system analyzed the results of this task and found that the analysis accuracy reached 92%, meeting the user's KPI requirements. The system recorded the configuration parameters and result data of this task for reference and optimization of subsequent similar tasks. If similar requests are encountered in future tasks, the system can dispatch the optimal resources more quickly and provide efficient services.
[0126] Figure 3 The flowchart of the embodiment of the service demand identification and on-demand AI service method based on the large language model is schematically shown, see Figure 3 As shown in the figure, the user initiates a service request to the network, inputs the service requirements and related data, and after being processed by the cognitive layer and the decision-making layer, the intelligent layer transmits the user data and AI service solutions to the execution layer according to the formulated service strategy, and calls, reuses or combines different AI models to output service results. For the data and models uploaded by users, the network will deploy and execute tasks based on these models and feed back the results to the users. At the same time, operators will regularly collect data and user feedback during the service process to discover network deficiencies and optimize the quality of AI services. The service requests initiated by users to the network include computer vision tasks, natural language processing tasks, sound recognition tasks, reinforcement learning tasks, and recommendation system tasks.
[0127] The present invention realizes personalized and efficient AI services by determining the deployed AI service network architecture and implementing an on-demand service solution for stitching new AI models, greatly improving the user experience and the optimal allocation of network resources. At the same time, it reduces the operating costs, enhances the reliability and intelligence level of the network, ensures the rapid response and high adaptability of the services, bringing significant economic benefits and market competitiveness to communication service providers, and laying a solid foundation for the sustainable development and innovation of future communication technologies. It is mainly reflected in the following aspects:
[0128] (1) Improve user satisfaction: By automatically extracting the communication KPIs described in the user's AI service request data, this solution can more accurately grasp the actual needs of users, thereby providing more suitable services and enhancing user satisfaction.
[0129] (2) Optimize network services: The deployed AI service network architecture and the new AI model, i.e., the on-demand service model, enable network services to be dynamically adjusted according to real-time changing demands, optimize resource allocation, and improve service efficiency and response speed.
[0130] (3) Enhance the intelligence level of the network: Using AI technology, the network can achieve autonomous decision-making and intelligent scheduling, reducing human intervention and improving the network's automation and intelligent management capabilities.
[0131] (4) Reduce operating costs: Through intelligent resource management and optimized AI model stitching strategies, the network can use resources more efficiently, reduce waste, and thus reduce operating costs.
[0132] (5) Improve resource utilization: The stitching and reuse mechanism of AI models can maximize the utilization of deployed model resources, improve resource utilization, and at the same time speed up the service delivery speed.
[0133] (6) Support diverse services: This solution supports the deployment and invocation of multiple AI models, including single model invocation, model reuse, combined use, and the upload and application of user-defined models, meeting diverse service requirements.
[0134] Based on the same inventive concept, as an implementation of the above-mentioned method for identifying service requirements based on large language models and on-demand AI services, an embodiment of the present invention also provides a device for identifying service requirements based on large language models and on-demand AI services. Figure 4 For the structural diagram of the device for identifying service requirements based on large language models and on-demand AI services in the embodiment of the present invention, see Figure 4 As shown, the device for identifying service requirements based on large language models and on-demand AI services may include:
[0135] The first determination module 401 is configured to determine the KPI information required by the user according to the AI service request data, the BERT model, the new service request, the KPIs in the knowledge base, and the entity values when receiving the AI service request data initiated by the user;
[0136] The second determination module 402 is configured to determine the deployed AI service network architecture by using the stitchable neural network and the AI service network architecture according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchy relationship, and the resource constraints;
[0137] The execution module 403 is configured to execute the user's task by using the deployed AI service network architecture to obtain the execution result corresponding to the user's task.
[0138] The first determination module 401 is specifically configured to perform text preprocessing on the AI service request data to obtain preprocessed data; use the preprocessed data to train the BERT model to obtain a trained BERT large language model; input the new service request into the trained BERT large language model so that the trained BERT large language model outputs the user's intention; search for the target KPI that matches the user's intention in the KPIs in the knowledge base, and extract the entity value corresponding to the target KPI through natural language processing technology; map the entity value to the corresponding target KPI to obtain the KPI information required by the user.
[0139] The first determination module 401 performs text preprocessing on the AI service request data to obtain preprocessed data, including: using the generative pre-training model to perform data augmentation on the AI service request data to obtain augmented data; cleaning the augmented data to obtain cleaned data, and the cleaned data is data without punctuation marks; using the BERT tokenizer to tokenize the cleaned data to obtain tokenized data; performing data annotation on the tokenized data to obtain preprocessed data.
[0140] The second determination module 402 is specifically configured to determine a new AI model by using the stitchable neural network according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchy relationship, and the resource constraints; determine the deployed AI service network architecture according to the new AI model and the AI service network architecture.
[0141] The second determination module 402 determines a new AI model according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchical relationship, and the resource constraints, including: selecting two AI models from the AI model library according to the KPI information required by the user; taking maximizing user satisfaction and on-demand service rate as the goal, and taking the service response time, the service availability, the AI model hierarchical relationship, and the resource constraints as the constraint conditions, establishing an optimization problem; using the dung beetle algorithm to solve the optimization problem to obtain the segmentation positions and deployment positions of the two AI models; segmenting and splicing the two AI models according to the stitchable neural network and the segmentation positions to obtain a new AI model.
[0142] The second determination module 402 determines the deployed AI service network architecture according to the new AI model and the AI service network architecture, including: using the deployment position to deploy the new AI model into the AI service network architecture to obtain the deployed AI service network architecture.
[0143] The second determination module 402 selects two AI models from the AI model library according to the KPI information required by the user, including: determining the corresponding tasks according to the KPI information required by the user, where the tasks include image classification tasks and sentiment analysis tasks; in the AI model library, selecting two AI models corresponding to the tasks, and the sizes of the two AI models are different.
[0144] It should be noted here that: the above description of the embodiments of the service demand recognition and on-demand AI service device based on the large language model is similar to the description of the embodiments of the service demand recognition and on-demand AI service method based on the large language model, and has beneficial effects similar to those of the embodiments of the service demand recognition and on-demand AI service method based on the large language model. For the technical details not disclosed in the embodiments of the service demand recognition and on-demand AI service device based on the large language model of the present invention, please refer to the description of the method embodiments of the present invention for understanding.
[0145] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A service demand identification and on-demand AI service method based on a large language model, characterized in that: include: When receiving AI service request data initiated by a user, determine the KPI information required by the user based on the AI service request data, the BERT model, the new service request, the KPI and entity values in the knowledge base; Determine the deployed AI service network architecture using the stitchable neural network and AI service network architecture based on the KPI information, AI model library, service response time, service availability, AI model hierarchical relationship and resource constraints required by the user; The deployed AI service network architecture is used to execute the user's task and obtain the execution result corresponding to the user's task.
2. The service demand identification and on-demand AI service method based on a large language model according to claim 1 is characterized in that: When receiving AI service request data initiated by a user, determining the KPI information required by the user according to the AI service request data, the BERT model, the new service request, the KPI and entity values in the knowledge base, including: Performing text preprocessing on the AI service request data to obtain preprocessed data; Using the preprocessed data, training the BERT model to obtain a trained BERT large language model; Inputting the new service request into the trained BERT large language model so that the trained BERT large language model outputs the user intention; Searching for a target KPI that matches the user's intention among the KPIs in the knowledge base, and extracting an entity value corresponding to the target KPI through natural language processing technology; The entity value is mapped to the corresponding target KPI to obtain the KPI information required by the user.
3. The service demand identification and on-demand AI service method based on a large language model according to claim 2 is characterized in that: The performing text preprocessing on the AI service request data to obtain preprocessed data includes: Using a generative pre-trained model, the AI service request data is enhanced to obtain enhanced data; Cleaning the enhanced data to obtain cleaned data, wherein the cleaned data is data that does not contain punctuation marks; Using a BERT word segmenter to segment the cleaned data, and standardizing the segmented data to obtain standardized data; The standardized data are annotated to obtain the preprocessed data.
4. The service demand identification and on-demand AI service method based on a large language model according to claim 1 is characterized in that: Determining the deployed AI service network architecture by using the stitchable neural network and AI service network architecture according to the KPI information, AI model library, service response time, service availability, AI model hierarchical relationship and resource limitation required by the user, including: Determine a new AI model using the stitchable neural network according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchy and the resource limitation; According to the new AI model and the AI service network architecture, the deployed AI service network architecture is determined.
5. The service demand identification and on-demand AI service method based on a large language model according to claim 4 is characterized in that: The determining a new AI model using the stitchable neural network according to the KPI information required by the user, the AI model library, the service response time, the service availability, the AI model hierarchy and the resource limitation includes: According to the KPI information required by the user, two AI models are selected from the AI model library; An optimization problem is established with the goal of maximizing user satisfaction and on-demand service rate, and with the service response time, the service availability, the AI model hierarchical relationship, and the resource limitation as constraints; The optimization problem is solved by using a dung beetle algorithm to obtain the segmentation position and deployment position of the two AI models; According to the stitchable neural network and the segmentation position, the two AI models are segmented and spliced to obtain the new AI model.
6. The service demand identification and on-demand AI service method based on a large language model according to claim 5 is characterized in that: The step of determining the deployed AI service network architecture according to the new AI model and the AI service network architecture includes: The new AI model is deployed into the AI service network architecture using the deployment location to obtain the deployed AI service network architecture.
7. The service demand identification and on-demand AI service method based on a large language model according to claim 5 is characterized in that: The expression of the optimization problem is: in, is user satisfaction, β is the stitching coefficient matrix, L is the location matrix of the AI model deployment layer, C1 is the constraint of the service response time, C2 is the constraint of the service availability, C3 is the constraint of the hierarchical relationship between the front and back AI models in stitching, C4 is the constraint of the deployment of a single AI model, C5 is the constraint of the stitching coefficient of the front model of the AI model, C6 is the constraint of the stitching coefficient matrix of the stitching model, C7 is the constraint of the stitching coefficient matrix of the single AI model deployment, C8 is the constraint of resource limitation, tq is the completion time of each new service request q, Dq is the preset deadline delay of each new service request q, is the availability of each new service request q, Aq is the preset availability requirement for each new service request q, The deployment location of the pre-model for each new service request q, The deployment location of the backend model for each new service request q, is the deployment location of a single AI model corresponding to the new service request q, is the stitching coefficient of the AI model pre-model associated with the new service request q, is the stitching coefficient of the AI model post-model associated with the new service request q, is the pre-stitching coefficient matrix of the stitching model associated with the new service request q, is the post-stitching coefficient matrix associated with the new service request q, is the stitching coefficient matrix of a single AI model deployment associated with the new service request q, For the front model, For the rear model, is a single model, Ci is the resource capacity, Ni is the number of nodes, mp is one of the two AI models, mr is the other of the two AI models, and i is the i-th network layer in the AI service network architecture.
8. The service demand identification and on-demand AI service method based on a large language model according to claim 5 is characterized in that: According to the KPI information required by the user, two AI models are selected from the AI model library, including: Determine corresponding tasks according to the KPI information required by the user, wherein the tasks include image classification tasks and sentiment analysis tasks; In the AI model library, the two AI models corresponding to the task are selected, and the two AI models have different sizes.
9. The service demand identification and on-demand AI service method based on a large language model according to claim 5 is characterized in that: The new service requests include computer vision task requests, natural language processing task requests, sound recognition task requests, reinforcement learning task requests and recommendation system task requests.
10. A service demand identification and on-demand AI service device based on a large language model, characterized in that: include: A first determination module is used to determine the KPI information required by the user based on the AI service request data, the BERT model, the new service request, the KPI and the entity value in the knowledge base when receiving the AI service request data initiated by the user; A second determination module is used to determine the deployed AI service network architecture using a stitchable neural network and an AI service network architecture according to the KPI information, AI model library, service response time, service availability, AI model hierarchy and resource constraints required by the user; The execution module is used to use the deployed AI service network architecture to execute the user's task and obtain the execution result corresponding to the user's task.