Plug-in-based AI digital human capability expansion method
Through the plug-in-based AI digital human ability expansion method, user profile is built using user information, relevant plug-ins are selected and preloaded, and plug-in functions are dynamically enabled, the problem of fixed traditional AI digital human ability is solved, and the flexible expansion and personalized customization of AI digital human ability is realized, and the user experience and interaction quality is improved.
Patent Information
- Application Number
- CN202510041423.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The capabilities of traditional AI digital people are preset fixed, and it is difficult to flexibly adapt to the needs of different users and changing task scenarios. The lack of dynamic expansion mechanisms limits its intelligent and personalized development.
Through the plug-in-based AI digital human capability expansion method, user profile is built using user information, relevant plug-ins are selected and preloaded, plug-in functions are enabled dynamically, AI digital human capabilities are expanded, and the ability matrix is output through the ability management system is used to customize a personalized task list.
It realizes flexible expansion and personalized customization of AI digital human capabilities, improves the quality of dynamic interaction, enhances the immersion and emotional resonance of user experience, and supports the in-depth penetration of AI digital humans in multiple fields of applications.
Smart Images

Figure CN120010948A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and AI digital human technology, and in particular to a plug-in-based AI digital human capability expansion method. Background Art
[0002] With the rapid development of artificial intelligence and virtual assistant technology, AI (Artificial Intelligence) digital humans have gradually become an important way of human-computer interaction. AI digital human technology breaks through the physical and time limitations of traditional human interaction, providing a virtual, intelligent and efficient interactive interface that can provide services and support when humans are not present immediately. AI digital humans can perform tasks without emotional bias, provide personalized and customized services, and gradually acquire interactive capabilities close to real humans through continuous learning and adaptation. This technology not only promotes the evolution of human-computer interaction, but also brings profound changes to industries such as medical care, education, entertainment, and finance, promotes the development of an intelligent society, and gives the digital world a more vivid and meaningful form of interaction, becoming a new field of technology.
[0003] However, the capabilities of traditional AI digital humans are often preset and fixed, making it difficult to flexibly adapt to the needs of different users and changing task scenarios. Although AI digital humans can perform certain standardized tasks, they often cannot provide personalized and efficient services when faced with complex and changing tasks. In addition, existing AI digital human systems usually lack a dynamic expansion mechanism, which makes it difficult for the system to quickly load and enable corresponding functions according to the different needs of users, limiting its intelligent and personalized development.
[0004] Plug-inization deconstructs the core capabilities of AI digital humans (such as speech recognition, sentiment analysis, natural language understanding and generation, virtual character behavior modeling and generation, etc.) into independent, loosely coupled functional units, giving the AI digital human system a high degree of flexibility and scalability. This design not only realizes the dynamic plug-in and unplug of various functional modules, allowing AI digital humans to be customized according to specific scenarios and needs, but also provides great convenience for cross-platform adaptation, thereby reducing the complexity of adaptation under diverse hardware, operating systems and application scenarios. In addition, this plug-in mode can also achieve high responsiveness in the functional evolution process of AI digital humans through automatic updates and iterations of intelligent plug-ins, quickly adapt to changes in user needs, and optimize functions based on real-time feedback. This method makes the interaction mode of AI digital humans more refined and personalized, and can express emotions, adjust behaviors and respond at multiple levels and dimensions, ultimately achieving an interactive effect closer to that of real humans, significantly improving the immersion and emotional resonance of user experience, and thus laying the foundation for the in-depth penetration of AI digital human technology in multi-field applications.
[0005] Based on the above background, there is an urgent need for an AI digital human capability expansion technology that can flexibly import plug-ins to improve the quality of dynamic interaction and provide customized services for users. Summary of the invention
[0006] The purpose of the present invention is to provide a plug-in-based AI digital human capability expansion method, which preliminarily builds a user portrait with the help of user information, preloads plug-ins and provides capability expansion strategies; the second purpose of the present invention is to provide a plug-in-based AI digital human capability management system.
[0007] The objective of the present invention is achieved through the following technical solutions:
[0008] A plug-in-based AI digital human capability expansion method includes the following steps:
[0009] Step S1: Extract the user information feature matrix based on the data submitted and interacted by the user, and construct a preliminary user portrait;
[0010] Step S2: Select relevant plug-ins based on the preliminary user profile, preload the plug-ins and automatically select the capability expansion strategy;
[0011] Step S3: Initialize the plug-in function, gradually enable the plug-in according to the capability expansion strategy and the user dialogue system, expand the capabilities of the AI digital human, and record the enabled function information;
[0012] Step S4: construct an AI digital human capability management system based on the enabled function information and output an AI digital human capability matrix;
[0013] Step S5: Utilize the document uploaded by the user to parse the user's task list, match the capability matrix, and customize the user's personalized task list.
[0014] Further, the step S1 specifically includes:
[0015] Step S101: Obtain the user's basic information input set X based on the user's personal information or the basic information filled in during the first interaction:
[0016] X={x1,x2,x3,x4,…,x n}
[0017] Among them, x1,x2,x3,x4,…,x n They represent the user's name, age, gender, position and other user information respectively, and n represents the number of information;
[0018] Step S102: Divide X into different types of data and process them separately. t ,x d ,xT ,x s Respectively represent categorical, numerical, date, and text data, x t ,x d ,x T ,x s ∈X; convert the above data into numerical values that can represent features, and combine them in order to generate a feature vector u containing user information features;
[0019] Step S103: Use the feature vector u to represent the general description of the user and construct a preliminary portrait of the user.
[0020] Further, the step S102 specifically includes:
[0021] (1) The collected categorical user information is converted into numerical representation using multi-hot encoding. For each category c∈C in the categorical variable C, C={c1,c2,c3,…,c m}, where m is the number of categories, for the i-th category c i The expression is:
[0022]
[0023] u l =[u1(x t ),u2(x t ),u3(x t ),…,u m (x t )]
[0024] where i∈[1,2,3,…,m], x t Represents the information input by the user of this category; concatenates the unique-hot encoding vectors of all categorical variables into a matrix u l ;
[0025] (2) Use the following formula to calculate the collected numerical data x d Perform standardization and input the data x that needs to be standardized d :
[0026]
[0027] Among them, μ, σ represent x d The mean and standard deviation of d represents the matrix after the normalization of all numerical data, and g represents the number of numerical data;
[0028] (3) The collected date data x T Convert it into time difference ΔT data and combine them to get the date data matrix:
[0029] ΔT1=|x T1 -T|
[0030] u T =[ΔT1, ΔT2, …, ΔT h ]
[0031]
[0032] Where T represents the base time set by the system, u T represents the matrix after all date data are converted, h represents the number of date data, and u is obtained after standardizing the date data T * ;
[0033] (4) Select the first text data x collected s1 , calculate the weight of each word in the text, when the word W k appears once in the text, then count(W k )=count(W k )+1, use each word in x s1 The number of occurrences in divided by x s1 The total number of words in the document is obtained, and the frequency of each word in the document is obtained. All words and calculated frequencies are stored in the corpus in the form of key-value pairs; the first Ω high-frequency words W are selected k Calculate its share x s1 The weights in construct the feature matrix u of each text s1 , k = 1, 2, ..., Ω; according to this method, the weight of high-frequency words in each text data is obtained to form the corresponding text data feature matrix u s , z represents the number of text data, u s =[u s1 ,u s2 ,…,u sz ];
[0034] (5) Concatenate all types of data matrices in order of type to obtain the feature matrix u:
[0035]
[0036] Among them, u represents the matrix containing all information characteristics of the user, and the number of all types of data is equal to the sum of all information data: n=m+g+h+z.
[0037] Further, the step S2 specifically includes:
[0038] Step S201: Based on the feature matrix u of the user's preliminary portrait, convert it into the user's current state space B, and define the action space A, where A represents the action of selecting a plug-in:
[0039]
[0040] a j ∈A
[0041] Among them, u i represents the features in u, Φ represents the threshold for selecting p features, U represents the feature matrix selected from u that is greater than the threshold Φ, μ U , σ U Respectively represent the mean and standard deviation of each feature, Vp represents the selection of the first p features from the selected features as the principal component matrix, and j represents the number of plug-ins;
[0042] Step S202: In the state space and action space, the fitness difference of the plug-in action is calculated step by step, and the optimization benefit brought by the evaluation of the state function of the currently selected plug-in combination is evaluated using the fitness difference:
[0043] ΔΠ=Π a+1 (b,a)-Π a (b,a)
[0044] Π a+1 -Π a =λ(β+μ(Π a (b',a'))-Π a (b,a))
[0045]
[0046] Among them, a+1 (b,a),Π a (b, a) represent the fitness values after and before updating action a in state s, respectively. λ represents the update step size. β represents the fitness value generated when executing action a. μ represents the fitness value parameter that will be increased by the next action. s' represents the next state entered after executing action a. a' represents the action selected in the next state. s q represents the state function at this time, q represents the round of selection, f(Π a q ) indicates that the qth round is the fitness value function for selecting the ath action, and κ indicates the importance of such state changes;
[0047] Step S203: Select s q A plug-in with a value greater than the threshold ν indicates that the plug-in is highly relevant to the user and is preloaded; an exchange operator is set to exchange or merge different parts of the existing operations, so that previously unconsidered areas can be explored in the solution space, the solution space is expanded, and the plug-in selection strategy is updated using the gradient descent method:
[0048]
[0049] in, represents the exchange operator in the qth round of selection, λ θ represents the step size of the selection strategy update, is the state function s q For the policy parameter θ a The partial derivative of q Represents the a Sensitivity to change.
[0050] Step S204: output the trained capability expansion strategy automatic selection model Ψ.
[0051] Further, the step S3 specifically includes:
[0052] Step S301: Initialize all available plug-ins and establish initial status and capability information for each plug-in, specifically including assigning an initial configuration to each plug-in:
[0053] a={id,name,state,description}
[0054] id=Category+Index
[0055] Where a represents a plug-in. The plug-in content includes the plug-in id, name, state value, and function description. The plug-in id is composed of a digital index based on the plug-in type Category and sequence number, and the leading zero is used to fill the gap. When state = 0, it means that the plug-in is not enabled, and when state = 1, it means that the plug-in is enabled.
[0056] Step S302: clustering the user conversation content using the text clustering model, dividing the user conversation content into categories η, performing strategy selection on the plug-in of category η using the capability extension strategy automatic selection model Ψ, and setting the plug-in status to 1;
[0057] Step S303: Identify the plug-in whose plug-in status is 1, execute the plug-in function, display the corresponding card according to the returned category value, and generate a vector set F containing all enabled functions.
[0058] Furthermore, the specific process of classifying the user conversation content in step S302 is as follows:
[0059] (1) Collect user conversation information and convert it into a numerical feature vector Y. Each conversation sample is represented as a feature vector y i ,Y=[y1,y2,…,y i ,…,y M ],Y∈R M×D, M represents the number of dialogue samples, D is the feature dimension of each sample, and R represents the feature vector matrix;
[0060] (2) Initialize the text clustering model parameters: set the number of clusters, that is, the number of clusters to η = 1, the convergence accuracy ε = 0.001, and the number of iterations o = 0;
[0061] (3) Initialize the prior probability ω i =1 / η;
[0062] (4) Given the prior probability, calculate the posterior probability that each sample y belongs to each cluster:
[0063]
[0064] in, represents the posterior probability that the Mth sample comes from the i-th cluster at the oth iteration, μ and σ represent the mean and variance respectively, e represents the exponential operation, and π represents pi;
[0065] (5) Update the parameter value of each cluster according to the posterior probability. The specific steps are as follows:
[0066]
[0067] (6) Determine whether the model has converged, and calculate whether the change of each parameter is less than the set convergence precision. If the change of any parameter is greater than the convergence precision, the number of iterations o = o + 1, and return to step (4); otherwise, go to step (7);
[0068] (7) At this time, η=η+1. If η≤l, go to step (3) to ensure that the level of the category is within level l;
[0069] (8) Calculate the silhouette coefficient of samples when η = 1, 2, 3, ..., l by comparing different values of η. The silhouette coefficient is the similarity between each sample and the samples in the same cluster divided by the similarity between each sample and the nearest cluster, so as to select the η with the largest silhouette coefficient as the optimal category;
[0070] (9) Return the posterior probability r of each dialogue i (M) and the category η.
[0071] Further, the step S4 specifically includes:
[0072] Step S401: construct an AI digital human capability management model Θ according to the enabled function vector set F, the model including: a feature vector embedding model, a convolutional neural network model, a multi-layer perceptron model, a feedforward neural network model, and a capability management model;
[0073] Step S402: Input the function vector into the AI digital human capability management model Θ, obtain the capability scores of the AI digital human for different tasks, and generate the AI digital human capability matrix D, where rows represent different tasks and columns represent different capabilities:
[0074]
[0075] Among them, H represents the number of tasks, I represents the number of capabilities, and d ij Represents the AI digital human's score on the jth ability of the i-th task.
[0076] Furthermore, the specific structure of the AI digital human capability management model θ constructed in step S401 is as follows:
[0077] The input end of the AI digital human capability management model Θ is used as the input end of the feature vector embedding model; the output end of the feature vector embedding model is connected to the input end of the convolutional neural network; the output end of the convolutional neural network is connected to the input end of the multilayer perceptron model; the output end of the multilayer perceptron model is connected to the input end of the feedforward neural network model; the output end of the feedforward neural network model is connected to the input end of the capability management model; the output end of the capability management model is used as the output end of the AI digital human capability management model Θ;
[0078] The feature vector embedding model includes: a fully connected layer, a Relu activation function layer and a batch processing layer; the input end of the feature vector embedding model is used as the input end of the fully connected layer; the output end of the fully connected layer is used as the input end of the Relu activation function layer; the output end of the Relu activation function layer is connected to the input end of the batch processing layer; the output end of the batch processing layer is used as the output end of the feature vector embedding model and is connected to the convolution layer of the convolutional neural network model;
[0079] The convolutional neural network model includes: a convolutional layer, a pooling layer and an activation function layer; the input end of the convolutional neural network model serves as the input end of the convolutional layer; the output end of the convolutional layer is connected to the input end of the pooling layer; the output end of the pooling layer is connected to the input end of the activation function layer; the output end of the activation function layer serves as the output end of the convolutional neural network model and is connected to the fully connected layer of the multi-layer perceptron;
[0080] The multi-layer perceptron model comprises: a fully connected layer, a Relu activation function layer and a Dropout layer; the input end of the multi-layer perceptron model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the Relu activation function layer; the output end of the Relu activation function layer is connected to the input end of the Dropout layer; the output end of the Dropout layer serves as the output end of the multi-layer perceptron model and is connected to the fully connected layer of the feedforward neural network model;
[0081] The feedforward neural network model includes: a fully connected layer, a feedforward neural network layer and a Dropout layer; the input end of the feedforward neural network model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the feedforward neural network layer; the output end of the feedforward neural network layer is connected to the input end of the Dropout layer; the output end of the Dropout layer serves as the output end of the feedforward neural network model and is connected to the fully connected layer of the capability management model;
[0082] The capability management model includes: a fully connected layer and a Softmax activation function layer; the input end of the capability management model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the Softmax activation function layer; and the Softmax activation function layer serves as the output end of the capability management model.
[0083] Furthermore, the scoring formula for calculating the ability corresponding to each task in step S402 is as follows:
[0084] E1=Relu(W1·F+b1)
[0085]
[0086] E4=Relu(W3·Mask(E3,ρ)+b4)
[0087] E5=Relu(W4·Mask(E4,ρ)+b5)
[0088] E6=Softmax(W5·E5+b6)
[0089] Wherein, E1 represents the output of the feature vector after Relu(·) activation, W1 represents the bias matrix of the fully connected layer, and b1 represents the bias parameter of the fully connected layer; E2 represents the output of the feature vector embedding model, δ represents the learnable scaling parameter, b2 represents the bias parameter of the batch layer, μ and σ represent the mean and variance, and ξ represents the parameter to prevent zero division errors; E3 represents the output of the convolutional layer, W2 and b3 represent the bias matrix and bias parameter of the convolutional layer, LeakyRelu(·) is the activation function, and len(·) represents the length function; E4 represents the output of the multi-layer perceptron, W3 and b4 represent the bias matrix and bias parameter of the multi-layer perceptron, Mask(·) represents the mask function, indicating whether neurons are discarded, and ρ represents the discard rate, which is usually between 0 and 1; E5 represents the output of the feedforward network layer, W4 and b5 represent the bias matrix and bias parameter of the feedforward network layer; E6 is the capability score of the task, W5 and b6 represent the bias matrix and bias parameter of the capability management module, and Softmax(·) is the activation function.
[0090] Further, the step S5 specifically includes:
[0091] Step S501: Standardize the task documents uploaded by the user and convert them into a parsable text format, remove redundant spaces, line breaks, and special characters that interfere with parsing, use periods as the separator for each sentence, and use line breaks as the separator for each task;
[0092] Step S502: Perform task keyword recognition on each sentence after preprocessing, use regular expressions to extract the complete task description, and parse the tasks into structured data S1 in order. After parsing all the tasks in the document, a task dictionary list S = [S1, S2, ..., S n ], n represents the number of tasks;
[0093] Step S503: Match the task dictionary list with the tasks in the capability matrix, calculate the sum of the capability values required for the task in the capability matrix as the score of the task, and output the user personalized task list in the order of the scores.
[0094] The beneficial effects of the present invention include:
[0095] (1) Through systematic feature engineering, different types of data are converted into unified numerical feature vectors to ensure that the converted data can accurately express the multi-dimensional characteristics of users, providing rich data support for subsequent capability expansion and personalized task customization, and ensuring that the expanded AI digital human services can better meet the real needs of users;
[0096] (2) AI digital humans select capability plug-ins that match user needs based on the user feature matrix and load the capability plug-ins in advance, so that AI digital humans have the ability to dynamically adjust in different situations, avoiding the subjectivity and fixedness of the capabilities given to traditional AI digital humans, and improving personalization and flexibility;
[0097] (3) The capability expansion strategy is dynamically adjusted and gradually enabled based on the user's immediate needs and conversations. This not only avoids loading too many unnecessary functions at one time, but also enables the required capabilities accurately according to the user's interaction mode and task requirements, thereby achieving on-demand expansion of AI digital human capabilities. It is not necessary to grant all capabilities, which improves the responsiveness and immediacy of the system.
[0098] (4) The AI digital human capability management system combines artificial intelligence technologies such as feature vector embedding models, convolutional layers, and pooling layers to effectively extract deep feature information from raw data and convert this information into a more intuitive and accurate capability representation. Through this deep feature expression, AI digital humans can obtain a more comprehensive task understanding and capability assessment, avoiding simple rule matching and improving the intelligence level of AI digital humans;
[0099] (5) By accepting the documents provided by users, tasks are automatically extracted from the documents, structured, intelligently analyzed, and associated with the capability matrix, thus realizing the automated processing from task input to task recommendation. This automated processing of business processes reduces manual intervention, improves the work efficiency of AI digital humans, and avoids the problem of inaccurate task matching caused by manual errors;
[0100] (6) By calculating the ability score of each task, the AI digital human can automatically adjust its ability allocation according to the current task requirements and give priority to its strongest ability in different tasks. This dynamic ability adjustment enables the AI digital human to adapt to complex working environments and optimize the scheduling of abilities in real time according to changes in tasks.
[0101] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description and the preceding claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0103] Figure 1 It is a schematic diagram of the overall workflow of the present invention;
[0104] Figure 2 This is a schematic diagram of a preliminary user portrait after feature extraction in the present invention;
[0105] Figure 3 This is a flow chart of classifying user conversation content using a Gaussian mixture model in the present invention;
[0106] Figure 4 This is a structural diagram of the AI digital human capability management system of the present invention;
[0107] Figure 5 This is a schematic diagram of the AI digital human capability matrix generated by the present invention using the data of the embodiment;
[0108] Figure 6 This is a diagram of the AI digital human capability management function interface of the present invention;
[0109] Figure 7 This is a diagram of the user task function interface analyzed by the present invention. DETAILED DESCRIPTION
[0110] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.
[0111] A plug-in-based AI digital human capability expansion method of the present invention, such as Figure 1 As shown, the following steps are included:
[0112] Step S1: Extract the user information feature matrix based on the data submitted and interacted by the user, and construct a preliminary user portrait;
[0113] Step S2: Select relevant plug-ins based on the preliminary user profile, preload the plug-ins and automatically select the capability expansion strategy;
[0114] Step S3: Initialize the plug-in function, gradually enable the plug-in according to the capability expansion strategy and the user dialogue system, expand the capabilities of the AI digital human, and record the enabled function information;
[0115] Step S4: construct an AI digital human capability management system based on the enabled function information and output an AI digital human capability matrix;
[0116] Step S5: Utilize the document uploaded by the user to parse the user's task list, match the capability matrix, and customize the user's personalized task list.
[0117] The specific steps of the above method will be further described below through a specific embodiment.
[0118] In this embodiment, step S1 specifically includes the following steps:
[0119] Step S101: Obtain the user's basic information input set X based on the user's personal information or the basic information filled in during the first interaction:
[0120] The input received by this embodiment is as follows: X = {"Name":"Yan","Age":35,"Gender":"Male","Position":"Senior Engineer","Joining Date="2015-08-15","Date of Birth="1987-09-15","Marriage Date="2015-08-20","Education":"Bachelor's Degree","Graduated School":"Tsinghua University","Salary":18000,"Cities Lived":["Beijing","Shanghai","Guangzhou"],"Hobbies":"Hobbies":["Programming","Running","Music","Watching Movies","Mountain Climbing"],"Last Business Trip Location":["Shanghai","Tokyo"],"Technology Stack":["Python","SQL"],"Years of Work Experience":10,"Family Members" :["Wife","Son","Parents","Brother"],"Signature":"There are only 10 kinds of people in the world, one who understands binary and one who doesn't.","Introduction":"I am a full-stack development engineer, focusing on in-depth exploration of programming and technology. In the past few years, I have accumulated rich development experience, especially in front-end and back-end development. Familiar with modern front-end frameworks such as React and Vue.js, I also have deep back-end development experience and am good at using frameworks such as Express and Django for server-side development. In my daily work, I am not only keen on solving technical problems and promoting the progress of projects, but also pay great attention to teamwork and experience sharing. In addition to programming, I also like running. Inspiration often emerges during running, which helps me have new inspirations in the fields of programming and technology."}
[0121] Step S102: Divide X into different types of data and process them separately. t ,x d ,x T ,x s Respectively represent categorical, numerical, date, and text data, x t ,x d ,x T ,x s ∈X; x t ={"Gender":"Male","Position":"Senior Engineer","Education":"Bachelor's degree","University graduated from":"Tsinghua University","Cities lived in":["Beijing","Shanghai","Guangzhou"],"Hobbies":["Programming","Running","Music","Watching movies","Mountain climbing"],"Last business trip location":["Shanghai"],"Technology stack":["Python","SQL"],"Family members":["Wife","Son","Parents","Brother"]};x d ={"age":35,"salary":18000,"years of work":10}; x T={"Joining Date="2015-08-15","Birth Date="1987-09-15",Marriage Date="2015-08-20"};x s ={"Personal Signature":"There are only 10 kinds of people in the world, one who understands binary and one who doesn't.","Personal Introduction":"I am a full-stack development engineer who focuses on in-depth exploration of programming and technology. In the past few years, I have accumulated rich development experience, especially in front-end and back-end development. Familiar with modern front-end frameworks such as React and Vue.js, I also have deep back-end development experience and am good at using frameworks such as Express and Django for server-side development. In my daily work, I am not only keen on solving technical problems and promoting the progress of projects, but also pay great attention to teamwork and experience sharing. In addition to programming, I also like running. Inspiration often emerges during running, which helps me have new inspiration in the fields of programming and technology."}; Convert the above data into numerical values that can represent features, and combine them in order to generate a feature vector u containing user information features. The specific steps are as follows:
[0122] (1) The collected categorical user information is converted into numerical representation using multi-hot encoding. For each category c∈C in the categorical variable C, C={c1,c2,c3,…,c9}, where c1~c9 represent gender, position, education, university, city lived in, hobbies, most recent business trip location, technology stack, and family members respectively; the specific types of each category are as follows: c1={male, female}, c2={senior engineer, junior engineer, intermediate engineer, project manager, senior manager}, c3={master, doctor, undergraduate, junior college}, c4={Tsinghua University, Peking University, Fudan University, Shanghai Jiaotong University}, c5={Beijing, Shenzhen, Shanghai, Guangzhou, Chengdu, Mianyang}, c6={programming, running, travel, music, reading, watching movies, mountain climbing}, c7={Shanghai, Tokyo, Beijing, Guangzhou}, c8={Python, Java, C++, JavaScript, Go, SQL}, c9={wife, son, parents, brothers, sisters, grandfather, grandmother, friend};
[0123] For the i-th category c i The expression is:
[0124]
[0125] u l =[u1(x t ),u2(x t ),u3(x t ),…,u m (x t )]
[0126] where i∈[1,2,3,…,m], x t Represents the information input by the user of this category; we get u1=[1,0],u2=[1,0,0,0,0],u3=[0,0,1,0],u4=[1,0,0,0],u5=[1,0,1,1,0,0],u6=[1,1,0,1,0,1,1],u7=[1,0,0,0],u8=[1,0,0,0,0,1],u9=[1,1,1,1,0,0,0,0];concatenate the one-hot encoding vectors of all categorical variables into a matrix u l =[1,0,1,0,0,0,0,0,0,1,0,1,0,0,0,1,0,1,1,0,0,1,1,0,1,0,1,1,1,0,0,0,1,0,0,0,0,1,1,1,1,0,0,0,0];
[0127] (2) Use the following formula to calculate the collected numerical data x d Standardization is performed, where age is limited to 0-100, salary is limited to 2000-50000, and working years are limited to 1-40. In this case, g = 3:
[0128]
[0129] Among them, μ, σ represent x d The mean and standard deviation of age x' can be obtained according to the formula: d1 =0.35, salary x' d2 =0.33, working years x' d3 =0.23, and the standardized matrix u is obtained d =[0.35,0.33,0.23];
[0130] (3) The collected date data x T Converted into time difference ΔT data, and combined to get a date data matrix. The system base time is 1900-1-1, and the base time is used in the following:
[0131]
[0132] u T =[ΔT1, ΔT2, …, ΔT h ]
[0133]
[0134] Among them, T represents the reference time set by the system. According to the formula, the ΔT of the employment date, birth date, and marriage date are 42210, 32033, 42215, and uT =[42210,32033,42215], the unit is day. In order to be consistent with other types of features, the date data is standardized to obtain u T * =[0.99,0,1];
[0135] (4) Select the first collected text data "personal signature", ignore the stop words, and get the top four words with the highest frequency, which are "world", "binary", "person", and "10 kinds", with probabilities of 0.11, 0.11, 0.11, and 0.11, respectively; select the second collected text data "personal introduction", ignore the stop words, and get the top four words with the highest frequency, which are "development", "programming", "technology", and "experience", with probabilities of 0.49, 0.37, 0.37, and 0.37, respectively; in addition, store all words and their corresponding probabilities in the corpus;
[0136] Comprehensively obtain the characteristics u of text data s =[0.11,0.11,0.11,0.11,0.49,0.37,0.37,0.37];
[0137] (5) Using the name as the identifier, concatenate all types of data matrices in order of type to obtain the feature matrix u:
[0138]
[0139] Where u is a vector containing all the information characteristics of the user. The number of all types of data is equal to the sum of all the information data: n=m+g+h+z. The feature vector of user “Yan” is [1,0,1,0,0,0,0,0,0,1,0,1,0,0,0,1,0,1,1,0,0,1,1,0,1,0,1,1,1,1,0,0,1,0,0,0,0,1,1,1,1,0,0,0,0,0,1,1,1,1,0,0,0,0,0.35,0.33,0.23,0.99,0,1,0.11,0.11,0.11,0.11,0.49,0.37,0.37,0.37];
[0140] Step S103: Use the feature vector u to represent the general description of the user and construct a preliminary profile of the user:
[0141] According to the feature vector, we can get that the gender of the user "Yan" is male, his position is senior engineer, his education is undergraduate, he graduated from Tsinghua University, he lives in Beijing, Shanghai, and Guangzhou, his hobbies are programming, running, music, watching movies, and mountaineering, and his most recent business trip was to Shanghai. His technology stack includes Python and SQL, and his family members include his wife, son, parents, and brothers. His age is standardized to 0.35, which means he is in the lower middle age group. His salary income is standardized to 0.33, which means he is in the lower middle income group. His years of work are standardized to 0.23, which means he has a shorter working experience. His employment time is relatively far from the benchmark time, his date of birth is relatively close to the benchmark time, and his marriage date is far from the benchmark time. The focus of his signature is on "world", "binary", "10 types", and "people". The focus of his personal introduction is on "development", "programming", "technology", and "experience". The user's initial portrait is as follows: Figure 2 shown.
[0142] Step S2 specifically includes the following steps:
[0143] Step S201: Based on the feature matrix u of the user's preliminary portrait, convert it into the user's current state space B, and define the action space A, where A represents the action of selecting a plug-in:
[0144]
[0145] a j ∈A
[0146] Among them, u i represents the features in u, Φ represents the threshold for selecting p features, U represents the feature matrix selected from u that is greater than the threshold Φ, μ U , σ U Respectively represent the mean and standard deviation of each feature, Vp represents the selection of the first p features from the selected features as the principal component matrix, and j represents the number of plug-ins;
[0147] Set Φ=0.25 and get the final 8-dimensional state space vector B=[0.92,-0.12,0.56,1.02,-0.73,0.43,-0.18,0.61].
[0148] Step S202: In the state space and action space, the fitness difference of the plug-in action is calculated step by step, and the optimization benefit brought by the evaluation of the state function of the currently selected plug-in combination is evaluated using the fitness difference:
[0149] ΔΠ=Π a+1 (b,a)-Π a (b,a)
[0150] Π a+1 -Π a=λ(β+μ(Π a (b',a'))-Π a (b,a))
[0151]
[0152] Among them, a+1 (b,a),Π a (b, a) represent the fitness values after and before updating action a in state s, respectively. λ represents the update step size. β represents the fitness value generated when executing action a. μ represents the fitness value parameter that will be increased by the next action. s' represents the next state entered after executing action a. a' represents the action selected in the next state. s q represents the state function at this time, q represents the round of selection, f(Π a q ) indicates that the qth round is the fitness value function for selecting the ath action, and κ indicates the importance of such state changes;
[0153] Step S203: Select s q A plug-in with a value greater than the threshold ν indicates that the plug-in is highly relevant to the user and is preloaded; an exchange operator is set to exchange or merge different parts of the existing operations, so that previously unconsidered areas can be explored in the solution space, the solution space is expanded, and the plug-in selection strategy is updated using the gradient descent method:
[0154]
[0155] in, represents the exchange operator in the qth round of selection, λ θ represents the step size of the selection strategy update, is the state function s q For the policy parameter θ a The partial derivative of q Represents the a Sensitivity to change.
[0156] Step S204: output the trained capability expansion strategy automatic selection model Ψ.
[0157] Step S3 specifically includes the following steps:
[0158] Step S301: Initialize all available plug-ins and establish initial status and capability information for each plug-in, specifically including assigning an initial configuration to each plug-in:
[0159] a1={"id":"study_01","name":"Study Suggestions","state":0,"description":"This plug-in provides suggestions on learning methods and strategies. Based on the user's learning progress, subjects, learning habits and other data, it provides personalized suggestions for optimizing learning plans, improving learning efficiency and increasing learning motivation. Functions include: analyzing learning habits, formulating personalized learning plans, recommending learning resources, regularly tracking learning progress and providing feedback on improvement plans."};
[0160] a2={"id":"emergency_01","name":"Emergency Escape Guide","state":0,"description":"This plugin provides users with escape advice in emergency situations. According to different emergency situations (such as fire, earthquake, violent attack, etc.), it provides detailed escape routes, emergency measures and safety precautions. Functions include: designing escape route maps for various emergency situations, providing emergency contact information, and prompting necessary emergency tools and behaviors."};
[0161] a3={"id":"drawing_01","name":"Smart Drawing","state":0,"description":"This plugin helps users automatically generate or optimize graphics through AI. It supports multiple graphic forms such as charts, illustrations, flowcharts, etc. Users can upload sketches or enter descriptions, and the plugin will generate professional images based on these inputs. Functions include: automatically generate charts, artistic style drawings, flowcharts and schematics, optimize design details, and adjust graphic style and content according to needs."};
[0162] a4={"id":"programming_01","name":"Programming Assistant","state":1,"description":"This plug-in provides users with programming assistance, including code generation, debugging, optimization suggestions and learning resources. It supports multiple programming languages, can understand the code intent and provide instant feedback. Functions include: automatic code completion, code error troubleshooting, performance optimization suggestions, programming skills recommendations and related tutorial resources."};
[0163] a5={"id":"teaching_01","name":"Teaching Plan","state":0,"description":"This plug-in designs personalized teaching plans based on user needs. It provides the best teaching plan, teaching method and evaluation method based on teaching content, target group and time arrangement. Functions include: providing teaching outlines according to user needs, formulating detailed teaching schedules, designing classroom activities and evaluation plans to ensure that the course content is comprehensive and meets the teaching objectives."};
[0164] a6={"id":"legal_01","name":"Legal Consulting","state":1,"description":"This plug-in provides users with legal consulting services, covering civil law, contract law, labor law, intellectual property and other fields. It provides concise legal opinions, legal document templates and litigation process guides. Functions include: quick answers to legal questions, generation of legal documents, case analysis and handling suggestions, and recommendations for courts and law firms."};
[0165] a7={"id":"leave_01","name":"Leave Application","state":0,"description":"This plugin helps users quickly fill out and submit leave applications. It supports multiple leave types, such as sick leave, personal leave, annual leave, etc. According to the regulations of the company or school, the application form is automatically generated and sent to the relevant reviewer. Functions include: automatically filling out the leave form, viewing the leave history, tracking the leave approval process, etc."};
[0166] a8={"id":"diet_01","name":"Dietary Health Analysis","state":0,"description":"This plug-in provides a dietary health analysis report based on the user's dietary records and health data. By analyzing the nutritional components in the diet, it provides healthy dietary suggestions and helps users improve their eating habits. Functions include: analyzing daily intake of calories, protein, fat and other nutrients, providing dietary plans and suggestions, tracking changes in health indicators and performing personalized dietary optimization."}.
[0167] When state=0, it means the plugin is not enabled, and when state=1, it means the plugin is enabled;
[0168] Step S302: clustering the user conversation content using a Gaussian mixture model to divide the user conversation content into categories η, such as Figure 3 As shown, the plug-in of category η uses the capability extension strategy to automatically select the model Ψ to select the strategy, and sets the plug-in state to 1;
[0169] Step S303: Identify the plug-in whose plug-in status is 1, execute the plug-in function, display the corresponding card according to the returned category value, and generate a vector set F containing all enabled functions.
[0170] (1) Collect user conversation information and convert it into a numerical feature vector Y. Each conversation sample is represented as a feature vector y i ,Y=[y1,y2,…,y i ,…,y M ],Y∈R M×D, M represents the number of dialogue samples, D is the feature dimension of each sample, and R represents the feature vector matrix;
[0171] (2) Initialize the parameters of the Gaussian mixture model: set the number of sub-Gaussian distributions, that is, the number of clusters, to η = 1, the convergence accuracy ε = 0.001, and the number of iterations o = 0;
[0172] (3) Initialize the prior probability ω i =1 / η;
[0173] (4) Given the prior probability, calculate the posterior probability that each sample y belongs to each sub-Gaussian distribution:
[0174]
[0175] in, represents the posterior probability that the Mth sample comes from the i-th sub-Gaussian distribution at the o-th iteration, μ and σ represent the mean and variance respectively, e represents the exponential operation, and π represents pi;
[0176] (5) Update the parameter value of each sub-Gaussian distribution according to the posterior probability. The specific steps are as follows:
[0177]
[0178] (6) Determine whether the model has converged, and calculate whether the change of each parameter is less than the set convergence precision. If the change of any parameter is greater than the convergence precision, the number of iterations o = o + 1, and return to step (4); otherwise, go to step (7);
[0179] (7) At this time, η = η + 1. If η ≤ 5, go to step (3) to ensure that the level of the category is within 5 levels;
[0180] (8) Calculate the silhouette coefficient of samples when η = 1, 2, 3, ..., l by comparing different values of η. The silhouette coefficient is the similarity between each sample and the samples in the same cluster divided by the similarity between each sample and the nearest cluster, so as to select the η with the largest silhouette coefficient as the optimal category;
[0181] (9) Return the posterior probability r of each dialogue i (M) and the category η.
[0182] Step S303: Identify the plug-in whose plug-in status value is 1, execute the plug-in function, and display the corresponding card according to the returned category value. When the return value of η is 1, text is returned; when the return value of η is 2, audio is returned; when the return value of η is 3, video is returned; when the return value of η is 4, a table is returned; when the return value of η is 5, an image is returned; and a vector set F containing all enabled functions is generated.
[0183] Step S4 specifically includes the following steps:
[0184] Step S401: According to the enabled function vector set F, an AI digital human capability management model Θ is constructed. The model includes: a feature vector embedding model, a convolutional neural network model, a multi-layer perceptron model, a feedforward neural network model, and a capability management model. The specific internal structure is as follows: Figure 4 As shown;
[0185] The input end of the AI digital human capability management model Θ is used as the input end of the feature vector embedding model; the output end of the feature vector embedding model is connected to the input end of the convolutional neural network; the output end of the convolutional neural network is connected to the input end of the multilayer perceptron model; the output end of the multilayer perceptron model is connected to the input end of the feedforward neural network model; the output end of the feedforward neural network model is connected to the input end of the capability management model; the output end of the capability management model is used as the output end of the AI digital human capability management model Θ;
[0186] The feature vector embedding model includes: a fully connected layer, a Relu activation function layer and a batch processing layer; the input end of the feature vector embedding model is used as the input end of the fully connected layer; the output end of the fully connected layer is used as the input end of the Relu activation function layer; the output end of the Relu activation function layer is connected to the input end of the batch processing layer; the output end of the batch processing layer is used as the output end of the feature vector embedding model and is connected to the convolution layer of the convolutional neural network model;
[0187] The convolutional neural network model includes: a convolutional layer, a pooling layer and an activation function layer; the input end of the convolutional neural network model serves as the input end of the convolutional layer; the output end of the convolutional layer is connected to the input end of the pooling layer; the output end of the pooling layer is connected to the input end of the activation function layer; the output end of the activation function layer serves as the output end of the convolutional neural network model and is connected to the fully connected layer of the multi-layer perceptron;
[0188] The multi-layer perceptron model comprises: a fully connected layer, a Relu activation function layer and a Dropout layer; the input end of the multi-layer perceptron model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the Relu activation function layer; the output end of the Relu activation function layer is connected to the input end of the Dropout layer; the output end of the Dropout layer serves as the output end of the multi-layer perceptron model and is connected to the fully connected layer of the feedforward neural network model;
[0189] The feedforward neural network model includes: a fully connected layer, a feedforward neural network layer and a Dropout layer; the input end of the feedforward neural network model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the feedforward neural network layer; the output end of the feedforward neural network layer is connected to the input end of the Dropout layer; the output end of the Dropout layer serves as the output end of the feedforward neural network model and is connected to the fully connected layer of the capability management model;
[0190] The capability management model includes: a fully connected layer and a Softmax activation function layer; the input end of the capability management model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the Softmax activation function layer; and the Softmax activation function layer serves as the output end of the capability management model.
[0191] Step S402: Input the function vector into the AI digital human capability management model DAM to obtain the capability scores of the AI digital human for different tasks. The rows correspond to different tasks and the columns correspond to different capabilities. The tasks are in the following order in rows: providing suggestions, designing solutions, data analysis, and information consultation; the capabilities are in the following order in columns: learning suggestions, emergency escape guides, smart drawing, programming assistants, teaching plans, legal consultations, leave applications, and dietary health analysis. Generate an AI digital human capability matrix as shown in the figure below: Figure 5 shown.
[0192] Furthermore, the scoring formula for calculating the ability corresponding to each task in step S402 is as follows:
[0193] E1=Relu(W1·F+b1)
[0194]
[0195] E4=Relu(W3·Mask(E3,ρ)+b4)
[0196] E5=Relu(W4·Mask(E4,ρ)+b5)
[0197] E6=Softmax(W5·E5+b6)
[0198] Wherein, E1 represents the output of the feature vector after Relu(·) activation, W1 represents the bias matrix of the fully connected layer, and b1 represents the bias parameter of the fully connected layer; E2 represents the output of the feature vector embedding model, δ represents the learnable scaling parameter, b2 represents the bias parameter of the batch layer, μ and σ represent the mean and variance, and ξ represents the parameter to prevent zero division errors; E3 represents the output of the convolutional layer, W2 and b3 represent the bias matrix and bias parameter of the convolutional layer, LeakyRelu(·) is the activation function, and len(·) represents the length function; E4 represents the output of the multi-layer perceptron, W3 and b4 represent the bias matrix and bias parameter of the multi-layer perceptron, Mask(·) represents the mask function, indicating whether neurons are discarded, and ρ represents the discard rate, which is usually between 0 and 1; E5 represents the output of the feedforward network layer, W4 and b5 represent the bias matrix and bias parameter of the feedforward network layer; E6 is the capability score of the task, W5 and b6 represent the bias matrix and bias parameter of the capability management module, and Softmax(·) is the activation function.
[0199] Step S5 specifically includes the following steps:
[0200] Step S501: Standardize the task documents uploaded by the user and convert them into a parsable text format, remove redundant spaces, line breaks, and special characters that interfere with parsing, use periods as the separator for each sentence, and use line breaks as the separator for each task;
[0201] Step S502: perform task keyword recognition on the preprocessed tasks, use regular expressions to extract complete task descriptions, and parse the tasks into structured data in sequence S1 = {"task_id":01, "task_description": provide learning suggestions for employees based on their programming abilities, task_features:{"programming assistant", "provide suggestions"}}, S2 = {"task_id":02, "task_description": I need to design a teaching plan for corporate law, task_features:{"legal consultation", "teaching plan"}}, S3 = {"task_id":03, "task_description": I want to draw a pie chart based on the health status of employees, task_features:{"diet health analysis", "intelligent drawing"}}, after parsing all the tasks in the document, a dictionary list S = [S1, S2, S3] of tasks is constructed;
[0202] Step S503: Match the task dictionary list with the tasks in the capability matrix, calculate the sum of the capability values required for the task in the capability matrix as the score of the task, and obtain the scores of the three tasks as 1.5, 1.6, and 1.2 respectively; output the user personalized task list [S2, S1, S3] in the order of the scores.
[0203] In this embodiment, the AI digital human capability management function interface is as shown in the figure Figure 6 As shown in the figure, the AI digital human processes user documents and decomposes the functional interface of user tasks. Figure 7 As shown in the figure, by calculating the ability score of the AI digital human plug-in, it is possible to ensure that the task matches the ability of the AI digital human and improve execution efficiency. The ability score can also dynamically adjust the difficulty of the task, support the gradual improvement of the ability of the AI digital human, optimize task allocation, and enhance the user experience and sense of participation. In addition, using plug-ins to expand the capabilities of AI digital humans can significantly improve their flexibility and adaptability, enabling them to dynamically load or modify specific functional modules according to different task requirements. This plug-in design not only improves the work efficiency and professionalism of AI digital humans, but also facilitates the maintenance and update of the system and reduces development costs. Through plug-ins, AI digital humans can be applied across fields, support needs in different scenarios, ensure their continuous innovation and expansion, and thus improve user experience and meet changing needs.
[0204] It should be appreciated that embodiments of the present invention may be implemented or enforced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method may be implemented in a computer program using standard programming techniques, including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may be implemented in an assembly or machine language. In any case, the language may be a compiled or interpreted language. In addition, the program may be run on a programmed ASIC for this purpose.
[0205] Furthermore, the operations of the processes described herein may be performed in any relevant order unless otherwise indicated herein or otherwise clearly contradicted by context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that may be executed by one or more processors.
[0206] Further, the method can be implemented in any type of computing platform that is operably connected to the relevant, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, a RAM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A plug-in-based AI digital human capability expansion method, characterized by: The method comprises the following steps: Step S1: Extract the user information feature matrix based on the data submitted and interacted by the user, and construct a preliminary user portrait; Step S2: Select relevant plug-ins based on the preliminary user profile, preload the plug-ins and automatically select the capability expansion strategy; Step S3: Initialize the plug-in function, gradually enable the plug-in according to the capability expansion strategy and the user dialogue system, expand the capabilities of the AI digital human, and record the enabled function information; Step S4: construct an AI digital human capability management system based on the enabled function information and output an AI digital human capability matrix; Step S5: Utilize the document uploaded by the user to parse the user's task list, match the capability matrix, and customize the user's personalized task list.
2. According to claim 1, a plug-in-based AI digital human capability expansion method is characterized by: The step S1 specifically includes: Step S101: Obtain the user's basic information input set X based on the personal information submitted by the user or the basic information filled in during the first interaction: X={x1,x2,x3,x4,…,x n } Among them, x1,x2,x3,x4,…,x n They represent the user's name, age, gender, position and other user information respectively, and n represents the number of information; Step S102: Divide X into different types of data and process them separately. t ,x d ,x T ,x s Respectively represent categorical, numeric, date, and text data, x t ,x d ,x T ,x s ∈X; convert the above data into numerical values that can represent features, and combine them in order to generate a feature vector u containing user information features; Step S103: Use the feature vector u to represent the general description of the user and construct a preliminary portrait of the user.
3. According to the plug-in-based AI digital human capability expansion method of claim 1, it is characterized by: The step S102 specifically includes: (1) The collected categorical user information is converted into numerical representation using multi-hot encoding. For each category c∈C in the categorical variable C, C={c1,c2,c3,…,c m }, where m is the number of categories, for the i-th category c i The expression is: u l =[u1(x t ),u2(x t ),u3(x t ),…,u m (x t )] where i∈[1,2,3,…,m], x t Represents the information input by the user of this category; concatenates the unique-hot encoding vectors of all categorical variables into a matrix u l ; (2) Use the following formula to calculate the collected numerical data x d Perform standardization and input the data x that needs to be standardized d : Among them, μ, σ represent x d The mean and standard deviation of d represents the matrix after the normalization of all numerical data, and g represents the number of numerical data; (3) The collected date data x T Convert it into time difference ΔT data and combine them to get the date data matrix: ΔT1=|x T1 -T| you T =[ΔT1,ΔT2,…,ΔT h ] Where T represents the base time set by the system, u T represents the matrix after all date data are converted, h represents the number of date data, and u is obtained after standardizing the date data T * ; (4) Select the first text data x collected s1 , calculate the weight of each word in the text, when the word W k appears once in the text, then count(W k )=count(W k )+1, use each word in x s1 The number of occurrences in divided by x s1 The total number of words in the document is obtained, and the frequency of each word in the document is obtained. All words and calculated frequencies are stored in the corpus in the form of key-value pairs; the first Ω high-frequency words W are selected k Calculate its share x s1 The weights in construct the feature matrix u of each text s1 , k = 1, 2, ..., Ω; according to this method, the weight of high-frequency words in each text data is obtained to form the corresponding text data feature matrix u s , z represents the number of text data, u s =[u s1 ,u s2 ,…,u sz ]; (5) Concatenate all types of data matrices in order of type to obtain the feature matrix u: Among them, u represents the matrix containing all information characteristics of the user, and the number of all types of data is equal to the sum of all information data: n=m+g+h+z.
4. According to claim 1, a plug-in-based AI digital human capability expansion method is characterized by: The step S2 specifically includes: Step S201: Based on the feature matrix u of the user's preliminary portrait, convert it into the user's current state space B, and define the action space A, where A represents the action of selecting a plug-in: a j ∈A Among them, u i represents the features in u, Φ represents the threshold for selecting p features, U represents the feature matrix selected from u that is greater than the threshold Φ, μ U , σ U Respectively represent the mean and standard deviation of each feature, Vp represents the selection of the first p features from the selected features as the principal component matrix, and j represents the number of plug-ins; Step S202: In the state space and action space, the fitness difference of the plug-in action is calculated step by step, and the optimization benefit brought by the evaluation of the state function of the currently selected plug-in combination is evaluated using the fitness difference: D.P.=P. a+1 (b,a)-P a (b,a) P a+1 -P a =λ(β+μ(Π a (b',a'))-P a (b,a)) Among them, a+1 (b,a),Π a (b, a) represent the fitness values after and before updating action a in state s, respectively. λ represents the update step size. β represents the fitness value generated when executing action a. μ represents the fitness value parameter that will be increased by the next action. s' represents the next state entered after executing action a. a' represents the action selected in the next state. s q represents the state function at this time, q represents the round of selection, f(Π a q ) indicates that the qth round is the fitness value function for selecting the ath action, and κ indicates the importance of such state changes; Step S203: Select s q A plug-in with a value greater than the threshold ν indicates that the plug-in is highly relevant to the user and is preloaded; an exchange operator is set to exchange or merge different parts of the existing operations, so that previously unconsidered areas can be explored in the solution space, the solution space is expanded, and the plug-in selection strategy is updated using the gradient descent method: in, represents the exchange operator in the qth round of selection, λ θ represents the step size of the selection strategy update, is the state function s q For the policy parameter θ a The partial derivative of q Represents the a Sensitivity to change. Step S204: output the trained capability expansion strategy automatic selection model Ψ.
5. The method for expanding the capabilities of an AI digital human based on a plug-in according to claim 1, characterized in that: The step S3 specifically includes: Step S301: Initialize all available plug-ins and establish initial status and capability information for each plug-in, specifically including assigning an initial configuration to each plug-in: a={id,name,state,description} id=Category+Index Where a represents a plug-in. The plug-in content includes the plug-in id, name, state value, and function description. The plug-in id is composed of a digital index based on the plug-in type Category and sequence number, and the leading zero is used to fill the gap. When state = 0, it means that the plug-in is not enabled, and when state = 1, it means that the plug-in is enabled. Step S302: clustering the user conversation content using the text clustering model, dividing the user conversation content into categories η, performing strategy selection on the plug-in of category η using the capability extension strategy automatic selection model Ψ, and setting the plug-in status to 1; Step S303: Identify the plug-in whose plug-in status is 1, execute the plug-in function, display the corresponding card according to the returned category value, and generate a vector set F containing all enabled functions.
6. The method for expanding the capabilities of an AI digital human based on a plug-in according to claim 5, characterized in that: The specific process of classifying the user conversation content in step S302 is as follows: (1) Collect user conversation information and convert it into a numerical feature vector Y. Each conversation sample is represented as a feature vector y i ,Y=[y1,y2,…,y i ,…,y M ],Y∈R M×D , M represents the number of dialogue samples, D is the feature dimension of each sample, and R represents the feature vector matrix; (2) Initialize the text clustering model parameters: set the number of clusters, that is, the number of clusters to η = 1, the convergence accuracy ε = 0.001, and the number of iterations o = 0; (3) Initialize the prior probability ω i =1 / η; (4) Given the prior probability, calculate the posterior probability that each sample y belongs to each cluster: in, represents the posterior probability that the Mth sample comes from the i-th cluster at the oth iteration, μ and σ represent the mean and variance respectively, e represents the exponential operation, and π represents pi; (5) Update the parameter value of each cluster according to the posterior probability. The specific steps are as follows: (6) Determine whether the model has converged, and calculate whether the change of each parameter is less than the set convergence precision. If the change of any parameter is greater than the convergence precision, the number of iterations o = o + 1, and return to step (4); otherwise, go to step (7); (7) At this time, η=η+1. If η≤l, go to step (3) to ensure that the level of the category is within level l; (8) Calculate the silhouette coefficient of samples when η = 1, 2, 3, ..., l by comparing different values of η. The silhouette coefficient is the similarity between each sample and the samples in the same cluster divided by the similarity between each sample and the nearest cluster, so as to select the η with the largest silhouette coefficient as the optimal category; (9) Return the posterior probability r of each dialogue i (M) and the category η.
7. The method for expanding the capabilities of an AI digital human based on a plug-in according to claim 1, characterized in that: The step S4 specifically includes: Step S401: construct an AI digital human capability management model Θ according to the enabled function vector set F, the model including: a feature vector embedding model, a convolutional neural network model, a multi-layer perceptron model, a feedforward neural network model, and a capability management model; Step S402: Input the function vector into the AI digital human capability management model Θ, obtain the capability scores of the AI digital human for different tasks, and generate the AI digital human capability matrix D, where rows represent different tasks and columns represent different capabilities: Among them, H represents the number of tasks, I represents the number of capabilities, and d ij Represents the AI digital human's score on the jth ability of the i-th task.
8. The method for expanding the capabilities of an AI digital human based on a plug-in according to claim 7, characterized in that: The specific structure of the AI digital human capability management model θ constructed in step S401 is as follows: The input end of the AI digital human capability management model Θ is used as the input end of the feature vector embedding model; the output end of the feature vector embedding model is connected to the input end of the convolutional neural network; the output end of the convolutional neural network is connected to the input end of the multilayer perceptron model; the output end of the multilayer perceptron model is connected to the input end of the feedforward neural network model; the output end of the feedforward neural network model is connected to the input end of the capability management model; the output end of the capability management model is used as the output end of the AI digital human capability management model Θ; The feature vector embedding model includes: a fully connected layer, a Relu activation function layer and a batch processing layer; the input end of the feature vector embedding model is used as the input end of the fully connected layer; the output end of the fully connected layer is used as the input end of the Relu activation function layer; the output end of the Relu activation function layer is connected to the input end of the batch processing layer; the output end of the batch processing layer is used as the output end of the feature vector embedding model and is connected to the convolution layer of the convolutional neural network model; The convolutional neural network model includes: a convolutional layer, a pooling layer and an activation function layer; the input end of the convolutional neural network model serves as the input end of the convolutional layer; the output end of the convolutional layer is connected to the input end of the pooling layer; the output end of the pooling layer is connected to the input end of the activation function layer; the output end of the activation function layer serves as the output end of the convolutional neural network model and is connected to the fully connected layer of the multi-layer perceptron; The multi-layer perceptron model comprises: a fully connected layer, a Relu activation function layer and a Dropout layer; the input end of the multi-layer perceptron model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the Relu activation function layer; the output end of the Relu activation function layer is connected to the input end of the Dropout layer; the output end of the Dropout layer serves as the output end of the multi-layer perceptron model and is connected to the fully connected layer of the feedforward neural network model; The feedforward neural network model includes: a fully connected layer, a feedforward neural network layer and a Dropout layer; the input end of the feedforward neural network model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the feedforward neural network layer; the output end of the feedforward neural network layer is connected to the input end of the Dropout layer; the output end of the Dropout layer serves as the output end of the feedforward neural network model and is connected to the fully connected layer of the capability management model; The capability management model includes: a fully connected layer and a Softmax activation function layer; the input end of the capability management model serves as the input end of the fully connected layer; the output end of the fully connected layer is connected to the input end of the Softmax activation function layer; and the Softmax activation function layer serves as the output end of the capability management model.
9. The method for expanding the capabilities of an AI digital human based on a plug-in according to claim 7, characterized in that: The scoring formula for calculating the ability corresponding to each task in step S402 is as follows: E1=Relu(W1·F+b1) E4=Relu(W3·Mask(E3,ρ)+b4) E5=Relu(W4·Mask(E4,ρ)+b5) E6=Softmax(W5·E5+b6) Wherein, E1 represents the output of the feature vector after Relu(·) activation, W1 represents the bias matrix of the fully connected layer, and b1 represents the bias parameter of the fully connected layer; E2 represents the output of the feature vector embedding model, δ represents the learnable scaling parameter, b2 represents the bias parameter of the batch layer, μ and σ represent the mean and variance, and ξ represents the parameter to prevent zero division errors; E3 represents the output of the convolutional layer, W2 and b3 represent the bias matrix and bias parameter of the convolutional layer, LeakyRelu(·) is the activation function, and len(·) represents the length function; E4 represents the output of the multi-layer perceptron, W3 and b4 represent the bias matrix and bias parameter of the multi-layer perceptron, Mask(·) represents the mask function, indicating whether neurons are discarded, and ρ represents the discard rate, which is usually between 0 and 1; E5 represents the output of the feedforward network layer, W4 and b5 represent the bias matrix and bias parameter of the feedforward network layer; E6 is the capability score of the task, W5 and b6 represent the bias matrix and bias parameter of the capability management module, and Softmax(·) is the activation function.
10. The plug-in-based AI digital human capability expansion method according to claim 1, characterized in that: The step S5 specifically includes: Step S501: Standardize the task documents uploaded by the user and convert them into a parsable text format, remove redundant spaces, line breaks, and special characters that interfere with parsing, use periods as the separator for each sentence, and use line breaks as the separator for each task; Step S502: Perform task keyword recognition on each sentence after preprocessing, use regular expressions to extract the complete task description, and parse the tasks into structured data S1 in order. After parsing all the tasks in the document, a task dictionary list S = [S1, S2, ..., S n ], n represents the number of tasks; Step S503: Match the task dictionary list with the tasks in the capability matrix, calculate the sum of the capability values required for the task in the capability matrix as the score of the task, and output the user personalized task list in the order of the scores.
Citation Information
Patent Citations
Digital interaction method and system based on artificial intelligence, and medium
CN117348736A
Task processing method and device, equipment and storage medium
CN119003023A
System and Method for Visual Art Streaming Runtime Platform
US20200167856A1
Artificial intelligence-based personalized financial recommendation assistant system and method
US20210303973A1