Multi-mode intelligent terminal interaction method oriented to working scene
By constructing a multimodal intelligent terminal interaction method, integrating multimodal data with a large language model, generating a knowledge base for all work scenarios, and training the large language model, the problem of low efficiency in traditional manual customer service systems is solved, and efficient and accurate user interaction and personalized services are achieved.
Patent Information
- Application Number
- CN202510870371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-28
AI Technical Summary
Traditional human customer service systems are inefficient in handling complex issues, unable to accurately respond to user needs, and rely on single-modal data, resulting in inaccurate information interaction and a poor user experience.
By constructing a multimodal intelligent terminal interaction method, combining a large language model, integrating multimodal data, performing feature extraction and cleaning, generating a knowledge base for the entire work scenario, and integrating time information to train a large language model for the work scenario, intelligent interaction of multimodal data is achieved.
It improves the accuracy and efficiency of interaction, reduces human intervention, lowers operating costs, provides personalized Q&A services, and enhances user experience and satisfaction.
Smart Images

Figure CN120849545A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of large language models, multimodal intelligent terminals, and digital humans, specifically to a multimodal intelligent terminal interaction method for work scenarios. This method helps users perform human-computer interaction on multimodal intelligent terminals based on different work scenarios, thereby reducing errors and costs associated with human customer service and improving the accuracy of user interaction. Background Technology
[0002] With the rapid development of artificial intelligence and deep learning technologies, more and more fields are beginning to apply smart terminals and large language models (such as GPT-3 / 4) to optimize workflows and improve productivity. Among these, multimodal interaction technology on smart terminals is gradually becoming an important research direction. While traditional human customer service systems can provide question-and-answer support to users to some extent, they rely on manual operation and are prone to errors in information transmission, low processing efficiency, and inability to accurately answer user questions. With the dramatic increase in information volume and the diversification of user needs, the limitations of human customer service are becoming increasingly apparent, especially when dealing with complex issues, where human customer service often struggles to quickly provide accurate answers based on customer needs.
[0003] Traditional interaction methods primarily rely on single-modal data input and output, such as text and voice. This approach is not only limited by context and background but also requires significant manual maintenance and constant adjustments to responses, leading to inefficiency and high costs. Especially in work scenarios, user needs are often highly context-dependent, requiring effective personalized responses based on specific work situations. Existing single-modal interaction systems cannot meet this need, resulting in imprecise information exchange and a poor user experience.
[0004] Therefore, how to improve the accuracy and efficiency of smart terminal interaction methods by constructing a knowledge base through multimodal input data has become a key research direction in the field of large-scale models and multimodal smart terminals. By combining large language models, deep reasoning can be performed based on real-time collected multimodal data, thereby providing users with more accurate and intelligent work scenario interaction support. At the same time, enabling intelligent perception and processing of work scenarios through smart terminals can greatly reduce the workload of human customer service, improve user interaction experience and satisfaction, reduce operating costs, and increase work efficiency. Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal intelligent terminal interaction method for work scenarios. By using a large language model of work scenarios, different types of user input are identified and classified, and the interaction results are output to the multimodal intelligent terminal.
[0006] A multimodal intelligent terminal interaction method for work scenarios includes the following steps:
[0007] S1: Use data extraction methods to integrate multimodal data according to different work scenarios to obtain knowledge bases collected for different work scenarios;
[0008] S2: Perform preprocessing, clean and standardize the knowledge bases collected from different work scenarios to obtain a full work scenario knowledge base;
[0009] S3: Combine the knowledge base of the entire work scenario with time dimension information to obtain a knowledge base of the entire work scenario that integrates time information;
[0010] S4: Perform feature learning on the full work scenario knowledge base that integrates time information, and use feature extraction methods to obtain a large language model of the work scenario;
[0011] S5: The large language model in the work scenario receives user questions input through the multimodal intelligent terminal, performs queries, and outputs answers to the multimodal intelligent terminal.
[0012] Furthermore, step S1 specifically includes:
[0013] S101: Use a data collector to acquire the work scenario array WS;
[0014] S102: Based on the work scenario array WS, use the data acquisition device to obtain the multimodal data list D respectively;
[0015] S103: Perform data processing on the multimodal data list D using a feature extraction method based on multimodal data;
[0016] S104: After feature extraction, the feature array of the multimodal data is integrated to obtain a knowledge base collected from different work scenarios, and the feature array of the multimodal data is integrated.
[0017] Furthermore, step S101 specifically includes:
[0018] The data acquisition device is used to obtain the work scene array WS, whose length is i, and the work scene array WS = {WS1,WS2,…,WS}. i The method for obtaining the work scenario array WS using a data collector is as follows:
[0019] WS = GetDiffWS()
[0020] Here, GetDiffWS() is a method to retrieve the array of work scenarios WS, where WS1 represents the first work scenario, WS2 represents the second work scenario, and WS... i This represents the i-th work scenario;
[0021] Furthermore, step S102 specifically includes:
[0022] Based on the work scenario array WS and the length i of the work scenario array, the data collector is invoked to obtain the multimodal data list D, where D = (D 1 D 2 ,…,D i ):
[0023] D i =RequireDataFromWS(WS i )
[0024] Where i represents the sequence number of the current work scenario, and D 1 D represents the multimodal data array for the first work scenario. 2 D represents the multimodal data array for the second work scenario. i Let D represent the multimodal data array for the i-th work scenario. i ={D1 i D2 i ,…,D j i}, j represents the multimodal data array D i The length of D1 i D represents the multimodal data array of the i-th work scenario. i The first stored multimodal data, D2 i D represents the multimodal data array of the i-th work scenario. i The second multimodal data stored, D j i D represents the multimodal data array of the i-th work scenario. i The j-th multimodal data is stored, and RequireDataFromWS() is the method to retrieve the multimodal data array;
[0025] Furthermore, step S103 specifically includes:
[0026] (1) Obtain a multimodal data array D from the multimodal data list D. i ;
[0027] (2) For the multimodal data array D i Obtain its multimodal data D1 i D2 i ,…,D j i ;
[0028] (3) Based on multimodal data D1 i D2 i ,…,D ji To determine the multimodal data type DT, the method for determining the multimodal data type is as follows:
[0029] DT = JudgeDT(D j i )
[0030] In this context, the multimodal data type DT is an integer, with a value range of 0 or 1. 0 indicates that the multimodal data is text data, and 1 indicates that the multimodal data is image data. The JudgeDT(·) method is used to determine the multimodal data type.
[0031] (4) Based on the determined multimodal data type DT, execute the feature extraction method based on multimodal data to obtain the multimodal data feature array DF. The data feature extraction method based on multimodal type is as follows:
[0032] DF j =ExecFE(D j i ,DT)
[0033] Wherein, the multimodal data feature array DF = {DF1, DF2, ..., DF} j}, j represents the multimodal data feature array DF and the multimodal data array D. i The length of DF1 represents the multimodal data array D for the i-th work scenario. i The data features of the first stored multimodal data, DF2 represents the multimodal data array D of the i-th working scenario. i The data features of the second multimodal data stored, DF j D represents the multimodal data array of the i-th work scenario. i The data features of the j-th multimodal data stored, ExecFE(·) is the data feature extraction method based on multimodal type;
[0034] (5) Repeat steps (1) to (4), appending the multimodal data feature array obtained in subsequent traversals to the end of the multimodal data feature array obtained in the previous traversal, until all multimodal data arrays have been traversed. The method for appending the multimodal data feature array DF obtained in subsequent traversals to the end of the previous traversal is as follows:
[0035] DF j =AppendEnd(DF j-1 )
[0036] Here, `AppendEnd()` is a method that appends the multimodal data feature array `DF` obtained from subsequent traversals to the end of the previous traversal. After the traversal ends, the length of `DF` is... getLen(·) is a function to calculate the length, where i represents the sequence number of the work scene;
[0037] Different feature extraction operations are performed depending on the multimodal type (DT):
[0038] If the multimodal type DT is 0, it indicates that the multimodal data D j i For text data, text TD is a set of characters, where TD = {TD1, TD2, ..., TD}. k}, where TD1 is the first character in the text TD, TD2 is the second character in the text TD, and TD... k Let k be the k-th character in the text TD, where k represents the length of the text TD; then, the text TD is used for feature extraction according to the following formula:
[0039] V = {V1, V2, ..., V} k}={Freq(TD1,TD),Freq(TD2,TD),...,Freq(TD k ,TD)}
[0040]
[0041] Where V is a word frequency table with length k, Freq(TD1, TD) is represented by V1, which represents the frequency of word TD1 in word frequency table V, and Freq(TD2, TD) is represented by V2, which represents the frequency of word TD2 in word frequency table V. k ,TD) use V k It indicates that the character TD represents... k The frequency of a word appearing in the word frequency table V, Freq(·) represents obtaining the TD. k The function of frequency of occurrence in text TD, ‖V‖ is the norm of the text frequency table, s represents the current loop number of the summation operation, δ is a hyperparameter set by the system to balance the possible errors in the norming and normalization process, and finally DF is assigned the value of the normalized feature of the text frequency table.
[0042] If the multimodal type DT is 1, it indicates that the multimodal data D j i Given image data, where image IM is a collection of pixels (PIX), and assuming the length of image IM is a and the width is b, then:
[0043]
[0044] Here, IM is composed of pixels, and IM is a row and b column determinant. The image is then processed to obtain image feature values. The subsequent processing steps are as follows:
[0045] First, using an n*n convolutional layer, RGB channel convolution kernels are used to extract the scale features of the image in RGB, resulting in n preliminary feature maps FI1, FI2, ..., FI at different levels. n ;
[0046] Then, in the preliminary feature maps FI1,FI2,…,FI n Using convolution, we obtain sub-feature values FE1, FE2, ..., FE at different depths. n ;
[0047] In the preliminary feature maps FI1, FI2, ..., FI n The convolution operation is performed on the top, as shown below:
[0048] FE1=(FI1*F)+w1
[0049] FE2=(FI2*F)+w2
[0050] ...
[0051] FE n =(FI) n *F)+w n
[0052] Where F represents the convolution operation with dimensions (a, b), w n This represents the depth value of the nth sub-feature;
[0053] Finally, the sub-eigenvalues FE1,FE2,…,FE n Feature fusion is performed to obtain the feature values FE of the image:
[0054]
[0055] Among them, max(·), min(·), and avg(·) methods respectively take n sub-feature values FE1, FE2, ..., FE n The maximum, minimum, and average values are determined by the method, where δ represents the hyperparameter set by the system; finally, DF is assigned as a feature of the image.
[0056] Furthermore, step S104 specifically includes:
[0057] First, obtain the feature array DF for the multimodal data;
[0058] Then, obtain the length of the multimodal data feature array DF and denote it as len.
[0059] Finally, feature integration is performed on the multimodal data feature array DF to obtain the knowledge base DK collected from different working scenarios. The feature integration method is as follows:
[0060]
[0061] The Process(·) method is a feature integration operation. The knowledge base DK collected from different work scenarios is a one-row, len-column matrix with len multimodal data feature dimensions. DK1 is the first knowledge in the knowledge base DK collected from different work scenarios, DK2 is the second knowledge in the knowledge base DK collected from different work scenarios, and DK... len The getL(·) method can be used to get the number of knowledge items in the knowledge base DK collected for different work scenarios.
[0062] Furthermore, step S2 specifically includes:
[0063] S201: Based on the knowledge base collected for different work scenarios, use relevant data cleaning methods to remove noisy data from the knowledge base;
[0064] S202: The cleaned knowledge base will be standardized to obtain a standardized knowledge base;
[0065] S203: Perform data fusion operations on the standardized knowledge base to obtain the full-work-scenario knowledge base WSK;
[0066] Furthermore, step S201 specifically includes:
[0067] First, identify and delete duplicate records to ensure the uniqueness of each characteristic, as follows:
[0068] For knowledge bases (DKs) collected for different work scenarios, there are:
[0069] DK = Trim(DK)
[0070] The Trim(·) method is used to identify and delete duplicate records.
[0071] Then, the missing features are filled using the mean imputation method;
[0072] Furthermore, excessively high outlier values are removed, as follows:
[0073] DK = DeleteHV(DK)
[0074] The DeleteHV(·) method is used to identify and delete excessively high abnormal feature values.
[0075] Finally, check the data quality to ensure that the cleaning process has not introduced any new errors;
[0076] Furthermore, step S202 specifically includes:
[0077] First, acquire the knowledge base DK1, DK2, ..., DK collected from different work scenarios. len ;
[0078] Then, logarithmic transformation is used to transform the knowledge DK1, DK2, ..., DK. len Standardization processes are performed, including:
[0079] DK len =ln(DK) len )
[0080] The ln(·) method is used to calculate DK1, DK2, ..., DK len The natural logarithm;
[0081] Finally, the standardized DK1, DK2, ..., DK len Reassign the value to DK;
[0082] Furthermore, step S203 specifically includes:
[0083] First, the knowledge base DK is trained independently using the WSRF and WSBoost methods, where:
[0084] ResRF = WSRF(DK)
[0085] ResBoost = WSBoost(DK)
[0086] Wherein, ResRF represents the set of prediction results obtained using the WSRF method, ResBoost represents the set of prediction results obtained using the WSBoost method, and WSRF(·) and WSBoost(·) are both prediction models;
[0087] Then, a weighted average is performed on ResRF and ResBoost as follows:
[0088] WSK = w ResRF *ResRF+w ResBoost *ResBoost
[0089] Among them, w ResRF For the weights of ResRF, w ResBoost These are the weights for ResBoost;
[0090] Finally, we obtain the WSK (Work Scenario Knowledge Base);
[0091] For the WSRF method, we have:
[0092]
[0093] The model contains a total of Z subtrees, where i represents the index of the current subtree, and w... i h is the weight of the i-th subtree. i (DK) represents the prediction result of the i-th subtree, indicating the support for the accuracy of knowledge DK;
[0094] For the WSBoost method, we have:
[0095]
[0096] Among them, w len is the weight of the len-th knowledge, reflecting the importance of that knowledge in the work scenario, and l(·) is the cross-entropy loss function. Let f be the predicted value of the len-th knowledge, t be the number of base models, and f t R is the output of the t-th base model, R(i) is the penalty term of the i-th work scenario rule, and Ω(·) represents the method for calculating the regularization term, which is used to prevent overfitting.
[0097] Furthermore, step S3 specifically includes:
[0098] S301: Obtain a multimodal data array D from the multimodal data list D. i For multimodal data array D i Obtain its multimodal data D1 i D2 i ,…,D j i From multimodal data D1 i D2 i ,…,D j i Obtain time information from the source;
[0099] S302: Based on time information, the knowledge base of the entire work scenario is integrated to obtain a knowledge base of the entire work scenario with integrated time information;
[0100] Furthermore, step S301 specifically includes:
[0101] Time information can be denoted as Time, and Time can be represented as:
[0102] Time = RequireTimeFD(D1) i D2 i ,…,D j i )
[0103] Among them, RequireTimeFD(·) is derived from multimodal data D1 i D2 i ,…,D j i Methods for obtaining time information in [the context of a mobile application];
[0104] Furthermore, step S302 specifically includes:
[0105] First, based on the full work scenario knowledge base WSK, a one-to-one correspondence is established with the time information Time, resulting in the full work scenario knowledge base OWSK sorted by time, where:
[0106] OWSK = OrderBT(WSK, Time)
[0107] The OrderBT(·) method sorts the entire work scenario knowledge base WSK according to the time information Time.
[0108] Then, the time interval GAP is set as a benchmark for subsequent correlation prediction;
[0109] Secondly, extract time series features (TSF) from each time interval. g Using the average value as a time series feature, a time series feature array TSF is constructed, where:
[0110] TSF = {TSF1, TSF2, ..., TSF} g} = GetFBT(OWSK, GAP)
[0111] The GetFBT(·) method retrieves time series features (TSFs) according to time intervals. g The method is as follows: g is the current time interval number, where TSF1 represents the time series feature corresponding to the first time interval, TSF2 represents the time series feature corresponding to the second time interval, and TSF... g This represents the time series feature corresponding to the g-th time interval;
[0112] Then, by weighting the time series features corresponding to different time windows, it can be expressed as:
[0113] FTSF = {w1*TSF1, w2*TSF2, ..., w g *TSF g}
[0114] Where FTSF is the weighted array of time series features, w1 is the weight of the time series feature corresponding to the first time interval, w2 is the weight of the time series feature corresponding to the second time interval, and w... gThe weights of the time series features corresponding to the g-th time interval are:
[0115] Finally, the FTSF is merged with the OWSK, a full work scenario knowledge base sorted by time, to obtain the full work scenario knowledge base TWSK with merged time information = FTSF ⊕ OWSK;
[0116] Time series features (TSF) are obtained according to time intervals. g The steps are as follows:
[0117] First, the OWSK knowledge base, which is sorted by time according to the time interval GAP, is divided into g sub-knowledge bases {OWSK1, OWSK2, ..., OWSK}. g};
[0118] Then, the sub-knowledge base {OWSK1,OWSK2,…,OWSK} g OWSK g Contains 1 to j different knowledge DK j Using the averaging method, the time series features (TSF) within each sub-knowledge base are calculated. g ,in:
[0119] Among them, DK j ∈OWSK g And n = SizeOf(OWSK) g )
[0120] The SizeOf(·) method is used to obtain the length of a sub-knowledge base segment.
[0121] Furthermore, step S4 specifically includes:
[0122] S401: Perform oversampling using a full-work-scene knowledge base that integrates time information to obtain a TWSK (Temporally Written Knowledge Base) of integrated time information. oversampled ;
[0123] S402: Fine-tune the model and output the WLLM (Work Scenario Large Language Model);
[0124] Furthermore, step S401 specifically includes:
[0125] The TWSK knowledge base for the entire working scenario, which incorporates time information, is oversampled using the oversampling method, as shown below:
[0126]
[0127] Among them, TWSK oversampled N is a knowledge base for the entire working scenario, formed by fusing time information after oversampling.minority TWSK represents the number of minority class samples, λ is the interpolation coefficient. i This represents the i-th data sample;
[0128] Furthermore, step S402 specifically includes:
[0129] Apply a work-scene attention mechanism based on the contextual information of the work scenario.
[0130] (1) TWSK, a knowledge base for the entire working scenario of fused time information after oversampling. oversampled Perform location encoding to obtain specific location encoding information, and map the input data into a vector representation of EMB;
[0131] (2) Perform weight matrix w on the vector representation EMB. QU w KE and w VA The multiplication operation yields a query, key, and value vector, denoted as QU, KE, and VA, respectively.
[0132] (3) Perform a full-connect operation on the query and key vectors based on the work scenario to obtain the model's attention to different work scenarios, where:
[0133]
[0134] Among them, S i To indicate the level of attention given to the i-th scene, Softmax(·) is the Softmax activation function, C represents the scene context information, and w i Let be the weight of the i-th scene, and DIM be the TWSK (Time-Based Knowledge Base) for the entire working scene after oversampling and fusion of temporal information. oversampled The dimension, KE T The transpose of the key vector KE is represented by the matrix;
[0135] (4) The degree of attention S to the i-th scene i Multiplying the vectors VA with the same value yields the output O = S of the fully connected layer. i *VA;
[0136] (5) Repeat steps (2) to (4) to adjust the weight matrix w. QU w KE and w VA This continues until the final output O converges;
[0137] (6) Output the WLLM (Work Scenario Large Language Model);
[0138] For each scenario SC i Extract its sequence number information to form a vector Cs;
[0139] scene data Perform a vector concatenation operation with Cs to obtain the scene context information C, as follows:
[0140] C = Concat(DK, Cs)
[0141] The Concat(·) function performs vector concatenation operations.
[0142] Furthermore, step S5 specifically includes:
[0143] S501: Use a data collector to acquire the question Q input by the user on the smart terminal, use the Working Scenario Large Language Model (WLLM) to answer the question, and obtain the output of the Working Scenario Large Language Model.
[0144] S502: Return the output of the large language model of the work scenario to the multimodal intelligent terminal to provide an external response;
[0145] Furthermore, step S501 specifically includes:
[0146] First, a data collector is used to obtain the user's question Q on the smart terminal, where:
[0147] Q = GetOpt()
[0148] GetOpt() is a method for obtaining user input on a smart terminal;
[0149] Then, the question Q is input into the Work Scenario Large Language Model (WLLM) for intelligent answering;
[0150] Finally, the output of the WLLM (Work Scenario Large Language Model) is obtained as follows:
[0151] Output = GetWLLM(Q)
[0152] GetWLLM(·) is a method for obtaining answers to questions using a large language model based on the work scenario.
[0153] The beneficial effects of this invention include:
[0154] (1) This invention significantly improves the interaction accuracy of smart terminals. Traditional interaction systems mostly rely on single-modal input and output data, such as plain text or voice. However, this invention combines multimodal data such as text and images, enabling the system to support more comprehensive and diverse input data. These different types of data can provide the system with richer contextual information and data content, thereby helping the smart terminal to understand the user's needs more accurately and provide more precise answers to meet the user's complex question-and-answer needs.
[0155] (2) This invention improves the accuracy and data integrity of the multimodal knowledge base by identifying and deleting duplicate records and abnormal feature values in knowledge bases for different work scenarios; it standardizes the data using a logarithmic transformation method and integrates WSRF and WSBoost prediction models to evaluate the accuracy of knowledge, ultimately generating a full-work-scenario knowledge base; the full-work-scenario knowledge base can be integrated with time information to increase the priority of recent key knowledge and maintain the priority focus on recent knowledge. Furthermore, this invention can dynamically adjust the interaction strategy between the multimodal terminal digital human and the user according to different work scenarios to meet the needs of different work scenarios;
[0156] (3) This invention integrates multimodal data according to work scenarios and constructs a knowledge base. It uses a neural network-based prediction model to improve the accuracy of knowledge and integrates time-dimensional information to obtain a full work scenario knowledge base with integrated time information. After being trained by a work scenario attention mechanism, the full work scenario knowledge base with integrated time information generates a large language model for work scenarios. The model can drive the digital human to provide accurate question answers based on user question input. In addition, this invention is applicable to different work scenarios, supports a wider variety of question types, makes the digital human's answers more accurate, and improves the user experience.
[0157] (4) This invention can reduce human intervention. Traditional customer service systems usually require human intervention to handle complex user issues, which not only increases operating costs but is also prone to errors or delays due to human factors. With the help of large language models and multimodal intelligent interaction technology, the system can automatically process and understand various types of user input data and complete most tasks without relying on human customer service. This automation method can not only improve work efficiency but also reduce errors that may occur when human intervention is involved, thereby reducing work costs and management difficulty.
[0158] (5) This invention can improve work efficiency. Traditional manual customer service systems usually rely on human judgment and processing, but the complexity and variability of work scenarios often result in lag in manual processing. By introducing a smart terminal based on a large language model, the system can understand user needs in different work scenarios in real time and accurately and provide timely responses. For example, in environments with large changes in work scenarios, the smart terminal can quickly switch its processing mode and automatically adjust its interaction method for the current work task to ensure that users receive timely and effective feedback.
[0159] (6) This invention can significantly reduce the operating costs of enterprises. Although human customer service is unavoidable in some situations (such as when the problem is extremely complex and difficult to understand), relying heavily on human customer service is very expensive and inefficient for enterprises. Through the multimodal intelligent interaction technology of this invention, the need for human customer service can be reduced without sacrificing service quality, thus greatly reducing labor costs. In addition, the efficient and accurate answers of the intelligent system can improve user satisfaction and loyalty, thereby further boosting enterprise efficiency.
[0160] (7) This invention can provide a better user experience. Through the understanding of different work scenarios by the smart terminal, the system can provide users with more personalized question and answer services. This customized service can not only meet the needs of different users, but also improve the overall user experience. As user needs continue to change, the system can adjust its interaction methods in real time and optimize the response content. This highly intelligent service makes users feel more convenient and satisfied when using the multimodal terminal digital human system.
[0161] Other advantages, objectives, and features of the present invention will be set forth in detail in the following description. Through a thorough study of the following text, those skilled in the art will be able to clearly recognize these advantages and features and gain valuable lessons from the practice of the invention, whose objectives and other advantages can be realized and embodied in the following description and the previously mentioned claims. Attached Figure Description
[0162] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:
[0163] Figure 1 This is a schematic diagram of a multimodal intelligent terminal interaction method for work scenarios according to the present invention;
[0164] Figure 2 This is an example diagram of the working scenario array in an embodiment of the present invention;
[0165] Figure 3 This is an example diagram of a multimodal data list according to an embodiment of the present invention;
[0166] Figure 4 This is an example diagram of a multimodal data feature array according to an embodiment of the present invention;
[0167] Figure 5 Example diagrams of the knowledge base collected for different working scenarios in embodiments of the present invention;
[0168] Figure 6 This is an example diagram of the knowledge base for all working scenarios in an embodiment of the present invention;
[0169] Figure 7 This is an example diagram of a knowledge base for all working scenarios that integrates time information, as described in an embodiment of the present invention.
[0170] Figure 8 This is an example diagram of a large language model in a working scenario according to an embodiment of the present invention;
[0171] Figure 9 This is an example diagram of a multimodal intelligent terminal digital human according to an embodiment of the present invention;
[0172] Figure 10 This is an example diagram of a multimodal intelligent terminal digital human answering e-commerce questions according to an embodiment of the present invention;
[0173] Figure 11 This is an example diagram of a multimodal intelligent terminal digital human answering virtual live broadcast questions according to an embodiment of the present invention. Detailed Implementation
[0174] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0175] like Figure 1 As shown, the present invention provides a multimodal intelligent terminal interaction method for work scenarios, comprising the following steps:
[0176] Step S1: Use data extraction methods to integrate multimodal data according to different work scenarios to obtain knowledge bases collected for different work scenarios;
[0177] In this embodiment, the data collector is a sub-functional module of the multimodal intelligent terminal interaction method for work scenarios, and the work scenario array is as follows: Figure 2 As shown, the multimodal data list is as follows: Figure 3 As shown, the feature array of multimodal data is as follows: Figure 4 As shown, the feature array of multimodal data is as follows: Figure 5 As shown.
[0178] In this embodiment, step S1 includes the following sub-steps:
[0179] Step S101: Use a data acquisition device to obtain the working scene array WS;
[0180] The data acquisition device is used to obtain the work scene array WS, whose length is i, and the work scene array WS = {WS1,WS2,…,WS}. i The method for obtaining the work scenario array WS using a data collector is as follows:
[0181] WS = GetDiffWS()
[0182] Here, GetDiffWS() is a method to retrieve the array of work scenarios WS, where WS1 represents the first work scenario, WS2 represents the second work scenario, and WS... i Let i represent the i-th work scenario.
[0183] Step S102: Based on the work scenario array WS, use the data acquisition device to obtain the multimodal data list D;
[0184] In this embodiment, Figure 2 The task data acquisition array WS is used as input, and the output is... Figure 3 The data acquisition unit obtains a list of multimodal data.
[0185] Based on the work scenario array WS and the length i of the work scenario array, the data collector is invoked to obtain the multimodal data list D, where D = (D 1 D 2 ,…,D i ):
[0186] D i =RequireDataFromWS(WS i )
[0187] Where i represents the sequence number of the current work scenario, and D 1 D represents the multimodal data array for the first work scenario. 2 D represents the multimodal data array for the second work scenario. i Let D represent the multimodal data array for the i-th work scenario. i ={D1 i D2 i ,…,D j i}, j represents the multimodal data array D i The length of D1 i D represents the multimodal data array of the i-th work scenario. i The first stored multimodal data, D2 i D represents the multimodal data array of the i-th work scenario. i The second multimodal data stored, D j i D represents the multimodal data array of the i-th work scenario. i The j-th multimodal data is stored, and RequireDataFromWS(·) is the method to obtain the multimodal data array.
[0188] Step S103: Process the multimodal data list D using a feature extraction method based on multimodal data;
[0189] In this embodiment, Figure 3 The multimodal data list is used as input, and the output is... Figure 4 The specific steps for creating the feature array (DF) for multimodal data are as follows:
[0190] (1) Obtain a multimodal data array D from the multimodal data list D. i ;
[0191] (2) For the multimodal data array D i Obtain its multimodal data D1 i D2 i ,…,D j i ;
[0192] (3) Based on multimodal data D1 i D2 i ,…,D j i To determine the multimodal data type DT, the method for determining the multimodal data type is as follows:
[0193] DT = JudgeDT(D j i )
[0194] In this context, the multimodal data type DT is an integer, with a value range of 0 or 1. 0 indicates that the multimodal data is text data, and 1 indicates that the multimodal data is image data. The JudgeDT(·) method is used to determine the multimodal data type.
[0195] (4) Based on the determined multimodal data type DT, execute the feature extraction method based on multimodal data to obtain the multimodal data feature array DF. The data feature extraction method based on multimodal type is as follows:
[0196] DF j =ExecFE(D j i ,DT)
[0197] Wherein, the multimodal data feature array DF = {DF1, DF2, ..., DF} j}, j represents the multimodal data feature array DF and the multimodal data array D. i The length of DF1 represents the multimodal data array D for the i-th work scenario. i The data features of the first stored multimodal data, DF2 represents the multimodal data array D of the i-th working scenario. iThe data features of the second multimodal data stored, DF j D represents the multimodal data array of the i-th work scenario. i The data features of the j-th multimodal data stored, ExecFE(·) is the data feature extraction method based on multimodal type;
[0198] (5) Repeat steps (1) to (4), appending the multimodal data feature array obtained in subsequent traversals to the end of the multimodal data feature array obtained in the previous traversal, until all multimodal data arrays have been traversed. The method for appending the multimodal data feature array DF obtained in subsequent traversals to the end of the previous traversal is as follows:
[0199] DF j =AppendEnd(DF j-1 )
[0200] Here, `AppendEnd()` is a method that appends the multimodal data feature array `DF` obtained from subsequent traversals to the end of the previous traversal. After the traversal ends, the length of `DF` is... getLen(·) is a function to calculate the length, where i represents the sequence number of the work scene;
[0201] Different feature extraction operations are performed depending on the multimodal type (DT):
[0202] If the multimodal type DT is 0, it indicates that the multimodal data D j i For text data, text TD is a set of characters, where TD = {TD1, TD2, ..., TD}. k}, where TD1 is the first character in the text TD, TD2 is the second character in the text TD, and TD... k Let k be the k-th character in the text TD, where k represents the length of the text TD; then, the text TD is used for feature extraction according to the following formula:
[0203] V = {V1, V2, ..., V} k}={Freq(TD1,TD),Freq(TD2,TD),...,Freq(TD k ,TD)}
[0204]
[0205] Where V is a word frequency table with length k, Freq(TD1, TD) is represented by V1, which represents the frequency of word TD1 in word frequency table V, and Freq(TD2, TD) is represented by V2, which represents the frequency of word TD2 in word frequency table V. k ,TD) use Vk It indicates that the character TD represents... k The frequency of a word appearing in the word frequency table V, Freq(·) represents obtaining the TD. k The function of frequency of occurrence in text TD, ‖V‖ is the norm of the text frequency table, s represents the current loop number of the summation operation, δ is a hyperparameter set by the system with a default value of 0.01, used to balance the possible errors in the norming and normalization process, and finally DF is assigned the feature after normalization of the text frequency table.
[0206] If the multimodal type DT is 1, it indicates that the multimodal data D j i Given image data, where image IM is a collection of pixels (PIX), and assuming the length of image IM is a and the width is b, then:
[0207]
[0208] Here, IM is composed of pixels, and IM is a row and b column determinant. The image is then processed to obtain image feature values. The subsequent processing steps are as follows:
[0209] First, a 3*3 convolutional layer is used to extract the scale features of the image in RGB using RGB channel convolutional kernels, resulting in three preliminary feature maps FI1, FI2, and FI3 at different levels.
[0210] Then, convolution is used on the initial feature maps FI1, FI2, FI3 to obtain sub-feature values FI1, FI2, FI3 at different depths.
[0211] Convolution operations are performed on the initial feature maps FI1, FI2, and FI3, as shown below:
[0212] FE1=(FI1*F)+w1
[0213] FE2=(FI2*F)+w2
[0214] FE3=(FI3*F)+w3
[0215] Where F is the convolution operation with size (a, b), and w3 represents the depth value of the third sub-feature;
[0216] Finally, the sub-feature values FE1, FE2, and FE3 are fused to obtain the image feature value FE:
[0217]
[0218] Among them, max(·), min(·), and avg(·) methods are used to obtain the maximum, minimum, and average values of the three sub-feature values FE1, FE2, and FE3, respectively. δ represents the hyperparameter set by the system, with a default value of 0.01. Finally, DF is assigned as a feature of the image.
[0219] Step S104: After feature extraction, the feature array of the multimodal data is integrated to obtain a knowledge base collected from different work scenarios, and the feature array of the multimodal data is integrated.
[0220] In this embodiment, Figure 4 The multimodal data feature array DF is used as input, and the output is... Figure 5 Knowledge base DK collected from different work scenarios;
[0221] First, obtain the feature array DF for the multimodal data;
[0222] Then, obtain the length of the multimodal data feature array DF and denote it as len.
[0223] Finally, feature integration is performed on the multimodal data feature array DF to obtain the knowledge base DK collected from different working scenarios. The feature integration method is as follows:
[0224]
[0225] The Process(·) method is a feature integration operation. The knowledge base DK collected from different work scenarios is a one-row, len-column matrix with len multimodal data feature dimensions. DK1 is the first knowledge in the knowledge base DK collected from different work scenarios, DK2 is the second knowledge in the knowledge base DK collected from different work scenarios, and DK... len The getL(·) method retrieves the number of knowledge items collected in the knowledge base DK for different work scenarios.
[0226] Step S2: Perform preprocessing, clean and standardize the knowledge bases collected from different work scenarios to obtain a full work scenario knowledge base;
[0227] In this embodiment, preprocessing techniques are used to clean and standardize the knowledge bases collected from different work scenarios, resulting in a comprehensive work scenario knowledge base, such as... Figure 6 As shown.
[0228] In this embodiment, S2 specifically includes the following steps:
[0229] Step S201: Based on the knowledge base collected for different work scenarios, use relevant data cleaning methods to clean out noisy data in the knowledge base;
[0230] First, identify and delete duplicate records to ensure the uniqueness of each characteristic, as follows:
[0231] For knowledge bases (DKs) collected for different work scenarios, there are:
[0232] DK = Trim(DK)
[0233] The Trim(·) method is used to identify and delete duplicate records.
[0234] Then, the missing features are filled using the mean imputation method;
[0235] Furthermore, excessively high outlier values are removed, as follows:
[0236] DK = DeleteHV(DK)
[0237] The DeleteHV(·) method is used to identify and delete excessively high abnormal feature values.
[0238] Finally, check the data quality to ensure that the cleaning process has not introduced any new errors.
[0239] Step S202: Standardize the cleaned knowledge base to obtain a standardized knowledge base.
[0240] First, acquire the knowledge base DK1, DK2, ..., DK collected from different work scenarios. len ;
[0241] Then, logarithmic transformation is used to transform the knowledge DK1, DK2, ..., DK. len Standardization processes are performed, including:
[0242] DK len =ln(DK) len )
[0243] The ln(·) method is used to calculate DK1, DK2, ..., DK len The natural logarithm;
[0244] Finally, the standardized DK1, DK2, ..., DK len Reassign the value to DK;
[0245] Step S203: Perform data fusion on the standardized knowledge base to obtain the full-work-scenario knowledge base WSK.
[0246] First, the knowledge base DK is trained independently using the WSRF and WSBoost methods, where:
[0247] ResRF = WSRF(DK)
[0248] ResBoost = WSBoost(DK)
[0249] Wherein, ResRF represents the set of prediction results obtained using the WSRF method, ResBoost represents the set of prediction results obtained using the WSBoost method, and WSRF(·) and WSBoost(·) are both prediction models;
[0250] Then, a weighted average is performed on ResRF and ResBoost as follows:
[0251] WSK = w ResRF *ResRF+w ResBoost *ResBoost
[0252] Among them, w ResRF For the weights of ResRF, w ResBoost These are the weights for ResBoost, with a default value of 0.5 for each weight.
[0253] Finally, we obtain the WSK (Work Scenario Knowledge Base);
[0254] For the WSRF method, we have:
[0255]
[0256] The model contains a total of Z subtrees, where i represents the index of the current subtree, and w... i h is the weight of the i-th subtree. i (DK) represents the prediction result of the i-th subtree, indicating the support for the accuracy of knowledge DK;
[0257] For the WSBoost method, we have:
[0258]
[0259] Among them, w len is the weight of the len-th knowledge, reflecting the importance of that knowledge in the work scenario, and l(·) is the cross-entropy loss function. Let f be the predicted value of the len-th knowledge, t be the number of base models, and f t R is the output of the t-th base model, R(i) is the penalty term of the i-th work scenario rule, and Ω(·) represents the method for calculating the regularization term, which is used to prevent overfitting.
[0260] Step S3: Combine the full work scenario knowledge base with time dimension information to obtain a full work scenario knowledge base that integrates time information;
[0261] In this embodiment, a full-work-scene knowledge base integrating time information is as follows: Figure 7 As shown.
[0262] In this embodiment, S3 specifically includes the following steps:
[0263] Step S301: Obtain a multimodal data array D from the multimodal data list D. i For multimodal data array D i Obtain its multimodal data D1 i D2 i ,…,D j i From multimodal data D1 i D2 i ,…,D j i Obtain time information from the middle.
[0264] Time information can be denoted as Time, and Time can be represented as:
[0265] Time = RequireTimeFD(D1) i D2 i ,…,D j i )
[0266] Among them, RequireTimeFD(·) is derived from multimodal data D1 i D2 i ,…,D j i Methods for obtaining time information.
[0267] Step S302: Based on time information, integrate the knowledge base of the entire work scenario to obtain a knowledge base of the entire work scenario with integrated time information.
[0268] First, based on the full work scenario knowledge base WSK, a one-to-one correspondence is established with the time information Time, resulting in the full work scenario knowledge base OWSK sorted by time, where:
[0269] OWSK = OrderBT(WSK, Time)
[0270] The OrderBT(·) method sorts the entire work scenario knowledge base WSK according to the time information Time.
[0271] Then, the time interval GAP is set as a benchmark for subsequent correlation prediction;
[0272] Secondly, extract time series features (TSF) from each time interval. gUsing the average value as a time series feature, a time series feature array TSF is constructed, where:
[0273] TSF = {TSF1, TSF2, ..., TSF} g} = GetFBT(OWSK, GAP)
[0274] The GetFBT(·) method retrieves time series features (TSFs) according to time intervals. g The method is as follows: g is the current time interval number, where TSF1 represents the time series feature corresponding to the first time interval, TSF2 represents the time series feature corresponding to the second time interval, and TSF... g This represents the time series feature corresponding to the g-th time interval;
[0275] Then, by weighting the time series features corresponding to different time windows, it can be expressed as:
[0276] FTSF = {w1*TSF1, w2*TSF2, ..., w g *TSF g}
[0277] Where FTSF is the weighted array of time series features, w1 is the weight of the time series feature corresponding to the first time interval, w2 is the weight of the time series feature corresponding to the second time interval, and w... g The weights of the time series features corresponding to the g-th time interval are:
[0278] Finally, the FTSF is merged with the OWSK, a full work scenario knowledge base sorted by time, to obtain the full work scenario knowledge base TWSK = FTSF ⊕ OWSK with merged time information.
[0279] Time series features (TSF) are obtained according to time intervals. g The steps are as follows:
[0280] First, the OWSK knowledge base, which is sorted by time according to the time interval GAP, is divided into g sub-knowledge bases {OWSK1, OWSK2, ..., OWSK}. g};
[0281] Then, the sub-knowledge base {OWSK1,OWSK2,…,OWSK} g OWSK g Contains 1 to j different knowledge DK j Using the averaging method, the time series features (TSF) within each sub-knowledge base are calculated. g ,in:
[0282] Among them, DKj ∈OWSK g And n = SizeOf(OWSK) g )
[0283] The SizeOf(·) method is used to obtain the length of a sub-knowledge base segment.
[0284] Step S4: Perform feature learning on the full work scenario knowledge base that integrates time information, and use feature extraction methods to obtain a large language model of the work scenario;
[0285] In this embodiment, the structure diagram of the large language model for the working scenario is as follows: Figure 8 As shown.
[0286] In this embodiment, S4 specifically includes the following steps:
[0287] Step S401: Perform oversampling using the full-work-scene knowledge base with fused time information to obtain the oversampled full-work-scene knowledge base TWSK with fused time information. oversampled ;
[0288] The TWSK knowledge base for the entire working scenario, which incorporates time information, is oversampled using the oversampling method, as shown below:
[0289]
[0290] Among them, TWSK oversampled N is a knowledge base for the entire working scenario, formed by fusing time information after oversampling. minority TWSK represents the number of minority class samples, λ is the interpolation coefficient. i This represents the i-th data sample.
[0291] Step S402: Fine-tune the model and output the large language model WLLM for the working scenario;
[0292] Apply a work-scene attention mechanism based on the contextual information of the work scenario.
[0293] (1) TWSK, a knowledge base for the entire working scenario of fused time information after oversampling. oversampled Perform location encoding to obtain specific location encoding information, and map the input data into a vector representation of EMB;
[0294] (2) Perform weight matrix w on the vector representation EMB. QU w KE and w VA The multiplication operation yields a query, key, and value vector, denoted as QU, KE, and VA, respectively.
[0295] (3) Perform a full-connect operation on the query and key vectors based on the work scenario to obtain the model's attention to different work scenarios, where:
[0296]
[0297] Among them, S i To indicate the level of attention given to the i-th scene, Softmax(·) is the Softmax activation function, C represents the scene context information, and w i Let be the weight of the i-th scene, and DIM be the TWSK (Time-Based Knowledge Base) for the entire working scene after oversampling and fusion of temporal information. oversampled The dimension, KE T The transpose of the key vector KE is represented by the matrix;
[0298] (4) The degree of attention S to the i-th scene i Multiplying the vectors VA with the same value yields the output O = S of the fully connected layer. i *VA;
[0299] (5) Repeat steps (2) to (4) to adjust the weight matrix w. QU w KE and w VA This continues until the final output O converges;
[0300] (6) Output the WLLM (Work Scenario Large Language Model);
[0301] For each scenario SC i Extract its sequence number information to form a vector Cs;
[0302] scene data Perform a vector concatenation operation with Cs to obtain the scene context information C, as follows:
[0303] C = Concat(DK, Cs)
[0304] The Concat(·) function performs vector concatenation operations.
[0305] Step S5: The large language model of the working scenario receives user questions input through the multimodal intelligent terminal, performs queries, and outputs answers to the multimodal intelligent terminal;
[0306] In this embodiment, to enhance the fun and interactive functions, according to Figure 2 The data task obtains the working scenario array WS to construct a multimodal intelligent terminal digital human. An example of a multimodal intelligent terminal digital human is shown below. Figure 9 As shown, an example of a multimodal intelligent terminal digital human answering e-commerce questions is as follows: Figure 10 As shown, an example of a multimodal intelligent terminal digital human answering questions in a virtual live broadcast is as follows. Figure 11As shown.
[0307] In this embodiment, S5 specifically includes the following steps:
[0308] Step S501: Use a data collector to acquire the question Q input by the user on the smart terminal, use the Working Scenario Large Language Model (WLLM) to answer the question, and obtain the output of the Working Scenario Large Language Model.
[0309] First, a data collector is used to obtain the user's question input on the smart terminal, where:
[0310] Q = GetOpt()
[0311] GetOpt() is a method for obtaining user input on a smart terminal;
[0312] Then, the question Q is input into the Work Scenario Large Language Model (WLLM) for intelligent answering;
[0313] Finally, the output of the WLLM (Work Scenario Large Language Model) is obtained as follows:
[0314] Output = GetWLLM(Q)
[0315] GetWLLM(·) is a method for obtaining answers to questions using a large language model based on the work scenario.
[0316] Step S502: Return the output of the large language model of the work scenario to the multimodal intelligent terminal to control the digital human to make a response.
[0317] In this embodiment, the work scenario, multimodal intelligent terminal, and large language model are combined. Figure 2 , Figure 3 All data comes from the data collector. Features are extracted and further analyzed from the work scenario array and multimodal data list obtained by the data collector. A work scenario attention mechanism is used to perform context learning on the knowledge base. The combination of these three methods improves the stability and security of the multimodal smart terminal and ultimately achieves good user experience.
[0318] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0319] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0320] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0321] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
[0322] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A multimodal intelligent terminal interaction method for work scenarios, characterized in that: Includes the following steps: Step S1: Use data extraction methods to integrate multimodal data according to different work scenarios to obtain knowledge bases collected for different work scenarios; Step S2: Perform preprocessing, clean and standardize the knowledge bases collected from different work scenarios to obtain a full work scenario knowledge base; Step S3: Combine the full work scenario knowledge base with time dimension information to obtain a full work scenario knowledge base that integrates time information; Step S4: Perform feature learning on the full work scenario knowledge base that integrates time information, and use feature extraction methods to obtain a large language model of the work scenario; Step S5: The large language model of the working scenario receives user questions input through the multimodal intelligent terminal, performs queries, and outputs answers to the multimodal intelligent terminal.
2. The multimodal intelligent terminal interaction method for work scenarios according to claim 1, characterized in that: Step S1 specifically includes: Step S101: Obtain the working scene array WS using a data acquisition device. The specific operation steps are as follows: The data acquisition device is used to obtain the work scene array WS, whose length is i, and the work scene array WS = {WS1,WS2,…,WS}. i The method for obtaining the work scenario array WS using a data collector is as follows: WS = GetDiffWS() Here, GetDiffWS() is a method to retrieve the array of work scenarios WS, where WS1 represents the first work scenario, WS2 represents the second work scenario, and WS... i This represents the i-th work scenario; Step S102: Based on the work scenario array WS, use a data acquisition device to obtain the multimodal data list D. The specific operation steps are as follows: Based on the work scenario array WS and the length i of the work scenario array, the data collector is invoked to obtain the multimodal data list D, where D = (D 1 D 2 ,…,D i ): D i =RequireDataFromWS(WS i ) Where i represents the sequence number of the current work scenario, and D 1 D represents the multimodal data array for the first work scenario. 2 D represents the multimodal data array for the second work scenario. i Let D represent the multimodal data array for the i-th work scenario. i ={D1 i D2 i ,…,D j i }, j represents the multimodal data array D i The length of D1 i D represents the multimodal data array of the i-th work scenario. i The first stored multimodal data, D2 i D represents the multimodal data array of the i-th work scenario. i The second multimodal data stored, D j i D represents the multimodal data array of the i-th work scenario. i The j-th multimodal data is stored, and RequireDataFromWS() is the method to retrieve the multimodal data array; Step S103: Process the multimodal data list D using a feature extraction method based on multimodal data. The specific steps are as follows: (1) Obtain a multimodal data array D from the multimodal data list D. i ; (2) For the multimodal data array D i Obtain its multimodal data D1 i D2 i ,…,D j i ; (3) Based on multimodal data D1 i D2 i ,…,D j i To determine the multimodal data type DT, the method for determining the multimodal data type is as follows: DT=JudgeDT(D j i ) In this context, the multimodal data type DT is an integer, with a value range of 0 or 1. 0 indicates that the multimodal data is text data, and 1 indicates that the multimodal data is image data. The JudgeDT(·) method is used to determine the multimodal data type. (4) Based on the determined multimodal data type DT, execute the feature extraction method based on multimodal data to obtain the multimodal data feature array DF. The data feature extraction method based on multimodal type is as follows: DF j =ExecFE(D j i ,DT) Wherein, the multimodal data feature array DF = {DF1, DF2, ..., DF} j }, j represents the multimodal data feature array DF and the multimodal data array D. i The length of DF1 represents the multimodal data array D for the i-th work scenario. i The data features of the first stored multimodal data, DF2 represents the multimodal data array D of the i-th working scenario. i The data features of the second multimodal data stored, DF j D represents the multimodal data array of the i-th work scenario. i The data features of the j-th multimodal data stored, ExecFE(·) is the data feature extraction method based on multimodal type; (5) Repeat steps (1) to (4), appending the multimodal data feature array obtained in subsequent traversals to the end of the multimodal data feature array obtained in the previous traversal, until all multimodal data arrays have been traversed. The method for appending the multimodal data feature array DF obtained in subsequent traversals to the end of the previous traversal is as follows: DF j =AppendEnd(DF j-1 ) Here, `AppendEnd()` is a method that appends the multimodal data feature array `DF` obtained from subsequent traversals to the end of the previous traversal. After the traversal ends, the length of `DF` is... getLen(·) is a function to calculate the length, where i represents the sequence number of the work scene; Step S104: After feature extraction, the multimodal data feature array is integrated to obtain a knowledge base collected from different work scenarios. The specific steps for integrating the multimodal data feature array are as follows: First, obtain the feature array DF for the multimodal data; Then, obtain the length of the multimodal data feature array DF and denote it as len. Finally, feature integration is performed on the multimodal data feature array DF to obtain the knowledge base DK collected from different working scenarios. The feature integration method is as follows: The Process(·) method is a feature integration operation. The knowledge base DK collected from different work scenarios is a one-row, len-column matrix with len multimodal data feature dimensions. DK1 is the first knowledge in the knowledge base DK collected from different work scenarios, DK2 is the second knowledge in the knowledge base DK collected from different work scenarios, and DK... len The getL(·) method retrieves the number of knowledge items collected in the knowledge base DK for different work scenarios.
3. The multimodal intelligent terminal interaction method for work scenarios according to claim 2, characterized in that: In step S103, the specific steps of the feature extraction method based on multimodal data are as follows: Different feature extraction operations are performed depending on the multimodal type (DT): If the multimodal type DT is 0, it indicates that the multimodal data D j i For text data, text TD is a set of characters, where TD = {TD1, TD2, ..., TD}. k }, where TD1 is the first character in the text TD, TD2 is the second character in the text TD, and TD... k Let k be the k-th character in the text TD, where k represents the length of the text TD; then, the text TD is used for feature extraction according to the following formula: V={V1,V2,…,V k }={Freq(TD1,TD),Freq(TD2,TD),...,Freq(TD k ,TD)} Where V is a word frequency table with length k, Freq(TD1, TD) is represented by V1, which represents the frequency of word TD1 in word frequency table V, and Freq(TD2, TD) is represented by V2, which represents the frequency of word TD2 in word frequency table V. k ,TD) use V k It indicates that the character TD represents... k The frequency of a word appearing in the word frequency table V, Freq(·) represents obtaining the TD. k The function of frequency of occurrence in text TD, ‖V‖ is the norm of the text frequency table, s represents the current loop number of the summation operation, δ is a hyperparameter set by the system to balance the possible errors in the norming and normalization process, and finally DF is assigned the value of the normalized feature of the text frequency table. If the multimodal type DT is 1, it indicates that the multimodal data D j i Given image data, where image IM is a collection of pixels (PIX), and assuming the length of image IM is a and the width is b, then: Here, IM is composed of pixels, and IM is a row and b column determinant. The image is then processed to obtain image feature values. The subsequent processing steps are as follows: First, using an n*n convolutional layer, RGB channel convolution kernels are used to extract the scale features of the image in RGB, resulting in n preliminary feature maps FI1, FI2, ..., FI at different levels. n ; Then, in the preliminary feature maps FI1,FI2,…,FI n Using convolution, we obtain sub-feature values FE1, FE2, ..., FE at different depths. n ; In the preliminary feature maps FI1, FI2, ..., FI n The convolution operation is performed on the top, as shown below: FE1=(FI1*F)+w1 FE2=(FI2*F)+w2 …… FE n =(FI n *F)+w n Where F represents the convolution operation with dimensions (a, b), w n This represents the depth value of the nth sub-feature; Finally, the sub-eigenvalues FE1,FE2,…,FE n Feature fusion is performed to obtain the feature values FE of the image: Among them, max(·), min(·), and avg(·) methods respectively take n sub-feature values FE1, FE2, ..., FE n The maximum, minimum, and average values are determined by the method, where δ represents the hyperparameter set by the system; finally, DF is assigned as a feature of the image.
4. The multimodal intelligent terminal interaction method for work scenarios according to claim 1, characterized in that: Step S2 specifically includes: Step S201: Based on the knowledge base collected for different work scenarios, use relevant data cleaning methods to remove noisy data from the knowledge base. The specific steps of data cleaning are as follows: First, identify and delete duplicate records to ensure the uniqueness of each characteristic, as follows: For knowledge bases (DKs) collected for different work scenarios, there are: DK = Trim(DK) The Trim(·) method is used to identify and delete duplicate records. Then, the missing features are filled using the average imputation method; Furthermore, excessively high outlier values are removed, as follows: DK = DeleteHV(DK) The DeleteHV(·) method is used to identify and delete excessively high abnormal feature values. Finally, check the data quality to ensure that the cleaning process has not introduced any new errors; Step S202: The cleaned knowledge base is then standardized to obtain a standardized knowledge base. The standardization process is as follows: First, acquire the knowledge base DK1, DK2, ..., DK collected from different work scenarios. len ; Then, logarithmic transformation is used to transform the knowledge DK1, DK2, ..., DK. len Standardization processes are performed, including: DK len =ln(DK len ) The ln(·) method is used to calculate DK1, DK2, ..., DK len The natural logarithm; Finally, the standardized DK1, DK2, ..., DK len Reassign the value to DK; Step S203: Perform data fusion on the standardized knowledge base to obtain the full-work-scenario knowledge base WSK. The data fusion process is as follows: First, the knowledge base DK is trained independently using the WSRF and WSBoost methods, where: ResRF = WSRF(DK) ResBoost = WSBoost(DK) Wherein, ResRF represents the set of prediction results obtained using the WSRF method, ResBoost represents the set of prediction results obtained using the WSBoost method, and WSRF(·) and WSBoost(·) are both prediction models; Then, a weighted average is performed on ResRF and ResBoost as follows: WSK=w ResRF *ResRF+w ResBoost *ResBoost Among them, w ResRF For the weights of ResRF, w ResBoost These are the weights for ResBoost; Finally, we obtained the WSK (Work Scenario Knowledge Base).
5. The multimodal intelligent terminal interaction method for work scenarios according to claim 4, characterized in that: In step S203, the prediction model is defined as: For the WSRF method, we have: The model contains a total of Z subtrees, where i represents the index of the current subtree, and w... i h is the weight of the i-th subtree. i (DK) represents the prediction result of the i-th subtree, indicating the support for the accuracy of knowledge DK; For the WSBoost method, we have: Among them, w len is the weight of the len-th knowledge, reflecting the importance of that knowledge in the work scenario, and l(·) is the cross-entropy loss function. Let f be the predicted value of the len-th knowledge, t be the number of base models, and f t R is the output of the t-th base model, R(i) is the penalty term of the i-th work scenario rule, and Ω(·) represents the method for calculating the regularization term, which is used to prevent overfitting.
6. The multimodal intelligent terminal interaction method for work scenarios according to claim 1, characterized in that: Step S3 specifically includes: Step S301: Obtain a multimodal data array D from the multimodal data list D. i For multimodal data array D i Obtain its multimodal data D1 i D2 i ,…,D j i From multimodal data D1 i D2 i ,…,D j i The process of obtaining time information is as follows: Time information can be denoted as Time, and Time can be represented as: Time=RequireTimeFD(D1 i ,D2 i ,…,D j i ) Among them, RequireTimeFD(·) is derived from multimodal data D1 i D2 i ,…,D j i Methods for obtaining time information in [the context of a sentence]; Step S302: Based on time information, the knowledge base of the entire work scenario is fused to obtain a knowledge base of the entire work scenario with fused time information. The specific steps for fusing the knowledge base of the entire work scenario based on time information are as follows: First, based on the full work scenario knowledge base WSK, a one-to-one correspondence is established with the time information Time, resulting in the full work scenario knowledge base OWSK sorted by time, where: OWSK = OrderBT(WSK, Time) The OrderBT(·) method sorts the entire work scenario knowledge base WSK according to the time information Time. Then, the time interval GAP is set as a benchmark for subsequent correlation prediction; Secondly, extract time series features (TSF) from each time interval. g Using the average value as a time series feature, a time series feature array TSF is constructed, where: TSF={TSF1,TSF2,…,TSF g }=GetFBT(OWSK,GAP) The GetFBT(·) method retrieves time series features (TSFs) according to time intervals. g The method is as follows: g is the current time interval number, where TSF1 represents the time series feature corresponding to the first time interval, TSF2 represents the time series feature corresponding to the second time interval, and TSF... g This represents the time series feature corresponding to the g-th time interval; Then, by weighting the time series features corresponding to different time windows, it can be expressed as: FTSF={w1*TSF1,w2*TSF2,…,w g *TSF g } Where FTSF is the weighted array of time series features, w1 is the weight of the time series feature corresponding to the first time interval, w2 is the weight of the time series feature corresponding to the second time interval, and w... g The weights of the time series features corresponding to the g-th time interval are: Finally, the FTSF is merged with the OWSK, a full work scenario knowledge base sorted by time, to obtain the full work scenario knowledge base TWSK = FTSF ⊕ OWSK with merged time information.
7. The multimodal intelligent terminal interaction method for work scenarios according to claim 6, characterized in that: In step S302, time series features (TSF) are obtained according to time intervals. g The steps are as follows: First, the OWSK knowledge base, which is sorted by time according to the time interval GAP, is divided into g sub-knowledge bases {OWSK1, OWSK2, ..., OWSK}. g }; Then, the sub-knowledge base {OWSK1,OWSK2,…,OWSK} g OWSK g Contains 1 to j different knowledge DK j Using the averaging method, the time series features (TSF) within each sub-knowledge base are calculated. g ,in: Among them, DK j ∈OWSK g And n = SizeOf(OWSK) g ) The SizeOf(·) method is used to obtain the length of a sub-knowledge base segment.
8. The multimodal intelligent terminal interaction method for work scenarios according to claim 1, characterized in that: Step S4 specifically includes: S401: Perform oversampling using a full-work-scene knowledge base that integrates time information to obtain a TWSK (Temporally Written Knowledge Base) of integrated time information. oversampled The specific steps of oversampling are as follows: The TWSK knowledge base for the entire working scenario, which incorporates time information, is oversampled using the oversampling method, as shown below: Among them, TWSK oversampled N is a knowledge base for the entire working scenario, formed by fusing time information after oversampling. minority TWSK represents the number of minority class samples, λ is the interpolation coefficient. i This represents the i-th data sample; Step S402: Fine-tune the model and output the large language model WLLM for the working scenario. The specific steps are as follows: Apply a work-scene attention mechanism based on the contextual information of the work scenario. (1) TWSK, a knowledge base for the entire working scenario of fused time information after oversampling. oversampled Perform location encoding to obtain specific location encoding information, and map the input data into a vector representation of EMB; (2) Perform weight matrix w on the vector representation EMB. QU 、w KE and w VA The multiplication operation yields a query, key, and value vector, denoted as QU, KE, and VA, respectively. (3) Perform a full-connect operation on the query and key vectors based on the work scenario to obtain the model's attention to different work scenarios, where: Among them, S i To indicate the level of attention given to the i-th scene, Softmax(·) is the Softmax activation function, C represents the scene context information, and w i Let be the weight of the i-th scene, and DIM be the TWSK (Time-Based Knowledge Base) for the entire working scene after oversampling and fusion of temporal information. oversampled The dimension, KE T The matrix representing the transpose of the key vector KE; (4) The degree of attention S to the i-th scene i Multiplying the vectors VA with the same value yields the output O = S of the fully connected layer. i *VA; (5) Repeat steps (2) to (4) to adjust the weight matrix w. QU 、w KE and w VA This continues until the final output O converges; (6) Output the large language model WLLM for the work scenario.
9. The multimodal intelligent terminal interaction method for work scenarios according to claim 1, characterized in that: In step S402, the specific steps for obtaining scene context information are as follows: For each scenario SC i Extract its sequence number information to form a vector Cs; scene data Perform a vector concatenation operation with Cs to obtain the scene context information C, as follows: C = Concat(DK, Cs) The Concat(·) function performs vector concatenation operations.
10. A multimodal intelligent terminal interaction method for work scenarios according to claim 1, characterized in that: Step S5 specifically includes: Step S501: Use a data collector to acquire the question Q input by the user on the smart terminal, use the Working Scenario Large Language Model (WLLM) to answer the question, and obtain the output of the Working Scenario Large Language Model. The specific operation of answering the question in the Working Scenario Large Language Model is as follows: First, a data collector is used to obtain the user's question Q on the smart terminal, where: Q = GetOpt() GetOpt() is a method for obtaining user input on a smart terminal; Then, the question Q is input into the Work Scenario Large Language Model (WLLM) for intelligent answering; Finally, the output of the WLLM (Work Scenario Large Language Model) is obtained as follows: Output = GetWLLM(Q) GetWLLM(·) is a method for obtaining answers to questions using a large language model based on the work scenario; Step S502: Return the output of the large language model of the work scenario to the multimodal intelligent terminal and provide an external response.