Large-model-based method and apparatus for acquiring multi-target prediction model, and large-model-based method and apparatus for predicting product data

By combining large language models and structured data to process user dialogue data and generate target prediction models, the problem of inaccurate user intention analysis in the existing technology is solved, the service quality and product conversion rate are improved, and the user complaint rate is reduced.

WO2025146228A1PCT designated stage Publication Date: 2025-07-10LINGXI TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/083191
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-19
Filing Date
2025-03-18
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The prior art cannot conduct comprehensive and accurate analysis of user intentions in application software services and telephone communication services, resulting in the inability to provide accurate high-quality service experience and the inability to effectively predict user satisfaction with the product.

Method used

Through the target large language model, multiple rounds of conversation data of users and intelligent customer service are processed, semantic information sequences are obtained, and structured data in business applications are combined and optimized using the expert network and task network in the initial prediction model to generate a target prediction model to achieve accurate prediction of user satisfaction.

Benefits of technology

It realizes accurate prediction of users' satisfaction with products, improves service quality and product conversion rate, and reduces user complaint rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025083191_10072025_PF_FP_ABST
    Figure CN2025083191_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of data prediction. Provided are a large-model-based method and apparatus for acquiring a multi-target prediction model, and a large-model-based method and apparatus for predicting product data. The large-model-based method for acquiring a multi-target prediction model comprises: using a target large language model to process multi-turn dialogue data between a user and an intelligent customer service, so as to acquire a semantic information sequence; acquiring a structural vector corresponding to structured data related to the user; fusing the semantic information sequence and the structural vector to acquire a dialogue semantic vector; inputting the structural vector into each expert network of an initial prediction model, so as to obtain a plurality of intermediate vectors; fusing the plurality of intermediate vectors and the dialogue semantic vector, and then inputting a fused vector into each task network of the initial prediction model, so as to acquire a prediction result for each task network; and on the basis of the prediction results and a loss function, optimizing the initial prediction model to acquire a target prediction model. The embodiments of the present disclosure can effectively improve the product conversion rate, and can provide high-quality service experience for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for obtaining multi-objective prediction model and product data prediction based on large model

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims priority to Chinese patent application number 2024103143797, filed with the Chinese Patent Office on March 19, 2024, entitled “Method and device for obtaining multi-objective prediction model and product data prediction based on large model”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0003] The present disclosure relates to the field of data prediction technology, and in particular to a method and device for obtaining a multi-objective prediction model and product data prediction based on a large model. Background Art

[0004] With the continuous development of artificial intelligence technology, intelligent customer robots are widely used in business areas related to product marketing.

[0005] Currently, product marketing service systems primarily consist of application software services and telephone communication services. To meet daily user needs, these two distinct service models are typically handled independently. Specifically, application software services determine user needs by analyzing behavioral characteristics, while telephone communication services determine user needs by analyzing the content of conversations. However, in practice, separate user intent analysis and targeting for application software and telephone communication services makes it impossible to fully predict user satisfaction with service content, nor can it provide users with a precise, high-quality service experience.

[0006] Therefore, how to provide a technical solution to accurately locate user intentions and improve service quality has become a technical problem that needs to be solved urgently.

[0007] Application Contents

[0008] The purpose of some embodiments of the present disclosure is to provide a method and device for obtaining a multi-objective prediction model and product data prediction based on a large model. Through the technical solutions of the embodiments of the present disclosure, it is possible to predict whether the user is satisfied with the product, accurately locate the user's intention, improve service quality and product conversion rate, and reduce user complaint rate.

[0009] Some embodiments of the present disclosure provide a method for obtaining a multi-objective prediction model based on a large model, including: using a target large language model to process multi-round conversation data between a user and an intelligent customer service representative to obtain a semantic information sequence; obtaining a structure vector corresponding to structured data related to the user, wherein the structured data is collected from a business application; fusing the semantic information sequence and the structure vector to obtain a conversation semantic vector; inputting the structure vector into each expert network of an initial prediction model to obtain a plurality of intermediate vectors; fusing the plurality of intermediate vectors and the conversation semantic vector and inputting them into each task network of the initial prediction model respectively to obtain prediction results of the respective task networks; optimizing the initial prediction model based on the prediction results and a loss function to obtain the target prediction model, wherein the target prediction model is used to predict product data, and the product data represents the user's satisfaction with the product.

[0010] Some embodiments of the present disclosure use a target large language model to obtain a semantic information sequence corresponding to conversation data. After obtaining a structural vector corresponding to the structured data, the two are fused to produce a conversation semantic vector. The conversation semantic information is combined with the structural vector, and inference and prediction are performed using multiple expert networks and multiple task networks in the initial prediction model. The initial prediction model is optimized to produce the final target prediction model. By fusing user conversation data with structured data to optimize the initial prediction model, some embodiments of the present disclosure can subsequently predict user satisfaction with a product, accurately pinpoint user intent, improve service quality and product conversion rates, and reduce user complaint rates.

[0011] In some embodiments, the use of the target large language model to process multi-round conversation data between the user and the intelligent customer service to obtain a semantic information sequence includes: segmenting the multi-round conversation data through a preset sliding window to obtain multiple conversation segments; using the target large language model to vectorize the multiple conversation segments to obtain multiple vectors, wherein the multiple vectors constitute the semantic information sequence.

[0012] Some embodiments of the present disclosure segment the conversation data through a preset sliding window and then use a target large language model to perform vectorization processing to obtain multiple vectors, thereby effectively acquiring the user's conversation data.

[0013] In some embodiments, before segmenting the multi-round conversation data using a preset sliding window to obtain multiple conversation segments, the method further includes:

[0014] Merge the consecutive sentences spoken by the same character in each round of dialogue data.

[0015] In some embodiments, the fusing of the semantic information sequence and the structural vector to obtain the conversation semantic vector includes: processing the semantic information sequence using a single-layer linear network to obtain a semantic information sequence dimensionally aligned with the structural vector; performing inner product calculation and weighted processing on the structural vector with each vector in the semantic information sequence dimensionally aligned with the structural vector to obtain the conversation semantic vector.

[0016] Some embodiments of the present disclosure can achieve effective data fusion by performing dimension alignment processing on the semantic information sequence and calculating the structure vector to obtain a complete conversation semantic vector.

[0017] In some embodiments, the initial prediction model includes multiple expert networks, multiple gating networks, and multiple task networks; wherein the number of the expert networks and the number of the task networks are determined according to the modeling objectives, and the gating network is used to perform weighted fusion on the data output by the multiple expert networks and the dialogue semantic vectors, and send the weighted fusion results to the corresponding task networks.

[0018] In some embodiments, the fusing of the multiple intermediate vectors and the conversation semantic vectors and inputting them into respective task networks of the initial prediction model to obtain prediction results of the respective task networks includes: using respective gating networks of the initial prediction model to perform weighted calculations on the multiple intermediate vectors and the conversation semantic vectors to obtain respective fused vectors; and inputting the respective fused vectors into the respective task networks to obtain prediction results of the respective task networks.

[0019] Some embodiments of the present disclosure process conversation semantic vectors through a gating network and a task network to obtain prediction results, providing data support for subsequent model optimization.

[0020] In some embodiments, the initial prediction model is optimized based on the prediction result and the loss function to obtain the target prediction model, including: inputting the prediction result into the loss function to obtain a loss value, wherein the weight value in the loss function is set by the product conversion rate and the product complaint rate; optimizing the initial prediction model using the loss value until the loss value meets a preset condition, and outputting the target prediction model; wherein the semantic information sequence remains unchanged during the optimization process of the initial prediction model.

[0021] Some embodiments of the present disclosure optimize the initial prediction model through the loss value output by the loss function to obtain the target prediction model, which can achieve accurate prediction of whether the user is satisfied with the product and improve the product marketing conversion rate.

[0022] Some embodiments of the present disclosure also provide a method for product data prediction based on a large model, comprising: using a target large language model to process multi-round conversation data between a user to be predicted and an intelligent customer service representative to obtain a semantic information sequence; obtaining a structural vector corresponding to structured data related to the user to be predicted, wherein the structured data includes: user information, and user business information related to the product; inputting the semantic information sequence and the structural vector into a target prediction model obtained by the method described in any method embodiment of the first aspect to obtain a predicted value of the product, wherein the predicted value represents the satisfaction of the user to be predicted with the product.

[0023] Some embodiments of the present disclosure also provide a device for obtaining a multi-objective prediction model based on a large model, including: an extraction module, configured to use a target large language model to process multiple rounds of conversation data between a user and an intelligent customer service representative to obtain a semantic information sequence; an acquisition module, configured to obtain a structure vector corresponding to structured data related to the user, wherein the structured data is collected from a business application; a fusion module, configured to fuse the semantic information sequence and the structure vector to obtain a conversation semantic vector; a processing module, configured to input the structure vector into each expert network of an initial prediction model to obtain a plurality of intermediate vectors; an output module, configured to fuse the plurality of intermediate vectors and the conversation semantic vector and input them into each task network of the initial prediction model respectively to obtain prediction results of the respective task networks; an optimization module, configured to optimize the initial prediction model based on the prediction results and a loss function to obtain the target prediction model, wherein the target prediction model is used to predict product data, and the product data represents the user's satisfaction with the product.

[0024] Some embodiments of the present disclosure also provide a device for product data prediction based on a large model, including: an extraction module, configured to use a target large language model to process multi-round conversation data between a user to be predicted and an intelligent customer service representative to obtain a semantic information sequence; an acquisition module, configured to obtain a structural vector corresponding to structured data related to the user to be predicted, wherein the structured data includes: user information and user business information related to the product; a prediction module, configured to input the semantic information sequence and the structural vector into a target prediction model obtained by the method described in any method embodiment of the first aspect to obtain a predicted value of the product, wherein the predicted value represents the satisfaction of the user to be predicted with the product.

[0025] Some embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.

[0026] Some embodiments of the present disclosure also provide an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor can implement the method described in any embodiment of the first aspect when executing the program.

[0027] Some embodiments of the present disclosure further provide a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of some embodiments of the present disclosure, the following briefly introduces the drawings required for use in some embodiments of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0029] FIG1 is a system diagram of obtaining a multi-objective prediction model based on a large model according to some embodiments of the present disclosure;

[0030] FIG2 is a flow chart of a method for obtaining a multi-objective prediction model based on a large model provided by some embodiments of the present disclosure;

[0031] FIG3 is a schematic diagram of the structure of an initial prediction model provided by some embodiments of the present disclosure;

[0032] FIG4 is a flowchart of a method for predicting product data based on a large model according to some embodiments of the present disclosure;

[0033] FIG5 is a second flowchart of a method for predicting product data based on a large model provided by some embodiments of the present disclosure;

[0034] FIG6 is a block diagram of an apparatus for obtaining a multi-objective prediction model based on a large model provided by some embodiments of the present disclosure;

[0035] FIG7 is a block diagram of a device for predicting product data based on a large model according to some embodiments of the present disclosure;

[0036] FIG8 is a schematic diagram of an electronic device provided by some embodiments of the present disclosure. DETAILED DESCRIPTION

[0037] The technical solutions in some embodiments of the present disclosure will be described below with reference to the accompanying drawings in some embodiments of the present disclosure.

[0038] It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. At the same time, in the description of this disclosure, the terms "first," "second," etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.

[0039] In related technologies, in telemarketing scenarios across various business sectors (such as banking, insurance, and credit), the business goal is to help users with business needs find suitable product solutions. The entire business service system consists of two components: the application software (APP) within the business scenario and the telephone communication scenario. The former focuses more on daily interactions with users, while the latter focuses on screening high-intent users for targeted reach. For example, in an APP scenario within a certain business sector, users enter their personal information when they first log in to the app. When selecting products, they do so within their personal account. Daily users also engage in various click-and-browse activities within the app. The app assesses the user's personal information level based on their historical product usage behavior. It also recommends products and articles of interest based on their past click-and-browse behavior within the app. In telephone communication scenarios, if the probability of a candidate user completing a transaction and the probability of a complaint regarding a particular product can be estimated offline in advance, it is possible to selectively reach these users, improving telemarketing efficiency and reducing customer complaint rates. For example, users with a high probability of completing a transaction and a low probability of complaining can be prioritized for reach. Based on the user's behavior and communication content in different scenarios, the corresponding user intent can be modeled separately. However, the same user usually has interactive behaviors or communication content with the same product in the above two scenarios. The independent analysis methods of the above scenarios cannot achieve a comprehensive and accurate analysis of user intent, and thus cannot accurately assess user satisfaction with the product.

[0040] In view of this, some embodiments of the present disclosure provide a method for obtaining a multi-target prediction model based on a large model, which extracts semantic information sequences from multiple rounds of conversation data between users and intelligent customer service through a target large language model, processes user-related structured data in business applications to obtain a structural vector; then fuses the two to obtain a conversation semantic vector; inputs the structural vector into multiple expert networks of the initial prediction model to obtain multiple intermediate vectors. Finally, the multiple intermediate vectors and the conversation semantic vector are fused and input into the task network to obtain the prediction results, so as to predict the initial prediction model through the prediction results and the loss function to obtain the target prediction model. The target prediction model can realize the prediction of the user's satisfaction with the product. Among them, the product data corresponding to the satisfaction level can be the product's transaction rate or complaint rate, etc., which is not specifically limited in the embodiments of the present disclosure.

[0041] The overall structure of a system for obtaining a multi-objective prediction model based on a large model provided by some embodiments of the present disclosure is exemplarily described below with reference to FIG1 .

[0042] As shown in Figure 1, some embodiments of the present disclosure provide a system for obtaining a multi-objective prediction model based on a large model. The system may include a terminal 100 and a prediction server 200. Terminal 100 may record multiple rounds of conversation data between a user and an intelligent customer service representative. Business applications on terminal 100 may also store structured user data. The prediction server 200 may obtain the multi-round conversation data and structured data from the terminal 100. The target large language model on the prediction server 200 then extracts semantic information sequences from the multi-round conversation data and processes the structured data to obtain a structure vector. The prediction server 200 then fuses the semantic information sequence and structure vector to obtain a conversation semantic vector. Finally, the final target prediction model is output by combining the structure vector, conversation semantic vector, and the expert network, gating network, and task network in the initial prediction model. The target prediction model can be used to predict user complaint rates and conversion rates (as a specific example of product data), thereby determining user satisfaction with the product or service, enabling precision marketing and improving user experience.

[0043] In some embodiments of the present disclosure, the terminal 100 may be a mobile terminal or a non-portable computer terminal, which is not specifically limited in the embodiments of the present disclosure.

[0044] In order to improve user satisfaction with products and services, the embodiments of the present disclosure can first predict user intentions through a target prediction model so as to provide customized marketing services in the future. Therefore, the implementation process of obtaining a multi-target prediction model based on a large model and performed by the prediction server 200 provided in some embodiments of the present disclosure is exemplified below in conjunction with Figure 2.

[0045] Please refer to FIG2 , which is a flow chart of a method for obtaining a multi-objective prediction model based on a large model provided by some embodiments of the present disclosure. The method for obtaining a multi-objective prediction model based on a large model may include:

[0046] S210: Use the target large language model to process multiple rounds of conversation data between the user and the intelligent customer service to obtain a semantic information sequence.

[0047] For example, in some embodiments of the present disclosure, after fine-tuning the initial large language model to obtain a target large language model that meets the needs of the business scenario, the target large language model can be used to extract the content of multiple rounds of conversations between the user and the intelligent customer service telephone to obtain a semantic information sequence.

[0048] In some embodiments of the present disclosure, S210 may include: segmenting the multi-round dialogue data through a preset sliding window to obtain multiple dialogue segments; using the target large language model to vectorize the multiple dialogue segments to obtain multiple vectors, wherein the multiple vectors constitute the semantic information sequence.

[0049] For example, in some embodiments of the present disclosure, a preset sliding window is used to segment multiple rounds of conversation data in a conversation text. Before segmentation, it is necessary to merge the consecutive sentences spoken by the same character in each round of conversation data. The speaker identity of the first sentence in each segmented conversation segment is customer service (as a specific example of intelligent customer service), and the speaker identity of the last sentence is the user. The window length of the preset sliding window can be adjusted according to the actual modeling effect. For example, the window length of the preset sliding window is: segmentation is performed with 3 rounds of conversation content as one segment. Each of the multiple segmented conversation segments is collected into an ordered array in chronological order. Finally, the business-based fine-tuning LLM (as a specific example of the target large language model) is used to extract the semantic information representation embedding vector (as a specific example of multiple vectors) of each conversation segment in the array to form a semantic information sequence.

[0050] S220: Obtain a structure vector corresponding to structured data related to the user, wherein the structured data is collected from a business application.

[0051] For example, in some embodiments of the present disclosure, after the user authorizes, the user's personal basic information, product usage information, personal rating information and other structured data can be collected through the business application in the business scenario. For example, in a credit scenario, the structured data in the financial credit APP can be the user's personal basic information, credit behavior information, risk control credit assessment, etc. After collection, the structured data can be processed to obtain the corresponding structure vector. Specifically, a target-attention structure similar to the DIN model is used: the embedding representation of structured data such as the user's personal basic information, credit behavior information, risk control credit assessment, etc. is used as the target vector (as a specific example of a structure vector).

[0052] S230: Fusing the semantic information sequence and the structural vector to obtain a dialogue semantic vector.

[0053] For example, in some embodiments of the present disclosure, the target vector and the semantic information sequence are fused and calculated to obtain a conversation semantic representation weighted based on user personalized information (as a specific example of a conversation semantic vector).

[0054] In some embodiments of the present disclosure, S230 may include: processing the semantic information sequence using a single-layer linear network to obtain a semantic information sequence aligned with the structural vector dimension; performing inner product calculation and weighted processing on the structural vector and each vector in the semantic information sequence aligned with the structural vector dimension to obtain the dialogue semantic vector.

[0055] For example, in some embodiments of the present disclosure, a single-layer linear network is used to reduce the dimensionality of the semantic information sequence, aligning its dimensions with those of the structured data embedding representation. Finally, the inner product is performed between the target vector and each vector in the semantic information sequence. The semantic information sequence is weighted according to the inner product to obtain a conversation semantic representation weighted by the user's personalized information.

[0056] S240: Input the structure vector into each expert network of the initial prediction model to obtain multiple intermediate vectors.

[0057] For example, in some embodiments of the present disclosure, the structure of the initial prediction model is shown in Figure 3. Figure 3 includes: multiple expert networks, namely Expert1, Expert2, and Expert3; multiple gating networks, namely Gate1 and Gate2; and multiple task networks, namely Task1 and Task2. The dashed line connecting Gate1 in Figure 3 indicates that Gate1 can perform a weighted fusion of the data output by Expert1, Expert2, and Expert3 and the conversation semantic vector, and then provide it to Task1. The dashed line connecting Gate2 in Figure 3 indicates that Gate1 can perform a weighted fusion of the data output by Expert1, Expert2, and Expert3 and the conversation semantic vector, and then provide it to Task2, to achieve modeling of different objectives. The number of gating networks and task networks is consistent with the number of modeling objectives; the number of expert networks can be set as needed, and the embodiments of the present disclosure are not limited to this. For example, in Figure 3 of the present disclosure, the initial prediction model is trained with the two modeling objectives of user product transaction rate and complaint rate as the modeling objectives to obtain the target prediction model. Specifically, the structure vector is input into Expert1, Expert2 and Expert3 in FIG3 respectively, and three intermediate vectors (as a specific example of multiple intermediate vectors) are output.

[0058] S250 , fusing the multiple intermediate vectors and the dialogue semantic vector and inputting them into each task network of the initial prediction model respectively to obtain prediction results of each task network.

[0059] For example, in some embodiments of the present disclosure, the three intermediate vectors obtained through the network structure of Figure 3 and the conversation semantic vector weighted based on user personalized information are fused and input into Task 1 and Task 2 for prediction, and the corresponding prediction results are output.

[0060] In some embodiments of the present disclosure, S250 may include: using each gating network of the initial prediction model to perform weighted calculations on the multiple intermediate vectors and the dialogue semantic vector to obtain each fusion vector; and inputting each fusion vector into each task network respectively to obtain prediction results of each task network.

[0061] For example, in some embodiments of the present disclosure, the fusion of the intermediate vector and the conversation semantic vector can be achieved through the Gate network of each target. Taking Gate1 as an example, the calculation formula of the fusion vector emb involved is as follows:

[0062] emb=emb conversation semantic vector·g0+embExpert1·g1+embExpert2·g2+embExpert3·g3

[0063] Among them, emb complete conversation semantics is the conversation semantic vector, embExpert1, embExpert2, and embExpert3 are the three intermediate vectors output by the expert network, and g0~g3 are the weight values ​​of Gate1.

[0064] Two fusion vectors can be obtained through Gate1, Gate2 and the above formula, and then the two fusion vectors are respectively input into the corresponding Task1 and Task2 for prediction to obtain two prediction results (as a specific example of the prediction results of each task network).

[0065] S260, optimizing the initial prediction model based on the prediction result and the loss function to obtain the target prediction model, wherein the target prediction model is used to predict product data, and the product data represents the user's satisfaction with the product.

[0066] For example, in some embodiments of the present disclosure, since the aforementioned disclosure simultaneously models two objectives, namely, whether a user completes a transaction and whether a user complains, both of which are binary classification objectives, a weighted cross-entropy loss is used. The initial prediction model is optimized by combining the two prediction results with the loss value output by the loss function to obtain the final target prediction model.

[0067] In some embodiments of the present disclosure, S260 may include: inputting the prediction result into the loss function to obtain a loss value, wherein the weight value in the loss function is set by the product conversion rate and the product complaint rate; using the loss value to optimize the initial prediction model until the loss value meets the preset conditions, and outputting the target prediction model; wherein the semantic information sequence remains unchanged during the optimization process of the initial prediction model.

[0068] For example, in some embodiments of the present disclosure, for a telephone communication scenario, the conversion rate of a product over a certain period of time in history is statistically calculated as A and the complaint rate is calculated as B (as a specific example of a weight value). Based on this, the following multi-objective loss function is designed: wherein pi is the prediction result obtained by Task 1 prediction of the product's conversion rate, and qi is the prediction result obtained by Task 2 prediction of the product's complaint rate. The initial prediction model in FIG3 is optimized by the loss, and when the loss meets the set threshold (as a specific example of a preset condition), the target prediction model is output. In order to prevent the complete conversation semantic representation extracted by the LLM from being over-generalized by the deep network and losing the memory of the conversation semantic information, the update of the embedding vector of the frozen semantic information sequence is carried out during the iterative training of the network parameters of the initial prediction model. That is, the semantic information sequence is only involved in the calculation and will not be updated during the iterative training process.

[0069] By simultaneously modeling both user conversion and complaint goals, we can leverage the learning capabilities of the conversion goal, which has more samples, and transfer them to the complaint goal, which has fewer samples, effectively reducing the complaint rate. It is understood that the modeling goals can be selected as needed, and the embodiments of this disclosure are not limited to this.

[0070] After completing the training of the above-mentioned target prediction model, the specific process of product data prediction provided by some embodiments of the present disclosure is exemplarily described below with reference to FIG4 .

[0071] In order to improve precision marketing and service quality for users, a target prediction model can be used to predict user satisfaction with a particular product. By predicting the probability of user complaints, customized services can be provided to users to improve user experience. Therefore, please refer to Figure 4, which is a flow chart of a method for product data prediction based on a large model provided by some embodiments of the present disclosure. This method for product data prediction based on a large model may include:

[0072] S410: Process the multi-round conversation data between the predicted user and the intelligent customer service using the target large language model to obtain a semantic information sequence.

[0073] For example, in some embodiments of the present disclosure, the collected conversation data between the user to be predicted and the intelligent customer service telephone is segmented and processed, and the corresponding semantic information sequence is extracted using the target large language model.

[0074] S420 , obtaining a structure vector corresponding to structured data related to the user to be predicted, wherein the structured data includes user information and user service information related to a product.

[0075] For example, in some embodiments of the present disclosure, user information and user service information of the user to be predicted after authorization are collected from the service application of the corresponding product to obtain the corresponding structure vector.

[0076] S430: Input the semantic information sequence and the structural vector into a target prediction model to obtain a predicted value of the product, wherein the predicted value represents the user's satisfaction with the product.

[0077] For example, in some embodiments of the present disclosure, the semantic information sequence and the structural vector are input into the target prediction model obtained by the method embodiments shown in Figures 2 and 3 to obtain a predicted value. This predicted value can represent the user's complaint rate (that is, the satisfaction level of the user to be predicted). When the complaint rate exceeds a set value, the user can be given special attention and customized services can be provided to the user to improve service quality and user experience and reduce the probability of complaints.

[0078] The specific process of product data prediction based on a large model provided by some embodiments of the present disclosure is exemplarily described below with reference to FIG5 .

[0079] Please refer to Figure 5, which is a flow chart of a method for product data prediction based on a large model provided by some embodiments of the present disclosure.

[0080] The above implementation process is described below by way of example.

[0081] S501, dividing the multi-round dialogue data into multiple dialogue segments using a preset sliding window to obtain multiple dialogue segments.

[0082] S502: Vectorize the multiple dialogue segments using the target large language model to obtain a semantic information sequence.

[0083] S503: Obtain a structure vector corresponding to the structured data related to the user.

[0084] S504: Process the semantic information sequence using a single-layer linear network to obtain a semantic information sequence aligned with the structural vector dimension.

[0085] S505 , performing inner product calculation and weighting processing on the structural vector and each vector in the semantic information sequence aligned with the structural vector dimension to obtain a dialogue semantic vector.

[0086] S506: Input the structure vector into each expert network of the initial prediction model to obtain multiple intermediate vectors.

[0087] S507 , using each gating network of the initial prediction model to perform weighted calculations on multiple intermediate vectors and the dialogue semantic vector to obtain each fusion vector.

[0088] S508: Input each fusion vector into each task network to obtain the prediction results of each task network.

[0089] S509: Input the prediction result into the loss function to obtain the loss value.

[0090] S510, optimizing the initial prediction model using the loss value until the loss value meets a preset condition, and then outputting the target prediction model.

[0091] S511: Input the semantic information sequence and structure vector related to the user to be predicted into the target prediction model to obtain the predicted value of the product.

[0092] It should be noted that the specific implementation process of S501 to S511 can refer to the method embodiment provided above, and in order to avoid repetition, the detailed description is appropriately omitted here.

[0093] Please refer to Figure 6, which shows a block diagram of the apparatus for obtaining a multi-objective prediction model based on a large model, provided in some embodiments of the present disclosure. It should be understood that the apparatus for obtaining a multi-objective prediction model based on a large model corresponds to the above-mentioned method embodiment and is capable of performing each step involved in the above-mentioned method embodiment. The specific functions of the apparatus for obtaining a multi-objective prediction model based on a large model can be found in the description above. To avoid repetition, a detailed description is omitted here.

[0094] The device for obtaining a multi-target prediction model based on a large model in FIG6 includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the device for obtaining a multi-target prediction model based on a large model. The device for obtaining a multi-target prediction model based on a large model includes: an extraction module 610, configured to use a target large language model to process multiple rounds of conversation data between a user and an intelligent customer service representative to obtain a semantic information sequence; an acquisition module 620, configured to obtain a structure vector corresponding to structured data related to the user, wherein the structured data is collected from a business application; a fusion module 630, configured to process the semantic information sequence and the structure vector. The structure vectors are fused to obtain the dialogue semantic vector; the processing module 640 is configured to input the structure vector into each expert network of the initial prediction model to obtain multiple intermediate vectors; the output module 650 is configured to fuse the multiple intermediate vectors and the dialogue semantic vector and input them into each task network of the initial prediction model respectively to obtain the prediction results of the each task network; the optimization module 660 is configured to optimize the initial prediction model based on the prediction results and the loss function to obtain the target prediction model, wherein the target prediction model is used to predict product data, and the product data represents the user's satisfaction with the product.

[0095] Please refer to Figure 7, which shows a block diagram of an apparatus for predicting product data based on a large model, as provided in some embodiments of the present disclosure. It should be understood that this apparatus for predicting product data based on a large model corresponds to the aforementioned method embodiment and is capable of executing each of the steps involved in the aforementioned method embodiment. The specific functions of this apparatus for predicting product data based on a large model can be found in the description above, and a detailed description is omitted here to avoid repetition.

[0096] The device for product data prediction based on a big model in Figure 7 includes at least one software functional module that can be stored in a memory in the form of software or firmware or solidified in the device for product data prediction based on a big model. The device for product data prediction based on a big model includes: an extraction module 710, configured to use a target big language model to process multiple rounds of conversation data between the user to be predicted and the intelligent customer service to obtain a semantic information sequence; an acquisition module 720, configured to obtain a structural vector corresponding to the structured data related to the user to be predicted, wherein the structured data includes: user information and user business information related to the product; a prediction module 730, configured to input the semantic information sequence and the structural vector into a target prediction model to obtain a predicted value of the product, wherein the predicted value represents the satisfaction of the user to be predicted with the product.

[0097] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0098] Some embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations corresponding to any of the above methods provided in the above embodiments.

[0099] Some embodiments of the present disclosure further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.

[0100] As shown in Figure 8, some embodiments of the present disclosure provide an electronic device 800, which includes: a memory 810, a processor 820, and a computer program stored in the memory 810 and executable on the processor 820, wherein the processor 820 reads the program from the memory 810 through a bus 830 and executes the program to implement a method as in any of the above embodiments.

[0101] The processor 820 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, the processor 820 can be a microprocessor.

[0102] The memory 810 can be configured to store instructions executed by the processor 820 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all of the functions of one or more modules described in the embodiments of the present disclosure. The processor 820 of the embodiment of the present disclosure can be configured to execute the instructions in the memory 810 to implement the method shown above. The memory 810 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memory known to those skilled in the art.

[0103] The above description is merely an embodiment of the present disclosure and is not intended to limit the scope of protection of the present disclosure. For those skilled in the art, the present disclosure may be subject to various modifications and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure. It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.

[0104] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

[0105] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element. Industrial Applicability

[0106] The disclosed embodiments provide a method and apparatus for obtaining a multi-objective prediction model and product data prediction based on a large model, which can predict whether a user is satisfied with a product, accurately locate user intentions, improve service quality and product conversion rate, and reduce user complaint rates.

Claims

1. A method for obtaining a multi-object prediction model based on a large model, characterized in that, Including: Processing multi-turn conversation data between a user and an intelligent customer service using a target large language model to obtain a semantic information sequence; Obtaining a structure vector corresponding to structured data related to the user, where the structured data is collected from a business application; Fusing the semantic information sequence and the structure vector to obtain a conversation semantic vector; Inputting the structure vector into each expert network of an initial prediction model to obtain multiple intermediate vectors; Fusing the multiple intermediate vectors and the conversation semantic vector and then respectively inputting them into each task network of the initial prediction model to obtain prediction results of each task network; Optimizing the initial prediction model based on the prediction results and a loss function to obtain a target prediction model, where the target prediction model is used to predict product data, and the product data represents the user's satisfaction with the product.

2. The method according to claim 1, wherein The processing the multi-turn conversation data between a user and an intelligent customer service using a target large language model to obtain a semantic information sequence includes: Segmenting the multi-turn conversation data through a preset sliding window to obtain multiple conversation segments; Using the target large language model to perform vectorization processing on the multiple conversation segments to obtain multiple vectors, where the multiple vectors constitute the semantic information sequence.

3. The method according to claim 2, wherein Before segmenting the multi-turn conversation data through a preset sliding window to obtain multiple conversation segments, the method further includes: Merging each sentence continuously spoken by the same role in each turn of conversation data.

4. The method according to claim 1 or 2, characterized in that The fusing the semantic information sequence and the structure vector to obtain a conversation semantic vector includes: Processing the semantic information sequence using a single-layer linear network to obtain a semantic information sequence aligned with the dimension of the structure vector; Performing inner product calculation and weighting processing on the structure vector and each vector in the semantic information sequence aligned with the dimension of the structure vector to obtain the conversation semantic vector.

5. The method according to claim 1, wherein The initial prediction model includes multiple expert networks, multiple gating networks, and multiple task networks; Among them, the number of expert networks and the number of task networks are determined according to the modeling objective, and the gating network is used to perform weighted fusion on the data output by the multiple expert networks and the conversation semantic vector and send the weighted fusion result to the corresponding task network.

6. The method according to claim 1 or 2, characterized in that The fusing the multiple intermediate vectors and the conversation semantic vector and then respectively inputting them into each task network of the initial prediction model to obtain prediction results of each task network includes: Using each gating network of the initial prediction model to perform weighted calculation on the multiple intermediate vectors and the conversation semantic vector to obtain respective fusion vectors; Respectively inputting the respective fusion vectors into each task network to obtain prediction results of each task network.

7. The method according to claim 1 or 2, characterized in that, The optimizing the initial prediction model based on the prediction results and a loss function to obtain the target prediction model includes: Inputting the prediction results into the loss function to obtain a loss value, where the weight value in the loss function is set through the product conversion rate and the product complaint rate; Optimize the initial prediction model using the loss value until the loss value meets a preset condition, and then output the target prediction model; wherein, the semantic information sequence remains unchanged during the optimization process of the initial prediction model.

8. A method for predicting product data based on a large model, characterized in that, It includes: Process the multi-round conversation data of the user to be predicted and the intelligent customer service using a target large language model to obtain a semantic information sequence; Obtain a structure vector corresponding to the structured data related to the user to be predicted, where the structured data includes: user information and user service information related to the product; Input the semantic information sequence and the structure vector into the target prediction model obtained by the method according to any one of claims 1 to 5 to obtain a prediction value of the product, where the prediction value represents the satisfaction degree of the user to be predicted with respect to the product.

9. An apparatus for obtaining a multi-object prediction model based on a large model, characterized in that It includes: An extraction module configured to process the multi-round conversation data of the user and the intelligent customer service using a target large language model to obtain a semantic information sequence; An acquisition module configured to obtain a structure vector corresponding to the structured data related to the user, where the structured data is collected from a business application; A fusion module configured to fuse the semantic information sequence and the structure vector to obtain a dialogue semantic vector; A processing module configured to input the structure vector into each expert network of the initial prediction model to obtain a plurality of intermediate vectors; An output module configured to fuse the plurality of intermediate vectors and the dialogue semantic vector and then input them into each task network of the initial prediction model respectively to obtain the prediction results of the respective task networks; An optimization module configured to optimize the initial prediction model based on the prediction results and a loss function to obtain a target prediction model, where the target prediction model is used to predict product data, and the product data represents the satisfaction degree of the user with respect to the product.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, wherein the computer program, when run by a processor, executes the method according to any one of claims 1-8.

11. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor, wherein the computer program, when run by the processor, executes the method according to any one of claims 1-8.

12. A computer program product, characterized in that, The computer program product includes a computer program, wherein the computer program, when run by a processor, executes the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Multimedia resource pushing method and device, model training method and device and storage medium

    CN114363671A

  • Training method of multi-task prediction model, and multi-task prediction method and device

    CN115423016A

  • Dialogue generation method fusing user interests and preferences

    CN115757707A

  • Information recommendation method and device, storage medium and computer equipment

    CN115774814A

  • Conversation generation method fusing basic knowledge and user information

    CN116010575A