Method and device for predicting user behavior information, storage medium and electronic equipment

By introducing cross-entropy loss and information entropy loss into the user profiling model, the problem of low confidence in the user profiling model is solved, and the prediction accuracy and user experience are improved.

CN120821847APending Publication Date: 2025-10-21INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511126723.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing user profiling models have low confidence levels, resulting in low prediction accuracy, especially when dealing with marginal users, where classification is inaccurate.

Method used

A target user profile model is trained using a target loss function composed of cross-entropy loss and information entropy loss. User behavior data is processed by vectorization, and model parameters are optimized within a deep learning framework to reduce model uncertainty.

Benefits of technology

It improved the accuracy of user profiling models, reduced misjudgments of marginal users, lowered resource consumption, enhanced user experience, and prevented user churn.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821847A_ABST
    Figure CN120821847A_ABST
Patent Text Reader

Abstract

The invention discloses a user behavior information prediction method and device, a storage medium and electronic equipment. The method relates to the field of financial science and technology, and comprises the following steps: collecting behavior data of a target user under the condition of obtaining authorization of the target user; vectorizing the behavior data to obtain vector representation of the behavior data; the vector representation of the behavior data is input into a target user portrait model for behavior information prediction, a prediction result is obtained, the target user portrait model is obtained through training of a target loss function composed of cross entropy loss and information entropy loss, and the prediction result is used for representing whether a target user has a target behavior or not. Through the method and the device, the problem of low prediction accuracy of the model caused by low confidence coefficient of the user portrait model in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of financial technology, and specifically, to a method and device for predicting user behavior information, a storage medium, and an electronic device. Background Art

[0002] User profiling is a broad category of tasks, which can be categorized into tasks such as educational background prediction and interest classification, depending on their objectives. Existing user profiling models use a classification threshold to determine the final result. Specifically, the model outputs a probability value. If the value is greater than the threshold, the user is considered a positive sample; otherwise, it is considered a negative sample. This simple threshold approach ignores the uncertainty of the model's predictions. When the model's confidence in the answer is low, a large number of marginal users will be misclassified, resulting in low model prediction accuracy.

[0003] Currently, no effective solution has been proposed to address the problem of low confidence in user portrait models in related technologies, which leads to low prediction accuracy of the models. Summary of the Invention

[0004] The main purpose of this application is to provide a method and device for predicting user behavior information, a storage medium and an electronic device to solve the problem in related technologies that the confidence of the user portrait model is low, resulting in low prediction accuracy of the model.

[0005] To achieve the above objectives, according to one aspect of the present application, a method for predicting user behavior information is provided. The method comprises: obtaining authorization from a target user, collecting behavior data of the target user; vectorizing the behavior data to obtain a vector representation of the behavior data; and inputting the vector representation of the behavior data into a target user profile model to predict behavior information and obtain a prediction result. The target user profile model is trained using a target loss function composed of a cross-entropy loss and an information entropy loss. The prediction result is used to indicate whether the target user has engaged in the target behavior.

[0006] Furthermore, a target user portrait model is generated through the following steps: obtaining training data, wherein the training data is obtained after preprocessing the user behavior data; performing word segmentation and word embedding processing on the training data to obtain a vector representation of the training data; training the initial network model based on the vector representation of the training data and the target loss function to obtain the target user portrait model.

[0007] Furthermore, the initial network model is trained based on the vector representation of the training data and the target loss function to obtain the target user portrait model, including: inputting the vector representation of the training data into the initial network model for processing to obtain the probability distribution of the predicted user portrait label; adjusting the model parameters based on the probability distribution and the target loss function until the predetermined convergence conditions are met to obtain the target user portrait model.

[0008] Furthermore, the vector representation of the training data is input into the initial network model for processing to obtain the probability distribution of the predicted user portrait label, including: inputting the vector representation of the training data into the encoder module of the initial network model for encoding to obtain the vector representation of the semantic space; inputting the vector representation of the semantic space into the classifier module of the initial network model for prediction to obtain the probability distribution.

[0009] Furthermore, adjusting the model parameters according to the probability distribution and the target loss function includes: calculating the cross entropy loss between the probability distribution and the true label, and calculating the information entropy loss of the probability distribution; obtaining the loss value of the target loss function based on the cross entropy loss and the information entropy loss; and updating the model parameters through the back propagation algorithm based on the loss value.

[0010] Furthermore, performing word segmentation and word embedding processing on the training data to obtain a vector representation of the training data includes: performing word segmentation processing on the training data to obtain a word-gram sequence corresponding to the training data; and performing word embedding processing on the word-gram sequence to obtain a vector representation of the training data.

[0011] Furthermore, word embedding processing is performed on the word-gram sequence to obtain a vector representation of the training data, including: mapping each word-gram in the word-gram sequence to a pre-trained word embedding matrix to obtain a word embedding vector for each word-gram; appending a corresponding position embedding vector to the word embedding vector of each word-gram to obtain a word embedding vector with position information; and concatenating the word embedding vectors with position information to obtain a vector representation of the training data.

[0012] To achieve the above-mentioned objectives, according to another aspect of the present application, a device for predicting user behavior information is provided. The device comprises: a first acquisition unit for collecting the behavior data of a target user with authorization from the target user; a first processing unit for vectorizing the behavior data to obtain a vector representation of the behavior data; and a second processing unit for inputting the vector representation of the behavior data into a target user profile model to predict behavior information and obtain a prediction result, wherein the target user profile model is trained using a target loss function composed of a cross entropy loss and an information entropy loss, and the prediction result is used to indicate whether the target user has engaged in the target behavior.

[0013] Furthermore, the device also includes the following units, which are used to generate a target user portrait model through the following steps: a second acquisition unit, which is used to acquire training data, wherein the training data is obtained after preprocessing the user behavior data; a third processing unit, which is used to perform word segmentation and word embedding processing on the training data to obtain a vector representation of the training data; and a fourth processing unit, which is used to train the initial network model based on the vector representation of the training data and the target loss function to obtain the target user portrait model.

[0014] Furthermore, the fourth processing unit includes: a first processing sub-unit, used to input the vector representation of the training data into the initial network model for processing to obtain the probability distribution of the predicted user portrait label; a second processing sub-unit, used to adjust the model parameters according to the probability distribution and the target loss function until the predetermined convergence conditions are met to obtain the target user portrait model.

[0015] Furthermore, the first processing sub-unit includes: an encoding module, which is used to input the vector representation of the training data into the encoder module of the initial network model for encoding to obtain the vector representation of the semantic space; a prediction module, which is used to input the vector representation of the semantic space into the classifier module of the initial network model for prediction to obtain the probability distribution.

[0016] Furthermore, the second processing subunit includes: a first calculation module, which is used to calculate the cross entropy loss between the probability distribution and the true label, and calculate the information entropy loss of the probability distribution; a second calculation module, which is used to calculate the loss value of the target loss function based on the cross entropy loss and the information entropy loss; an update module, which is used to update the model parameters through the back propagation algorithm based on the loss value.

[0017] Furthermore, the third processing unit includes: a third processing sub-unit, used to perform word segmentation processing on the training data to obtain a word unit sequence corresponding to the training data; and a fourth processing sub-unit, used to perform word embedding processing on the word unit sequence to obtain a vector representation of the training data.

[0018] Furthermore, the fourth processing sub-unit includes: a first processing module, used to map each word in the word sequence to a pre-trained word embedding matrix to obtain a word embedding vector for each word; a second processing module, used to append a corresponding position embedding vector to the word embedding vector of each word to obtain a word embedding vector with position information; a third processing module, used to splice the word embedding vectors with position information to obtain a vector representation of the training data.

[0019] According to another aspect of an embodiment of the present invention, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein any one of the above-mentioned user behavior information prediction methods is executed when the program is running.

[0020] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, which stores a program, wherein when the program is running, the device where the storage medium is located is controlled to execute any of the above-mentioned user behavior information prediction methods.

[0021] In an embodiment of the present application, the following steps are adopted: with the authorization of the target user, the behavior data of the target user is collected; the behavior data is vectorized to obtain a vector representation of the behavior data; the vector representation of the behavior data is input into the target user portrait model to predict the behavior information and obtain a prediction result, wherein the target user portrait model is trained using a target loss function composed of a cross entropy loss and an information entropy loss, and the prediction result is used to characterize whether the target user has the target behavior. This solves the technical problem in the related art that the confidence of the user portrait model is low, resulting in low prediction accuracy of the model. In this solution, by introducing information entropy into the loss function modeling, the uncertainty of the model can be reduced, and the accuracy of the user portrait model can be improved, thereby improving the model's classification accuracy for marginal users, reducing resource consumption, improving user experience, and avoiding user churn. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0023] Figure 1 A hardware structure block diagram of a computer terminal for implementing a method for predicting user behavior information is shown;

[0024] Figure 2 is a flowchart of a method for predicting user behavior information provided in an embodiment of the present application;

[0025] Figure 3 This is a flowchart of building a user portrait model according to an embodiment of the present application;

[0026] Figure 4 This is a flowchart of a user portrait model training method according to an embodiment of the present application;

[0027] Figure 5 is a schematic diagram of a device for predicting user behavior information according to an embodiment of the present application;

[0028] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation portals for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.

[0032] Example 1

[0033] According to an embodiment of the present application, an embodiment of a method for predicting user behavior information is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0034] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1The hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for predicting user behavior information is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0035] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for predicting user behavior information in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned method for predicting user behavior information. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0037] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0038] The display may be a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0039] Under the above operating environment, this application provides Figure 2 The prediction method of user behavior information shown. Figure 2 Flowchart of a method for predicting user behavior information according to the first embodiment of the present application. The method for predicting user behavior information includes:

[0040] Step S201: After obtaining authorization from the target user, collect the target user's behavior data.

[0041] Optionally, behavioral data includes, but is not limited to, users' browsing history, search history, and click history. This data forms the basis for building user profiles. By analyzing these behavioral patterns, users' needs, preferences, and potential financial activity intentions can be inferred. This step emphasizes compliance and privacy protection principles during the data collection process, ensuring that all operations are conducted in compliance with regulations and with the user's informed consent. Users can review the purpose of data usage in real time through the authorization interface and have the right to withdraw authorization or delete data at any time. Upon withdrawal of authorization, the system will terminate the relevant data processing within 24 hours.

[0042] Step S202: vectorize the behavior data to obtain a vector representation of the behavior data.

[0043] Optionally, the vectorization process includes at least data preprocessing, word segmentation, and word embedding. For example, assuming that the user's historical search information is a text string of length L, the string is first preprocessed, including data filtering, to filter out low-quality words, invalid characters, etc. Then, the text is segmented. For example, the query statement entered by the user is: What is the deposit interest rate of Bank A? The text is segmented using a word segmentation algorithm and converted into the following character sequence: ["Bank A", "deposit", "interest rate", "is", "how much"]. Then, word embedding is performed to convert the character sequence into a dense vector representation, that is, to obtain a vector representation of the behavioral data.

[0044] In step S203, the vector representation of the behavior data is input into the target user portrait model to predict the behavior information and obtain a prediction result, wherein the target user portrait model is trained using a target loss function composed of cross entropy loss and information entropy loss, and the prediction result is used to characterize whether the target user has performed the target behavior.

[0045] Optionally, the vector representation of the behavioral data is input into the target user portrait model for behavioral information prediction, wherein the target user portrait model is trained through the joint optimization of cross entropy loss and information entropy loss under the deep learning framework. The cross entropy loss ensures that the label probability distribution predicted by the model is as consistent as possible with the actual label, and the information entropy loss controls the uncertainty of the predicted probability distribution. The combination of the two significantly reduces the uncertainty of the model and enhances the accuracy of the model.

[0046] For example, suppose we want to predict whether a user will take out a loan. We collect the user's search history on the bank's website, including keywords such as "loan interest rate" and "mortgage calculator." After preprocessing and word segmentation, these keywords are converted into a series of tokens. Then, using word embedding technology, these tokens are converted into vectors and input into the target user profile model. The model predicts a probability of 0.85 for the user to take out a loan. Simultaneously, by calculating information entropy, the model can assess the uncertainty of this prediction. If the information entropy is low, the model's confidence level is high, and the prediction of the user's loan behavior is reliable. Conversely, if the information entropy is high, even a predicted probability of 0.85 indicates that the model is uncertain about the prediction.

[0047] In summary, we collect user-authorized behavioral data, such as browsing history, purchase history, and search keywords on the platform, and convert it into vector form for processing by machine learning models. This vectorized behavioral data captures the complexity and nuances of user behavior, thereby more accurately reflecting user characteristics. The vectors are then fed into a user profiling model, which is trained using a deep learning framework through a joint optimization of cross-entropy and information entropy losses. The cross-entropy loss ensures that the model's predicted label probability distribution is as consistent as possible with the actual label, while the information entropy loss controls the uncertainty of the predicted probability distribution. This combination significantly reduces model uncertainty and enhances model accuracy. This addresses the problem of traditional user profiling models inaccurately classifying marginal users. By reducing misclassification of marginal users, resource consumption is reduced, the user experience is improved, and potential user churn is avoided.

[0048] Optionally, in the method for predicting user behavior information provided in an embodiment of the present application, a target user portrait model is generated by the following steps: obtaining training data, wherein the training data is obtained after preprocessing the user behavior data; performing word segmentation and word embedding processing on the training data to obtain a vector representation of the training data; training the initial network model based on the vector representation of the training data and the target loss function to obtain a target user portrait model.

[0049] In an optional embodiment, in the process of generating a user portrait model, it is first necessary to collect a large amount of user behavior data as a training set. These data may include various types such as text and images, so preprocessing is required, such as cleaning, standardization and other steps to eliminate noise and improve data quality. Then, the text data is segmented to divide the continuous text into independent vocabulary units. For example, "I like shopping at night" is divided into ["I", "like", "at", "night", "shopping"]. Then, through word embedding technology, each word is converted into a high-dimensional vector, that is, a vector representation of the training data is obtained. Then, the target loss function is used to guide the model training so that the model can learn the correlation between the user portrait label and the behavior data, thereby improving the prediction ability of the model.

[0050] Optionally, the objective loss function is as follows:

[0051]

[0052] in, is the information entropy, H(p,q) is the cross entropy, and the formula is as follows:

[0053]

[0054] Where q is the probability distribution of the model output and p is the one-hot encoded vector of the target label.

[0055] It is important to note that, without increasing the cost of model training and inference, this approach explicitly models model uncertainty in an end-to-end manner. This involves modeling the confidence level of the results into the user profile classification model, encouraging the model to reduce its classification information entropy while ensuring correct classification. This significantly reduces model uncertainty and enhances the accuracy of the user profile model. This addresses the issues of non-robust user classification and the anomalous classification of marginal users in user profile models based on maximum likelihood estimation and implemented using natural language processing.

[0056] Optionally, in the user behavior information prediction method provided in the embodiment of the present application, the training data is segmented and word embedded to obtain a vector representation of the training data, including: performing word segmentation on the training data to obtain a word sequence corresponding to the training data; performing word embedding on the word sequence to obtain a vector representation of the training data.

[0057] In an optional embodiment, a text segmentation algorithm is applied to the training data to segment it into a series of words or subwords. A word embedding process is then performed to convert the segmented words or subwords into dense vectors of fixed dimension, forming a vector representation of the training data. For example, for the Chinese text "The weather is really good today," after word segmentation, it becomes ["today," "weather," "real," "good"]. Next, the word embedding step converts these word units into fixed-length vectors. These vectors not only contain the semantic information of the word units themselves, but also take into account the relative position of the word units in the text, achieved by appending positional embedding vectors.

[0058] The original text is decomposed into a series of tokens through word segmentation processing, and the tokens are converted into vector representations through word embedding processing. This helps the model understand and process natural language, enabling it to capture sentence structure and contextual relationships, and ensure the smooth progress of model training.

[0059] Optionally, in the method for predicting user behavior information provided in an embodiment of the present application, word embedding processing is performed on the word-gram sequence to obtain a vector representation of the training data, including: mapping each word-gram in the word-gram sequence to a pre-trained word embedding matrix to obtain a word embedding vector for each word-gram; appending a corresponding position embedding vector to the word embedding vector of each word-gram to obtain a word embedding vector with position information; and splicing the word embedding vectors with position information to obtain a vector representation of the training data.

[0060] In an optional embodiment, word embeddings consist of two parts: word embeddings and position embeddings. Word embeddings convert word-gram sequences into vectors representing word meanings, while position embeddings represent the positional order information between different word-grams. First, each word-gram in the word-gram sequence is mapped to a pre-trained word embedding matrix to obtain a word embedding vector for each word-gram. Then, the corresponding position embedding vector is appended to the word embedding vector of each word-gram to obtain a word embedding vector with positional information. These word embedding vectors with positional information are then concatenated to obtain a vector representation of the training data.

[0061] Optionally, a pre-trained word embedding matrix is ​​typically trained on a large corpus. Each word in the matrix has a corresponding vector, which mathematically reflects the semantic similarity of the words. Positional embedding vectors are added to enhance the model's understanding of sentence structure.

[0062] By concatenating the position embedding vector with the word embedding vector, the model can simultaneously take into account the semantics and position information of the word unit, improving the model's semantic understanding and expression capabilities, thereby improving the overall prediction performance.

[0063] Optionally, in the method for predicting user behavior information provided in an embodiment of the present application, the initial network model is trained based on the vector representation of the training data and the target loss function to obtain the target user portrait model, including: inputting the vector representation of the training data into the initial network model for processing to obtain the probability distribution of the predicted user portrait label; adjusting the model parameters based on the probability distribution and the target loss function until the predetermined convergence conditions are met to obtain the target user portrait model.

[0064] In an optional embodiment, during the network model training phase, the model first receives vectorized training data, and outputs the predicted probability distribution of user portrait labels through forward propagation of a multi-layer neural network. For example, for a two-classification problem, the model outputs a probability of 0.8 for the user to perform a certain behavior and a probability of 0.2 for not performing the behavior. Then, the performance of the model is comprehensively evaluated by calculating the gap between the predicted probability distribution and the actual label, i.e., the cross entropy loss, and the uncertainty of the predicted probability distribution itself, i.e., the information entropy loss. Based on these two losses, the model adjusts the weights and biases in the network through a backpropagation algorithm to minimize the loss function. This process is iterated continuously until the changes in the model parameters tend to stabilize, that is, the convergence conditions are met.

[0065] The cross entropy loss ensures that the label probability distribution predicted by the model is as consistent as possible with the actual label, and the information entropy loss controls the uncertainty of the predicted probability distribution. The combination of the two significantly reduces the uncertainty of the model and enhances the accuracy of the model.

[0066] Optionally, in the method for predicting user behavior information provided in an embodiment of the present application, the vector representation of the training data is input into the initial network model for processing to obtain the probability distribution of the predicted user portrait label, including: inputting the vector representation of the training data into the encoder module of the initial network model for encoding to obtain the vector representation of the semantic space; inputting the vector representation of the semantic space into the classifier module of the initial network model for prediction to obtain the probability distribution.

[0067] In an optional embodiment, the model includes an encoder module and a classifier module, wherein the encoder module can adopt an existing natural language processing model structure to convert a sentence into a vector representation of the semantic space, that is, the vector representation of the training data is input into the encoder module of the initial network model for encoding, and a vector representation of the semantic space can be obtained. The classifier module can be a multi-layer feedforward network, which receives the output of the encoder module and returns a probability value to represent the probability that the sample is a positive sample (positive sample: such as whether the user has a deposit or loan demand), that is, the vector representation of the semantic space is input into the classifier module of the initial network model for prediction to obtain a probability distribution.

[0068] Through the combination of encoders and classifiers, the model can flexibly adapt to different task requirements, enhancing the adaptability and scalability of the model.

[0069] Optionally, in the method for predicting user behavior information provided in an embodiment of the present application, adjusting the model parameters based on the probability distribution and the target loss function includes: calculating the cross entropy loss between the probability distribution and the true label, and calculating the information entropy loss of the probability distribution; obtaining the loss value of the target loss function based on the cross entropy loss and the information entropy loss; and updating the model parameters based on the loss value through the back propagation algorithm.

[0070] In an optional embodiment, the objective loss function is as follows:

[0071]

[0072] in, is the information entropy, H(p,q) is the cross entropy, and the formula is as follows:

[0073]

[0074] Where q is the probability distribution of the model output and p is the one-hot encoded vector of the target label.

[0075] The model's performance is comprehensively evaluated by calculating the gap between the predicted probability distribution and the actual label, namely the cross-entropy loss, and the uncertainty of the predicted probability distribution itself, namely the information entropy loss. Based on these two losses, the model uses a backpropagation algorithm to return the loss value to the specific parameters of each network and optimize each part of the network based on this loss.

[0076] By simultaneously optimizing cross entropy loss and information entropy loss, the model can improve prediction accuracy while reducing prediction uncertainty, making the model more robust when processing complex and ambiguous user behavior data.

[0077] In an optional embodiment, Figure 3This is a flowchart of building a user portrait model according to an embodiment of the present application. Figure 3 As shown, relevant text information is collected from users' historical search and browsing history, and other behavioral data is preprocessed through data filtering, including the removal of duplicate, low-quality, and abnormal data. The processed data is then used as training data, which is then segmented to obtain a series of tokens. This token sequence is then converted into a dense vector representation using word embedding technology. This vector representation is then input into an encoder for main model training, where it is converted into a vector representation in the semantic space. The encoder's output serves as the input to a classifier, which predicts the probability of a user profile label (such as deposit willingness or loan demand). Training is optimized jointly using cross-entropy loss and information entropy loss. Based on these two losses, the model uses a backpropagation algorithm to return the loss value to the specific parameters of each network. Based on this loss, each part of the network is optimized to obtain a trained user profile model, which can then be used for model inference, output classification results, and be applied to various business scenarios.

[0078] In an optional embodiment, Figure 4 This is a flowchart of training a user portrait model according to an embodiment of the present application. Figure 4 As shown, user behavior data is used as training data. For example, suppose a user searches for "What is the deposit interest rate of Bank A" on Bank A's website. The length of this text string is L. The text is segmented using a word segmentation algorithm and converted into the following character sequence: ["Bank A", "deposit", "interest rate", "is", "how much"]. Then, through word embedding, the text string is converted into a vector representation of shape L*H, where H represents the size of the model's hidden layer. The encoder converts it into a vector representation in the semantic space. The encoder output is also an L*H vector representation. The classifier receives the L*H vector representation output by the encoder and returns an L*1 probability value, which is used to represent the probability that the sample is a positive sample (positive sample: for example, whether the user has a deposit demand). Based on cross-entropy loss and information entropy loss, the model returns the loss value to the specific parameters of each network through the backpropagation algorithm and optimizes each part of the network based on this loss.

[0079] The method for predicting user behavior information provided in the embodiment of the present application collects user-authorized behavior data, such as the user's browsing history, purchase history, search keywords, etc. on the platform, and converts it into a vector form for processing by a machine learning model. After vectorization, these behavior data can capture the complexity and subtle differences of user behavior, thereby more accurately reflecting user characteristics. Subsequently, the vector is input into the user portrait model, which is trained through the joint optimization of cross entropy loss and information entropy loss under the deep learning framework. Cross entropy loss ensures that the label probability distribution predicted by the model is as consistent as possible with the actual label, while information entropy loss controls the uncertainty of the predicted probability distribution. The combination of the two significantly reduces the uncertainty of the model and enhances the accuracy of the model. It solves the problem of inaccurate classification of traditional user portrait models when facing marginal users. By reducing the misjudgment of marginal users, it reduces resource consumption, improves user experience, and avoids potential user loss.

[0080] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0081] Example 2

[0082] The present application also provides a device for predicting user behavior information. It should be noted that the device for predicting user behavior information in the present application can be used to execute the method for predicting user behavior information provided in the present application. The following describes the device for predicting user behavior information provided in the present application.

[0083] According to an embodiment of the present application, a device for predicting user behavior information for implementing the above-mentioned method for predicting user behavior information is also provided. Figure 5 As shown, the device includes: a first acquisition unit 501, a first processing unit 502, and a second processing unit 503.

[0084] The first acquisition unit 501 is configured to collect the behavior data of the target user after obtaining authorization from the target user;

[0085] A first processing unit 502 is configured to perform vectorization processing on the behavior data to obtain a vector representation of the behavior data;

[0086] The second processing unit 503 is used to input the vector representation of the behavior data into the target user portrait model to predict the behavior information and obtain a prediction result, wherein the target user portrait model is trained using a target loss function composed of cross entropy loss and information entropy loss, and the prediction result is used to characterize whether the target user has the target behavior.

[0087] The user behavior information prediction device provided in the embodiment of the present application collects the behavior data of the target user through the first acquisition unit 501 with the authorization of the target user; the first processing unit 502 vectorizes the behavior data to obtain the vector representation of the behavior data; the second processing unit 503 inputs the vector representation of the behavior data into the target user portrait model to predict the behavior information and obtain a prediction result, wherein the target user portrait model is trained using a target loss function composed of cross entropy loss and information entropy loss, and the prediction result is used to characterize whether the target user has the target behavior. It solves the technical problem in the related art that the confidence of the user portrait model is low, resulting in low prediction accuracy of the model. In this solution, by introducing information entropy into the loss function modeling, the uncertainty of the model can be reduced, and the accuracy of the user portrait model can be improved, thereby improving the classification accuracy of the model for marginal users, reducing resource consumption, improving user experience, and avoiding user churn.

[0088] Optionally, in the user behavior information prediction device provided in the embodiment of the present application, the device also includes the following units, which are used to generate a target user portrait model through the following steps: a second acquisition unit, which is used to acquire training data, wherein the training data is obtained after preprocessing the user behavior data; a third processing unit, which is used to perform word segmentation and word embedding processing on the training data to obtain a vector representation of the training data; and a fourth processing unit, which is used to train the initial network model based on the vector representation of the training data and the target loss function to obtain a target user portrait model.

[0089] Optionally, in the user behavior information prediction device provided in the embodiment of the present application, the fourth processing unit includes: a first processing sub-unit, used to input the vector representation of the training data into the initial network model for processing to obtain the probability distribution of the predicted user portrait label; a second processing sub-unit, used to adjust the model parameters according to the probability distribution and the target loss function until the predetermined convergence conditions are met to obtain the target user portrait model.

[0090] Optionally, in the user behavior information prediction device provided in the embodiment of the present application, the first processing sub-unit includes: an encoding module, which is used to input the vector representation of the training data into the encoder module of the initial network model for encoding to obtain the vector representation of the semantic space; and a prediction module, which is used to input the vector representation of the semantic space into the classifier module of the initial network model for prediction to obtain the probability distribution.

[0091] Optionally, in the user behavior information prediction device provided in the embodiment of the present application, the second processing sub-unit includes: a first calculation module, used to calculate the cross entropy loss between the probability distribution and the true label, and calculate the information entropy loss of the probability distribution; a second calculation module, used to calculate the loss value of the target loss function based on the cross entropy loss and the information entropy loss; an update module, used to update the model parameters through the back propagation algorithm based on the loss value.

[0092] Optionally, in the user behavior information prediction device provided in the embodiment of the present application, the third processing unit includes: a third processing sub-unit, used to perform word segmentation processing on the training data to obtain a word sequence corresponding to the training data; and a fourth processing sub-unit, used to perform word embedding processing on the word sequence to obtain a vector representation of the training data.

[0093] Optionally, in the user behavior information prediction device provided in the embodiment of the present application, the fourth processing sub-unit includes: a first processing module, used to map each word in the word sequence to a pre-trained word embedding matrix to obtain a word embedding vector for each word; a second processing module, used to append a corresponding position embedding vector to the word embedding vector of each word to obtain a word embedding vector with position information; and a third processing module, used to splice the word embedding vectors with position information to obtain a vector representation of the training data.

[0094] It should be noted that the first acquisition unit 501, the first processing unit 502, and the second processing unit 503 described above correspond to steps S201 to S203 in Example 1. The examples and application scenarios implemented by the three units and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and can be run in the computer terminal 10 provided in Example 1.

[0095] Example 3

[0096] An embodiment of the present application may provide an electronic device, Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 Only one is shown) processor 602, memory 604, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0097] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0098] The processor can call the information and applications stored in the memory through the transmission device to perform the following steps: with the authorization of the target user, collect the behavior data of the target user; vectorize the behavior data to obtain a vector representation of the behavior data; input the vector representation of the behavior data into the target user portrait model to predict the behavior information and obtain a prediction result, wherein the target user portrait model is trained using a target loss function composed of cross entropy loss and information entropy loss, and the prediction result is used to characterize whether the target user has performed the target behavior.

[0099] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain training data, where the training data is obtained after preprocessing the user behavior data; perform word segmentation and word embedding processing on the training data to obtain a vector representation of the training data; train the initial network model based on the vector representation of the training data and the target loss function to obtain a target user portrait model.

[0100] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: input the vector representation of the training data into the initial network model for processing to obtain the probability distribution of the predicted user portrait label; adjust the model parameters according to the probability distribution and the target loss function until the predetermined convergence conditions are met to obtain the target user portrait model.

[0101] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: input the vector representation of the training data into the encoder module of the initial network model for encoding to obtain the vector representation of the semantic space; input the vector representation of the semantic space into the classifier module of the initial network model for prediction to obtain the probability distribution.

[0102] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: calculate the cross entropy loss between the probability distribution and the true label, and calculate the information entropy loss of the probability distribution; calculate the loss value of the target loss function based on the cross entropy loss and the information entropy loss; based on the loss value, update the model parameters through the back propagation algorithm.

[0103] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: perform word segmentation processing on the training data to obtain a word unit sequence corresponding to the training data; perform word embedding processing on the word unit sequence to obtain a vector representation of the training data.

[0104] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: map each word in the word sequence to the pre-trained word embedding matrix to obtain the word embedding vector of each word; append the corresponding position embedding vector to the word embedding vector of each word to obtain a word embedding vector with position information; splice the word embedding vectors with position information to obtain a vector representation of the training data.

[0105] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), or a PAD. Figure 6 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 6 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 6 Different configurations shown.

[0106] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0107] Example 4

[0108] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the user behavior information prediction method provided in the first embodiment.

[0109] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0110] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of the method for predicting user behavior information.

[0111] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0112] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0113] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0114] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0115] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0116] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0117] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for predicting user behavior information, characterized in that: include: With authorization from the target user, collect the target user's behavior data; performing vectorization processing on the behavior data to obtain a vector representation of the behavior data; The vector representation of the behavior data is input into the target user portrait model to predict the behavior information to obtain a prediction result, wherein the target user portrait model is trained using a target loss function composed of cross entropy loss and information entropy loss, and the prediction result is used to characterize whether the target user has performed the target behavior.

2. The method according to claim 1, characterized in that Generate the target user portrait model through the following steps: Acquiring training data, wherein the training data is obtained by preprocessing user behavior data; Performing word segmentation and word embedding processing on the training data to obtain a vector representation of the training data; The initial network model is trained based on the vector representation of the training data and the target loss function to obtain the target user portrait model.

3. The method according to claim 2, characterized in that The initial network model is trained according to the vector representation of the training data and the target loss function to obtain the target user portrait model, including: Inputting the vector representation of the training data into the initial network model for processing to obtain a probability distribution of predicted user profile labels; The model parameters are adjusted according to the probability distribution and the target loss function until a predetermined convergence condition is met to obtain the target user portrait model.

4. The method according to claim 3, characterized in that Inputting the vector representation of the training data into the initial network model for processing to obtain the probability distribution of the predicted user profile label includes: Inputting the vector representation of the training data into the encoder module of the initial network model for encoding to obtain a vector representation of the semantic space; The vector representation of the semantic space is input into the classifier module of the initial network model for prediction to obtain the probability distribution.

5. The method according to claim 3, characterized in that Adjusting the model parameters according to the probability distribution and the target loss function includes: Calculating the cross entropy loss between the probability distribution and the true label, and calculating the information entropy loss of the probability distribution; Calculating a loss value of the target loss function based on the cross entropy loss and the information entropy loss; Based on the loss value, the model parameters are updated through a back-propagation algorithm.

6. The method according to claim 2, characterized in that Performing word segmentation and word embedding processing on the training data to obtain a vector representation of the training data includes: Performing word segmentation processing on the training data to obtain a word unit sequence corresponding to the training data; Perform word embedding processing on the word sequence to obtain a vector representation of the training data.

7. The method according to claim 6, characterized in that Performing word embedding processing on the word sequence to obtain a vector representation of the training data includes: Mapping each word in the word-unit sequence to a pre-trained word embedding matrix to obtain a word embedding vector for each word; Adding the corresponding position embedding vector to the word embedding vector of each word element to obtain a word embedding vector with position information; The word embedding vectors with position information are concatenated to obtain a vector representation of the training data.

8. A device for predicting user behavior information, characterized in that: include: A first acquisition unit is configured to collect behavior data of a target user upon obtaining authorization from the target user; a first processing unit, configured to perform vectorization processing on the behavior data to obtain a vector representation of the behavior data; The second processing unit is used to input the vector representation of the behavior data into the target user portrait model to predict the behavior information and obtain a prediction result, wherein the target user portrait model is trained using a target loss function composed of cross entropy loss and information entropy loss, and the prediction result is used to characterize whether the target user has performed the target behavior.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the user behavior information prediction method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: a memory storing an executable program; A processor is used to run the program, wherein the program, when running, executes the method for predicting user behavior information according to any one of claims 1 to 7.

Citation Information

Cited By

  • User portrait modeling reasoning method and device, computer equipment and storage medium

    CN121860676A