Method and system for generating live broadcast room population portrait based on big data
By acquiring, cleaning and analyzing the data of live broadcast users, and using deep learning to generate multi-scale fusion feature vectors, the problem of live broadcast platforms generating crowd portraits is solved, and the user experience and service accuracy is improved.
Patent Information
- Application Number
- CN202311202178.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-09-18
AI Technical Summary
When the live broadcast platform processes and analyzes complex user data, it is difficult to accurately generate portraits of people in the live broadcast room, and it is impossible to fully understand the user's interests, needs and behavior patterns.
By obtaining the basic information, interactive data and consumption data of live broadcast users, data cleaning and structured representation are carried out, semantic analysis is used using deep learning technology to generate multi-scale fusion correlation feature vectors for user information, and finally generating population portraits based on adaptive information granular clustering.
It has achieved a more comprehensive understanding of users' interests, needs and behavior patterns, and improved the personalized service capabilities and user experience of the live broadcast platform.
Smart Images

Figure CN118916731B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of generating population portraits, and more specifically, to a method and system for generating a population portrait of a live broadcast room based on big data. Background Art
[0002] The live broadcast industry has developed vigorously, and the number of users on live broadcast platforms is huge. Generating a population portrait of a live broadcast room can better understand and comprehend live broadcast users, so as to provide services such as personalized content recommendations, precise advertisement placements, and refined user management. Through the population portrait, the live broadcast platform can better grasp the needs and interests of users, provide live broadcast content that meets user preferences, and improve user experience and retention rate.
[0003] However, a large amount of data left during the operation of the live broadcast platform is usually miscellaneous and disorderly. How to process and analyze this data and accurately generate a population portrait of the live broadcast room is an important technical problem. Summary of the Invention
[0004] In view of this, the present disclosure proposes a method and system for generating a population portrait of a live broadcast room based on big data, which can more comprehensively understand the interests, needs, and behavior patterns of users.
[0005] According to one aspect of the present disclosure, there is provided a method for generating a population portrait of a live broadcast room based on big data, which includes:
[0006] Obtaining basic information, interaction data, and consumption data of the live broadcast user to be analyzed;
[0007] Performing data cleaning and structured representation on the basic information, interaction data, and consumption data of the live broadcast user to be analyzed to obtain a user basic information coding vector, a user interaction coding vector, and a user consumption coding vector;
[0008] Performing semantic analysis on the user basic information coding vector, the user interaction coding vector, and the user consumption coding vector to obtain a user information multi-scale fusion correlation feature vector; and
[0009] Generating a population portrait based on the user information multi-scale fusion correlation feature vector.
[0010] According to another aspect of the present disclosure, there is provided a system for generating a population portrait of a live broadcast room based on big data, which includes:
[0011] A data acquisition module for obtaining basic information, interaction data, and consumption data of the live broadcast user to be analyzed;
[0012] A data cleaning and structured representation module for cleaning and structurally representing the basic information, interaction data, and consumption data of the analyzed live user to obtain a user basic information encoding vector, a user interaction encoding vector, and a user consumption encoding vector;
[0013] A semantic analysis module for performing semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion correlation feature vector; and
[0014] A population portrait generation module for generating a population portrait based on the user information multi-scale fusion correlation feature vector.
[0015] According to an embodiment of the present disclosure, it first obtains the basic information, interaction data, and consumption data of the analyzed live user. Then, it cleans and structurally represents the basic information, interaction data, and consumption data of the analyzed live user to obtain a user basic information encoding vector, a user interaction encoding vector, and a user consumption encoding vector. Next, it performs semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion correlation feature vector. Finally, it generates a population portrait based on the user information multi-scale fusion correlation feature vector. In this way, the user's interests, needs, and behavior patterns can be understood more comprehensively. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings included in the specification and constituting a part of the specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.
[0017] Figure 1 A flowchart showing a method for generating a population portrait of a live broadcast room based on big data according to an embodiment of the present disclosure.
[0018] Figure 2 A schematic structural diagram showing a method for generating a population portrait of a live broadcast room based on big data according to an embodiment of the present disclosure.
[0019] Figure 3 A flowchart showing a sub-step S130 of a method for generating a population portrait of a live broadcast room based on big data according to an embodiment of the present disclosure.
[0020] Figure 4 A block diagram showing a system for generating a population portrait of a live broadcast room based on big data according to an embodiment of the present disclosure.
[0021] Figure 5 An application scenario diagram showing a method for generating a population portrait of a live broadcast room based on big data according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts also belong to the scope of protection of the present disclosure.
[0023] As used in the present disclosure and the claims, unless the context clearly dictates otherwise, words such as "a," "an," "one," and / or "the" are not specifically intended to refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" merely indicate the inclusion of the steps and elements that have been expressly identified, and these steps and elements do not constitute an exclusive listing. A method or apparatus may also contain other steps or elements.
[0024] The various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the drawings. Like reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings need not be drawn to scale unless otherwise specified.
[0025] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0026] To address the above technical problems, the technical concept of the present disclosure is to utilize deep learning technology and combine the basic information, interaction data, and consumption data of users to achieve the intelligent generation of the crowd portrait in the live broadcast room.
[0027] Based on this, Figure 1 The flowchart of the method for generating a crowd portrait in a live broadcast room based on big data according to an embodiment of the present disclosure is shown. Figure 2 The schematic structural diagram of the method for generating a crowd portrait in a live broadcast room based on big data according to an embodiment of the present disclosure is shown. As Figure 1 and Figure 2As shown, the method for generating a live broadcast audience portrait based on big data according to an embodiment of the present disclosure includes the steps of: S110, obtaining the basic information, interaction data, and consumption data of the live broadcast user to be analyzed; S120, performing data cleaning and structured representation on the basic information, interaction data, and consumption data of the live broadcast user to be analyzed to obtain a user basic information encoding vector, a user interaction encoding vector, and a user consumption encoding vector; S130, performing semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion association feature vector; and S140, generating an audience portrait based on the user information multi-scale fusion association feature vector.
[0028] Specifically, in the technical solution of the present disclosure, first, obtain the basic information, interaction data, and consumption data of the live broadcast user to be analyzed; and perform data cleaning and structured representation on the basic information, interaction data, and consumption data of the live broadcast user to be analyzed to obtain a user basic information encoding vector, a user interaction encoding vector, and a user consumption encoding vector. It should be understood that the basic information, interaction data, and consumption data of the live broadcast user provide information in multiple dimensions, including the user's personal characteristics, behavior preferences, and consumption habits, etc. By comprehensively analyzing this information, the user's interests, needs, and behavior patterns can be understood more comprehensively.
[0029] Then, perform semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion association feature vector. That is, mine the feature distribution of the user's interests and behavior patterns from the basic information, interaction data, and consumption data of the live broadcast user to be analyzed.
[0030] In a specific example of the present disclosure, the encoding process of performing semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion association feature vector includes: first, passing the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a user information global association feature extractor based on a fully connected layer to obtain a user information global association feature vector; at the same time, passing the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a user information context semantic association feature extractor based on a transformer to obtain a user information context semantic association feature vector; subsequently, calculating the mutual gain between the user information global association feature vector and the user information context semantic association feature vector to obtain a user information multi-scale fusion association feature vector.
[0031] Correspondingly, as Figure 3As shown, semantic analysis is performed on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion association feature vector, including: S131, extracting the global association features of user information among the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information global association feature vector; S132, extracting the context semantic association features among the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information context semantic association feature vector; and S133, fusing the user information global association feature vector and the user information context semantic association feature vector to obtain the user information multi-scale fusion association feature vector. It should be understood that the generation of the user information multi-scale fusion association feature vector is carried out through the three steps of S131, S132, and S133. In step S131, the association features among the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector are extracted. These features may include the basic attributes of the user, the interaction behaviors of the user on social media, and the consumption behaviors of the user, etc. The extracted association features can help understand the overall behavior patterns and preferences of the user. In step S132, the context semantic association features among the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector are extracted. The context semantic association features may include the behavior characteristics of the user in different environments, such as the behavior changes of the user at different times, locations, or social circles. These features can help understand the context relevance of the user's behavior. In step S133, the user information global association feature vector and the user information context semantic association feature vector are fused together to generate the user information multi-scale fusion association feature vector. This vector combines the overall behavior patterns and preferences of the user and the context relevance of the user's behavior, providing a more comprehensive representation of user information. This multi-scale fusion association feature vector can be used in subsequent user analysis, personalized recommendation, and other tasks.
[0032] More specifically, in step S131, extracting the global user information association features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a global user information association feature vector includes: passing the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a global user information association feature extractor based on a fully connected layer to obtain the global user information association feature vector. It is worth mentioning that the fully connected layer is one of the most commonly used layer types in deep neural networks. Its role is to connect each neuron in the input to each neuron in the output, forming a fully connected network structure. The main role of the fully connected layer is to learn the complex non-linear relationships between the input data and encode these relationships as feature representations. In the global user information association feature extractor, the fully connected layer is used to combine and fuse the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to generate a global association feature vector of user information. Through the fully connected layer, the network can learn the weights and relationships between different input features, thereby capturing higher-level abstract features. These abstract features can better represent the overall behavior patterns and preferences of users, helping the model better understand the characteristics and behaviors of users. The fully connected layer usually includes weight parameters and bias parameters, and these parameters will be learned and updated through the backpropagation algorithm during the training process to minimize the prediction error of the model as much as possible. The output of the fully connected layer can be used as the input for the next layer or as the final output result for different tasks and application scenarios.
[0033] More specifically, in step S132, the context semantic association features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector are extracted to obtain a user information context semantic association feature vector, including: passing the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a Transformer-based user information context semantic association feature extractor to obtain the user information context semantic association feature vector. It is worth mentioning that in the user information context semantic association feature extractor, a Transformer-based method is used to extract the context semantic association feature vector of user information. The Transformer refers to the Transformer model, which is a neural network structure based on the self-attention mechanism. The main role of the Transformer is to encode and model the input sequence, capturing the context relationship and semantic information between elements in the sequence. In the user information context semantic association feature extractor, the Transformer is used to process the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to extract the context semantic association features between them. Through the self-attention mechanism of the Transformer, the model can globally focus on and interact with the input vectors, thereby capturing the relevance and importance between different elements. The Transformer gradually models and extracts features from the input through multiple stacked self-attention layers and feed-forward neural network layers to form a context semantic association feature vector. The advantage of the Transformer is its ability to handle long sequences and model global dependencies, while having good parallel computing performance. It can learn complex patterns and relationships in the input sequence, extract richer context semantic association features, and help better understand the context relevance of user behavior. These feature vectors can be used for subsequent tasks.
[0034] More specifically, in step S133, fusing the global associated feature vector of the user information and the context semantic associated feature vector of the user information to obtain the multi-scale fused associated feature vector of the user information includes: calculating the mutual gain between the global associated feature vector of the user information and the context semantic associated feature vector of the user information to obtain the multi-scale fused associated feature vector of the user information. It is worth mentioning that mutual gain refers to the degree of gain provided by different feature vectors in the fused feature by comparing the correlation and complementarity between different feature vectors in multi-scale feature fusion. In this step, mutual gain is used to calculate the complementarity and correlation between the global associated feature vector of the user information and the context semantic associated feature vector of the user information to obtain the multi-scale fused associated feature vector of the user information. Specifically, mutual gain can be calculated in different ways, and common methods include mutual information, correlation coefficient, etc. These methods can measure the degree of correlation and complementarity between two feature vectors. By calculating the mutual gain, we can understand the contributions and importance of the global associated feature vector of the user information and the context semantic associated feature vector of the user information in multi-scale fusion. This helps to determine how to appropriately fuse these feature vectors to obtain a more comprehensive and accurate multi-scale fused associated feature vector of the user information.
[0035] Furthermore, passing the multi-scale fused associated feature vector of the user information through an AIGC-based population portrait generator to generate a population portrait. Correspondingly, in step S140, generating a population portrait based on the multi-scale fused associated feature vector of the user information includes: passing the multi-scale fused associated feature vector of the user information through an AIGC-based population portrait generator to generate a population portrait. It is worth mentioning that AIGC (Adaptive Information Granularity Clustering) is a method based on adaptive information granularity clustering for generating population portraits. AIGC can take the multi-scale fused associated feature vector of the user information as input and divide users into different groups or categories through the method of adaptive information granularity clustering. It analyzes the similarities and differences between feature vectors and clusters users with similar features together to form different population groups. During the clustering process, AIGC will adaptively adjust the information granularity of clustering according to the characteristics and distribution of the data to adapt to different data features, so that it can better capture the subtle differences and common features between user groups. The population portrait generated by AIGC can provide an in-depth understanding and description of the user group.
[0036] Furthermore, in the technical solution of the present disclosure, the method for generating a live broadcast room population portrait based on big data is characterized in that it further includes a training step: training the user information global association feature extractor based on the fully connected layer, the user information context semantic association feature extractor based on the transformer, and the population portrait generator based on AIGC. It should be understood that the training step plays a key role in the method for generating a live broadcast room population portrait based on big data. By training the user information global association feature extractor based on the fully connected layer, the user information context semantic association feature extractor based on the transformer, and the population portrait generator based on AIGC, the following objectives can be achieved: 1. Parameter optimization: The training step can optimize the parameters of the model, enabling the feature extractor and the population portrait generator to extract and generate user information features more accurately. Through repeated iteration and optimization on the training data, the model can learn better feature representations and generation capabilities. 2. Improved accuracy: Through training, the feature extractor and the population portrait generator can learn patterns and rules that are more adaptable to the actual data, thereby improving the accuracy of the generated population portrait and better reflecting the characteristics and behaviors of users. 3. Enhanced generalization ability: The training step can help the model learn more general feature representations and generation methods, enabling it to perform well on unseen data, which improves the model's generalization ability and enables it to adapt to different live broadcast rooms and different users. Through the training step, the model can be continuously improved and optimized to better adapt to the actual application scenario and generate more accurate and comprehensive population portraits.
[0037] Among them, more specifically, the training steps include: obtaining training data, where the training data includes the training basic information, training interaction data, and training consumption data of the live user to be analyzed, as well as the true value of the population portrait; performing data cleaning and structured representation on the training basic information, training interaction data, and training consumption data of the live user to be analyzed to obtain a training user basic information coding vector, a training user interaction coding vector, and a training user consumption coding vector; passing the training user basic information coding vector, the training user interaction coding vector, and the training user consumption coding vector through the training user information global association feature extractor based on the fully connected layer to obtain a training user information global association feature vector; passing the training user basic information coding vector, the training user interaction coding vector, and the training user consumption coding vector through the user information context semantic association feature extractor based on the Transformer to obtain a training user information context semantic association feature vector; calculating the mutual gain between the training user information global association feature vector and the training user information context semantic association feature vector to obtain a training user information multi-scale fusion association feature vector; optimizing the feature distribution of the training user information multi-scale fusion association feature vector to obtain an optimized training user information multi-scale fusion association feature vector; passing the optimized training user information multi-scale fusion association feature vector through the population portrait generator based on AIGC to obtain a training generated population portrait; calculating the cross-entropy function value between the training generated population portrait and the true value of the population portrait; and using the cross-entropy function value as the loss function value to train the user information global association feature extractor based on the fully connected layer, the user information context semantic association feature extractor based on the Transformer, and the population portrait generator based on AIGC.
[0038] In the technical solution of the present disclosure, since the training user information global association feature vector and the training user information context semantic association feature vector are obtained from the training user basic information encoding vector, the training user interaction encoding vector, and the training user consumption encoding vector through a user information global association feature extractor based on a fully connected layer and a user information context semantic association feature extractor based on a transformer, respectively, the training user information global association feature vector and the training user information context semantic association feature vector will respectively have semantic feature representations corresponding to the local feature distribution scales of each initial encoding vector. Thus, by calculating the mutual gain between the training user information global association feature vector and the training user information context semantic association feature vector, the training user information multi-scale fusion association feature vector will also have semantic feature representations corresponding to the local feature distribution scales of each initial encoding vector. In this way, when the training user information multi-scale fusion association feature vector passes through the AIGC-based population portrait generator, it will also perform scale heuristic distribution probability density mapping based on the local feature distribution scale, thereby obtaining the training generated population portrait. However, considering that under the local feature distribution scale of the training user information multi-scale fusion association feature vector, it will simultaneously contain the global association semantic features of the vector value granularity of the initial vector and the context association features of the vector granularity, that is, it will contain mixed granularity semantic feature representations under the local feature distribution scale, which will reduce the training efficiency of the AIGC-based population portrait generator.
[0039] Based on this, during the training process, when the applicant of the present disclosure generates a population portrait by passing the training user information multi-scale fusion association feature vector through the AIGC-based population portrait generator, semantic information uniform activation of feature rank expression is performed on the training user information multi-scale fusion association feature vector.
[0040] Correspondingly, in a specific example, optimizing the feature distribution of the training user information multi-scale fusion association feature vector to obtain an optimized training user information multi-scale fusion association feature vector includes: optimizing the feature distribution of the training user information multi-scale fusion association feature vector with the following optimization formula to obtain the optimized training user information multi-scale fusion association feature vector; where the optimization formula is:
[0041]
[0042] where V is the training user information multi-scale fusion association feature vector, v iis the i-th eigenvalue of the multi-scale fusion associated feature vector of the training user information, ||V||2 represents the bi-norm of the multi-scale fusion associated feature vector of the training user information, log is the logarithmic function with base 2, and α is the weight hyperparameter, v′ i It is the i-th eigenvalue of the multi-scale fusion associated feature vector of the optimized training user information.
[0043] Here, considering the feature distribution mapping of the feature distribution of the multi-scale fusion associated feature vector V of the training user information from the high-dimensional feature space to the generated probability density mapping space, different mapping modes will be presented at different feature distribution levels based on the mixed granularity semantic features, resulting in the inability of the scale-heuristic mapping strategy to obtain the optimal efficiency. Therefore, the rank expression semantic information based on the feature vector norm is homogenized instead of the scale for feature matching, which can activate similar feature rank expressions in a similar manner and reduce the correlation between feature rank expressions with large differences, thereby solving the problem of low efficiency of probability expression mapping of the feature distribution of the multi-scale fusion associated feature vector V of the training user information under different spatial rank expressions, and improving the training efficiency of the AIGC-based crowd portrait generator.
[0044] In summary, the method for generating crowd portraits in a live broadcast room based on big data according to the embodiment of the present disclosure can more comprehensively understand the interests, needs and behavior patterns of users.
[0045] Figure 4 FIG. 1 is a block diagram of a live broadcast crowd portrait generation system 100 based on big data according to an embodiment of the present disclosure. Figure 4 As shown, according to the big data-based live broadcast room crowd portrait generation system 100 of the embodiment of the present disclosure, it includes: a data acquisition module 110, used to obtain the basic information, interaction data and consumption data of the analyzed live broadcast user; a data cleaning and structured representation module 120, used to perform data cleaning and structured representation on the basic information, interaction data and consumption data of the analyzed live broadcast user to obtain a user basic information coding vector, a user interaction coding vector and a user consumption coding vector; a semantic analysis module 130, used to perform semantic analysis on the user basic information coding vector, the user interaction coding vector and the user consumption coding vector to obtain a multi-scale fusion associated feature vector of user information; and a crowd portrait generation module 140, used to generate a crowd portrait based on the multi-scale fusion associated feature vector of user information.
[0046] In a possible implementation, the semantic analysis module 130 includes: a global association feature extraction unit, configured to extract the global user information association feature between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a global user information association feature vector; a semantic association feature extraction unit, configured to extract the context semantic association feature between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a context semantic association feature vector of the user information; and a fusion unit, configured to fuse the global user information association feature vector and the context semantic association feature vector of the user information to obtain the multi-scale fusion association feature vector of the user information.
[0047] Here, those skilled in the art can understand that the specific functions and operations of each unit and module in the above live broadcast room population portrait generation system 100 based on big data have been introduced in detail in the description of the above-mentioned Figures 1 to 3 live broadcast room population portrait generation method based on big data, and therefore, the repeated description thereof will be omitted.
[0048] As described above, the live broadcast room population portrait generation system 100 based on the embodiments of the present disclosure can be implemented in various wireless terminals, such as a server having an algorithm for generating a live broadcast room population portrait based on big data. In a possible implementation, the live broadcast room population portrait generation system 100 based on the embodiments of the present disclosure can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the live broadcast room population portrait generation system 100 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the live broadcast room population portrait generation system 100 can also be one of the many hardware modules of the wireless terminal.
[0049] Alternatively, in another example, the live broadcast room population portrait generation system 100 and the wireless terminal can also be separate devices, and the live broadcast room population portrait generation system 100 can be connected to the wireless terminal through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.
[0050] Figure 5 The application scenario diagram of the live broadcast room population portrait generation method according to the embodiments of the present disclosure is shown. As Figure 5 shown, in this application scenario, first, the basic information of the live user to be analyzed (for example, Figure 5 D1 shown in Figure 5 ), interaction data (for example,Figure 5 as shown in D3), and then input the basic information, interaction data, and consumption data of the analyzed live user into a server (e.g., Figure 5 as shown in S) that deploys an algorithm for generating a crowd portrait of a live broadcast room based on big data. Among them, the server can use the algorithm for generating a crowd portrait of a live broadcast room based on big data to process the basic information, interaction data, and consumption data of the analyzed live user to generate a crowd portrait.
[0051] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, the segment of a program, or the part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0052] The various embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for generating a crowd portrait of a live broadcast room based on big data, characterized in that, Including: Obtain the basic information, interaction data, and consumption data of the live user to be analyzed; Perform data cleaning and structured representation on the basic information, interaction data, and consumption data of the live user to be analyzed to obtain a user basic information encoding vector, a user interaction encoding vector, and a user consumption encoding vector; Perform semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion correlation feature vector; And Generate a population portrait based on the user information multi-scale fusion correlation feature vector; Among them, performing semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion correlation feature vector includes: Extract the global user information correlation features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a global user information correlation feature vector; Extract the context semantic correlation features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a context semantic correlation feature vector of user information; and Fuse the global user information correlation feature vector and the context semantic correlation feature vector of user information to obtain the user information multi-scale fusion correlation feature vector; Among them, extracting the global user information correlation features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a global user information correlation feature vector includes: Pass the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a global user information correlation feature extractor based on a fully connected layer to obtain the global user information correlation feature vector; Among them, extracting the context semantic correlation features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a context semantic correlation feature vector of user information includes: Pass the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a context semantic correlation feature extractor of user information based on a Transformer to obtain the context semantic correlation feature vector of user information.
2. The method for generating a live broadcast room population portrait based on big data according to claim 1, wherein Fusing the global user information correlation feature vector and the context semantic correlation feature vector of user information to obtain the user information multi-scale fusion correlation feature vector includes: Calculate the mutual gain between the global user information correlation feature vector and the context semantic correlation feature vector of user information to obtain the user information multi-scale fusion correlation feature vector.
3. The method for generating a live broadcast room population portrait based on big data according to claim 2, wherein Generating a population portrait based on the user information multi-scale fusion correlation feature vector includes: Pass the user information multi-scale fusion correlation feature vector through a population portrait generator based on AIGC to generate a population portrait.
4. The method for generating a live broadcast room population portrait based on big data according to claim 3, characterized in that It also includes a training step: training the user information global association feature extractor based on the fully connected layer, the user information context semantic association feature extractor based on the Transformer, and the crowd portrait generator based on AIGC; Among them, the training step includes: Obtaining training data, where the training data includes the training basic information, training interaction data, and training consumption data of the live user to be analyzed, as well as the true value of the crowd portrait; Performing data cleaning and structured representation on the training basic information, training interaction data, and training consumption data of the live user to be analyzed to obtain a training user basic information encoding vector, a training user interaction encoding vector, and a training user consumption encoding vector; Passing the training user basic information encoding vector, the training user interaction encoding vector, and the training user consumption encoding vector through the training user information global association feature extractor based on the fully connected layer to obtain a training user information global association feature vector; Passing the training user basic information encoding vector, the training user interaction encoding vector, and the training user consumption encoding vector through the user information context semantic association feature extractor based on the Transformer to obtain a training user information context semantic association feature vector; Calculating the mutual gain between the training user information global association feature vector and the training user information context semantic association feature vector to obtain a training user information multi-scale fusion association feature vector; Performing feature distribution optimization on the training user information multi-scale fusion association feature vector to obtain an optimized training user information multi-scale fusion association feature vector; Passing the optimized training user information multi-scale fusion association feature vector through the crowd portrait generator based on AIGC to obtain a training generated crowd portrait; Calculating the cross-entropy function value between the training generated crowd portrait and the true value of the crowd portrait; and Using the cross-entropy function value as the loss function value to train the user information global association feature extractor based on the fully connected layer, the user information context semantic association feature extractor based on the Transformer, and the crowd portrait generator based on AIGC.
5. The method for generating a live broadcast room population portrait based on big data according to claim 4, characterized in that Performing feature distribution optimization on the training user information multi-scale fusion association feature vector to obtain an optimized training user information multi-scale fusion association feature vector, including: Performing feature distribution optimization on the training user information multi-scale fusion association feature vector with the following optimization formula to obtain the optimized training user information multi-scale fusion association feature vector; Among them, the optimization formula is: where V is the multi-scale fusion correlation feature vector of the training user information, and v i is the i-th eigenvalue of the multi-scale fusion correlation feature vector of the training user information, ||V||2 represents the two-norm of the multi-scale fusion correlation feature vector of the training user information, log is the logarithmic function with base 2, and α is the weight hyperparameter, v′ i is the i-th eigenvalue of the optimized multi-scale fusion correlation feature vector of the training user information.
6. A live broadcast room population portrait generation system based on big data, characterized in that, Including: A data acquisition module for acquiring the basic information, interaction data, and consumption data of the live user to be analyzed; A data cleaning and structured representation module for performing data cleaning and structured representation on the basic information, interaction data, and consumption data of the live user to be analyzed to obtain a user basic information encoding vector, a user interaction encoding vector, and a user consumption encoding vector; A semantic analysis module for performing semantic analysis on the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information multi-scale fusion association feature vector; and A population portrait generation module for generating a population portrait based on the user information multi-scale fusion association feature vector; wherein, the semantic analysis module includes: A global association feature extraction unit for extracting the user information global association features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information global association feature vector; A semantic association feature extraction unit for extracting the context semantic association features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information context semantic association feature vector; and A fusion unit for fusing the user information global association feature vector and the user information context semantic association feature vector to obtain the user information multi-scale fusion association feature vector; wherein, extracting the user information global association features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information global association feature vector includes: Passing the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a user information global association feature extractor based on a fully connected layer to obtain the user information global association feature vector; wherein, extracting the context semantic association features between the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector to obtain a user information context semantic association feature vector includes: Passing the user basic information encoding vector, the user interaction encoding vector, and the user consumption encoding vector through a user information context semantic association feature extractor based on a transformer to obtain the user information context semantic association feature vector.
Citation Information
Patent Citations
Data processing method for directional analysis and prediction of e-commerce platform
CN115879972A
User portrait construction method and system based on feature fusion of improved LDA
CN116385037A