Language model enhancement method and system based on live-streaming room user behavior network
By building a live broadcast room user behavior network for entity embedding representation training and language model enhancement training, the problem of large errors in user language analysis in live broadcast scenarios is solved, more accurate user behavior and text expression analysis is achieved, and the application effect of the language model is improved.
Patent Information
- Application Number
- PCT/CN2024/135021
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-19
AI Technical Summary
In live broadcast scenarios that are more focused on entertainment, the user's text expression is arbitrary, the text is short, and the information expressed is fragmented. It is difficult for conventional language models to accurately analyze, resulting in insufficient understanding of the user and the live broadcast room content and deviating the language analysis effect.
By constructing a language model enhancement method based on the live broadcast room user behavior network, it includes building a user behavior network based on the live broadcast room information, user identification information, user behavior information and user text information, conducting entity embedding representation training, output text vectors, and constructing similar text data sets based on the text vectors for language model enhancement training.
This method can improve the analysis accuracy of the language model in live broadcast scenarios, enhance the accurate analysis ability of user behavior and text expression, improve the content understanding ability of the language model, and improve the application effect.
Smart Images

Figure CN2024135021_19062025_PF_FP_ABST
Abstract
Description
A language model enhancement method and system based on live broadcast room user behavior network
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 15, 2023, with application number 202311739893.7, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a method and system for enhancing a language model based on a user behavior network in a live broadcast room. Background Art
[0003] As live streaming applications continue to develop and offer increasingly diverse features, users are increasingly engaging in more ways within the platform. Users can access their favorite live streaming rooms, tip the host, take the microphone, send comments, and interact with other room members. To better understand user characteristics, a common approach is to use language models to analyze user speech and text.
[0004] However, in entertainment-focused livestreams, users often express themselves in casual, short, fragmented text, and a lot of noise. Conventional language models struggle to accurately analyze these texts, resulting in an inadequate understanding of the user and the livestream content, and inaccurate language analysis. Summary of the Invention
[0005] The embodiments of the present application provide a language model enhancement method and system based on a live broadcast room user behavior network, which can improve the accuracy of language model analysis and solve the technical problem of large errors in user language analysis in live broadcast scenarios.
[0006] In a first aspect, an embodiment of the present application provides a method for enhancing a language model based on a live broadcast room user behavior network, comprising:
[0007] Construct a user behavior network based on live broadcast room information, user identification information, user behavior information and user text information;
[0008] Entity embedding representation training is performed based on the user behavior network, and the text vector corresponding to the user text information is output;
[0009] A similar text dataset corresponding to the user's text information is constructed based on the text vector, and language model enhancement training is performed based on the similar text dataset.
[0010] In a second aspect, an embodiment of the present application provides a language model enhancement system based on a live broadcast room user behavior network, comprising:
[0011] A network construction module is configured to construct a user behavior network based on live broadcast room information, user identification information, user behavior information, and user text information;
[0012] An entity embedding module is configured to train entity embedding representations based on the user behavior network and output a text vector corresponding to the user's text information;
[0013] The model enhancement module is configured to construct a similar text dataset corresponding to the user text information based on the text vector, and perform language model enhancement training based on the similar text dataset.
[0014] In a third aspect, an embodiment of the present application provides a language model enhancement device based on a live broadcast room user behavior network, comprising:
[0015] memory and one or more processors;
[0016] The memory is configured to store one or more programs;
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the language model enhancement method based on the live broadcast room user behavior network as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a computer processor, are configured to execute the language model enhancement method based on the live broadcast room user behavior network as described in the first aspect.
[0019] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed on a computer or processor, the computer or processor executes the language model enhancement method based on the live broadcast room user behavior network as described in the first aspect.
[0020] The embodiment of the present application constructs a user behavior network based on live broadcast room information, user identification information, user behavior information and user text information; performs entity embedding representation training based on the user behavior network and outputs a text vector corresponding to the user text information; constructs a similar text dataset corresponding to the user text information based on the text vector, and performs language model enhancement training based on the similar text dataset. By adopting the above technical means, by constructing a user behavior network for entity embedding representation training, similar text datasets can be accurately mined based on text vector representations, and similar samples from different sources can be expanded. In this way, language model training can be performed, which can enhance the content understanding ability of the language model in different contextual scenarios, perform accurate language analysis of user behavior and text expressions, and improve the application effect of the language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] FIG1 is a flow chart of a method for enhancing a language model based on a live broadcast room user behavior network according to an embodiment of the present application;
[0022] FIG2 is a schematic diagram of the structure of a user behavior network in an embodiment of the present application;
[0023] FIG3 is a flow chart of language model enhancement training in an embodiment of the present application;
[0024] FIG4 is a schematic diagram of an enhanced language model application in an embodiment of the present application;
[0025] FIG5 is a structural diagram of a language model enhancement system based on a live broadcast room user behavior network provided by an embodiment of the present application;
[0026] Figure 6 is a structural diagram of a language model enhancement device based on a live broadcast room user behavior network provided by an embodiment of the present application. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only parts related to the present application, not all of the contents, are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0028] The language model enhancement method based on the live broadcast room user behavior network provided in this application aims to accurately mine similar samples from different sources through text vector representation, so as to train the language model's content understanding ability in different contextual scenarios.
[0029] In live streaming scenarios, users tip hosts, use microphones and public screens in live streaming rooms, interact with other room members, obtain various top-up tasks and coupons issued by the platform, and use them for on-device top-ups and gift giving. These actions all generate corresponding text data. The live streaming server collects various types of text data generated in live streaming rooms, analyzes users' language expressions using language models, and then applies the analysis results accordingly.
[0030] Generally speaking, to improve the analysis accuracy of language models, text posted by users within live streaming applications is collected and trained on the language models to enhance the accuracy of language analysis in live streaming scenarios. However, in entertainment-oriented live streaming scenarios, users' textual expressions are casual, short, fragmented, and filled with a large amount of noise. Conventional language models struggle to accurately analyze these texts, resulting in an inadequate understanding of the users and the content of the live streaming room, and skewed language analysis results. Using this noisy live streaming room text data for conventional language model training actually fails to achieve better results, as the noise in the text can degrade the model's performance. Therefore, it is difficult to easily characterize user interests and understand users and the events occurring within the live streaming room through text analysis, which poses a challenge to the efficient operation of the live streaming room and users.
[0031] Based on this, a language model enhancement method based on a live broadcast room user behavior network is provided in an embodiment of the present application to solve the technical problem of user language analysis errors in live broadcast scenarios.
[0032] Example:
[0033] FIG1 shows a flow chart of a method for enhancing a language model based on a live broadcast room user behavior network provided by an embodiment of the present application. The method for enhancing a language model based on a live broadcast room user behavior network provided in this embodiment can be performed by a language model enhancing device based on a live broadcast room user behavior network. The language model enhancing device based on a live broadcast room user behavior network can be implemented by software and / or hardware. The language model enhancing device based on a live broadcast room user behavior network can be composed of two or more physical entities, or it can be composed of one physical entity. Generally speaking, the language model enhancing device based on a live broadcast room user behavior network can be a computing device such as a live broadcast service-end device and a server host.
[0034] The following description will be made by taking the language model enhancement device based on the live broadcast room user behavior network as an example to perform the language model enhancement method based on the live broadcast room user behavior network. Referring to Figure 1, the language model enhancement method based on the live broadcast room user behavior network specifically includes:
[0035] S110: Construct a user behavior network based on the live broadcast room information, user identification information, user behavior information, and user text information.
[0036] When performing language model enhancement, the embodiment of the present application establishes a user behavior network through user behavior in a live broadcast scenario, and mines similar text data sets in the live broadcast scenario based on the user behavior network to perform model enhancement training.
[0037] The user behavior network is a heterogeneous network. A heterogeneous network is composed of multiple types of nodes, such as users, live streaming rooms, and different user behaviors. Because different node types lead to different types of edges, it is called a heterogeneous network. This user behavior network can be used to identify contextual relationships between different nodes for entity embedding training.
[0038] Referring to Figure 2, a schematic diagram of the structure of the user behavior network of an embodiment of the present application is provided. Based on the different types of nodes that constitute the user behavior network, the corresponding live broadcast room information is collected. Among them, the collected live broadcast room information includes live broadcast room information (such as room 1, room 2, room 3, etc.), user identification information (such as user 1, user 2, user 3, etc.), user behavior information (such as speaking, following, entering, rewarding) and user text information (text 1, text 2 and text 3). By collecting the above types of information, a corresponding user behavior network can be constructed.
[0039] It is understandable that for user text information in different scenarios, similar text information can be clustered by combining the contextual relationship of user behavior. In this way, some noisy and difficult-to-understand texts can be clustered with similar text information through contextual relationships, and this part of information can be easily analyzed and understood based on the clustering results. Based on this, the embodiment of the present application constructs the user behavior network, collects user text information, and also collects the live broadcast room information of different types of nodes under the corresponding user behavior according to the contextual relationship to construct the user behavior network, so as to accurately cluster user text information.
[0040] It should be noted that in actual applications, corresponding types of live broadcast room information can also be collected based on various contextual relationships of user text information in live broadcast scenarios for use in constructing user behavior networks, thereby achieving more accurate and comprehensive user text information clustering.
[0041] S120: Perform entity embedding representation training based on the user behavior network, and output a text vector corresponding to the user text information.
[0042] Furthermore, based on the above user behavior network, the embodiment of the present application uses entity embedding representation training to determine text vectors representing user text information. Based on the determined text vectors, accurate similar text clustering can be performed.
[0043] Among them, entity embedding representation training based on user behavior network includes:
[0044] The live broadcast room information, user identification information, user behavior information and user text information are taken as entities, and entity embedding representation training is performed based on the contextual relationship between entities to obtain the embedding vectors representing each entity in the user behavior network.
[0045] In the embodiment of the present application, live broadcast room information, user identification information, user behavior information and user text information are used as entities. The entity embedding representation training adopts the random walk method to sample the network structure, and then learns the embedding vector representation of different entities in the network, and determines the embedding vector representing each entity. During the training process, the corresponding array is used to assign the initial embedding vector of each entity, and then the similarity of the embedding vector is calculated according to the context relationship, and the loss function is calculated according to the similarity. Based on the loss function, the embedding vector can be adjusted, and then the loss function is recalculated, and so on, until the loss converges. For each entity, the above method is used to train and determine its embedding vector. Taking the user behavior of "user 1-enter-room 1" as an example, the embedding vector representing "user 1" plus the embedding vector representing "enter" should be equal to the embedding vector representing "room 1". Similarly, by limiting the loss function of each entity through the above context relationship, the embedding vector that accurately represents each entity can be trained.
[0046] Optionally, for each entity in the user behavior network, the entity relationship transition probability of the entity to its neighboring entities is calculated based on the contextual relationship. Both the entity and its neighboring entities can be from the user behavior network. A neighboring entity is an entity that is directly connected to the entity. For example, if entity A is directly connected to entity B, entity B can be called a neighboring entity of entity A.
[0047] For each entity, the embodiment of the present application can determine the entity relationship transition probability from the entity to its adjacent entities. The entity relationship transition probability is composed of the entity relationship ratio and the reverse triple probability. Among them, the entity relationship ratio is used to characterize the ratio of any entity relationship from the entity to the adjacent entity in all entity relationships, and the reverse triple probability is used to characterize the importance of any entity relationship in all triples. In this way, the transition probability of an entity to its adjacent entity through each entity relationship can be determined based on the entity relationship transition probability.
[0048] Among them, when calculating the entity relationship transfer probability of the entity transferring to the adjacent entity of the entity, for each entity in the user behavior network, the specified entity relationship between the entity and the adjacent entity of the entity is obtained; for the specified entity relationship in all the obtained entity relationships, the ratio of the specified entity relationship in the entity relationship between the entity and the adjacent entity of the entity is determined; the number of times the specified entity relationship appears in the triples of the user behavior network is counted; based on the counted number of times and the number of triples, the reverse triple probability corresponding to the specified entity relationship is determined; based on the specified entity relationship ratio and the reverse triple probability, the specified entity relationship transfer probability is obtained.
[0049] According to the entity relationship transition probability of the target entity and the preset jump step number of the target entity, all reference entities corresponding to the target entity are determined. Here, based on the entity relationship transition probability obtained above, the embodiment of the present application aims to use the entity relationship transition probability to determine the corresponding reference entity for the target entity. The reference entity can be an entity describing the above-mentioned target entity generated by random walk, that is, the reference entity can not only be an adjacent entity directly connected to the target entity, but also an entity indirectly connected to the target entity, for example: entity A is directly connected to entity B, entity B is directly connected to entity C, and entity A and entity C are not directly connected, then entity A and entity C are indirectly connected through entity B, in which case entity C can be called the reference node of entity A. In specific operations, the reference entity corresponding to the target entity can be determined by setting a preset jump step number, for example: setting the jump step number to 1, then the adjacent node directly connected to the entity is used as the reference entity; setting the jump step number to 2, then taking the entity as the starting point, the entity corresponding to the jump two steps is used as the reference entity, and so on.
[0050] Afterwards, the embedding vector of the target entity is calculated based on the target entity and all reference entities corresponding to the target entity; the embedding vector reflects the entity relationship between the target entity and all reference entities. Here, in the embodiment of the present application, an embedding vector can be used to represent the entity. Since in the user behavior network, the above-mentioned entities may be described in text form, for the original obtained data, in order to facilitate computer processing, it is usually necessary to convert it into a vector representation, that is, to encode the entity into a vector space, so that each entity is represented by a vector in the vector space. For the initial vectorized representation of the original obtained entity, that is, mapping the entity into the vector space, common methods or models can be selected, such as existing semantic mapping methods, etc., which are not limited here. It is precisely because the current vector mapping of entities cannot fully reflect the association between entities, therefore, the embodiment of the present application determines the reference entity corresponding to the entity, performs calculations or multiple rounds of iterative calculations, so that the calculated entity vector can be integrated or reflect the vector characteristics of the reference entity, so that the original vector representation of the entity is optimized.
[0051] It should be noted that there are many ways to implement entity embedding training. The embodiments of this application do not impose fixed restrictions on the specific entity embedding vector representation method, so they will not be elaborated here.
[0052] Furthermore, outputting a text vector corresponding to the user text information includes: determining an embedding vector of an entity corresponding to the user text information in the user behavior network as the text vector of the user text information.
[0053] Afterwards, based on the determined embedding vector, an embedding vector representing the user text information is selected as the text vector of the user text information. It is understandable that since many user text information in live broadcast scenes are in noise and difficult to understand, the embodiment of the present application uses the contextual relationship of the user behavior network to perform entity embedding representation training on various entities. Therefore, the user text information with normal expressions and user text information with noise expressions with similar semantics can be represented with similar embedding vectors with the help of the contextual relationship. Even if there is a large difference between the two user text information due to the influence of noise, the two can be represented with similar embedding vectors with the help of the contextual relationship of the user behavior network, so as to facilitate subsequent accurate text clustering. As for the embedding vectors of other entities, they are only used for the calculation of text vectors, and this part of the embedding vector does not need to be output in actual applications.
[0054] S130: construct a similar text dataset corresponding to the user text information according to the text vector, and perform language model enhancement training based on the similar text dataset.
[0055] Finally, based on the determined text vectors, adaptive text clustering can be performed to cluster semantically similar user text messages with normal expressions and noisy expressions together for language model enhancement training. The similar text dataset can be constructed using a nearest neighbor matching algorithm to calculate similarity between the determined text vectors. Similar text vectors are then found and used to construct the dataset.
[0056] Among them, a similar text dataset corresponding to the user's text information is constructed based on the text vector, including:
[0057] A vector similarity comparison is performed based on the text vectors of the user text information, and the corresponding user text information is selected based on the vector similarity comparison results to construct a similar text dataset.
[0058] It can be understood that if two text vectors have the same vector similarity, even if the actual text representations of the corresponding user text information are different, such as the presence of noise, reversed order, synonyms, homophones, and other different text representations, the text vectors derived from the entity embedding vector representation described above can still achieve a certain degree of similarity between the text vectors of two user text information with similar contextual semantics. Therefore, all text vectors with similarity within a set ratio can be clustered to obtain a similar text dataset.
[0059] Further, referring to FIG3 , language model enhancement training is performed based on a similar text dataset, including:
[0060] S1301. Obtain word segmentation language samples and mask training samples based on a similar text dataset;
[0061] S1302: Input the word segmentation language samples and the mask training samples into the pre-trained language model, and train the language model through the multi-head self-attention network layer.
[0062] Specifically, language models include the BERT model, which uses a Transformer encoder. Due to its self-attention mechanism, the upper and lower layers of the model are directly interconnected. While OpenAI GPT uses a Transformer decoder, it is a restricted Transformer structure requiring left-to-right processing. ELMo uses a bidirectional LSTM, which, while bidirectional, simply concatenates two unidirectional LSTM layers at the top level. Only the BERT model is truly bidirectional in all sentence feature extraction layers, capable of simultaneously capturing the semantic information of the entire sentence context. The language training samples mentioned in the embodiments of this application are the corresponding text training samples, which are continuously trained to produce the corresponding language model. In language model training, the standard left-to-right prediction of the next word is no longer used as the target task. Instead, two new tasks are proposed. The first task, called MLM, randomly blocks 15% of the words in the input word sequence, and the task is to predict these blocked words. Compared to the traditional language model prediction objective function, MLM can predict the probability of a word based on the entire context of these blocked words, rather than just unidirectional information. For example, if the input words are "today, is, a, nice, day," during pre-training, randomly mask a word, such as "is," and then use the contextual information "today," "a," "nice," and "day" to train the model for prediction. This process is repeated to ultimately learn the model parameters.
[0063] Traditional language models don't consider the relationships between sentences. To enable the model to learn these relationships, a second objective task is to predict the next sentence. This is essentially a binary classification problem: 50% of the time, the input is the concatenation of one sentence and the next, with a positive label. The other 50% of the time, the input is the concatenation of one sentence and a random sentence other than the next, with a negative label. The objective function for the entire training process is to perform maximum likelihood learning on these two types of samples.
[0064] By using the above-mentioned language model training method, the pre-trained language model is further enhanced and trained using a similar text dataset, thereby obtaining the ability to understand and analyze text with various noises and difficult to understand. The pre-trained language model can also be trained based on conventional training samples. The embodiments of this application do not impose fixed restrictions on the pre-trained language model and will not be elaborated here.
[0065] In this way, by combining various user behaviors with language model training, the language model's understanding of contexts such as user behavior in different live broadcast rooms can be enhanced.
[0066] Optionally, after performing language model enhancement training based on a similar text dataset, the following is further included:
[0067] Obtain live broadcast room text data, input the live broadcast room text data into the language model to obtain language analysis results, and based on the language analysis results, perform similar user clustering, add content tags to the anchor room, or mine user live broadcast rooms.
[0068] Referring to Figure 4, in actual applications, for various types of difficult-to-understand user text information in live broadcast scenarios, after language model enhancement training is performed using the above-mentioned language model enhancement method, similar user text information can be found in the live broadcast scenario, and then users can be clustered based on similar user text information. Similarly, by analyzing the user text information of the live broadcast room, the content type of the live broadcast room can be understood, and different content type labels can be added to the live broadcast room. In addition, based on the user text information of different live broadcast rooms and combined with the text information of the current user, the live broadcast rooms that the user may be interested in can be analyzed and accurately recommended. In this way, with the help of an enhanced language model, through accurate analysis and understanding of user text information, the application effect of the language model in the live broadcast scenario can be improved, and the live broadcast service and operation and maintenance effects can be improved.
[0069] In the above, a user behavior network is constructed based on live broadcast room information, user identification information, user behavior information, and user text information; entity embedding representation training is performed based on the user behavior network to output text vectors corresponding to the user text information; similar text datasets corresponding to the user text information are constructed based on the text vectors, and language model enhancement training is performed based on the similar text datasets. Using the above technical means, by constructing a user behavior network for entity embedding representation training, similar text datasets can be accurately mined based on text vector representations, and similar samples from different sources can be expanded. This is used to train a language model, which can enhance the language model's ability to understand content in different contextual scenarios, conduct accurate language analysis of user behavior and text expressions, and improve the application effect of the language model.
[0070] Based on the above embodiment, FIG5 is a structural diagram of a language model enhancement system based on a live broadcast room user behavior network provided by this application. Referring to FIG5, the language model enhancement system based on a live broadcast room user behavior network provided by this embodiment specifically includes: a network construction module 21, an entity embedding module 22, and a model enhancement module 23.
[0071] The network construction module 21 is configured to construct a user behavior network based on live broadcast room information, user identification information, user behavior information and user text information;
[0072] The entity embedding module 22 is configured to perform entity embedding representation training based on the user behavior network and output a text vector corresponding to the user text information;
[0073] The model enhancement module 23 is configured to construct a similar text dataset corresponding to the user text information according to the text vector, and perform language model enhancement training based on the similar text dataset.
[0074] Specifically, entity embedding representation training is performed based on the user behavior network, including:
[0075] The live broadcast room information, user identification information, user behavior information and user text information are taken as entities, and entity embedding representation training is performed based on the contextual relationship between entities to obtain the embedding vectors representing each entity in the user behavior network.
[0076] Among them, the text vector corresponding to the user's text information is output, including:
[0077] Determine the embedding vector of the entity corresponding to the user text information in the user behavior network as the text vector of the user text information.
[0078] Specifically, a similar text dataset corresponding to the user's text information is constructed based on the text vector, including:
[0079] A vector similarity comparison is performed based on the text vectors of the user text information, and the corresponding user text information is selected based on the vector similarity comparison results to construct a similar text dataset.
[0080] Specifically, language model enhancement training based on similar text datasets includes:
[0081] Obtain word segmentation language samples and mask training samples based on similar text datasets;
[0082] The word segmentation language samples and mask training samples are input into the pre-trained language model, and the language model is trained through the multi-head self-attention network layer.
[0083] After language model enhancement training based on similar text datasets, it also includes:
[0084] Obtain live broadcast room text data, input the live broadcast room text data into the language model to obtain language analysis results, and based on the language analysis results, perform similar user clustering, add content tags to the anchor room, or mine user live broadcast rooms.
[0085] In the above, a user behavior network is constructed based on live broadcast room information, user identification information, user behavior information, and user text information; entity embedding representation training is performed based on the user behavior network to output text vectors corresponding to the user text information; similar text datasets corresponding to the user text information are constructed based on the text vectors, and language model enhancement training is performed based on the similar text datasets. Using the above technical means, by constructing a user behavior network for entity embedding representation training, similar text datasets can be accurately mined based on text vector representations, and similar samples from different sources can be expanded. This is used to train a language model, which can enhance the language model's ability to understand content in different contextual scenarios, conduct accurate language analysis of user behavior and text expressions, and improve the application effect of the language model.
[0086] The language model enhancement system based on the live broadcast room user behavior network provided in the embodiment of the present application can be configured to execute the language model enhancement method based on the live broadcast room user behavior network provided in the above embodiment, and has corresponding functions and beneficial effects.
[0087] Based on the above practical example, the embodiment of the present application also provides a language model enhancement device based on the live broadcast room user behavior network. Referring to Figure 6, the language model enhancement device based on the live broadcast room user behavior network includes: a processor 31, a memory 32, a communication module 33, an input device 34 and an output device 35. The memory, as a computer-readable storage medium, can be configured to store software programs, computer executable programs and modules, such as the program instructions / modules corresponding to the language model enhancement method based on the live broadcast room user behavior network described in any embodiment of the present application (for example, the network construction module, entity embedding module and model enhancement module in the language model enhancement system based on the live broadcast room user behavior network). The communication module is configured to perform data transmission. The processor executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory, that is, realizing the above-mentioned language model enhancement method based on the live broadcast room user behavior network. The input device can be configured to receive input digital or character information, and generate key signal input related to the user settings and function control of the device. The output device may include a display device such as a display screen. The above-mentioned language model enhancement device based on the live broadcast room user behavior network can be configured to execute the language model enhancement method based on the live broadcast room user behavior network provided in the above-mentioned embodiment, and has corresponding functions and beneficial effects.
[0088] On the basis of the above embodiments, the embodiments of the present application further provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed by a computer processor, are configured to execute a language model enhancement method based on a live broadcast room user behavior network. The storage medium can be any of various types of memory devices or storage devices. Of course, the computer-executable instructions of the computer-readable storage medium provided in the embodiments of the present application are not limited to the language model enhancement method based on a live broadcast room user behavior network as described above, and can also execute the relevant operations in the language model enhancement method based on a live broadcast room user behavior network provided in any embodiment of the present application.
[0089] On the basis of the above embodiments, the embodiments of the present application also provide a computer program product. The essence of the technical solution of the present application or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes a number of instructions for enabling a computer device, a mobile terminal or a processor therein to execute all or part of the steps of the language model enhancement method based on the live broadcast room user behavior network described in each embodiment of the present application.
Claims
1. A language model enhancement method based on a live broadcast room user behavior network, wherein: include: Construct a user behavior network based on live broadcast room information, user identification information, user behavior information and user text information; Perform entity embedding representation training based on the user behavior network, and output a text vector corresponding to the user text information; A similar text data set corresponding to the user text information is constructed according to the text vector, and language model enhancement training is performed based on the similar text data set.
2. The language model enhancement method based on the live broadcast room user behavior network according to claim 1, wherein: The entity embedding representation training based on the user behavior network includes: The live broadcast room information, the user identification information, the user behavior information and the user text information are taken as entities, and entity embedding representation training is performed based on the contextual relationship between entities to obtain an embedding vector representing each entity in the user behavior network.
3. The language model enhancement method based on the live broadcast room user behavior network according to claim 2, wherein: The output corresponds to the text vector of the user text information, including: Determine an embedding vector of an entity corresponding to the user text information in the user behavior network as a text vector of the user text information.
4. The language model enhancement method based on the live broadcast room user behavior network according to claim 1, wherein: The step of constructing a similar text data set corresponding to the user text information according to the text vector includes: A vector similarity comparison is performed based on the text vectors of the user text information, and the corresponding user text information is selected based on the vector similarity comparison result to construct a similar text data set.
5. The language model enhancement method based on the live broadcast room user behavior network according to claim 1, wherein: The language model enhancement training based on the similar text data set includes: Acquire word segmentation language samples and mask training samples based on the similar text dataset; The word segmentation language sample and the mask training sample are input into a pre-trained language model, and the language model is trained through a multi-head self-attention network layer.
6. The language model enhancement method based on the live broadcast room user behavior network according to claim 5, wherein: After performing language model enhancement training based on the similar text dataset, the method further includes: Acquire live broadcast room text data, input the live broadcast room text data into the language model to obtain a language analysis result, and perform similar user clustering, anchor room content tag addition, or user live broadcast room mining based on the language analysis result.
7. A language model enhancement system based on a live broadcast room user behavior network, wherein: include: A network construction module, configured to construct a user behavior network based on live broadcast room information, user identification information, user behavior information and user text information; An entity embedding module, configured to perform entity embedding representation training based on the user behavior network and output a text vector corresponding to the user text information; The model enhancement module is configured to construct a similar text data set corresponding to the user text information according to the text vector, and perform language model enhancement training based on the similar text data set.
8. A language model enhancement device based on a live broadcast room user behavior network, wherein: include: memory and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the language model enhancement method based on the live broadcast room user behavior network as described in any one of claims 1-6.
9. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions, which, when executed by a computer processor, are configured to execute the language model enhancement method based on the live broadcast room user behavior network as described in any one of claims 1-6.
10. A computer program product, wherein: The computer program product includes instructions, which, when executed on a computer or a processor, enable the computer or the processor to execute the language model enhancement method based on the live broadcast room user behavior network as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Language model training method, copywriting generation method and related equipment
CN114048289A
Multi-modal named entity identification method and system based on cross-modal feature enhancement network
CN117057352A
Language model enhancement method and system based on live broadcast room user behavior network
CN117709339A
Domain terminology expansion by relevancy
US20180197530A1