Dialogue text vector representation method and computer readable storage medium

By processing dialogue text into sentences and performing multi-layer embedding operations, and combining a mask matrix to determine the role sentence vector matrix, this method solves the problem that existing methods do not consider dialogue scenario information, thereby improving the accuracy of dialogue text vector representation and the accuracy of topic classification.

CN119066198BActive Publication Date: 2026-05-22WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WEBANK (CHINA)
Filing Date
2024-08-08
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing methods for representing dialogue text sentences fail to adequately consider the information of the dialogue scenario, resulting in low accuracy in subsequent topic classification.

Method used

By segmenting the target dialogue text into sentences and inputting it into a large language model for sentence embedding, role embedding, and dialogue turn embedding, and combining the mask matrix to determine the role sentence vector matrix, a deep vector representation of the dialogue text is achieved.

Benefits of technology

It enhances the connections between dialogues between different characters and the sub-levels between dialogue rounds, improves the accuracy of subsequent topic classification, covers both the individuality and commonalities of characters, reflects the differences in identity and status, and improves the grounding of vector representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119066198B_ABST
    Figure CN119066198B_ABST
Patent Text Reader

Abstract

The application discloses a dialogue text vector representation method and a computer readable storage medium. The method comprises the following steps: processing a target dialogue text by sentence to obtain a plurality of sentences, inputting the plurality of sentences into a first large language model to obtain a plurality of sentence vectors, performing sentence embedding operation, role embedding operation and dialogue round embedding operation on the plurality of sentence vectors to obtain a plurality of sentence embedding sentence vectors, a plurality of role embedding sentence vectors and a plurality of dialogue round embedding sentence vectors, adding the plurality of sentence embedding sentence vectors, the plurality of role embedding sentence vectors and the plurality of dialogue round embedding sentence vectors to obtain a first sentence vector matrix, inputting the first sentence vector matrix into a second large language model to obtain a second sentence vector matrix, obtaining a mask matrix of each role in at least two roles to obtain at least one mask matrix, determining a sentence vector matrix of at least one role according to the second sentence vector matrix and the at least one mask matrix, and determining a target dialogue text sentence vector representation according to the sentence vector matrix of the at least one role.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method for representing dialogue text vectors and a computer-readable storage medium. Background Technology

[0002] Currently, banking services are gradually shifting from in-person, offline branch services to online platforms, resulting in a large volume of telephone transactions and online chat conversations. These call recordings can be converted into conversational text corpora using Automatic Sentence Recognition (ASR). NLP techniques can then be used to analyze customer needs and staff behavior, and to categorize conversation topics, thus better meeting customer needs and providing better service. The technological foundation for analyzing these conversational texts lies in better vector representation of the dialogue sentences. Currently, existing methods for representing conversational text sentences (such as contrastive algorithms and BERT) do not consider the context of the conversation, leading to lower accuracy in subsequent topic classification. Summary of the Invention

[0003] This application provides a method for representing dialogue text vectors and a computer-readable storage medium, which can deeply consider the information of the dialogue scenario to improve the accuracy of subsequent topic classification.

[0004] In a first aspect, embodiments of this application provide a method for representing dialogue text vectors, the method comprising:

[0005] Obtain target dialogue text, wherein the target dialogue text includes at least two characters and at least one round of dialogue;

[0006] The target dialogue text is segmented into sentences to obtain multiple sentences;

[0007] The multiple sentences are input into the first large language model to obtain multiple sentence vectors;

[0008] Perform sentence embedding, role embedding, and dialogue turn embedding operations on the multiple sentence vectors to obtain multiple sentence embedding sentence vectors, multiple role embedding sentence vectors, and multiple dialogue turn embedding sentence vectors;

[0009] The first sentence vector matrix is ​​obtained by adding the sentence vectors embedded in the multiple sentences, the sentence vectors embedded in the multiple roles, and the sentence vectors embedded in the multiple dialogue rounds.

[0010] Input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix;

[0011] Obtain the mask matrix for each of the at least two roles to obtain at least one mask matrix;

[0012] The sentence vector matrix of at least one character is determined based on the second sentence vector matrix and the at least one mask matrix;

[0013] The sentence vector representation of the target dialogue text is determined based on the sentence vector matrix of the at least one character.

[0014] Secondly, embodiments of this application provide a dialog text vector representation apparatus, the apparatus comprising: an acquisition unit, a sentence segmentation unit, an input unit, an embedding unit, and a determination unit, wherein,

[0015] The acquisition unit is used to acquire target dialogue text, which includes at least two characters and at least one round of dialogue.

[0016] The sentence segmentation unit is used to segment the target dialogue text into multiple sentences;

[0017] The input unit is used to input the multiple sentences into the first large language model to obtain multiple sentence vectors;

[0018] The embedding unit is used to perform sentence embedding, role embedding, and dialogue turn embedding operations on the multiple sentence vectors to obtain multiple sentence embedding sentence vectors, multiple role embedding sentence vectors, and multiple dialogue turn embedding sentence vectors; the multiple sentence embedding sentence vectors, the multiple role embedding sentence vectors, and the multiple dialogue turn embedding sentence vectors are added together to obtain a first sentence vector matrix;

[0019] The input unit is also used to input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix;

[0020] The acquisition unit is also used to acquire the mask matrix of each of the at least two roles, to obtain at least one mask matrix;

[0021] The determining unit is configured to determine the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix; and to determine the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one character.

[0022] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing the steps in the first aspect of embodiments of this application.

[0023] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of embodiments of this application.

[0024] Fifthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of embodiments of this application. The computer program product may be a software installation package.

[0025] Implementing the embodiments of this application has the following beneficial effects:

[0026] The dialogue text vector representation method and computer-readable storage medium described in this application involve obtaining target dialogue text, which includes at least two roles and at least one round of dialogue. The target dialogue text is segmented into sentences to obtain multiple sentences. These sentences are input into a first language model to obtain multiple sentence vectors. Sentence embedding, role embedding, and round-of-dialogue embedding operations are performed on these sentence vectors to obtain multiple sentence-embedded sentence vectors, multiple role-embedded sentence vectors, and multiple round-of-dialogue embedding sentence vectors. These vectors are then summed to obtain a first sentence vector matrix. This first sentence vector matrix is ​​input into a second language model to obtain a second sentence vector matrix. A mask matrix is ​​obtained for each of the at least two roles, resulting in at least one mask matrix. Based on the second sentence vector matrix... A vector matrix and at least one mask matrix determine the sentence vector matrix of at least one character. Based on the sentence vector matrix of at least one character, the sentence vector representation of the target dialogue text is determined. In this way, considering the context can enhance the relationship between different characters' dialogues and guide the sub-levels between dialogue rounds, so that each round of dialogue can be smoothly connected, and the vector representation is also smoother. Characters represent identities, and different identities correspond to different social statuses. Deeply exploring the emotions and status between characters, different rounds correspond to the switching of different dialogue sub-topics, that is, information related to various dimensions of the dialogue topic can be mined from different dimensions, so that the final sentence vector representation of the dialogue text is more in line with the actual application scenario of the topic, and also covers the personality and commonality of characters, as well as the differences between different statuses and identities, making the vector representation more down-to-earth and helping to improve the accuracy of subsequent topic classification. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1A This is a flowchart illustrating a method for representing dialogue text vectors according to an embodiment of this application;

[0029] Figure 1B This is another flowchart illustrating a method for representing dialogue text vectors provided in an embodiment of this application;

[0030] Figure 2 This is a flowchart illustrating another method for representing dialog text vectors provided in an embodiment of this application;

[0031] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0032] Figure 4 This is a block diagram of the functional units of a dialog text vector representation device provided in an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0034] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0035] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0036] The following is a brief introduction to the relevant terminology used in this application.

[0037] The electronic devices described in this application embodiment may include smartphones (such as Android phones, iOS phones, Windows Phones, etc.), tablet computers, PDAs, dashcams, servers, laptops, mobile internet devices (MIDs) or wearable devices (such as smartwatches, Bluetooth headsets), etc. The above are merely examples and not exhaustive, including but not limited to the above electronic devices. The server may be an outsourced server, cloud server, edge server, etc., which are not limited here.

[0038] Natural Language Processing (NLP) enables computers to understand, generate, and process natural languages, such as English and Chinese, which are languages ​​used by humans in daily life. Tasks include sentiment analysis, machine translation, named entity recognition, speech recognition, and question answering systems.

[0039] Transformer: A deep learning model based on the self-attention mechanism. It is widely used in many other tasks in natural language processing, such as text classification, sentiment analysis, and question answering systems.

[0040] Word embedding is a technique that maps discrete data (such as words, categories, etc.) to continuous vectors. In the field of Natural Language Processing (NLP), word embeddings are typically used to convert words or phrases into fixed-length vector representations so that computers can better understand and process natural language data.

[0041] Contrastive learning is an unsupervised learning method that aims to learn useful feature representations from large amounts of unlabeled data. Its core idea is to automatically construct similar and dissimilar instances, training a model such that similar instances are close in distance in the feature space, while dissimilar instances are far apart. This approach encourages the model to learn high-level features capable of distinguishing different categories of data.

[0042] Automatic Speech Recognition (ASR) is a technology that converts human speech signals into computer-readable text. It is an important branch of artificial intelligence and natural language processing, and is widely used in scenarios such as intelligent assistants, voice search, customer service automation, and accessibility technology.

[0043] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on Transformers. In BERT, `[CLS]` and `[SEP]` are special tokens used to indicate the beginning and end of a sentence, or to separate multiple sentences.

[0044] Among related technologies, with the rapid development of deep learning, deep learning-based Natural Language Processing (NLP) technology has demonstrated significant advantages over traditional rule-based and syntactic analysis methods in dialogue systems, information retrieval, text mining, and sentiment analysis. Text vectorization is the foundation of NLP algorithms. Vector representation methods aim to capture and represent semantic information by mapping text into a continuous, low-dimensional vector space. After obtaining the vector representation of the text, it can be input into downstream models that calculate similarity and text classification to complete tasks such as information retrieval and sentiment recognition.

[0045] In practical implementation, let's take an example, suppose we have the following dialogue case D:

[0046] Customer Service: Hello, this is XX Bank customer service.

[0047] Customer: What gifts can I redeem with my points? I'd like to know.

[0048] Customer service: Points can be redeemed for various gifts, such as travel packages, movie tickets, etc. You can log in to your online banking to view the redemption list and choose the gift that suits you.

[0049] Customer: A travel package? That sounds good, I'll have to check it out!

[0050] In related technologies, the contrast algorithm (Dial2vec) can be used to represent the dialogue text vectors. Dial2vec has achieved good results on many validation datasets in clustering, dialogue retrieval, and dialogue semantic similarity ranking. However, because this method treats the entire dialogue as a single sample input, and the model's time complexity is O(nd^2), Dial2vec is limited to a length of 512. However, in real-world banking customer service scenarios, many dialogues exceed 512 characters, especially in bank telemarketing and debt collection scenarios, where the dialogue text length often exceeds 512 characters. Therefore, this limitation necessitates truncation of the dialogue text in many scenarios using Dial2vec. Furthermore, training a large amount of text with a length of 512 characters is time-consuming and places high demands on hardware resources.

[0051] To address the shortcomings of related technologies, embodiments of this application provide a method for representing dialogue text vectors, which may include the following steps:

[0052] Obtain target dialogue text, wherein the target dialogue text includes at least two characters and at least one round of dialogue;

[0053] The target dialogue text is segmented into sentences to obtain multiple sentences;

[0054] The multiple sentences are input into the first large language model to obtain multiple sentence vectors;

[0055] Perform sentence embedding, role embedding, and dialogue turn embedding operations on the multiple sentence vectors to obtain multiple sentence embedding sentence vectors, multiple role embedding sentence vectors, and multiple dialogue turn embedding sentence vectors;

[0056] The first sentence vector matrix is ​​obtained by adding the sentence vectors embedded in the multiple sentences, the sentence vectors embedded in the multiple roles, and the sentence vectors embedded in the multiple dialogue rounds.

[0057] Input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix;

[0058] Obtain the mask matrix for each of the at least two roles to obtain at least one mask matrix;

[0059] The sentence vector matrix of at least one character is determined based on the second sentence vector matrix and the at least one mask matrix;

[0060] The sentence vector representation of the target dialogue text is determined based on the sentence vector matrix of the at least one character.

[0061] In this embodiment, a target dialogue text is obtained, comprising at least two roles and at least one round of dialogue. The target dialogue text is processed into sentences to obtain multiple sentences. These multiple sentences are input into a first large language model to obtain multiple sentence vectors. Sentence embedding, role embedding, and dialogue round embedding operations are performed on these multiple sentence vectors to obtain multiple sentence-embedded sentence vectors, multiple role-embedded sentence vectors, and multiple dialogue round-embedded sentence vectors. These vectors are then added together to obtain a first sentence vector matrix. This first sentence vector matrix is ​​input into a second large language model to obtain a second sentence vector matrix. A mask matrix is ​​obtained for each of the at least two roles to obtain at least one mask matrix. The sentence vector matrix for at least one role is determined based on the second sentence vector matrix and the at least one mask matrix. The target dialogue text sentence vector representation is determined by the matrix. In this way, considering the context can enhance the connection between different roles in the dialogue and guide the sub-level between dialogue rounds, so that each round of dialogue can be smoothly connected, and the vector representation is also smoother. Roles represent identities, and different identities correspond to different social statuses. The deep mining of the emotions and status between roles, and the switching of different dialogue sub-topics in different rounds, can mine various dimensions of information related to the dialogue topic from different dimensions, so that the final dialogue text sentence vector representation is more in line with the actual application scenario of the topic, and also covers the personality and commonality of the roles, as well as the differences between different statuses and identities, making the vector representation more down-to-earth and helping to improve the accuracy of subsequent topic classification. At the same time, it solves the problem of input length and resource problem (computational resources) of the comparison algorithm (Dial2vec) in the paper, which directly encodes the entire dialogue into word-token.

[0062] The embodiments of this application will be described in detail below.

[0063] Please see Figure 1A , Figure 1A This is a flowchart illustrating a dialog text vector representation method provided in an embodiment of this application. As shown in the figure, this dialog text vector representation method includes:

[0064] 101. Obtain the target dialogue text, wherein the target dialogue text includes at least two characters and at least one round of dialogue.

[0065] In this embodiment of the application, the target dialogue text can be the dialogue text converted from a call recording. Specifically, a call recording can be obtained, and the call recording can be translated into dialogue text to obtain the target dialogue text.

[0066] The target dialogue text may include at least two characters and may include at least one round of dialogue.

[0067] 102. The target dialogue text is segmented into sentences to obtain multiple sentences.

[0068] In this embodiment of the application, the target dialogue text can be processed by sentence segmentation, that is, a dialogue can be split into multiple sentences to obtain multiple sentences.

[0069] 103. Input the multiple sentences into the first large language model to obtain multiple sentence vectors.

[0070] In this embodiment, the first major language model can be preset or set by system default. The first major language model may include at least one of the following: Generative Pre-Trained (GPT) model, BERT model, t5 model, Roberta model, etc., without limitation.

[0071] In practice, multiple sentences can be input into the first language model in parallel. Specifically, after inputting sentences in parallel, multiple sentence vectors are obtained. In this way, the sentence vectors in the dialogue can be learned at the word-token level.

[0072] 104. Perform sentence embedding, role embedding, and dialogue turn embedding operations on the multiple sentence vectors to obtain multiple sentence embedding sentence vectors, multiple role embedding sentence vectors, and multiple dialogue turn embedding sentence vectors.

[0073] In this embodiment, multiple sentence vectors can be embedded using sentence embedding, role embedding, and dialogue turn embedding operations to obtain multiple sentence-embedded sentence vectors, multiple role-embedded sentence vectors, and multiple dialogue turn embedded sentence vectors. Multiple sentence-embedded sentence vectors can enhance the connections between different roles by considering the context, deeply explore the emotions and status between roles by using multiple role-embedded sentence vectors, and guide the sub-levels between dialogue turns by using multiple dialogue turn embedded sentence vectors, allowing each turn of dialogue to be smoothly connected. In other words, considering the context enhances the connections between different roles and guides the sub-levels between dialogue turns, enabling a smoother connection between each turn of dialogue. The vector representation is also smoother. Roles represent identities, and different identities correspond to different social statuses. Deeply exploring the emotions and status between roles, and the switching of different dialogue sub-topics in different turns, allows for the mining of various dimensions related to the dialogue topic from different dimensions. This makes the final dialogue text sentence vector representation more consistent with the actual application scenario of the topic, and also covers the individuality and commonalities of roles, as well as the differences between different statuses and identities. This makes the vector representation more grounded and helps improve the accuracy of subsequent topic classification.

[0074] 105. Add the multiple sentence embedding vectors, the multiple role embedding vectors, and the multiple dialogue turn embedding vectors together to obtain the first sentence vector matrix.

[0075] In this embodiment, multiple sentence embeddings, multiple role embeddings, and multiple dialogue round embeddings are summed to obtain the first sentence vector matrix. This approach, considering context, enhances the connections between different roles in the dialogue and guides the sub-levels between dialogue rounds, allowing for a smoother connection between each round. The vector representation is also smoother. Roles represent identities; different identities correspond to different social statuses. Deeply exploring the emotions and statuses between roles, and the switching of different dialogue sub-topics in different rounds, allows for the extraction of various dimensions related to the dialogue theme. This makes the final dialogue text sentence vector representation more consistent with the actual application scenario of the theme, encompassing both the individuality and commonalities of the roles, as well as the differences between different statuses and identities. This makes the vector representation more grounded and helps improve the accuracy of subsequent theme classification. For example, the same sentence can have different meanings depending on the context in which it is spoken, or the meaning can differ depending on the speaker.

[0076] 106. Input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix.

[0077] The second language model can be preset or set by the system default, while the first language model can be the same as or different from the second language model. The second language model can include at least one of the following: Generative Pre-Trained (GPT) model, BERT model, t5 model, Roberta model, etc., without limitation.

[0078] In practice, the first sentence vector matrix can be input into the second language model to obtain the second sentence vector matrix. In this way, the sentence vectors in the dialogue can be learned at the sentence-token level.

[0079] The second language model can be the same as or different from the first language model.

[0080] 107. Obtain the mask matrix of each of the at least two roles to obtain at least one mask matrix.

[0081] In this embodiment of the application, the mask matrix of each of the at least two roles can be obtained to obtain at least one mask matrix. That is, in order to obtain the sentence vector matrix corresponding to different roles, the mask matrix of each of the at least two roles can be obtained to obtain at least one mask matrix, and the sentence vector matrix corresponding to different roles can be obtained through the mask matrix.

[0082] 108. Determine the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix.

[0083] In the specific implementation, each role corresponds to a mask matrix. Each role's mask matrix is ​​used to extract the chat content and related features corresponding to that role. The sentence vector matrix of at least one role can be determined based on the second sentence vector matrix and at least one mask matrix. In this way, the chat content of each role can be obtained based on the dialogue order to generate its corresponding sentence vector matrix.

[0084] 109. Determine the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one role.

[0085] In practice, the sentence vector representation of the target dialogue text can be determined based on the sentence vector matrix of at least one character. In this way, the sentence vector representation of the dialogue text of each character can be obtained. Of course, the sentence vector representation of the dialogue text of multiple characters can also be obtained.

[0086] In some possible examples, step 109 above, determining the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one character, may include the following steps:

[0087] 91. Determine the self-interaction sentence vector matrix and the role interaction sentence vector matrix for each role based on the sentence vector matrix of at least one role;

[0088] 92. Determine the first sentence vector representation of each of the at least one characters based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of the at least one character;

[0089] 93. Determine the second sentence vector representation of each character in the at least one character based on the first sentence vector representation of each character and the first sentence vector matrix;

[0090] 94. Determine the target dialogue text sentence vector representation based on the second sentence vector representation of each of the at least one role.

[0091] In this context, the self-interaction sentence vector matrix of each character can be understood as the result of the interaction between the sentence vector matrix of that character and the sentence vector matrix of that character, and the character interaction sentence vector matrix can be understood as the result of the interaction between the sentence vector matrix of that character and the sentence vector matrices of other characters.

[0092] In practice, the self-interaction sentence vector matrix and the role interaction sentence vector matrix of each role can be determined based on the sentence vector matrix of at least one role. The self-interaction sentence vector matrix can be used to repeatedly explore the identity and status of the role, and the role attributes of the role can be enhanced by deduce the emotional changes between similar roles. Correspondingly, the role interaction sentence vector matrix can contain the attitude of the role towards other roles, showing the differences between identities, as well as other humble attributes, and deeply explore the identity and status attributes of the role. Moreover, by deduce the emotional changes between different roles, the role vector representation becomes more grounded.

[0093] Furthermore, the first sentence vector representation of each character in at least one character can be determined based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of at least one character. The second sentence vector representation of each character in at least one character can be determined based on the first sentence vector representation of each character in at least one character and the first sentence vector matrix. The sentence vector representation of the target dialogue text can be determined based on the second sentence vector representation of each character in at least one character.

[0094] Furthermore, in some possible examples, step 91 above, determining the self-interaction sentence vector matrix and the character interaction sentence vector matrix for each character based on the sentence vector matrix of the at least one character, may include the following steps:

[0095] 911. Obtain the first transpose matrix of the sentence vector matrix of the first character, wherein the first character is any one of the at least one characters;

[0096] 912. Determine the self-interaction sentence vector matrix of the first character based on the sentence vector matrix of the first character and the first transpose matrix;

[0097] 913. Determine the second transpose of the sentence vector matrix of the second role, wherein the second role is any role other than the first role among the at least one role;

[0098] 914. Determine the character interaction sentence vector matrix of the first character based on the sentence vector matrix of the first character and the second transpose matrix.

[0099] In specific implementation, taking the first role as an example, the first role is any one of at least one roles. The first transpose matrix of the sentence vector matrix of the first role can be obtained. The self-interaction sentence vector matrix of the first role can be determined based on the sentence vector matrix of the first role and the first transpose matrix. Through the self-interaction sentence vector matrix, the identity and status of this role can be repeatedly explored. Moreover, by deducing the emotional changes between similar roles, the role attributes of this role can be enhanced.

[0100] Furthermore, taking the second role as an example, the second role is any role other than the first role among at least one roles. The second transpose of the sentence vector matrix of the second role can be determined. Then, the role interaction sentence vector matrix of the first role can be determined based on the sentence vector matrix of the first role and the second transpose matrix. In this way, the role interaction sentence vector matrix can contain the attitude of the role towards other roles, showing the differences between identities and other humble attributes, deeply exploring the identity and status attributes of the roles, and making the role vector representation more down-to-earth by deducing the emotional changes between different roles.

[0101] To illustrate, let's take a target dialogue text containing two characters, r1 and r2. To obtain interaction information, the implementation performs a dot product operation between sentence-level characters, which yields the self-interaction matrix C. r1 And the role-interaction matrix C between different roles r1′ And calculate the interaction matrix for each interlocutor.

[0102] In the specific implementation, the interaction matrix can be obtained for character r1:

[0103] C r1 =H r1 (H r1 ) T

[0104] C r1′ =H r1 (H r2 ) T

[0105] Among them, C r1 ∈R r1*r1 C r1′ =R r1*r2 C r1 H represents the self-interaction sentence vector matrix of character r1. r1 The sentence vector matrix representing character r1, (H r1 ) T C represents the transpose of the sentence vector matrix of character r1. r1′ The vector matrix representing the character interaction sentences of character r1, (H r2 ) T This represents the transpose of the sentence vector matrix of character r2.

[0106] Accordingly, for character r2, the interaction matrix can be obtained as follows:

[0107] C r2 =H r2 (Hr2 ) T

[0108] C r2′ =H r2 (H r1 ) T

[0109] Among them, C r2 ∈R r2*r2 C r2′ =R r2*r1 C r2 H represents the self-interaction sentence vector matrix of character r2. r2 The sentence vector matrix representing character r2, (H r2 ) T C represents the transpose of the sentence vector matrix of character r2. r2′ The vector matrix representing the character interaction sentences of character r2, (H r1 ) T This represents the transpose of the sentence vector matrix of character r1.

[0110] Furthermore, in some possible examples, step 92 above, determining the first sentence vector representation of each of the at least one character based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of the at least one character, may include the following steps:

[0111] 921. Determine the self-interaction sentence vector representation of the second role based on the self-interaction sentence vector matrix of the second role and the sentence vector matrix of the second role;

[0112] 922. Determine the character sentence vector representation of the second character based on the character interaction sentence vector matrix of the second character and the sentence vector matrix of the third character; the third character is any character other than the second character among the at least one character;

[0113] 923. Determine the first sentence vector representation of the second role based on the self-interactive sentence vector representation and the role sentence vector representation.

[0114] In this embodiment, taking the second role as an example, the second role is any one of at least one roles. The self-interaction sentence vector representation of the second role can be determined based on the self-interaction sentence vector matrix of the second role and the sentence vector matrix of the second role. Then, the third role is any one of at least one roles other than the second role. The role sentence vector representation of the second role is determined based on the role interaction sentence vector matrix of the second role and the sentence vector matrix of the third role. The first sentence vector representation of the second role is determined based on the self-interaction sentence vector representation and the role sentence vector representation. In specific implementation, the first sentence vector representation of the second role can include the self-interaction sentence vector representation and the role sentence vector representation of the second role. The self-interaction can repeatedly explore the identity and status of the role, and enhance the role attributes of the role by deducing the emotional changes between similar roles. The interaction between corresponding roles can include the attitude of the role towards other roles, showing the differences between identities, as well as other humble attributes, deeply exploring the identity and status attributes of the role, and making the role vector representation more down-to-earth by deducing the emotional changes between different roles.

[0115] In practice, when the target dialogue text includes two roles, the third role can be the first role; when the target dialogue text includes three or more roles, the third role can be any role other than the first or second role.

[0116] To illustrate, let's take a target dialogue text that includes two characters, r1 and r2. Then, we can obtain two different sentence vector representations for each character.

[0117] In the specific implementation, for character r1:

[0118] E r1 =C r1 H r1

[0119] E r1′ =C r1′ H r2

[0120] Among them, E r1 ∈R r1*n E r1 ′∈R r1*n , where C r1 H represents the self-interaction sentence vector matrix of character r1. r1 E represents the sentence vector matrix of character r1. r1 H represents the self-interaction sentence vector representation of character r1. r2 E represents the sentence vector matrix of character r2. r1′ The sentence vector representation of role r1.

[0121] Furthermore, regarding character r2:

[0122] E r2 =C r2 H r2

[0123] E r2′ =C r2′ H r1

[0124] Among them, E r2 ∈R r2*n E r2 ′∈R r2*n , where C r2 H represents the self-interaction sentence vector matrix of character r2. r2 E represents the sentence vector matrix of character r2. r2 H represents the self-interaction sentence vector representation of role r2. r1 E represents the sentence vector matrix of character r1. r2′ The sentence vector representation of role r2.

[0125] Furthermore, for each role, a vector E representing the information combined with self-interaction information can be obtained. r1 E r2 And a representation vector E that combines information about interactions between different roles. r1′ E r2′ .

[0126] Furthermore, in some possible examples, step 93 above, determining the second sentence vector representation of each of the at least one character based on the first sentence vector representation of each character and the first sentence vector matrix, may include the following steps:

[0127] 931. Perform a weighted operation on the self-interactive sentence vector representation and the role sentence vector representation to obtain the intermediate sentence vector representation;

[0128] 932. Concatenate the first sentence vector matrix and the intermediate sentence vector representation to obtain the second sentence vector representation of the second character.

[0129] In this embodiment, the weights of the self-interaction sentence vector representation and the role sentence vector representation can be obtained, resulting in multiple weights. The sum of these multiple weights is 1. Then, the weights, along with the self-interaction sentence vector representation and the role sentence vector representation, are used to perform a weighted operation to obtain the intermediate sentence vector representation. Next, the first sentence vector matrix and the intermediate sentence vector representation are concatenated to obtain the second sentence vector representation of the second role. The self-interaction can repeatedly explore the identity and status of the role, and enhance the role attributes of the role by deducing the emotional changes between similar roles. The interaction between corresponding roles can include the role's attitude towards other roles, showing the differences between identities, as well as other humble attributes, deeply exploring the identity and status attributes of the role. Moreover, by deducing the emotional changes between different roles, the role vector representation becomes more grounded.

[0130] For example, let's take a target dialogue text that includes two characters, r1 and r2. Taking r1 as an example, the details are as follows:

[0131] f i r1 =H i ||E i

[0132]

[0133] Where || denotes vector concatenation, θ is a parameter used to control the proportion of self-interaction and inter-role interaction in the final dialogue text sentence vector information, and f i r1 ∈R 1*2d , i∈[1,r1]. By controlling the parameter θ, E i It can combine semantic information from self-interaction and interaction between roles simultaneously, or it can utilize these two types of information separately, such as E. i =E r1 ′(θ=0), or E i =E r1 (θ=1). E r1 E represents the self-interaction sentence vector representation of character r1. r1′ E represents the sentence vector representation of character r1. i H represents the vector representation of the middle sentence of character r1. i It can represent the word-token level sentence vector representation of role r1.

[0134] Similarly, taking character r2 as an example, the details are as follows:

[0135] f i r2 =H i ||Ei

[0136]

[0137] Where || denotes vector concatenation, θ is a parameter used to control the proportion of self-interaction and inter-role interaction in the final dialogue text sentence vector information, and f i r2 ∈R 2*2d , i∈[[1, r2]. By controlling the parameter θ, E i It can combine semantic information from self-interaction and interaction between roles simultaneously, or it can utilize these two types of information separately, such as E. i =E r1′ (θ=0), or E i =E r1 (θ=1). E r1 E represents the self-interaction sentence vector representation of character r1. r2 E represents the self-interactive sentence vector representation of character r2. r2′ H represents the sentence vector representation of character r2. i It can represent the word-token level sentence vector representation of role r2.

[0138] In this example, sentence vectors from both the word-token and sentence-token levels are learned for representation. The two vectors are then concatenated, preserving both the basic semantic information from a large corpus of publicly available text and learning semantic information specific to the dialogue, such as context, turn order, and dialogue character interactions. This approach overcomes the limitations of contrastive algorithms (Dial2vec), which directly encode the entire dialogue into word-tokens, incurring input length and computational resource issues.

[0139] In some possible examples, step 108 above, determining the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix, may include the following steps:

[0140] 81. Determine the third transpose of the first mask matrix corresponding to the first role; the first mask matrix is ​​any one of the at least one mask matrices; the first role is one of the at least one roles;

[0141] 82. Perform a dot product operation on the second sentence vector matrix and the third transpose matrix to obtain the sentence vector matrix of the first character.

[0142] In this embodiment of the application, the third transpose of the first mask matrix corresponding to the first role is determined. The first mask matrix is ​​any one of the at least one mask matrix, and the first role is one of the at least one roles. Next, the dot product operation is performed on the second sentence vector matrix and the third transpose matrix to obtain the sentence vector matrix of the first role. In this way, the dialogue content of each role can be obtained based on the dialogue order to generate its corresponding sentence vector matrix.

[0143] To illustrate, let's take a target dialogue text with two characters as an example. In order to obtain the sentence vector matrix corresponding to different characters, a mask matrix m is generated for each of the two speakers. r1 and m r2 , m r1 The i-th element in the matrix has a value of 1 when the input token comes from r1, and a value of 0 otherwise. The sentence vector matrix for different roles can be obtained by calculating using the following formula:

[0144] H r1 =H c ⊙(m r1 ) T

[0145] H r2 =H e ⊙(m r2 ) T

[0146] Here, ⊙ denotes element-wise multiplication of matrices. If two characters in a dialogue, r1 speaks r1 sentences and r2 speaks r2 sentences, then the sentence vectors H for each character are obtained. r1 ∈R r1*d H r2 ∈R r2*d H r1 H represents the sentence vector matrix of character r1. r2 The sentence vector matrix representing character r2, m r1 This represents the mask matrix for character r1, (m r1 ) T m represents the transpose of the mask matrix of character r1. r2 This represents the mask matrix for character r2, (m r2 ) T H represents the transpose of the mask matrix of character r2. e ∈R k*d For sentence-token level semantic learning of all sentences in the dialogue, the corresponding initialization vector (i.e., H) is obtained. e (Corresponding to the vector matrix in the second sentence above).

[0147] In some possible examples, the following steps may also be included:

[0148] A1. Use the target dialogue text as a positive sample;

[0149] A2. Randomly sample and replace sentences in the target dialogue text to obtain negative samples;

[0150] A3. Train the first large language model and / or the second large language model based on the positive samples, the negative samples, and a preset loss function.

[0151] In this embodiment of the application, the preset loss function can be preset or set by system default.

[0152] In practice, the target dialogue text can be used as a positive sample. Sentences in the target dialogue text can be randomly sampled and replaced to obtain negative samples. Then, the first language model can be trained based on the positive and negative samples and the preset loss function. In this way, the generalization ability of the first language model can be enhanced.

[0153] Of course, in a specific implementation, the target dialogue text can be used as a positive sample, and sentences in the target dialogue text can be randomly sampled and replaced to obtain negative samples. Then, the second language model can be trained based on the positive and negative samples and a preset loss function. In this way, the generalization ability of the second language model can be enhanced.

[0154] In some possible examples, the following steps may also be included:

[0155] B1. Obtain the audio file of the target dialogue text;

[0156] B2. Sample the audio files to obtain the corpus content;

[0157] B3. Sample dialogue sentences and dialogue scene expansion prompts based on the corpus content, input the dialogue sentences and dialogue scene expansion prompts into the first large language model to expand the dialogue corpus as training samples, so as to enhance the generalization ability of the first large language model.

[0158] In this embodiment, an audio file of the target dialogue text can be obtained. That is, the audio file contains not only text information but also corresponding sound information. The audio file is sampled to obtain the corpus content. The corpus content may also include the user's voice characteristics. Then, dialogue sentences and dialogue scene expansion prompts are sampled according to the corpus content. The dialogue sentences and dialogue scene expansion prompts are input into the first language model to expand the dialogue corpus. Such dialogue corpus not only deeply matches the user's identity but also deeply matches the actual context. This dialogue corpus, as a training sample, can enhance the generalization ability of the first language model.

[0159] In specific implementation, taking the first major language model as a generative language model as an example, we can obtain the audio file of the target dialogue text, then sample the audio file to obtain the corpus content, and then sample dialogue sentences and dialogue scene expansion prompts based on the corpus content, and input them together into the generative language model to expand the dialogue corpus as training samples to enhance the generalization ability of the model. That is, by using the generative language model to generate dialogue corpus to enhance the dialogue data, it is beneficial to improve the generalization ability of dialogue sentence vectors.

[0160] Of course, taking the second largest language model as a generative language model as an example, we can also obtain the audio file of the target dialogue text, then sample the audio file to obtain the corpus content, and then sample dialogue sentences and dialogue scenarios to expand the prompt words based on the corpus content. These are then input into the generative language model to expand the dialogue corpus as training samples to enhance the generalization ability of the model. That is, by using the generative language model to generate dialogue corpus to enhance the dialogue data, it is beneficial to improve the generalization ability of dialogue sentence vectors.

[0161] In some possible examples, the following steps may also be included:

[0162] The target dialogue text sentence vector representation is input into a preset topic model to obtain the target topic classification label.

[0163] In this embodiment of the application, the preset topic model can be preset or defaulted to by the system. For example, the preset topic model may include at least one of the following: large language model, neural network model, etc., which are not limited here.

[0164] In practice, since the target dialogue text sentence vector reflects the characteristics of the topic to a certain extent, the target dialogue text sentence vector representation can be input into the preset topic model to obtain the target topic classification label, and the entire dialogue can be represented by vectors. Multiple types of dialogue vector representation methods are provided for dialogue topic classification, which solves the problem of needing to classify conversation topics in practical applications. In this way, topic classification can be achieved quickly and accurately.

[0165] For example, in this embodiment of the application, let symbol D represent a dialogue. Part of D comes from the dialogue recording translated by ASR (ASR dialogue), and the other part is generated by sampling the corpus in the ASR dialogue and using a generative large language model such as GPT to generate a portion of the corpus to expand the dialogue training corpus (GPT dialogue) and increase the generalization ability of the model.

[0166] Where D = {S1, S2, ..., S} k Let S = {u1, u2, ..., u}, where S represents a sentence and k represents the number of sentences in dialogue D. n} represents a sentence, u represents a character or word, sentence S contains n characters or n words, and M represents the basic pre-trained large language model used, such as BERT, t5, Roberta, etc. By inputting the relevant information of the word-token level and sentence-token level of the dialogue scenario into the pre-trained basic large model, the basic vectors of word-token level and sentence-token level can be obtained.

[0167] Compared to Dial2vec, which inputs the entire dialogue as a single sample into model M, limiting the total dialogue length to 512 characters, this approach inputs all sentences from the dialogue in parallel. After obtaining the vector for each sentence, it uses the dialogue context, character interaction information, and other dialogue vectors for training. If each sentence is limited to 128 characters and the number of dialogue rounds is also limited to 128, then it can train on texts with a dialogue length of 128 * 128 = 16384 characters, solving the 512-character dialogue length problem of Dial2vec. Furthermore, since the time complexity of model M is O(nd^2) (where 0 represents complexity, n represents sentence length, and d represents the length of the hidden layer), the input length in this embodiment is only 1 / 4 that of Dial2vec, thus reducing the training time complexity. Moreover, it captures semantic interaction patterns between speakers, including their own contextual semantic interactions and the semantic interactions between speaking characters, resulting in a more comprehensive semantic learning of the dialogue text.

[0168] M is a pre-trained basic model used in the algorithm. The pre-trained model is trained based on a large corpus and can be used for ordinary text vectorization. In the dialogue scenario, the role of the basic pre-trained model is to obtain the initial sentence vectors and then learn and improve the vectorization according to the dialogue scenario.

[0169] Furthermore, let h = M(S), h ∈ R n*d Where h is the hidden layer output of model M, d is the hidden layer length of model M, and n is the sentence length. The following introduction will use BERT as the basic model.

[0170] After inputting sentences in parallel, each sentence can be obtained. in, Let be the vector of a word in the sentence. The vector of the special character cls is often used directly as a sentence vector, but the effect is not good.

[0171] Let H 1 ∈R n*dThe algorithm's sentence vector matrix, representing the first sentence of the dialogue, is used to obtain H. 1 You can use text-CNN for convolutional pooling, or directly extract the average value (avg) to obtain the vector for each sentence. The vector corresponding to all sentences in the dialogue is H∈R. k*d The sentence vector is obtained as a word-token layer.

[0172] After obtaining the vector of each sentence in the dialogue, it is used as input to the sentence-level model embedding. For role embedding, r i ∈R 1*d , i∈[1,k], turn embedding, t i ∈R 1 *d , i∈[1,k]. Compared to Dial2vec, sentence position embedding is no longer performed because word position encoding has already been input during word-token processing. For sentence-token processing, the tum embedding already reflects the sentence's position information in the dialogue. Therefore, when learning the contextual semantics of sentences in the dialogue, the input embedding consists of sentence embedding, turn embedding, and role embedding. After adding these three types of embeddings together, the sentence vector that integrates dialogue role and turn information is h. i +r i +t i h i +r i +t i ∈R 1*d Similarly, input model After (the second largest language model), for all sentences, we can obtain... To obtain the sentence vector matrices corresponding to different roles, a mask matrix m is generated for each of the two interlocutors. r1 and m r2 , m r1 The i-th element in the array has a value of 1 when the input token comes from r1, and a value of 0 otherwise.

[0173] The sentence vector matrix for different roles can be obtained by calculating using the following formula.

[0174] H r1 =He ⊙(m r1 ) T

[0175] H r2 =H e ⊙(m r2 ) T

[0176] Here, ⊙ denotes element-wise multiplication of matrices. If two characters in a dialogue, r1 speaks r1 sentences and r2 speaks r2 sentences, then the sentence vectors H for each character are obtained. r1 ∈R r1*d H r2 ∈R r2*d H r1 H represents the sentence vector matrix of character r1. r2 The sentence vector matrix representing character r2, m r1 This represents the mask matrix for character r1, (m r1 ) T m represents the transpose of the mask matrix of character r1. r2 This represents the mask matrix for character r2, (m r2 ) T H represents the transpose of the mask matrix of character r2. e ∈R k*d For each sentence in the dialogue, an initialization vector (i.e., H) is learned at the sentence-token level semantic level. e (Corresponding to the vector matrix in the second sentence above).

[0177] For character r1, the interaction matrix can be obtained:

[0178] C r1 =H r1 (H r1 ) T

[0179] C r1′ =H r1 (H r2 ) T

[0180] Among them, C r1 ∈R r1*r1 C r1′ =R r 1 *r2 .

[0181] Similarly, for character r2, the interaction matrix can be obtained:

[0182] C r2 =H r2 (H r2 )T

[0183] C r2′ =H r2 (H r1 ) T

[0184] Among them, C r2 ∈R r2*r2 C r2′ =R r2*r1 .

[0185] Then, the two roles each obtain two different sentence vector representations as follows:

[0186] E r1 =C r1 H r1

[0187] E r1′ =C r1′ H r2

[0188] Among them, E r1 ∈R r1*n E r1′ ∈R r1*n .

[0189] For character r2:

[0190] E r2 =C r2 H r2

[0191] E r2′ =C r2′ H r1

[0192] Among them, E r2 ∈R r2*n E r2′ ∈R r2*n .

[0193] For each role, a vector E is obtained that combines self-interaction information. r1 E r2 And a representation vector E that combines information about the interactions between different roles. r1′ E r2′ .

[0194] The word-token-level sentence vector H∈R learned from word-token-level semantics is fused together. k*n The final sentence vector can be expressed as:

[0195] f i r1 =H t ||E i

[0196]

[0197] Where || denotes vector concatenation, θ is a parameter used to control the proportion of self-interaction and inter-role interaction in the final dialogue text sentence vector information, and f i r1 ∈R 1*2d , i∈[1, r1]. By controlling the parameter θ, E i It can combine semantic information from self-interaction and interaction between roles simultaneously, or it can utilize these two types of information separately, such as E. i =E r1′ (θ=0), or E i =E r1 (θ=1), similarly, as follows:

[0198] f i r2 =H i ||E i

[0199]

[0200] Furthermore, we can obtain the vector f. i r2 ∈R 2*2d , i∈[1,r2].

[0201] Finally, by constructing positive and negative sample pairs, with sentences from the original dialogue as positive samples and randomly sampled sentences from the dialogue as negative samples, and using NT-Xent loSs as the loss function L, a large language model is trained. The specific loss function L is as follows:

[0202]

[0203] Where N is the total number of positive samples, and M is the total number of training sample pairs (f) corresponding to a given positive sample. j f j ′), j∈[1,M], (f i f i ′) represents a pair of samples consisting of a positive sample and a semantically similar sample. The neural network model is trained by minimizing the loss function.

[0204] Furthermore, after obtaining the sentence vector representations in the dialogue, various tasks such as clustering, similarity calculation, intent recognition, sentiment analysis, and vector retrieval can be performed.

[0205] Furthermore, in practical applications, to perform topic labeling for automatic dialogue, the vectors learned during training can be used to represent the entire dialogue segment by fusing all sentence vectors from the sentence-attention fusion process, thus obtaining the vector representation df of the entire dialogue for character r1. r1 ∈R 1*2d Vectorization df of the entire dialogue for character r2 r2 ∈R 1*2d Dialogue vector df∈R 1*2d The calculation method is as follows:

[0206]

[0207] Among them, the feedforward neural network (FFN) is a two-layer feedforward network with the activation function ReLU, γ i γ is the weight corresponding to each sentence vector in the fusion dialogue, where i represents any sentence in the entire dialogue. i That is, the weight of the i-th sentence in the dialogue text, FFN(f i Let r1 represent the number of times character r1 speaks and r2 represent the number of times character r2 speaks. By deeply integrating the dialogue scenario and considering the context, the relationship between different characters' dialogues can be enhanced. The weight of each sentence is deeply evaluated, and the vector representation is smoother. It runs through the entire dialogue. The characters represent identities, and different identities correspond to different social statuses. It deeply explores the emotions and status between characters. Different rounds correspond to the switching of different sub-topics of the dialogue. That is, it can explore various dimensions of information related to the dialogue theme from different dimensions, so that the final dialogue text sentence vector representation is more in line with the actual application scenario of the theme. It also covers the personality and commonality of the characters, as well as the differences between different statuses and identities, making the vector representation more down-to-earth and helping to improve the accuracy of subsequent theme classification.

[0208] Correspondingly, df can be obtained r1 df r2 The details are as follows:

[0209]

[0210] Where i represents any sentence spoken by character r1, and γ i FFN(f) represents the weight of the i-th sentence in the partial dialogue text corresponding to character r1. i) represents the result of the vector representation of the i-th sentence, and r1 represents the number of times character r1 speaks. Deep integration with the dialogue scenario and consideration of context can enhance the relationship between different characters' dialogues, deeply evaluate the weight of each sentence of a specified character, and make the vector representation smoother. It runs through the entire dialogue, which helps to deeply explore the characteristics of the characters and improve the accuracy of subsequent topic classification.

[0211] Then, correspondingly, df can be obtained. r2 The details are as follows:

[0212]

[0213] Where i represents any sentence spoken by character r1, and γ i FFN(f) represents the weight of the i-th sentence in the partial dialogue text corresponding to character r2. i ) represents the result of the vector representation of the i-th sentence, and r2 represents the number of times character r2 speaks. By deeply integrating the dialogue scene and considering the context, the relationship between different characters' dialogues can be enhanced. The weight of each sentence of a specified character can be deeply evaluated, and the vector representation is also smoother. It runs through the entire dialogue, which helps to deeply explore the characteristics of the characters and improve the accuracy of subsequent topic classification.

[0214] Next, we can use these three different types of dialogue vectors to label the entire dialogue as a topic, classify the dialogue as a topic, and input the vectors transformed from the entire dialogue into the topic model to obtain the corresponding topic classification labels.

[0215] Let me give another example, such as Figure 1B As shown, a recording is obtained, and ASR is performed to obtain the translated dialogue text. Then, dialogue text is generated based on GPT. The translated dialogue text and the generated dialogue text can be input into the sentence input model M to obtain word vectors. On the one hand, sentence vectors are obtained based on word vectors. On the other hand, the initial dialogue sentence vectors are obtained and dialogue information embedding (sentence embedding, turn embedding, role embedding) is added. Then, role embedding (self-interaction) and role embedding (interaction between roles) are performed. The sentence vectors and interaction results are concatenated and the concatenated result is input into FFN for training to improve the model's capabilities.

[0216] The dialogue text vector representation method described in this application involves obtaining a target dialogue text, which includes at least two roles and at least one round of dialogue. The target dialogue text is then segmented into sentences to obtain multiple sentences. These sentences are input into a first language model to obtain multiple sentence vectors. Sentence embedding, role embedding, and round-of-dialogue embedding operations are performed on these sentence vectors to obtain multiple sentence-embedded sentence vectors, multiple role-embedded sentence vectors, and multiple round-of-dialogue embedding sentence vectors. These vectors are then summed to obtain a first sentence vector matrix. This first sentence vector matrix is ​​input into a second language model to obtain a second sentence vector matrix. A mask matrix is ​​obtained for each of the at least two roles, resulting in at least one mask matrix. Based on the second sentence vector matrix and... At least one mask matrix determines the sentence vector matrix of at least one character. Based on the sentence vector matrix of at least one character, the sentence vector representation of the target dialogue text is determined. In this way, considering the context can enhance the relationship between different characters' dialogues and guide the sub-levels between dialogue rounds, so that each round of dialogue can be smoothly connected, and the vector representation is also smoother. Characters represent identities, and different identities correspond to different social statuses. Deeply exploring the emotions and status between characters, different rounds correspond to the switching of different dialogue sub-topics, that is, information related to various dimensions of the dialogue topic can be mined from different dimensions, so that the final sentence vector representation of the dialogue text is more in line with the actual application scenario of the topic, and also covers the personality and commonality of characters, as well as the differences between different statuses and identities, making the vector representation more grounded and helping to improve the accuracy of subsequent topic classification.

[0217] For those consistent with the above, please refer to Figure 2 , Figure 2 This is a flowchart illustrating another dialog text vector representation method provided in this application embodiment. As shown in the figure, this dialog text vector representation method includes:

[0218] 201. Obtain the target dialogue text, which includes at least two characters and at least one round of dialogue.

[0219] 202. The target dialogue text is segmented into sentences to obtain multiple sentences.

[0220] 203. Input the multiple sentences into the first large language model to obtain multiple sentence vectors.

[0221] 204. Perform sentence embedding, role embedding, and dialogue turn embedding operations on the multiple sentence vectors to obtain multiple sentence embedding sentence vectors, multiple role embedding sentence vectors, and multiple dialogue turn embedding sentence vectors.

[0222] 205. Add the multiple sentence embedding vectors, the multiple role embedding vectors, and the multiple dialogue round embedding vectors together to obtain the first sentence vector matrix.

[0223] 206. Input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix.

[0224] 207. Obtain the mask matrix of each of the at least two roles to obtain at least one mask matrix.

[0225] 208. Determine the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix.

[0226] 209. Determine the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one role.

[0227] 210. Input the target dialogue text sentence vector representation into a preset topic model to obtain the target topic classification label.

[0228] For a detailed description of steps 201-210 above, please refer to [link to relevant documentation]. Figure 1A The steps involved in the described method for representing dialogue text vectors will not be elaborated upon here.

[0229] The dialogue text vector representation method described in this application involves: acquiring a target dialogue text, which includes at least two roles and at least one round of dialogue; processing the target dialogue text into sentences to obtain multiple sentences; inputting these multiple sentences into a first language model to obtain multiple sentence vectors; performing sentence embedding, role embedding, and round-of-dialogue embedding operations on these multiple sentence vectors to obtain multiple sentence-embedded sentence vectors, multiple role-embedded sentence vectors, and multiple round-of-dialogue embedding sentence vectors; summing these multiple sentence-embedded sentence vectors to obtain a first sentence vector matrix; inputting this first sentence vector matrix into a second language model to obtain a second sentence vector matrix; obtaining a mask matrix for each of the at least two roles to obtain at least one mask matrix; and determining at least one role based on the second sentence vector matrix and the at least one mask matrix. The sentence vector matrix is ​​used to determine the sentence vector representation of the target dialogue text based on the sentence vector matrix of at least one character. The sentence vector representation of the target dialogue text is then input into a preset topic model to obtain the target topic classification label. In this way, considering the context can enhance the relationship between different characters' dialogues and guide the sub-levels between dialogue rounds, so that each round of dialogue can be smoothly connected, and the vector representation is also smoother. Characters represent identities, and different identities correspond to different social statuses. The deep exploration of the emotions and status between characters, and the switching of different sub-topics of dialogues in different rounds, can mine various dimensions of information related to the dialogue topic from different dimensions, so that the final sentence vector representation of the dialogue text is more in line with the actual application scenario of the topic, and also covers the personality and commonality of characters, as well as the differences between different statuses and identities, making the vector representation more down-to-earth and helping to improve the accuracy of subsequent topic classification.

[0230] Consistent with the above embodiments, please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs. The one or more programs are stored in the memory and configured to be executed by the processor. In this embodiment, the program includes instructions for performing the following steps:

[0231] Obtain the target dialogue text, which includes at least two characters and at least one round of dialogue.

[0232] The target dialogue text is segmented into sentences to obtain multiple sentences;

[0233] The multiple sentences are input into the first large language model to obtain multiple sentence vectors;

[0234] Perform sentence embedding, role embedding, and dialogue turn embedding operations on the multiple sentence vectors to obtain multiple sentence embedding sentence vectors, multiple role embedding sentence vectors, and multiple dialogue turn embedding sentence vectors;

[0235] The first sentence vector matrix is ​​obtained by adding the sentence vectors embedded in the multiple sentences, the sentence vectors embedded in the multiple roles, and the sentence vectors embedded in the multiple dialogue rounds.

[0236] Input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix;

[0237] Obtain the mask matrix for each of the at least two roles to obtain at least one mask matrix;

[0238] The sentence vector matrix of at least one character is determined based on the second sentence vector matrix and the at least one mask matrix;

[0239] The sentence vector representation of the target dialogue text is determined based on the sentence vector matrix of the at least one character.

[0240] In some possible examples, in determining the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one character, the above procedure includes instructions for performing the following steps:

[0241] The self-interaction sentence vector matrix and the role interaction sentence vector matrix of each role are determined based on the sentence vector matrix of the at least one role.

[0242] The first sentence vector representation of each of the at least one characters is determined based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of the at least one character.

[0243] The second sentence vector representation of each of the at least one characters is determined based on the first sentence vector representation of each character and the first sentence vector matrix;

[0244] The target dialogue text sentence vector representation is determined based on the second sentence vector representation of each of the at least one role.

[0245] In some possible examples, in determining the self-interaction sentence vector matrix and the role interaction sentence vector matrix for each role based on the sentence vector matrix of the at least one role, the above procedure includes instructions for performing the following steps:

[0246] Obtain the first transpose of the sentence vector matrix of the first character, where the first character is any one of the at least one characters;

[0247] The self-interaction sentence vector matrix of the first character is determined based on the sentence vector matrix of the first character and the first transpose matrix;

[0248] Determine the second transpose of the sentence vector matrix of the second role, wherein the second role is any role other than the first role among the at least one roles;

[0249] The character interaction sentence vector matrix of the first character is determined based on the sentence vector matrix of the first character and the second transpose matrix.

[0250] In some possible examples, in determining the first sentence vector representation of each of the at least one character based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of the at least one character, the above procedure includes instructions for performing the following steps:

[0251] The self-interaction sentence vector representation of the second role is determined based on the self-interaction sentence vector matrix of the second role and the sentence vector matrix of the second role;

[0252] The character sentence vector representation of the second character is determined based on the character interaction sentence vector matrix of the second character and the sentence vector matrix of the third character; the third character is any character other than the second character among the at least one character.

[0253] The first sentence vector representation of the second role is determined based on the self-interactive sentence vector representation and the role sentence vector representation.

[0254] In some possible examples, regarding the determination of a second sentence vector representation for each of the at least one characters based on a first sentence vector representation for each character and the first sentence vector matrix, the above procedure includes instructions for performing the following steps:

[0255] The intermediate sentence vector representation is obtained by performing a weighted operation on the self-interaction sentence vector representation and the role sentence vector representation.

[0256] By concatenating the first sentence vector matrix and the middle sentence vector representation, the second sentence vector representation of the second character is obtained.

[0257] In some possible examples, regarding the determination of the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix, the above procedure includes instructions for performing the following steps:

[0258] Determine the third transpose of the first mask matrix corresponding to the first role; the first mask matrix is ​​any one of the at least one mask matrices; the first role is one of the at least one roles;

[0259] The sentence vector matrix of the first character is obtained by performing a dot product operation on the second sentence vector matrix and the third transpose matrix.

[0260] In some possible examples, the above procedure also includes instructions for performing the following steps:

[0261] The target dialogue text is used as a positive sample.

[0262] Negative samples are obtained by randomly sampling and replacing sentences in the target dialogue text.

[0263] The first large language model and / or the second large language model are trained based on the positive samples, the negative samples, and a preset loss function.

[0264] In some possible examples, the above procedure also includes instructions for performing the following steps:

[0265] Obtain the audio file of the target dialogue text;

[0266] The audio files are sampled to obtain the corpus content;

[0267] Based on the corpus content, sample dialogue sentences and dialogue scene expansion prompts, input the dialogue sentences and dialogue scene expansion prompts into the first large language model to expand the dialogue corpus as training samples, so as to enhance the generalization ability of the first large language model.

[0268] In some possible examples, the above procedure also includes instructions for performing the following steps:

[0269] The target dialogue text sentence vector representation is input into a preset topic model to obtain the target topic classification label.

[0270] The electronic device described in this application acquires target dialogue text, which includes at least two roles and at least one round of dialogue. It then segments the target dialogue text into sentences to obtain multiple sentences, inputs these sentences into a first language model to obtain multiple sentence vectors, performs sentence embedding, role embedding, and dialogue round embedding operations on these sentence vectors to obtain multiple sentence-embedded sentence vectors, multiple role-embedded sentence vectors, and multiple dialogue round-embedded sentence vectors, adds these vectors together to obtain a first sentence vector matrix, inputs this first sentence vector matrix into a second language model to obtain a second sentence vector matrix, acquires a mask matrix for each of the at least two roles to obtain at least one mask matrix, and then uses the second sentence vector matrix and at least one... Each mask matrix determines the sentence vector matrix of at least one character. Based on the sentence vector matrix of at least one character, the sentence vector representation of the target dialogue text is determined. In this way, considering the context can enhance the connection between different characters' dialogues and guide the sub-levels between dialogue rounds, so that each round of dialogue can be smoothly connected, and the vector representation is also smoother. Characters represent identities, and different identities correspond to different social statuses. By deeply exploring the emotions and status between characters, and the switching of different dialogue sub-topics in different rounds, information related to the dialogue topic can be mined from different dimensions. This makes the final sentence vector representation of the dialogue text more in line with the actual application scenario of the topic, and also covers the personality and commonality of characters, as well as the differences between different statuses and identities, making the vector representation more grounded and helping to improve the accuracy of subsequent topic classification.

[0271] Figure 4 This is a functional unit block diagram of a dialogue text vector representation device 400 according to an embodiment of this application. The dialogue text vector representation device 400 includes: an acquisition unit 401, a sentence segmentation unit 402, an input unit 403, an embedding unit 404, and a determination unit 405, wherein...

[0272] The acquisition unit 401 is used to acquire target dialogue text, the target dialogue text includes at least two characters, and the target dialogue text includes at least one round of dialogue.

[0273] The sentence segmentation unit 402 is used to segment the target dialogue text into multiple sentences.

[0274] The input unit 403 is used to input the multiple sentences into the first large language model to obtain multiple sentence vectors;

[0275] The embedding unit 404 is used to perform sentence embedding operation, role embedding operation and dialogue turn embedding operation on the multiple sentence vectors to obtain multiple sentence embedded sentence vectors, multiple role embedded sentence vectors and multiple dialogue turn embedded sentence vectors; and to add the multiple sentence embedded sentence vectors, the multiple role embedded sentence vectors and the multiple dialogue turn embedded sentence vectors to obtain a first sentence vector matrix;

[0276] The input unit 403 is also used to input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix;

[0277] The acquisition unit 401 is also used to acquire the mask matrix of each of the at least two roles, to obtain at least one mask matrix;

[0278] The determining unit 405 is used to determine the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix; and to determine the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one character.

[0279] In some possible examples, in determining the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one character, the determining unit 405 is specifically used for:

[0280] The self-interaction sentence vector matrix and the role interaction sentence vector matrix of each role are determined based on the sentence vector matrix of the at least one role.

[0281] The first sentence vector representation of each of the at least one characters is determined based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of the at least one character.

[0282] The second sentence vector representation of each of the at least one characters is determined based on the first sentence vector representation of each character and the first sentence vector matrix;

[0283] The target dialogue text sentence vector representation is determined based on the second sentence vector representation of each of the at least one role.

[0284] In some possible examples, in determining the self-interaction sentence vector matrix and the character interaction sentence vector matrix for each character based on the sentence vector matrix of the at least one character, the determining unit 405 is specifically used for:

[0285] Obtain the first transpose of the sentence vector matrix of the first character, where the first character is any one of the at least one characters;

[0286] The self-interaction sentence vector matrix of the first character is determined based on the sentence vector matrix of the first character and the first transpose matrix;

[0287] Determine the second transpose of the sentence vector matrix of the second role, wherein the second role is any role other than the first role among the at least one roles;

[0288] The character interaction sentence vector matrix of the first character is determined based on the sentence vector matrix of the first character and the second transpose matrix.

[0289] In some possible examples, in determining the first sentence vector representation of each of the at least one character based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character and the sentence vector matrix of the at least one character, the determining unit 405 is specifically used for:

[0290] The self-interaction sentence vector representation of the second role is determined based on the self-interaction sentence vector matrix of the second role and the sentence vector matrix of the second role;

[0291] The character sentence vector representation of the second character is determined based on the character interaction sentence vector matrix of the second character and the sentence vector matrix of the third character; the third character is any character other than the second character among the at least one character.

[0292] The first sentence vector representation of the second role is determined based on the self-interactive sentence vector representation and the role sentence vector representation.

[0293] In some possible examples, in determining the second sentence vector representation of each of the at least one characters based on the first sentence vector representation of each character and the first sentence vector matrix, the determining unit 405 is specifically used for:

[0294] The intermediate sentence vector representation is obtained by performing a weighted operation on the self-interaction sentence vector representation and the role sentence vector representation.

[0295] By concatenating the first sentence vector matrix and the middle sentence vector representation, the second sentence vector representation of the second character is obtained.

[0296] In some possible examples, in determining the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix, the determining unit 405 is specifically used for:

[0297] Determine the third transpose of the first mask matrix corresponding to the first role; the first mask matrix is ​​any one of the at least one mask matrices; the first role is one of the at least one roles;

[0298] The sentence vector matrix of the first character is obtained by performing a dot product operation on the second sentence vector matrix and the third transpose matrix.

[0299] In some possible examples, the dialogue text vector representation device 400 is also specifically used for:

[0300] The target dialogue text is used as a positive sample.

[0301] Negative samples are obtained by randomly sampling and replacing sentences in the target dialogue text.

[0302] The first large language model and / or the second large language model are trained based on the positive samples, the negative samples, and a preset loss function.

[0303] In some possible examples, the dialogue text vector representation device 400 is also specifically used for:

[0304] Obtain the audio file of the target dialogue text;

[0305] The audio files are sampled to obtain the corpus content;

[0306] Based on the corpus content, sample dialogue sentences and dialogue scene expansion prompts, input the dialogue sentences and dialogue scene expansion prompts into the first large language model to expand the dialogue corpus as training samples, so as to enhance the generalization ability of the first large language model.

[0307] In some possible examples, the dialogue text vector representation device 400 is also specifically used for:

[0308] The target dialogue text sentence vector representation is input into a preset topic model to obtain the target topic classification label.

[0309] The dialogue text vector representation apparatus described in this application acquires a target dialogue text, which includes at least two roles and at least one round of dialogue. It then segments the target dialogue text into sentences to obtain multiple sentences, inputs these sentences into a first language model to obtain multiple sentence vectors, performs sentence embedding, role embedding, and round-of-dialogue embedding operations on these sentence vectors to obtain multiple sentence-embedded sentence vectors, multiple role-embedded sentence vectors, and multiple round-of-dialogue embedding sentence vectors, adds these vectors together to obtain a first sentence vector matrix, inputs this first sentence vector matrix into a second language model to obtain a second sentence vector matrix, acquires a mask matrix for each of the at least two roles to obtain at least one mask matrix, and then applies the second sentence vector matrix and... At least one mask matrix determines the sentence vector matrix of at least one character. Based on the sentence vector matrix of at least one character, the sentence vector representation of the target dialogue text is determined. In this way, considering the context can enhance the relationship between different characters' dialogues and guide the sub-levels between dialogue rounds, so that each round of dialogue can be smoothly connected, and the vector representation is also smoother. Characters represent identities, and different identities correspond to different social statuses. Deeply exploring the emotions and status between characters, different rounds correspond to the switching of different dialogue sub-topics, that is, information related to various dimensions of the dialogue topic can be mined from different dimensions, so that the final sentence vector representation of the dialogue text is more in line with the actual application scenario of the topic, and also covers the personality and commonality of characters, as well as the differences between different statuses and identities, making the vector representation more grounded and helping to improve the accuracy of subsequent topic classification.

[0310] It is understood that the functions of each program module of the dialogue text vector representation device in this embodiment can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can be referred to the relevant descriptions in the above method embodiments, and will not be repeated here.

[0311] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.

[0312] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.

[0313] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0314] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0315] In the embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical or other forms.

[0316] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0317] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0318] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0319] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0320] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for representing dialogue text vectors, characterized in that, The method includes: Obtain the target dialogue text, which includes at least two characters and at least one round of dialogue. The target dialogue text is segmented into sentences to obtain multiple sentences; The multiple sentences are input into the first large language model to obtain multiple sentence vectors; Perform sentence embedding, role embedding, and dialogue turn embedding operations on the multiple sentence vectors to obtain multiple sentence embedding sentence vectors, multiple role embedding sentence vectors, and multiple dialogue turn embedding sentence vectors; The first sentence vector matrix is ​​obtained by adding the sentence vectors embedded in the multiple sentences, the sentence vectors embedded in the multiple roles, and the sentence vectors embedded in the multiple dialogue rounds. Input the first sentence vector matrix into the second language model to obtain the second sentence vector matrix; Obtain the mask matrix for each of the at least two roles to obtain at least one mask matrix; The sentence vector matrix of at least one character is determined based on the second sentence vector matrix and the at least one mask matrix; The sentence vector representation of the target dialogue text is determined based on the sentence vector matrix of the at least one character; Determining the sentence vector representation of the target dialogue text based on the sentence vector matrix of the at least one character includes: The self-interaction sentence vector matrix and the role interaction sentence vector matrix of each role are determined based on the sentence vector matrix of the at least one role. The first sentence vector representation of each of the at least one characters is determined based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of the at least one character. The second sentence vector representation of each of the at least one characters is determined based on the first sentence vector representation of each character and the first sentence vector matrix; The target dialogue text sentence vector representation is determined based on the second sentence vector representation of each of the at least one role.

2. The method according to claim 1, characterized in that, The step of determining the self-interaction sentence vector matrix and the role interaction sentence vector matrix for each role based on the sentence vector matrix of the at least one role includes: Obtain the first transpose of the sentence vector matrix of the first character, where the first character is any one of the at least one characters; The self-interaction sentence vector matrix of the first character is determined based on the sentence vector matrix of the first character and the first transpose matrix; Determine the second transpose of the sentence vector matrix of the second role, wherein the second role is any role other than the first role among the at least one roles; The character interaction sentence vector matrix of the first character is determined based on the sentence vector matrix of the first character and the second transpose matrix.

3. The method according to claim 2, characterized in that, The step of determining the first sentence vector representation of each of the at least one character based on the self-interaction sentence vector matrix and the character interaction sentence vector matrix of each character, as well as the sentence vector matrix of the at least one character, includes: The self-interaction sentence vector representation of the second role is determined based on the self-interaction sentence vector matrix of the second role and the sentence vector matrix of the second role; The character sentence vector representation of the second character is determined based on the character interaction sentence vector matrix of the second character and the sentence vector matrix of the third character; the third character is any character other than the second character among the at least one character. The first sentence vector representation of the second role is determined based on the self-interactive sentence vector representation and the role sentence vector representation.

4. The method according to claim 3, characterized in that, The step of determining the second sentence vector representation of each of the at least one characters based on the first sentence vector representation of each character and the first sentence vector matrix includes: The intermediate sentence vector representation is obtained by performing a weighted operation on the self-interaction sentence vector representation and the role sentence vector representation. By concatenating the first sentence vector matrix and the middle sentence vector representation, the second sentence vector representation of the second character is obtained.

5. The method according to any one of claims 1-4, characterized in that, The step of determining the sentence vector matrix of at least one character based on the second sentence vector matrix and the at least one mask matrix includes: Determine the third transpose of the first mask matrix corresponding to the first role; the first mask matrix is ​​any one of the at least one mask matrices; the first role is one of the at least one roles; The sentence vector matrix of the first character is obtained by performing a dot product operation on the second sentence vector matrix and the third transpose matrix.

6. The method according to claim 5, characterized in that, The method further includes: The target dialogue text is used as a positive sample. Negative samples are obtained by randomly sampling and replacing sentences in the target dialogue text. The first large language model and / or the second large language model are trained based on the positive samples, the negative samples, and a preset loss function.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtain the audio file of the target dialogue text; The audio files are sampled to obtain the corpus content; Based on the corpus content, sample dialogue sentences and dialogue scene expansion prompts, input the dialogue sentences and dialogue scene expansion prompts into the first large language model to expand the dialogue corpus as training samples, so as to enhance the generalization ability of the first large language model.

8. The method according to any one of claims 1-6, characterized in that, The method further includes: The target dialogue text sentence vector representation is input into a preset topic model to obtain the target topic classification label.

9. A computer-readable storage medium, characterized in that, A computer program for storing electronic data interchange is provided, wherein the computer program causes a computer to perform the method as described in any one of claims 1-8.