Information Processing Method, Apparatus, Computer Device, and Storage Medium

By obtaining the content information of multimedia data and the global description information of target comment information, identifying the reply type and obtaining matching reply strategies, and generating reply information using deep learning models, solving the problem of lack of flexibility and accuracy of reply information in the prior art, and achieving higher accuracy and diversity of reply information generation.

CN114329005BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111133730.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2025-07-18
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

The existing multimedia data comment information reply method is generated based on the reply template, resulting in the lack of flexibility and low accuracy of reply information.

Method used

By obtaining the content information of multimedia data and the global description information of the target comment information, identifying the reply type and obtaining matching reply strategies, differentiated reply strategies are used to generate reply information, and deep learning models such as comment reply models and generation models are used to generate reply information.

Benefits of technology

It improves the accuracy and diversity of reply information, enhances the attractiveness of interaction between multimedia data and users, and improves user stickiness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329005B_ABST
    Figure CN114329005B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an information processing method, apparatus, computer device, and storage medium. The method includes: obtaining content information of multimedia data and target comment information for the multimedia data; obtaining global description information, where the global description information is used to describe the information semantics of the content information and the target comment information; using the global description information to identify a reply type for the target comment information, and obtaining a reply strategy matching the reply type; generating a reply message for the target comment information according to the reply strategy and based on the content information of the multimedia data, and outputting the reply message, which can improve the accuracy of the generated reply message.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an information processing method, apparatus, computer device, and storage medium. Background Art

[0002] With the continuous in-depth development of computer technology, more and more multimedia data is being released. Then, in order to enhance the attractiveness of the released multimedia data to users, corresponding reply information can be generated for the comment information of the multimedia data, so as to form an interaction and achieve the goal of attracting users. Most of the existing ways to generate reply information for comment information are based on reply templates. However, the way of generating reply information based on reply templates makes the reply information lack flexibility, and the reply templates need to be updated regularly, resulting in relatively low accuracy of the generated reply information. Therefore, how to flexibly generate reply information with relatively high accuracy has become a current research hotspot. Summary of the Invention

[0003] Embodiments of the present invention provide an information processing method, apparatus, computer device, and storage medium, which can improve the accuracy of the generated reply information.

[0004] On the one hand, embodiments of the present invention provide an information processing method, including:

[0005] Obtain the content information of the multimedia data and the target comment information for the multimedia data;

[0006] Obtain global description information, where the global description information is used to describe the information semantics of the content information and the target comment information;

[0007] Use the global description information to identify the reply type for the target comment information, and obtain a reply strategy that matches the reply type;

[0008] Generate the reply information for the target comment information according to the reply strategy and based on the content information of the multimedia data, and output the reply information.

[0009] On the other hand, embodiments of the present invention provide an information processing apparatus, including:

[0010] An obtaining unit, configured to obtain the content information of the multimedia data and the target comment information for the multimedia data;

[0011] The obtaining unit is further configured to obtain global description information, where the global description information is used to describe the information semantics of the content information and the target comment information;

[0012] A processing unit, configured to identify a reply type for the target comment information by using the global description information, and obtain a reply strategy matching the reply type;

[0013] The processing unit is further configured to generate a reply message for the target comment information according to the reply strategy and based on the content information of the multimedia data, and output the reply message.

[0014] In another aspect, an embodiment of the present invention provides a computer device, including a processor, an input device, an output device, and a memory, where the processor, the input device, the output device, and the memory are interconnected. The memory is configured to store a computer program for supporting the computer device to execute the above method. The computer program includes program instructions, and the processor is configured to call the program instructions to perform the following steps:

[0015] Obtain the content information of the multimedia data and the target comment information for the multimedia data;

[0016] Obtain global description information, where the global description information is used to describe the information semantics of the content information and the target comment information;

[0017] Identify a reply type for the target comment information by using the global description information, and obtain a reply strategy matching the reply type;

[0018] Generate a reply message for the target comment information according to the reply strategy and based on the content information of the multimedia data, and output the reply message.

[0019] In another aspect, an embodiment of the present invention provides a computer-readable storage medium. Program instructions are stored in the computer-readable storage medium. When the program instructions are executed by a processor, the program instructions are used to execute the information processing method as described in the first aspect.

[0020] In the embodiments of the present application, when a computer device needs to generate a reply message for the comment information of multimedia data, it will obtain the content information of the multimedia data, the target comment information, and the global description information used to describe the information semantics of the content information and the target comment information. Further, the computer device can use the global description information to identify the reply type for the target comment information. Thus, the computer device can generate the reply type of the target comment information according to the reply strategy matching the reply type and the content information of the multimedia data, enabling the computer device to implement different reply strategies based on the differences in the information types of the comment information to generate reply messages for comment information of different information types. Moreover, based on the differences in the information types of the comment information, different reply strategies are used to generate reply messages, enhancing the diversity in the process of generating reply messages. In addition, since the content information of the multimedia data is fully considered in the process of generating reply messages, the rationality and accuracy of the reply messages generated by the computer device can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic diagram of a model call relationship provided by an embodiment of the present invention;

[0023] Figure 2 It is a schematic flowchart of an information processing method provided by an embodiment of the present invention;

[0024] Figure 3 It is a schematic flowchart of an information processing method provided by an embodiment of the present invention;

[0025] Figure 4a It is a schematic diagram of a comment reply model provided by an embodiment of the present invention;

[0026] Figure 4b It is a schematic diagram of a generation model provided by an embodiment of the present invention;

[0027] Figure 5 It is a schematic block diagram of an information processing device provided by an embodiment of the present invention;

[0028] Figure 6 It is a schematic block diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] An embodiment of the present application proposes an information processing method, enabling a computer device to generate a reply message for comment information (such as target comment information) in multimedia data. According to the different information types of the target comment information to be replied, different reply strategies are adopted to generate the reply message. That is to say, when the computer device generates a reply message for the target comment information, the generated reply message is closely related to the target comment information to be replied, thereby effectively improving the accuracy of the computer device in generating reply messages for comment information. Among them, the multimedia data includes data integrating various media forms such as text, sound, and image. The multimedia data can be, for example, video data or audio data. In the embodiment of the present application, the multimedia data is mainly taken as video data for illustration, and the same can be referred to for other multimedia data in the embodiment of the present application; in addition, the comment information includes subjective or objective elaboration information sent by users for the multimedia data. The comment information can be question information sent by users, or can also be descriptive information related (or unrelated) to the multimedia data, etc. It can be understood that the number of comment information for the multimedia data can be one or more, and the number of comment information that each user can send to the multimedia data can also be one or more. The target comment information proposed in the embodiment of the present application can be any one of one or more comment information received by the multimedia data. Among them, the comment information of the multimedia data can include one or more of text information, expression information, and audio-visual information. The target comment information mentioned in the embodiment of the present application is for the text information in the comment information. When the target comment information includes expression information and / or audio-visual information, the expression information and / or audio-visual information can be text-converted to obtain text comment information. Among them, the computer device can be implemented as various types of terminal devices such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (such as mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), in-vehicle terminals, etc. Or, the computer device can also be a server. In the embodiment of the present application, the computer device is mainly taken as a terminal device for exemplary illustration.

[0030] In one embodiment, based on the semantics expressed by the target comment information, the target comment information can be divided into different information types. Among them, if the semantics expressed by the target comment information is related to objective questions such as asking for the source or asking about a person, or is related to the content discussion of multimedia data, personal attitude comments, etc., then the computer device can determine that the information type of the target comment information is an objective fact type; if the semantics expressed by the target comment information is irrelevant to the above semantics, then the computer device can consider that the information type of the target comment information is a general type. Then, when the computer device determines that the information type of the target comment information is an objective fact type, the computer device will adopt a response strategy associated with the objective fact type to generate response information. The response strategy associated with the objective fact type is a strategy that uses open-domain technology and constructs response information based on the content information of multimedia data in combination with the target comment information. Among them, the open-domain technology is an intelligent question-and-answer technology based on deep learning. Then it can be understood that if the response information is constructed by the computer device using the response strategy associated with the objective fact type, the content of the multimedia data and the semantics of the target comment information will be fully considered, thereby improving the accuracy of the constructed response information. If the computer device determines that the information type of the target comment information is a general type, the computer device can generate corresponding response information for the target comment information by combining the plot type of the multimedia data and the content information of the multimedia data. Generating corresponding response information based on the comment information sent by the user helps to improve the user stickiness between the multimedia data and the user, thereby improving user satisfaction. In addition, since the computer device can adopt different response strategies to generate response information when generating corresponding response information for the target comment information, the diversity of the computer device in generating response information is also improved.

[0031] In one embodiment, different response strategies adopted by the computer device based on different information types of the target comment information to be replied are implemented by calling different models. When the computer device determines that the information type of the target comment information is an objective fact type, the computer device can call a comment reply model to generate response information. If the computer device determines that the information type of the target comment information is a general type, the computer device generates response information by calling a generation model. Among them, both the comment reply model and the generation model are network models obtained based on deep learning. The comment reply model and the generation model can adopt, for example Figure 1Connect in the manner shown. When the computer device determines that a comment reply needs to be made to the multimedia data, after determining the target comment information to be replied, it can first input the target comment information into the comment reply model. Then, based on the processing of the target comment information by the comment reply model, it can be determined whether the target comment information is of the objective fact type. In one embodiment, if, after the computer device inputs the target comment information into the comment reply model, a reply information for the target comment information is obtained, it indicates that the information type of the target comment information is of the objective fact type. Otherwise, the comment reply model can further input the target comment information into the generation model, so that the generation model can generate corresponding reply information for the target comment information. That is to say, when the computer device generates corresponding reply information for the target comment information, it processes the target comment information by calling the comment reply model to determine the information type of the target comment information. Then, based on the information type of the target comment information, it can be determined whether to directly call the present comment reply model to generate corresponding reply information for the target comment information, or to call the generation model to generate corresponding reply information for the target comment information.

[0032] Please refer to Figure 2 , which is a schematic flowchart of an information processing method proposed in an embodiment of the present application. The information processing method can be executed by the above computer device, such as Figure 2 shown, the method may include:

[0033] S201, obtain the content information of the multimedia data and the target comment information for the multimedia data.

[0034] S202, obtain the global description information, where the global description information is used to describe the information semantics of the content information and the target comment information.

[0035] In steps S201 and S202, the multimedia data includes videos, audios, etc. that are published to a certain platform (or application). After the multimedia data is published, in order to enhance the attraction of the multimedia data to users, or to enhance the user stickiness of the user on the publishing platform of the multimedia data, the computer device can, after publishing the multimedia data to the corresponding platform, generate corresponding reply information for the comment information for the multimedia data. Based on the computer device generating corresponding reply information for the comment information, it can guide the user who publishes the comment information to further discuss around the multimedia data. Based on this discussion, it can further attract other users, which helps to enhance the attraction of the multimedia data to users, and thus can enhance the user stickiness of the corresponding publishing platform of the multimedia data.

[0036] In one embodiment, if a computer device determines from the comment information included in multimedia data that a corresponding reply message needs to be generated for target comment information, the computer device may first obtain the content information of the multimedia data and the target comment information. Among them, the content information of the multimedia data includes the text information in the multimedia data. When the multimedia data is a video, the text information included in the content information of the video may include: text extracted from the image frames of the video based on Optical Character Recognition (OCR) technology, text converted from the speech of the video based on Automated Speech Recognition (ASR) technology. In addition, the text information may further include text information such as the title of the video, topic tags, creation type, operation type, existing comments, and replies. The target comment information may be any comment information submitted by any user. In the embodiments of the present application, the manner in which the computer device selects the target comment information from one or more comment information of the multimedia data is not limited. In addition, the content information of the multimedia data further includes the image information of the multimedia data. Among them, when the multimedia data is a video, the image information of the video may include some or all of the video frames in the video. In one embodiment, the text information may further include the creator information of the target comment information, such as creator identification (ID), profile, level, user portrait, and other information.

[0037] In one embodiment, based on the computer device's acquisition of the content information and target comment information of the multimedia data, the computer device will further obtain the global description information for the content information and the target comment information. Among them, the global description information is used to describe the content information and the target comment information. That is to say, the global description information includes the semantics included in the content information and the semantics of the target comment information. Then it can be understood that by obtaining the global description information, the computer device can fully consider the semantics of the target comment information and the semantics of the content information in the process of generating the reply message for the target comment information, so as to ensure the rationality and accuracy of the generated reply message. Among them, when the computer device obtains the global description information, it may first construct a first coding sequence based on the coding sequence of the text information included in the content information and the coding sequence of the target comment information, and then based on the word vectors of each token in the first coding sequence, and in combination with the self-attention mechanism, obtain the global description information. Among them, when the computer device obtains the global description information based on the word vectors of each token in the first coding sequence and the self-attention mechanism, it can be determined by calculating the word vectors of each token in the first coding sequence and the correlation relationship between the word vectors.

[0038] After the computer device obtains the content information of the multimedia data and the target comment information, and after obtaining the global description information, it can determine the reply type for the target comment information based on the global description information, and thus adopt the corresponding reply strategy to generate the reply information, that is, it turns to execute step S203.

[0039] S203. Use the global description information to identify the reply type for the target comment information, and obtain the reply strategy that matches the reply type.

[0040] S204. According to the reply strategy, and based on the content information of the multimedia data, generate the reply information for the target comment information, and output the reply information.

[0041] In step S203 and step S204, when the computer device identifies the reply type for the target comment information based on the global description information, it is determined by the information type of the target comment information determined by the computer device based on the global description information. Among them, the information type of the comment information includes objective fact type and general type. The comment information of the objective fact type refers to the comment information for which the corresponding answer can be found through the comment information and the multimedia data. The comment information of the objective fact type can be, for example, the inquiry information about the character included in the multimedia data, etc. The general comment information refers to the comment information for which no specific answer can be found. In the embodiments of the present application, any comment information other than the comment information of the objective fact type can be used as the general comment information. In one embodiment, since one information type is associated with one reply type, then, after the computer device determines the information type of the target comment information, it also obtains the reply type of the target comment information. Based on the determination of the reply type by the computer device, the reply strategy that matches the reply type can be obtained, and the reply information can be generated based on the obtained reply strategy.

[0042] In one embodiment, the process of the computer device obtaining the reply strategy that matches the reply type based on the determined reply type to generate the reply information is the process of calling different models and generating the reply information. Among them, if the reply type determined by the computer device is associated with the information type of the objective fact type, the computer device will generate the reply information by calling the comment reply model. If the reply type determined by the computer device is associated with the information type of the general type, the computer device generates the reply information by calling the generation model.

[0043] When a computer device generates a reply message by invoking a comment reply model, the comment reply model is generated based on a first encoding sequence obtained by encoding text information and target comment information included in content information. In a specific implementation, when the computer device invokes the comment reply model to generate a reply message for the target comment information based on the first encoding sequence, it will separately invoke two linear layers in the comment reply model to respectively identify each word vector in the first encoding sequence, so as to determine the probability that each word vector in the first encoding sequence is the starting position of the reply message, and the probability that each word vector in the first encoding sequence is the ending position of the reply message. Further, the computer device can select some or all of the encoding sequences from the first encoding sequence for decoding based on the probabilities corresponding to the starting position and the ending position of each word vector in the first encoding sequence, so as to obtain a reply message for the target comment information. If the computer device determines that it is necessary to invoke a generation model to generate a reply message, then when generating the reply message, the generation model is obtained by decoding a second encoding sequence generated from the text information and image information included in the content information. In one embodiment, when the computer device invokes the generation model and decodes the reply message based on the second encoding sequence, it will perform an attention operation according to the second encoding sequence to obtain the reply message. Since the attention operation is performed when invoking the generation model to generate the reply message, it also makes the computer device reasonable in generating the reply message according to the second encoding sequence.

[0044] In an embodiment of the present application, when the computer device needs to generate a reply message for the comment information of multimedia data, it will obtain the content information and target comment information of the multimedia data, as well as global description information used to describe the information semantics of the content information and target comment information. Further, the computer device can use the global description information to identify the reply type for the target comment information, so that the computer device can generate the reply type for the target comment information according to the reply strategy matching the reply type and the content information of the multimedia data, enabling the computer device to implement different reply strategies based on the differences in the information types of the comment information to generate reply messages for comment information of different information types. And based on the differences in the information types of the comment information, different reply strategies are used to generate reply messages, which improves the diversity in the process of generating reply messages. In addition, since the content information of the multimedia data is fully considered in the process of generating the reply message, the rationality and accuracy of the reply message generated by the computer device can be improved.

[0045] Please refer to Figure 3 , which is a schematic flowchart of an information processing method proposed in an embodiment of the present application. The information processing method can be executed by the above computer device, such asFigure 3 As shown, the method may include:

[0046] S301, obtaining the content information of the multimedia data and the target comment information for the multimedia data, where the content information includes the text information of the multimedia data.

[0047] The content information of the multimedia data obtained by the computer device at least includes the text information of the multimedia data. This text information includes the text extracted from the multimedia data, such as the ocr, asr, title, etc. of the multimedia data. In addition, the text information may also include the text related to the multimedia data, such as the topic tag, creation type, etc. of the multimedia data. In one embodiment, the content information may further include the image information of the multimedia data. The image information may include some or all of the image frames extracted from the multimedia data. The target comment information obtained by the computer device may be any one selected from all the comment information included in the multimedia data. Then, based on the computer device's acquisition of the content information and target comment information of the multimedia data, the computer device obtains the global description information for the text information and the target comment information. It should be noted that the text information included in the content information of the multimedia data obtained by the computer device also includes the comment information with existing replies in the multimedia data and the corresponding replies. Thus, when the computer device generates the reply information for the target comment information, it can be generated in combination with the comment information with existing replies, and further improve the rationality of the reply information generated by the computer device.

[0048] After the computer device obtains the content information and target comment information of the multimedia data, it can perform word segmentation processing on the text information in the content information and the target comment information respectively, and then obtain the global description information for the text information and the target comment information based on the semantics of each word segment.

[0049] S302, obtaining the first coding sequence, where the first coding sequence includes the word vectors of each word segment in the text information and the word vectors of each word segment in the target comment information.

[0050] S303, calculating the similarity between any word segment and any other word segment according to the word vector of any word segment and the word vector of any other word segment, and performing weighted processing on the word vector of each word segment based on the similarity.

[0051] S304, using the vector sequence composed of the weighted word vectors as the global description information for the text information and the target comment information.

[0052] In steps S302 to S304, when the computer device obtains the first coding sequence, after obtaining the text information of the multimedia data and the target comment information, it can perform word segmentation on the target comment information to obtain a word segmentation sequence corresponding to the target comment information, and perform word segmentation on the text information of the multimedia data (such as text information such as ocr, asr, title, topic tags, creation type, operation type, existing comments, and replies) to obtain a word segmentation sequence of the text information; furthermore, the computer device can splice the word segmentation sequence of the target comment information and the word segmentation sequence of the text information to obtain a target splicing sequence. Among them, after obtaining the target splicing sequence, the computer device will also add a separator between the word segmentation sequence of the target comment information and the word segmentation sequence of the text information, and add a start character at the starting position of the target splicing sequence. The separator can be [SEP], and the start character is [CLS]. Then, if the word segmentation sequence of the target comment information is sequence A (assumed to be X1X2X3), and the word segmentation sequence of the text information is sequence B (assumed to be Y1Y2Y3), then, after the computer device splices the word segmentation sequence of the target comment information and the word segmentation sequence of the text information, and adds a separator and a start character, the obtained target splicing sequence is CLSX1X2X3SEPY1Y2Y3.

[0053] In one embodiment, when the computer device splices the token sequence of the target comment information and the token sequence of the text information to obtain the target splicing sequence, the computer device can also obtain the sequence length of the token sequence of the target comment information and the sequence length of the token sequence of the text information. Here, the sequence length of the token sequence of the target comment information is the sequence length of the above A sequence X1X2X3, and the sequence length of the token sequence of the text information is the sequence length of the above B sequence Y1Y2Y3. After the computer device respectively obtains the sequence length of the token sequence of the target comment information and the sequence length of the token sequence of the text information, it can further determine the sum of the sequence lengths between the sequence length of the token sequence of the target comment information and the sequence length of the token sequence of the text information, which is the sum of the lengths between the above A sequence X1X2X3 and B sequence Y1Y2Y3. Then, when the sum of the lengths is less than or equal to the length threshold, the computer device can use the sequence directly spliced from the token sequence of the target comment information and the token sequence of the text information as the target splicing sequence; if the sum of the lengths is greater than the length threshold, the computer device can perform sequence segmentation on the token sequence of the text information based on the sequence length of the token sequence of the target comment information and the length threshold, and splice each segmented sequence with the token sequence of the target comment information respectively, and each obtained splicing sequence is the target splicing sequence. In one embodiment, the length threshold can be 512, for example. Then, as described above, if the computer device determines that the sum of the lengths between the A sequence X1X2X3 and the B sequence Y1Y2Y3 is less than or equal to 512, then the computer device determines that the obtained target splicing sequence is CLSX1X2X3SEP Y1Y2Y3; in another implementation, if the computer device determines that the sum of the lengths between the A sequence X1X2X3 and the B sequence Y1Y2Y3 is greater than 512, and if the computer device determines that the sequence length of the A sequence X1X2X3 is 300, then when the computer device performs sequence segmentation on the B sequence Y1Y2Y3, it can segment the B sequence Y1Y2Y3 according to the subsequence length of 212 based on the length threshold 512. If the B sequence Y1Y2Y3 is segmented into the sequence Y1Y2 (whose corresponding sequence length is 212) and the sequence Y3 (whose corresponding sequence length is 106), then the target splicing sequences finally obtained by the computer device include: CLSX1X2X3SEP Y1Y2, and CLS X1X2X3SEP Y3.

[0054] After obtaining the target splicing sequence, the computer device can perform encoding processing on each token in the target splicing sequence to obtain a first encoding sequence, so that the first encoding sequence includes the word vectors of each token in the text information of the multimedia data and the word vectors of each token in the target comment information. After the computer device obtains the first encoding sequence and determines the word vectors of each token in the text information and the word vectors of each token in the target comment information in the first encoding sequence, it can adopt the self-attention mechanism to enable each token to obtain global description information from multiple perspectives, thereby deeply understanding the semantics of the text information and the target comment information. In one embodiment, when the computer device adopts the self-attention mechanism to obtain global description information, it can be determined according to the word vector of any token in the first encoding sequence and the word vector of any other token. In a specific implementation, the computer device can calculate the similarity between the word vector of any token and the word vector of any other token, and this similarity can be understood as the attention score for any token and any other token, so that the computer device can perform weighted processing on the word vector of the corresponding token according to the attention score between any token and any other token, and use the weighted word vector as the global description information. It can be understood that based on the acquisition of global description information, the overall semantic understanding of the text information and the target comment information of the multimedia data by the computer device can be effectively improved. Then, the rationality of the reply information generated subsequently for the target comment information can also be improved. In one embodiment, after the computer device obtains the global description information, it can use the global description information as the starting character of the first encoding sequence and use the target character pair to indicate the global description information that is the starting character, where the target character can be the character CLS. Then it can be understood that the computer device can obtain the global description information by reading the starting character CLS of the first encoding sequence. After the computer device obtains the global description information, it can use the global description information to identify the reply type for the target comment information and adopt a reply strategy matching the reply type to generate reply information, that is, then execute step S305.

[0055] In one embodiment, when the computer device encodes each word segment in the target splicing sequence to obtain a first encoded sequence, the computer device can directly perform vector conversion on the corresponding word segments in the target splicing sequence based on the mapping relationship between the word segments and vectors, so as to obtain the first encoded sequence. Alternatively, in the process of encoding each word segment in the target splicing sequence to obtain the first encoded sequence, the computer device can also introduce other word segment information into the word segment vector when encoding each word segment in the target splicing sequence. In a specific implementation, when the computer device obtains the word segment vector of each word segment, it can also generate a position vector and / or a type vector for each word segment vector, so as to improve the information richness of each word segment vector of the first encoded sequence obtained by the computer device, that is, the computer device realizes introducing more word segment information of each word segment in the target splicing sequence during the encoding process. Based on the introduction of more information, the computer device can improve the accuracy of obtaining the global description information based on the first encoded sequence obtained by encoding. In one embodiment, the position vector is used to indicate the position of the corresponding word segment in the target splicing sequence corresponding to the word segment vector, and the type vector is used to indicate what kind of text information (or target comment information) the word segment corresponding to the word segment vector is, such as indicating whether the word segment corresponding to the word segment vector is ocr, asr, title, topic label, creation type, operation type, existing comment and corresponding reply, or which one of the target comment information.

[0056] S305, identify the reply type for the target comment information using the global description information, and obtain a reply strategy matching the reply type.

[0057] In one embodiment, the first encoded sequence can be generated by calling a comment reply model. The comment reply model is a model trained using deep learning. The first encoded sequence is obtained by encoding the text information of the multimedia data included in the content information and the target comment information by the encoder of the comment reply model. Among them, the model structure of the comment reply model can be as Figure 4a shown. The comment reply model includes an encoder, a discriminant network, and a decoder. As Figure 4a shown, the discriminant network and the decoder are respectively connected to the encoder. The encoder is used to encode the target splicing sequence composed of the text information of the multimedia data and the target comment information, and obtain the first encoded sequence. The discriminant network is a linear layer (such as Figure 4aa linear layer marked by 40 is used to determine the information type (or the reply type for the target comment information) for the target comment information according to the first encoded sequence output by the encoder, and the decoder generates a reply information based on the discrimination result of the discrimination network and the first encoded sequence obtained by the encoder. The first encoded sequence is obtained after inputting the text information and the target comment information into an Embedding network (a kind of serialization network) for serialization. The Embedding network can be an independent network in the comment reply model, or can also be built into the encoder of the comment reply model.

[0058] Then, when the computer device obtains the first encoded sequence through the encoder, the global description information can be recorded as the starting character included in the first encoded sequence. When the computer device identifies the reply type for the target comment information by using the global description information, it can call the discrimination network in the comment reply model to identify and process the global description information to obtain a type discrimination score. The discrimination network can first identify and process the global description information to obtain an initial score of a numerical type. Then, the discrimination network can call the sigmoid function (a kind of smoothing function) to convert the initial score into a probability value between 0 and 1, and the probability value is used to indicate the probability that the target comment information can find an accurate answer based on the first encoded sequence. In one embodiment, the converted probability value is the type discrimination score. After the computer device obtains the type discrimination score for the target comment information, it can determine the information type of the target comment information according to the type discrimination score. In one embodiment, after generating the first encoded sequence, the encoder can directly input the first encoded sequence into the discrimination network, so that the discrimination network obtains the global description information by identifying the starting character of the first encoded sequence, and determines the reply type for the target comment information based on the global description information. Or, after generating the first encoded sequence, the encoder can also only input the global description information recorded in the starting character of the first encoded sequence into the discrimination network, so that the discrimination network identifies the reply type for the target comment information based on the global description information.

[0059] In one embodiment, since an information type is associated with a reply type, and the information type includes objective fact type or general type, then, based on the determined information type of the target comment information, the computer device can obtain the corresponding reply type for the target comment information. That is to say, the reply type determined by the computer device for the target comment information includes the reply type for the objective fact type of the target comment information and the reply type for the general type of the target comment information. After the computer device determines the reply type for the target comment information, it can obtain a reply strategy that matches the determined reply type. Among them, when the computer device obtains a reply strategy that matches the reply type, if the computer device determines that the type discrimination score determined based on the discrimination network is greater than or equal to the preset score threshold, then the computer device determines that the accurate and correct answer corresponding to the target comment information can be obtained based on the first coding sequence. Then, the computer device can use the reply type associated with the objective fact type as the reply type for the target comment information and use the strategy indicating information reply through the comment reply model as the reply strategy, so as to further call the comment reply model (i.e., the decoder included in the comment reply information) to generate reply information. In another implementation manner, if the type discrimination score determined by the computer device based on the discrimination network is less than the preset score threshold, it means that the computer device cannot find the answer corresponding to the target comment information based on the first coding sequence. Then, the computer device can use the reply type associated with the general type as the reply type for the target comment information and use the strategy indicating information reply through the generation model as the reply strategy, so as to call the generation model to generate reply information.

[0060] In one embodiment, the encoder in the comment reply model can be Albert (A Lite BERT, a lightweight word vector encoding model) to improve the training convergence speed and effect. The discrimination network is a linear network layer, and the decoder is a bilinear network layer. Among them, the above-mentioned comment reply model is a trained model, and the training process of the comment reply model can be obtained by jointly training the encoder, discrimination network, and decoder, or it can also be obtained by separately training the encoder, discrimination network, and decoder in the comment reply model. Among them, when separately training the comment reply model, the cross-entropy loss function can be used to train the discrimination network in the comment reply model.

[0061] S306, according to the reply strategy and based on the content information of the multimedia data, generate the reply information of the target comment information and output the reply information.

[0062] When a computer device generates a reply message according to the obtained reply policy, if the reply policy indicates that the reply message is generated by calling a comment reply model, the computer device can generate the reply message by calling the decoder included in the comment reply model. Among them, as Figure 4a shown, the decoder is composed of two linear layers, that is, the decoder in the comment reply model includes a first linear layer and a second linear layer. In one embodiment, one of the two linear layers constituting the decoder is used to perform mapping processing on all the word vectors included in the first encoded sequence to obtain the score (or probability) of each word vector in the first encoded sequence as the start position of the reply. The other linear layer is also used to map the word vectors in the first encoded sequence, and further obtain the score (or probability) of each word vector in the first encoded sequence as the end position of the reply. That is to say, when the computer device generates a reply message for the target comment message according to the reply policy and based on the content information of the multimedia data, it can use the first linear layer in the decoder of the comment reply model to perform recognition processing on the first encoded sequence obtained from the encoder to determine the probability that each segmented word vector constituting the first encoded sequence is the start position of the reply message, and, use the second linear layer in the decoder of the comment reply model to perform recognition processing on the first encoded sequence obtained from the encoder to determine the probability that each segmented word vector constituting the first encoded sequence is the end position of the reply message; then further, the computer device can intercept a partial encoded sequence from the first encoded sequence according to the probabilities of each word vector in the first encoded sequence being the start position of the reply and being the end position of the reply, and perform decoding processing on the partial encoded sequence to obtain the reply message of the target comment message.

[0063] In one embodiment, when the decoder determines the probabilities that each word vector in the first encoded sequence corresponds to a start position (or an end position) respectively, the obtained probability values are transformed by the sigmoid function, are values in the range of 0 to 1, and are independent of each other. Similarly, when training the decoder, cross-entropy can also be used as the loss function for training. Then, when the computer device extracts a partial encoded sequence from the first encoded sequence according to the probabilities that each word vector in the first encoded sequence is the start position of the reply and the end position of the reply respectively, the computer device can determine the start and end positions that meet the reply restriction conditions based on the above-determined probabilities, and then can sort the start and end positions that meet the reply restriction conditions, and select the word vectors corresponding to the start and end positions with the maximum combined probability for generating reply information. In a specific implementation, the computer device can select a reply start position and a reply end position that meet the reply restriction conditions from the first encoded sequence, and the reply restriction conditions include one or both of the following: the reply end position is greater than the reply start position, and the length of the encoded sequence determined based on the selected reply start position and reply end position is greater than a length threshold; further, the computer device can select the word vector with the maximum combined probability as the reply start position and the word vector with the maximum combined probability as the reply end position according to the probability of the word vector corresponding to the selected reply start position and the probability of the word vector corresponding to the corresponding reply end position, and use the selected word vectors and the word vectors between the selected word vectors as the partial encoded sequence. For example, if the first encoded sequence determined by the computer device is CLSX1X2X3SEP Y1Y2Y3, and if the word vectors corresponding to the start positions that meet the restriction conditions determined by the computer device include X1 and X3, and the word vectors corresponding to the end positions include Y1 and Y2, then, based on the principle of maximizing the corresponding selection combination probability, the computer device determines that the word vectors with the maximum combined probability are X1 and Y2, and the computer device can determine that the partial encoded sequence selected from the first encoded sequence is CLS X1X2X3SEP Y1Y2, and can obtain the reply information based on the decoding of the partial encoded sequence.

[0064] In one embodiment, if the reply policy indicates that the computer device will generate reply information by invoking a generation model, then when the computer device generates a reply information for the target comment information according to the reply policy and based on the content information of the multimedia data, the computer device may invoke the generation model to encode the image information of the multimedia data included in the content information and the tokenization sequence of the text information of the multimedia data included in the content information, to obtain a second encoding sequence, where the second encoding sequence includes the word vectors of each token in the text information; furthermore, the generation model may be invoked to add corresponding global description information to each word vector in the second encoding sequence, to obtain a new second encoding sequence. Then, after obtaining the new second encoding sequence, the computer device may obtain the target episode label corresponding to the multimedia data, and generate a reply information for the target comment information by using the new second encoding sequence and the target episode label. In one embodiment, when the computer device invokes the generation model to encode the image information of the multimedia data included in the content information and the tokenization sequence of the text information of the multimedia data included in the content information to obtain a second encoding sequence, it may first obtain the image vectors of each image in the multimedia data and the sequence length of the tokenization sequence; thus, when the sequence length is less than or equal to the length threshold, the image vectors may be added to the starting character of the tokenization sequence, such as the above [CLS], to obtain a second encoding sequence; in addition, when the sequence length is greater than the length threshold, the computer device may perform sequence segmentation on the tokenization sequence based on the length threshold, and add image vectors at the starting character of each segmented sequence to obtain new segmented sequences, and each new segmented sequence obtained is the second encoding sequence. Similarly, the length threshold may be the same as or different from the length threshold involved in the process of determining the first encoding sequence, that is, the length threshold may also be 512 or the like.

[0065] In one embodiment, since the text information of the multimedia data includes video titles, tags, OCR, ASR, existing comments and replies, etc., when the computer device obtains the token sequence of the text information, it can first tokenize the video titles, tags, OCR, ASR, existing comments and replies respectively, and then use a delimiter (such as the above [SEP]) to separate and splice the tokens from the same source, so as to obtain the token sequence of the text information. After the computer device obtains the token sequence of the text information, it can encode each token in the token sequence, and form a second encoding sequence from the word vectors of each token. Similarly, in order to enrich the semantics of the word vectors in the second encoding sequence, after generating the second encoding sequence, corresponding position vectors and type vectors can be added to each word vector in the second encoding sequence. After the computer device obtains the image vector of the multimedia data, it can first perform a linear transformation on the image vector to make the image vector and the word vectors in the second encoding sequence have the same vector dimension, and then the linearly transformed image can be concatenated and added to the starting character of the second encoding sequence. In addition, the computer device can also perform digit supplementation based on the length threshold. For example, when the computer device determines that the sequence length of the token sequence is less than the length threshold, it can add a placeholder to the token sequence so that the sequence length of the token sequence after adding the placeholder is equal to the length threshold, where the placeholder can be [PAD] for example.

[0066] In one embodiment, the generation model is a recurrent neural network structure, such as Figure 4b shown, the generation model includes an encoder (such as Figure 4b the network labeled 41 in Figure 4b ), an episode classification network (such as Figure 4bThe network marked by 43 in the figure), where the encoder is used to generate the above-mentioned second encoded sequence. The encoder adopts a transformer encoder structure (a type of encoding structure) to transform the input information (such as the text information and image information of the above-mentioned multimedia data) through multiple layers of attention operations (or self-attention operations), and transform the original vector (i.e., the vector sequence, such as the above-mentioned second encoded sequence) into a vector representation with abstract semantics. Each vector in this input vector representation sequence contains both the semantic information of the words and the current context information. Among them, the attention operation is executed by a self-attention module. The self-attention operation calculates the similarity between every two vectors, normalizes the similarity between each vector and all other vectors into weights whose sum is 1, then uses the obtained weights to perform weighted summation with other vectors to obtain the current vector, and further obtains the global description information of each token in the encoded sequence. Then, after the self-attention operation, each vector can obtain the information carried by other vectors, so that each vector can obtain the global description information related to itself, so that the generation model can obtain corresponding grammar, semantics, etc.

[0067] The episode classification network is used to enable the computer device to obtain the target episode label corresponding to the multimedia data. Among them, the episode classification network takes the vector output by the encoder (such as any word vector in the first encoded sequence) as input, and performs a dot product between any word vector in the first encoded sequence and all episode type vectors. At this time, each episode type gets a score, and after passing through the sigmoid function, a probability value in the range of 0-1 is obtained. This probability value represents the probability that the multimedia data belongs to this episode type. After obtaining the probability values of all episode types, select the episode type greater than the preset threshold as the target episode type for output. Among them, during the training of the episode classification network, a corresponding vector will be randomly generated for each episode type for training, and during the training, the type label will be trained using a multi-hot vector (a type of multi-classification vector). Among them, the multi-hot vector is a vector in which the corresponding position is set to 1 if it contains a certain episode type, and 0 otherwise.

[0068] When the computer device obtains the target plot label of the multimedia data by using the trained plot classification network, it can first obtain the label vector corresponding to any plot label, perform vector operations on the new second coding sequence and any label vector to obtain the matching degree between the new second coding sequence and any label vector, and then select the plot label corresponding to the label vector with the corresponding matching degree greater than or equal to the matching degree threshold as the target plot label of the multimedia data. Then, after the computer device obtains the target plot label, it can call the generation network in the generation model and generate a reply message of the target comment information by using the new second coding sequence and the target plot label. In one embodiment, the network structure of the generation network in the generation model adopts a transformer decoder structure (a decoding structure), and its input includes all vector sequences output by the encoder (i.e., the second coding sequence) and the target plot label (label vector) output by the plot classification network. The initial state of the decoder respectively adopts the above plot type label vector to generate comments related to the plot type of the video (multimedia data). Therefore, multiple comments related to different plot types can be generated for the multimedia data. The vector sequence output by the encoder is used for the attention operation of the decoder, which is convenient for the decoder to obtain more relevant information and make the comment generation result more reasonable and smooth.

[0069] When the computer device generates a response message based on the second coding sequence and the target plot label, the specific decoding process of the generation network is to first select multiple video plot vectors (plot type labels) with the highest probabilities, and start the encoding on the decoder side with the selected video plot vectors as the initial vectors of the transformer decoder. The macroscopic decoding steps are to predict the first response token (token) based on the initial vector, predict the second token based on the initial vector and the first token, and so on until the predicted token is the end symbol or the prediction length reaches the threshold. The specific steps are to first perform a linear mapping on the obtained video plot vector, then calculate the similarity between the mapped vector and the output vector of the transformer encoder, and use the similarity score and the encoder vector to calculate the vector related to this video plot. Then, by multiplying this vector with the vocabulary matrix, the probability distribution of the first token is obtained, and the k tokens with the highest probabilities are selected as the first token respectively, and input into the transformer encoder together with the plot vector to predict the second token, and so on. The final output is similar to a tree structure. Each route from the root node to the leaf node of the tree is a comment, and the score of each comment is the product of the probabilities of each token on the path. Finally, the k comments with the highest scores are selected. Then, after the model prediction is completed, the computer device can select the top k comments from each video plot type and output them as the response to the target comment information.

[0070] That is to say, the computer device can predict the i-th response token based on the label vector corresponding to the target plot label and the new second coding sequence; i≥1 and is a positive integer; furthermore, the (i + 1)-th response token can be predicted using the i-th response token. Then, when the obtained (i + j)-th response token meets the prediction termination condition, the computer device can generate a response message to the target comment information based on the predicted (i + j) response tokens; j≥1 and is a positive integer. When the computer device predicts the i-th response token based on the label vector corresponding to the target plot label and the new second coding sequence, the computer device can perform a mapping process on the label vector of the target plot label to obtain a mapped vector of the label vector, and calculate the similarity between the new second coding sequence and the mapped vector; thus, a label vector of the plot label associated with the target plot label can be generated according to the similarity and the new second coding sequence; furthermore, based on the generated label vector of the associated plot label and the vocabulary matrix, the probability distribution of each token in the vocabulary matrix being selected as the i-th response token can be generated, and one or more tokens can be selected from the vocabulary matrix as the i-th response token based on the probability distribution. Furthermore, the above-mentioned tree structure can be deduced, and the final response message can be selected and output.

[0071] In an embodiment of the present application, after a computer device obtains the content information of multimedia data and target comment information, it can obtain a first coding sequence based on the text information of the multimedia data included in the content information and the target comment information. After the computer device obtains the first coding sequence, it can obtain global description information for the text information and the target comment information according to the similarity between the word vectors of any two words in the first coding sequence, so that the computer device can determine different reply strategies for the target comment information based on the global description information. For objective fact-based comment information, the corresponding reply can be determined from the text information of the multimedia data based on a comment reply model, greatly improving the accuracy of objective fact-based comment information. For general comment information, a deep generation model can be used and combined with the type label of the multimedia data to generate reply information, so as to improve the rationality and accuracy of the reply information generated by the computer device. In addition, since the computer device refers to existing comments and replies when generating reply information, and refers to relevant information of the multimedia data in different modalities when constructing the coding sequence, the accuracy and diversity of the computer device when generating corresponding reply information based on the coding sequence are also improved.

[0072] Based on the description of the above information processing method embodiment, an embodiment of the present invention also proposes an information processing device, which can be a computer program (including program code) running in the above computer device. The information processing device can be used to execute as Figure 2 and Figure 3 the information processing method described, please refer to Figure 5 , and the information processing device includes: an acquisition unit 501 and a processing unit 502.

[0073] The acquisition unit 501 is used to acquire the content information of the multimedia data and the target comment information for the multimedia data;

[0074] The acquisition unit 501 is further used to acquire global description information, and the global description information is used to describe the information semantics of the content information and the target comment information;

[0075] The processing unit 502 is used to identify the reply type for the target comment information by using the global description information, and acquire a reply strategy matching the reply type;

[0076] The processing unit 502 is further used to generate a reply information for the target comment information according to the reply strategy and based on the content information of the multimedia data, and output the reply information.

[0077] In one embodiment, the obtaining unit 501 is specifically configured to:

[0078] Obtain a first encoding sequence, where the first encoding sequence includes word vectors of each word segment in the text information and word vectors of each word segment in the target comment information;

[0079] Calculate the similarity between any word segment and any other word segment according to the word vectors of any word segment and the word vectors of any other word segment, and perform weighted processing on the word vectors of each word segment based on the similarity;

[0080] Use the vector sequence composed of the weighted word vectors as the global description information for the text information and the target comment information.

[0081] In one embodiment, the obtaining unit 501 is specifically configured to:

[0082] Perform word segmentation on the target comment information to obtain a word segmentation sequence corresponding to the target comment information, and perform word segmentation on the text information to obtain a word segmentation sequence of the text information;

[0083] Concatenate the word segmentation sequence of the target comment information and the word segmentation sequence of the text information to obtain a target concatenated sequence;

[0084] Perform encoding processing on each word segment in the target concatenated sequence to obtain a first encoding sequence.

[0085] In one embodiment, the global description information is used as the starting character of the first encoding sequence. The first encoding sequence is generated by calling a comment reply model, and the comment reply model is a model trained by deep learning. The first encoding sequence is obtained by encoding the text information of the multimedia data included in the content information and the target comment information by the encoder of the comment reply model, and the encoder is connected to a discriminant network; the processing unit 502 is specifically configured to:

[0086] Call the discriminant network in the comment reply model to perform recognition processing on the global description information to obtain a type discrimination score; the global description information is obtained by the encoder to obtain the starting character of the first encoding sequence;

[0087] Determine the information type of the target comment information according to the type discrimination score, where one information type is associated with one reply type, and the information type includes at least one of an objective fact type and a general type.

[0088] In one embodiment, if the reply policy indicates that information reply is to be performed by invoking a comment reply model, the comment reply model further includes a decoder connected to the encoder. The encoder is configured to encode the text information of the multimedia data included in the content information and the target comment information to obtain a first encoded sequence. The decoder includes a first linear layer and a second linear layer. The processing unit 502 is specifically configured to:

[0089] Use the first linear layer in the decoder of the comment reply model to perform recognition processing on the first encoded sequence obtained from the encoder, and determine the probability that the word vectors corresponding to the respective word segments constituting the first encoded sequence are the starting positions of the reply information.

[0090] Use the second linear layer in the decoder of the comment reply model to perform recognition processing on the first encoded sequence obtained from the encoder, and determine the probability that the word vectors corresponding to the respective word segments constituting the first encoded sequence are the ending positions of the reply information.

[0091] According to the probabilities that the respective word vectors in the first encoded sequence are the starting position of the reply and the ending position of the reply, intercept a partial encoded sequence from the first encoded sequence, and perform decoding processing on the partial encoded sequence to obtain the reply information for the target comment information.

[0092] In one embodiment, if the reply policy indicates that the reply information is to be generated by invoking a generation model; the processing unit 602 is specifically configured to:

[0093] Invoke the generation model to perform encoding processing on the image information of the multimedia data included in the content information and the word segment sequence of the text information of the multimedia data included in the content information, to obtain a second encoded sequence; the second encoded sequence includes the word vectors of the respective word segments in the text information.

[0094] Invoke the generation model to add corresponding global description information to the respective word vectors in the second encoded sequence, to obtain a new second encoded sequence.

[0095] Obtain the target plot label corresponding to the multimedia data, and use the new second encoded sequence and the target plot label to generate the reply information for the target comment information.

[0096] In an embodiment of the present application, when the obtaining unit 501 needs to generate a reply message for the comment message of the multimedia data, it will obtain the content information of the multimedia data, the target comment message, and the global description information used to describe the information semantics of the content information and the target comment message. Further, the processing unit 502 can use the global description information to identify the reply type for the target comment message. Thus, the processing unit 502 can generate the reply message for the target comment message according to the reply strategy matching the reply type and the content information of the multimedia data. It can implement generating the reply messages for comment messages of different information types by using differentiated reply strategies based on the differences in the information types of the comment messages. And by using differentiated reply strategies to generate reply messages based on the different information types of the comment messages, the diversity in the process of generating reply messages is improved. In addition, since the content information of the multimedia data is fully considered in the process of generating the reply message, the rationality and accuracy of the generated reply message can be improved.

[0097] Please refer to Figure 6 , which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present invention. As Figure 6 shown, the computer device in this embodiment may include: one or more processors 601; one or more input devices 602, one or more output devices 603, and a memory 604. The above-mentioned processors 601, input devices 602, output devices 603, and memory 604 are connected through a bus 605. The memory 604 is used to store a computer program, and the computer program includes program instructions. The processor 601 is used to execute the program instructions stored in the memory 604.

[0098] The memory 604 may include a volatile memory, such as a random-access memory (RAM); the memory 604 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the memory 604 may further include a combination of the above types of memories.

[0099] The processor 601 may be a central processing unit (CPU). The processor 601 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), etc. The PLD may be a field-programmable gate array (FPGA), a generic array logic (GAL), etc. The processor 601 may also be a combination of the above structures.

[0100] In an embodiment of the present invention, the memory 604 is used to store a computer program, the computer program includes program instructions, and the processor 601 is used to execute the program instructions stored in the memory 604 to implement the corresponding method steps as described above Figure 2 and Figure 3 above.

[0101] In one embodiment, the processor 601 is configured to call the program instructions to execute:

[0102] Obtain the content information of the multimedia data and the target comment information for the multimedia data;

[0103] Obtain global description information, where the global description information is used to describe the information semantics of the content information and the target comment information;

[0104] Use the global description information to identify the reply type for the target comment information, and obtain a reply strategy that matches the reply type;

[0105] Generate a reply message for the target comment information according to the reply strategy and based on the content information of the multimedia data, and output the reply message.

[0106] An embodiment of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above as Figure 2 or Figure 3The method embodiments shown. Among them, the computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0107] What is disclosed above is only partial embodiments of the present invention. Of course, the scope of rights of the present invention cannot be limited by this. Those of ordinary skill in the art can understand the entire or partial processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. An information processing method, characterized in that, Including: Obtaining content information of multimedia data and target comment information for the multimedia data; Obtaining global description information, where the global description information is used to describe the information semantics of the content information and the target comment information; Using the global description information to identify the reply type for the target comment information and obtaining a reply strategy that matches the reply type; Generating a reply message for the target comment information according to the reply strategy and based on the content information of the multimedia data, and outputting the reply message; if the reply strategy indicates information reply by calling a generation model; The generating a reply message for the target comment information according to the reply strategy and based on the content information of the multimedia data includes: calling the generation model to perform encoding processing on the image information of the multimedia data included in the content information and the word segmentation sequence of the text information of the multimedia data included in the content information, to obtain a second encoding sequence; the second encoding sequence includes word vectors of each word in the text information; Calling the generation model to add corresponding global description information to each word vector in the second encoding sequence to obtain a new second encoding sequence; obtaining a target plot label corresponding to the multimedia data, and generating a reply message for the target comment information using the new second encoding sequence and the target plot label.

2. The method according to claim 1, wherein The content information includes the text information of the multimedia data; the obtaining global description information includes: Obtaining a first encoding sequence, where the first encoding sequence includes word vectors of each word in the text information and word vectors of each word in the target comment information; Calculating the similarity between any word vector and any other word vector, and performing weighted processing on the word vector of each word based on the similarity; Taking the vector sequence composed of the weighted word vectors as the global description information for the text information and the target comment information.

3. The method according to claim 2, characterized in that, The obtaining the first encoding sequence includes: Performing word segmentation on the target comment information to obtain a word segmentation sequence corresponding to the target comment information, and performing word segmentation on the text information to obtain a word segmentation sequence of the text information; Performing sequence splicing on the word segmentation sequence of the target comment information and the word segmentation sequence of the text information to obtain a target splicing sequence; Performing encoding processing on each word in the target splicing sequence to obtain a first encoding sequence.

4. The method according to claim 3, wherein The performing sequence splicing on the word segmentation sequence of the target comment information and the word segmentation sequence of the text information to obtain a target splicing sequence includes: Obtaining the sequence length of the word segmentation sequence of the target comment information and the sequence length of the word segmentation sequence of the text information; If the sum of the sequence lengths between the word segmentation sequence of the target comment information and the word segmentation sequence of the text information is less than or equal to a length threshold, taking the sequence directly obtained by splicing the word segmentation sequence of the target comment information and the word segmentation sequence of the text information as the target splicing sequence; When the length is greater than the length threshold, based on the sequence length of the word segmentation sequence of the target comment information and the length threshold, the word segmentation sequence of the text information is segmented, and each segmented sequence is respectively concatenated with the word segmentation sequence of the target comment information, and each concatenated sequence obtained is the target concatenated sequence.

5. The method according to claim 3, wherein The method further includes: Adding a separator between the word segmentation sequence of the target comment information and the word segmentation sequence of the text information; And adding global description information as a starting character at the starting position of the target concatenated sequence.

6. The method according to claim 1, wherein The global description information is used as the starting character of the first encoded sequence. The first encoded sequence is generated by calling a comment reply model, which is a model trained by deep learning. The first encoded sequence is obtained by encoding the text information of the multimedia data included in the content information and the target comment information by the encoder of the comment reply model. The encoder is connected to a discriminant network; Identifying the reply type for the target comment information by using the global description information includes: Calling the discriminant network in the comment reply model to perform identification processing on the global description information to obtain a type discrimination score; the global description information is obtained by the encoder to obtain the starting character of the first encoded sequence; According to the type discrimination score, determine the information type of the target comment information, where one information type is associated with one reply type, and the information type includes at least one of objective fact type and general type.

7. The method according to claim 6, wherein Obtaining a reply strategy that matches the reply type includes: When the type discrimination score is greater than or equal to a preset score threshold, taking the reply type associated with the objective fact type as the reply type of the target comment information, and taking the strategy for indicating information reply through the comment reply model as the reply strategy; When the type discrimination score is less than the preset score threshold, taking the reply type associated with the general type as the reply type of the target comment information, and taking the strategy for indicating information reply through the generation model as the reply strategy.

8. The method according to claim 1, characterized in that, If the reply strategy indicates information reply by calling the comment reply model, the comment reply model further includes a decoder, and the decoder is connected to the encoder. The encoder is used to encode the text information of the multimedia data included in the content information and the target comment information to obtain a first encoded sequence. The decoder includes a first linear layer and a second linear layer; Generating the reply information of the target comment information according to the reply strategy and based on the content information of the multimedia data includes: Using the first linear layer in the decoder of the comment reply model to perform identification processing on the first encoded sequence obtained from the encoder, and determining the probability that each word segment constituting the first encoded sequence corresponds to the word vector at the starting position of the reply information; Using the second linear layer in the decoder of the described comment reply model, perform recognition processing on the first encoded sequence obtained from the encoder to determine the probability that each token constituting the first encoded sequence corresponds to the end position of the reply information; According to the probabilities that each token vector in the first encoded sequence is the start position of the reply and the end position of the reply respectively, extract a partial encoded sequence from the first encoded sequence, and perform decoding processing on the partial encoded sequence to obtain the reply information of the target comment information.

9. The method according to claim 8, wherein The extracting a partial encoded sequence from the first encoded sequence according to the probabilities that each token vector in the first encoded sequence is the start position of the reply and the end position of the reply respectively includes: Selecting a start position and an end position of the reply that satisfy the reply restriction conditions from the first encoded sequence, where the reply restriction conditions include one or both of the following: the end position of the reply is greater than the start position of the reply, and the length of the encoded sequence determined based on the selected start position and end position of the reply is greater than the length threshold; According to the probability of the token vector corresponding to the selected start position of the reply and the probability of the token vector corresponding to the corresponding end position of the reply, select the token vector with the largest combined probability as the token vector for the start position of the reply and the token vector for the end position of the reply, and use the selected token vectors and the token vectors between the selected token vectors as the partial encoded sequence.

10. The method according to claim 1, characterized in that, The calling the generation model to perform encoding processing on the image information of the multimedia data included in the content information and the token sequence of the text information of the multimedia data included in the content information to obtain a second encoded sequence includes: Obtaining the image vectors of each image in the multimedia data and the sequence length of the token sequence; When the sequence length is less than or equal to the length threshold, adding the image vector to the start character of the token sequence to obtain a second encoded sequence; When the sequence length is greater than the length threshold, perform sequence segmentation on the token sequence based on the length threshold, and add the image vector at the start character of each segmented sequence to obtain a new segmented sequence, and each obtained new segmented sequence is the second encoded sequence.

11. The method according to claim 1, wherein The obtaining the target plot label corresponding to the multimedia data includes: Obtaining a label vector corresponding to any plot label; Performing vector operation on the new second encoded sequence and any of the label vectors to obtain the matching degree between the new second encoded sequence and the any label vector; Selecting the plot label corresponding to the label vector with a corresponding matching degree greater than or equal to the matching degree threshold as the target plot label of the multimedia data.

12. The method according to claim 1, wherein The generating the reply information of the target comment information using the new second encoded sequence and the target plot label includes: Predicting the i-th reply token based on the label vector corresponding to the target plot label and the new second encoded sequence; i≥1 and is a positive integer; Predicting the (i + 1)-th reply token using the i-th reply token; When the obtained (i + j)-th reply word segmentation meets the prediction termination condition, a reply message for the target comment message is generated based on the predicted (i + j) reply word segmentations; j ≥ 1 and is a positive integer.

13. The method according to claim 12, wherein The predicting the i-th reply word segmentation based on the label vector corresponding to the target plot label and the new second encoding sequence includes: Performing a mapping process on the label vector of the target plot label to obtain a mapped vector of the label vector, and calculating the similarity between the new second encoding sequence and the mapped vector; Generating a label vector of a plot label associated with the target plot label according to the similarity and the new second encoding sequence; Based on the generated label vector of the associated plot label and the vocabulary matrix, generating a probability distribution of each word segmentation in the vocabulary matrix being selected as the i-th reply word segmentation, and selecting one or more word segmentations from the vocabulary matrix as the i-th reply word segmentation based on the probability distribution.

14. An information processing apparatus, characterized in that, including: An acquisition unit, configured to acquire content information of multimedia data and a target comment message for the multimedia data; The acquisition unit is further configured to acquire global description information, where the global description information is used to describe the information semantics of the content information and the target comment message; A processing unit, configured to use the global description information to identify a reply type for the target comment message and obtain a reply strategy matching the reply type; The processing unit is further configured to, according to the reply strategy, generate a reply message for the target comment message based on the content information of the multimedia data, and output the reply message; if the reply strategy indicates information reply by invoking a generation model; The generating a reply message for the target comment message according to the reply strategy and based on the content information of the multimedia data includes: invoking the generation model to perform encoding processing on the image information of the multimedia data included in the content information and the word segmentation sequence of the text information of the multimedia data included in the content information, to obtain a second encoding sequence; the second encoding sequence includes word vectors of each word segmentation in the text information; Invoking the generation model to add corresponding global description information to each word vector in the second encoding sequence to obtain a new second encoding sequence; obtaining a target plot label corresponding to the multimedia data, and generating a reply message for the target comment message by using the new second encoding sequence and the target plot label.

15. A computer device, characterized in that, including a processor, an input device, an output device, and a memory, where the processor, the input device, the output device, and the memory are interconnected, and the memory is configured to store a computer program, the computer program includes program instructions, and the processor is configured to invoke the program instructions to execute the method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 13.

17. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, and the computer program or the computer instructions, when executed by a processor, are used to implement the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method and device for cross-type conversation, apparatus, and computer-readable storage medium

    CN109033223A

  • Conversation generation method and device, video comment method and device, equipment and storage medium

    CN111625660A

  • Text sentiment analysis method, electronic device and storage medium

    CN113434682A