Question and answer interaction method and device, medium, product and equipment

By finding and displaying highlighted barrage in video comments, the problem of interrupting the search for answers during movie viewing in the prior art is solved, and the user experience is improved.

CN120343349APending Publication Date: 2025-07-18MIGU CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510546411.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When users have questions during the video viewing process, the existing technology needs to interrupt the movie viewing process and search for answers, which will affect the movie viewing effect.

Method used

Avoid interrupting the movie viewing process by finding the target answer comments in the existing comments of the target video and displaying them in a highlighted barrage when they exist.

Benefits of technology

It realizes answering user questions without interrupting the movie viewing process, improving user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343349A_ABST
    Figure CN120343349A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer interaction method and device, a medium, a product and equipment, and the method comprises the steps: searching whether a target answer comment exists in the existing comments of a target video or not based on a question related to the target video; and displaying the target content corresponding to the target answer comment under the condition that the target answer comment exists, so that whether the target answer comment exists in the existing comments of the target video or not can be directly searched after the user puts forward a question related to the target video, and the target answer comment is displayed in a highlight bullet screen manner when the target answer comment is searched, so that the user experience is improved. Therefore, questions proposed by the user in the film watching process can be answered while the film watching process of the user is prevented from being interrupted, so that the film watching experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular, to a question-answering interaction method, device, medium, product, and equipment. Background Art

[0002] Currently, when a user has a question about the content of a video during the video viewing process, the user usually needs to search for an answer by searching for the question in a common search engine.

[0003] However, this way of finding answers usually requires interrupting the user's video viewing process (for example, pausing the video and switching to the search interface to search), thereby affecting the viewing effect. Summary of the Invention

[0004] To solve the above technical problems, embodiments of this application propose a question-answering interaction method, device, medium, product, and equipment, which can answer the questions raised by the user during the video viewing process without interrupting the user's video viewing process, and improve the user's video viewing experience.

[0005] In a first aspect, embodiments of this application provide a question-answering interaction method, including:

[0006] Based on a question related to a target video, search for a target answer comment in the existing comments of the target video;

[0007] In the case where the target answer comment exists, display the target content corresponding to the target answer comment.

[0008] Further, before searching for a target answer comment in the existing comments of the target video, the method further includes:

[0009] Obtain a user instruction and convert the user instruction into semantic text;

[0010] In the case where the semantic text is determined to be an interrogative sentence, determine whether the semantic text is related to the target video. If it is related, use the semantic text as the question;

[0011] Among them, determining whether the semantic text is related to the target video includes any one of the following:

[0012] Identify the semantic text to obtain a first keyword, perform word matching on the first keyword in a preset keyword library, and determine whether the semantic text is related to the target video according to the word matching result, where the keyword library is used to store keywords corresponding to the target video;

[0013] Use a pre-trained question filtering model to screen the semantic text, and determine whether the semantic text is relevant to the target video according to the screening result, where the question filtering model is trained from a first pre-trained model based on the first text content of the target video and the text content of other videos, and the other videos are different from the target video.

[0014] Further, based on the question related to the target video, check whether there is a target answer comment in the existing comments of the target video, including:

[0015] Input the question into a pre-trained language model to generate an answer, obtaining a target answer, where the language model includes a first language model and / or a second language model, the first language model is trained from a second pre-trained model based on the second text content of the target video, and the second language model includes a large language model;

[0016] Perform text similarity matching between the existing comments and the target answer, and based on the text similarity matching result, check whether there is the target answer comment in the existing comments.

[0017] Further, the performing text similarity matching between the existing comments and the target answer, and based on the text similarity matching result, checking whether there is the target answer comment in the existing comments, includes:

[0018] For any first comment in the existing comments whose publication time is within the target time range, if the text similarity between the first comment and the target answer meets the preset similarity condition, then determine the first comment as a second comment, where the target time range is determined by the time when the question is raised;

[0019] When the number of the second comments is 1, use the second comment as the target answer comment;

[0020] When the number of the second comments is multiple, evaluate the score of each second comment based on the barrage information, publication time and / or keyword offset position corresponding to each second comment, and find the target answer comment from each of the second comments based on the score;

[0021] Wherein, the keyword offset position indicates the position of a second keyword in the corresponding second comment, and the second keyword is determined by the target answer.

[0022] Further, the bullet screen information includes at least one of the following: the amount of bullet screens in the target video when the second comment is displayed, the number of bullet screen likes for the second comment, the number of bullet screen comments for the second comment, and the bullet screen similarity between the second comment and other bullet screens displayed simultaneously therewith;

[0023] When the number of characters of the target answer is less than the preset character number threshold, the second keywords are the words that meet the preset keyword type among the words obtained by performing word segmentation on the target answer;

[0024] When the number of characters of the target answer is not less than the preset character number threshold, the second keywords are the words that meet the preset keyword type among the words obtained by performing word segmentation on the refined answer, and the refined answer is obtained by generating a text summary of the target answer.

[0025] Further, the method further includes:

[0026] In the case where there is no comment on the target answer, in the target video displayed on the user device that poses the question, the target answer is displayed in a highlighted bullet screen form and maintained for a preset duration, and after reaching the preset duration, the display of the target answer is stopped.

[0027] After displaying the target answer or the target content, an evaluation interface is displayed, and a user operation acting on the evaluation interface is obtained, where the user operation indicates an evaluation score of the user for the displayed target answer or the displayed target content.

[0028] In a second aspect, an embodiment of the present application provides a question-answer interaction device, including:

[0029] A search module, configured to search whether there is a target answer comment in the existing comments of the target video based on a question related to the target video;

[0030] A target content display module, configured to display target content corresponding to the target answer comment in the case where there is the target answer comment.

[0031] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the question-answer interaction method described in any one of the above are implemented.

[0032] In a fourth aspect, an embodiment of the present application provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the question-answer interaction method described in any one of the above are implemented.

[0033] Fifth aspect, an embodiment of the present application provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the question-answering interaction method described in any one of the above are implemented.

[0034] In summary, the embodiments of the present application at least have the following beneficial effects:

[0035] By adopting the embodiment of the present application, based on a question related to the target video, it is checked whether there is a target answer comment in the existing comments of the target video; in the case where the target answer comment exists, the target content corresponding to the target answer comment is displayed. After the user asks a question related to the target video, it can directly check whether there is a target answer comment in the existing comments of the target video, and when found, it is displayed in the form of a highlighted bullet screen, thus avoiding interrupting the user's movie-watching process while also answering the questions raised by the user during the movie-watching process to improve the user's movie-watching experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flowchart of the question-answering interaction method provided by the embodiment of the present application;

[0037] Figure 2 is a schematic diagram of the named entity recognition model provided by the embodiment of the present application;

[0038] Figure 3 is a schematic diagram when displaying bullet screens provided by the embodiment of the present application;

[0039] Figure 4 is a schematic diagram of the question-answering interaction method provided by the embodiment of the present application;

[0040] Figure 5 is a schematic diagram of the structure of the question-answering interaction device provided by the embodiment of the present application;

[0041] Figure 6 is a schematic diagram of the structure of the computer device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0043] In the description of this application, the terms "first", "second", "third", etc. are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "a plurality" is two or more. In the description of this application, the term "comprising" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "according to" means "at least partially according to". The term "an embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments".

[0044] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0045] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0046] In a first aspect, by way of example, the embodiments provided in the first aspect of this application can be applied to a terminal device. For ease of understanding, the following takes the terminal device as the execution subject for explanation.

[0047] See Figure 1 , which shows a schematic flowchart of a question-and-answer interaction method provided by an embodiment of this application. The method includes:

[0048] S103, based on a question related to a target video, search in the existing comments of the target video to find whether there is a target answer comment;

[0049] It should be noted that the target video in this embodiment may include the video that is currently being played or paused on the terminal device, and the problem related to the target video may refer to a question raised about the video content at a certain time point of the target video (for example, a question about the plot corresponding to the current time point, a question about the overall plot, a question about the character personality in the video, etc.).

[0050] It should be noted that the way to raise this question in this embodiment may be to input the text expression of the question in the question input interface awakened on the terminal device, or the terminal device collects the question voice data emitted by the user through a microphone (for example, a microphone wake-up instruction for collecting this question can be set) and converts it into this question, or the terminal device collects the body language image data of the user (such as gesture instruction image data) through a camera and converts it into this question. It should be understood that the microphone, camera, etc. in this embodiment need to be called only after the user has pre-authorized and allowed their use.

[0051] It should be noted that the existing comments on the target video in this embodiment may include bullet comments and / or comments in the comment area of the target video, etc.

[0052] In one example, the target answer comment may be the top N comments with the highest similarity to the question among the existing comments, where N is an integer greater than or equal to 1. The determination of the similarity may include: converting the target answer comment into a first feature vector, converting the existing comment into a second feature vector, and calculating the cosine similarity between the first feature vector and the second feature vector to determine the similarity; or, jointly inputting the target answer comment and the existing comment into a large model for similarity determination to obtain the similarity.

[0053] S104, in the case where there is the target answer comment, display the target content corresponding to the target answer comment. Among them, the target content may include the highlighted bullet comment of the target answer comment.

[0054] It should be noted that the target content to be displayed in this embodiment is the content determined by the target answer comment. For example, when the target answer comment is not in the bullet comments, at least some of the text in the target answer comment can be converted into highlighted and / or enlarged font and / or bold font bullet comments and displayed in the bullet comment area; when the target answer comment exists in the bullet comments, at least some of the text in the corresponding bullet comment can be highlighted and / or the font can be enlarged and / or made bold. In some existing technologies, after receiving a question, a large model may be automatically used or the corresponding answer may be generated through background search, but the machine-generated answers are usually stereotyped and also affect the communication between real users. In view of this, the text part of the target content displayed in this embodiment is all derived from the comments of real users, which can improve the communication experience between real users and avoid the selected answers being as rigid as machine answers.

[0055] In addition, the target content may not be limited to the complete text content of the target answer comment. For example, in some cases, some target answer comments may be too long, so technologies such as large models can be used to generate a text summary of the target answer comment with too long a length according to the question, and then the more concise target content generated can be converted into highlighted and / or enlarged font and / or bold font bullet comments and displayed in the bullet comment area. In this way, the target content related to the question can be presented to the user in a targeted manner, improving the viewing experience.

[0056] In an alternative embodiment, referring to Figure 1 , before searching for the target answer comment in the existing comments of the target video, the method further includes:

[0057] S101, obtaining a user instruction and converting the user instruction into semantic text;

[0058] It should be noted that the user instruction in this embodiment may include the user's voice instruction, the text instruction input by the user, and / or the user gesture instruction, etc. Exemplarily, taking the case where the user instruction is a voice instruction as an example, in this way, the voice instruction can be obtained through the microphone of the terminal device, and the well-known speech-to-text technology can be used to convert the voice instruction into the corresponding semantic text. Among them, how to call the microphone to obtain the voice instruction can be preset for the terminal device, for example, when detecting the set voice command of the specified user, detecting the operation allowing the microphone to be called on the display interface of the terminal device, etc.

[0059] S102, in the case where it is determined that the semantic text is an interrogative sentence, determining whether the semantic text is related to the target video. If it is related, the semantic text is used as the question;

[0060] In one example, determining whether a semantic text is an interrogative sentence may include at least one of the following:

[0061] 1. Determine whether the last symbol of the semantic text is a question mark. If so, determine that the semantic text is an interrogative sentence;

[0062] 2. Determine whether the last character of the semantic text is "ma". If so, determine that the semantic text is an interrogative sentence;

[0063] 3. Determine whether the first character or word of the semantic text is a preset interrogative word. If so, determine that the semantic text is an interrogative sentence, where the preset interrogative word includes any one of the following: who, what, which, why, how, how to, how many, several, whether, is it possible, no wonder;

[0064] 4. Use a regular expression to check whether the semantic text contains a preset interrogative word. If so, determine that the semantic text is an interrogative sentence; where a regular expression (Regular Expression) is usually abbreviated as regex or regexp, which is a tool in computer science used to match character combinations in a string. It consists of a series of characters and special symbols, used to describe the pattern to be matched when searching text. Regular expressions can be used to find, replace, verify, or extract specific string patterns.

[0065] Among them, determining whether the semantic text is related to the target video includes any one of the following:

[0066] Identify the semantic text to obtain a first keyword, perform a word match in a preset keyword library based on the first keyword, and determine whether the semantic text is related to the target video according to the word match result, where the keyword library is used to store the keywords corresponding to the target video;

[0067] It should be noted that the first keyword in this embodiment may include a named entity word identified and extracted from the semantic text. It should be understood that a named entity is a person's name, an organization's name, a place name, and all other entities identified by a name. A more extensive entity also includes numbers, dates, currencies, addresses, etc. The types of named entity words in this embodiment may include at least one of a person's name, a person's alias, a place name, an organization's name, an organization's alias, an event name, an item name, and a feature name. Correspondingly, the keywords stored in the keyword library are the named entity words corresponding to the target video.

[0068] In one example, when the first keyword is a named entity word, identifying the semantic text to obtain a first keyword may include:

[0069] Input the semantic text into a pre-trained named entity recognition model to obtain the first keyword output by the named entity recognition model;

[0070] Among them, the named entity recognition model can be trained by using the named entity tags annotated in the plot texts of multiple TV series (including the target video) as the training set and the validation set. Exemplarily, see Figure 2 , and the recognition model to be trained can be constructed based on the BiLSTM (Bidirectional Long short-term memory) model and the CRF (Conditional Random Fields) model. In this way, the named entity recognition model can be used to extract named entities in the semantic text.

[0071] Combined with the above example, it should be noted that at least one of the following types of parameters of the BiLSTM model in this embodiment can be configured by the user: the size of the total number of words, the control dictionary indicating the correspondence between tags and ids, the word embedding dimension (i.e., the input_size of the LSTM input layer), the hidden layer vector dimension, the number of neural network layers, the number of batches, the maximum length limit of the statement, and the tag size (corresponding to the matrix width of the final output score of the BiLSTM).

[0072] Combined with the above example, it should be understood that in this embodiment, the BiLSTM can be used. Compared with the unidirectional LSTM (Long short-term memory) model that can only capture the information transmitted from front to back, the bidirectional BiLSTM can capture both forward and reverse information at the same time, making the utilization of text information more comprehensive and the effect better. Exemplarily, in this embodiment, a linear layer can also be added after the final output layer of the BiLSTM. This linear layer is used to project the hidden layer output result generated by the BiLSTM into an interval with a certain preset expression tag feature meaning, and finally combined with the CRF, and the transition matrix is optimized through the Viterbi algorithm (or other dynamic programming algorithms), so as to finally predict the named entity word related to the target video.

[0073] In one example, performing word matching on the preset keyword library based on the first keyword, and judging whether the semantic text is related to the target video according to the word matching result may include:

[0074] Respectively determine the person name, organization name, and event name included in the first keyword;

[0075] Determine whether keywords corresponding to the person name, the organization name, and the event name can be respectively matched in the keyword library. If they can be matched, it is determined that the semantic text is relevant to the target video;

[0076] If they cannot be matched, based on the first keyword, a search engine is called to retrieve in the full text of the plot corresponding to the target video, so as to determine whether the semantic text is relevant to the target video according to the retrieval results.

[0077] In another example, keywords with a similarity higher than a preset similarity threshold to the first keyword can also be directly matched in the keyword library. When the number of keywords with a similarity higher than the preset similarity threshold is greater than the preset quantity threshold, it can be determined that the semantic text is relevant to the target video.

[0078] Use a pre-trained question filtering model to screen the semantic text, so as to determine whether the semantic text is relevant to the target video according to the screening results, where the question filtering model is trained on a first pre-training model based on the first text content of the target video and the text content of other videos, and the other videos are different from the target video.

[0079] It should be noted that the first text content in this embodiment may refer to the first sentence of the first preset number of sentences randomly extracted from the plot text of the target video, and the first sentence contains the plot text content of the named entity and is used as a positive training sample; the text content of the other video may refer to the second sentence of the second preset number of sentences randomly extracted from the plot text of the other video, and the second sentence contains the plot text content of the named entity and is used as a negative training sample.

[0080] In an example, the first pre-training model can generate text word vectors based on the Bert-Large-Chinese pre-training model, and on this basis, build a deep learning fully connected network layer as a fine-tuning model layer. The number of input and output nodes is 1024 and 2 respectively. 1024 is the word embedding dimension data output by the BERT model, and 2 is the result data. The corresponding meanings are: 1 is a question related to this drama, and 0 is not a question related to this drama. BERT is a pre-trained language model based on the Transformer architecture. The BERT model performs well in multiple natural language processing tasks, including text classification, named entity recognition, and question answering systems, etc. Based on its bidirectional encoding ability, it can utilize all the context information in the text, thus having very strong capabilities in semantic understanding.

[0081] In an alternative implementation manner, the method of finding whether there is a target answer comment in the existing comments of the target video based on the question related to the target video includes:

[0082] Input the problem into a pre-trained language model to generate a target answer. Among them, the language model includes a first language model and / or a second language model. The first language model is obtained by training a second pre-trained model based on the second text content of the target video. The second language model includes a large language model. Exemplarily, the second pre-trained model may include Chinese-Macbert-Base, which is a BERT model for text error correction trained under Chinese corpora.

[0083] Perform text similarity matching between the existing comments and the target answer, and based on the text similarity matching result, check whether there is a target answer comment in the existing comments.

[0084] In one example, the training method of the first language model may include: using the second text content of the target video (including the main plot summary of the target video, the plot text of each episode, the introduction of characters in the drama, and / or the relationship between tasks in the drama, etc.) as training corpus, and fine-tuning the second pre-trained model through a preset Python library (such as the Transformers library of Hugging Face) so that the fine-tuned second pre-trained model can answer given questions. Among them, the fine-tuning training can be implemented through the TrainingArguments class and the Trainer class (these two classes are the core components in the Transformers library of Hugging Face for model training, evaluation, and prediction).

[0085] In some cases, when the language model is the first language model, the first language model can be called to generate an answer to the input question to obtain a target answer. In this way, a target answer that is more matched with the target video can be obtained.

[0086] In some cases, when the language model is the second language model, the second language model can be called to generate an answer to the input question to obtain a target answer. In this way, the powerful semantic understanding ability of the large language model can be utilized without the need to train a model for the target video, and it can be applied to some videos with a small number of viewers or videos with incomplete plot texts to reduce the training cost.

[0087] In some cases, when the language model includes the first language model and the second language model, the first language model can be called first to generate an answer to the input question, and the second language model can be called to modify the first answer generated by the first language model to make the language in the first answer smoother.

[0088] In one example, performing text similarity matching on the existing comments and the target answer, and based on the text similarity matching result, searching for the target answer comment in the existing comments may include: converting the existing comments and the target answer into corresponding text feature vectors respectively, calculating the cosine similarity between the text feature vectors corresponding to the existing comments and the target answer respectively as the text similarity, and determining whether the maximum text similarity is greater than a preset text similarity threshold. If so, it is determined that there is a target answer comment. In this way, the existing comment corresponding to the maximum text similarity can be further used as the target answer comment.

[0089] In an alternative embodiment, performing text similarity matching on the existing comments and the target answer, and based on the text similarity matching result, searching for the target answer comment in the existing comments includes:

[0090] For any first comment among the existing comments whose publication time is within the target time range, if the text similarity between the first comment and the target answer meets the preset similarity condition, the first comment is determined as the second comment, where the target time range is determined by the time when the question is raised;

[0091] When the number of the second comments is 1, the second comment is used as the target answer comment;

[0092] When the number of the second comments is multiple, based on the barrage information, publication time, and / or keyword offset position corresponding to each second comment, evaluating the score of each second comment, and based on the score, searching for the target answer comment among the second comments;

[0093] Wherein, the keyword offset position indicates the position of the second keyword in the corresponding second comment, and the second keyword is determined by the target answer.

[0094] In one example, the text similarity meeting the preset similarity condition may include that the text similarity is greater than a preset text similarity threshold. In some cases, the preset text similarity threshold can be 60%. In addition, the similarity condition may also indicate that the text similarity is within a preset similarity range.

[0095] In one example, the text similarity between the first comment and the target answer can be measured by BERT-Whiteing (i.e., first obtaining word vectors through BERT, then performing vector whitening and eigenvalue decomposition) and the cosine of the cosine space vector. Here, if the text vectors of the first comment and the target answer are obtained through BERT, this embodiment can be used to solve the problem that the semantic similarity cannot be directly measured by the cosine similarity of the cosine space vector due to the anisotropy in the original BERT word vectors.

[0096] In some cases, the time when the problem is proposed in this embodiment may include the first proposed time and / or the second proposed time, and the publication time may include the first publication time and / or the second publication time, where:

[0097] The first proposed time can be understood as the year, month, day, hour, minute, and second when the problem is proposed. Correspondingly, the first publication time of the existing comment at this time should be understood as the year, month, day, hour, minute, and second when the corresponding existing comment is published. In this way, the selected first comment can be closer to the viewing time of the user who proposed the problem, and the real-time performance is better.

[0098] The second proposed time can be understood as the specific time point when the problem is proposed in the target video. Correspondingly, the second publication time of the existing comment at this time should be understood as a specific time point when the target video being watched by the relevant user is playing when the corresponding existing comment is published. For example, when the problem is proposed / the corresponding existing comment is published, the target video is playing to the 40th second, then the proposed time / publication time is the 40th second. In this way, the correlation between the selected first comment and the plot indicated by the proposed time can be greatly improved.

[0099] In an example combining the above situations, the target time range can indicate the same time point as the proposed time, or a preset duration range before the proposed time, or a preset duration range before and after the proposed time.

[0100] In one example, based on the barrage information, publication time, and / or keyword offset position corresponding to each second comment, evaluating the score of each second comment, and finding the target answer comment among each second comment based on the score may include: using a pre-configured scoring table to respectively determine the preset scores of the barrage information, publication time, and / or keyword offset position, and performing weighted summation on the preset scores to obtain the score of each second comment.

[0101] In an optional embodiment, the bullet screen information includes at least one of the following: the amount of bullet screens in the target video when the second comment is displayed, the number of bullet screen likes for the second comment, the number of bullet screen comments for the second comment, and the bullet screen similarity between the second comment and other simultaneously displayed bullet screens;

[0102] When the number of words in the target answer is less than the preset word count threshold, the second keyword is a word that meets the preset keyword type among the words obtained by performing word segmentation on the target answer;

[0103] When the number of words in the target answer is not less than the preset word count threshold, the second keyword is a word that meets the preset keyword type among the words obtained by performing word segmentation on the refined answer, and the refined answer is obtained by generating a text summary of the target answer.

[0104] Exemplarily, the preset word count threshold can be 10; the preset keyword type can include at least one of the following: nouns, adjectives, verbs; text summary generation can be performed through the mengzi-t5-base model.

[0105] It should be noted that the other bullet screens simultaneously displayed with the second comment in this embodiment refer to all other bullet screens displayed in the target video during the process from when the bullet screen corresponding to the second comment starts to be displayed until it stops being displayed in the target video.

[0106] The following is an example of how to use the bullet screen information, publication time, and / or keyword offset position corresponding to the second comment to evaluate the score of each second comment. It should be understood that before the evaluation, each second comment needs to be initialized to the same initial score:

[0107] 1. When evaluating using bullet screen information:

[0108] The bullet screen information may include the size of the bullet screens on the current screen when the second comment is displayed. Here, the score can be determined according to the size of the bullet screens on the current screen. If the bullet screen volume is large, the bullet screen score of the short text in the second comment is higher; if the bullet screen volume is small, the bullet screen score of the long text in the second comment is higher, and the increased score is calculated and added to the initial bullet screen score. Among them, the judgment of long text and short text can include: processing the second comment through a normalization function such as sigmod to obtain the text length of the second comment, and taking the text with a text length greater than the text length threshold as long text, otherwise as short text.

[0109] The bullet screen information may include the number of comments and / or likes of the second comment bullet screen. Here, the corresponding comment score and / or like score can be configured according to the numerical size of the number of comments and / or likes.

[0110] The barrage information may also include the similarity between the second comment barrage and other barrages displayed simultaneously. Here, the lower the similarity, the higher the additional value score. In this way, it is possible to prevent similar barrages from appearing simultaneously in the answer barrage area and the non-answer barrage area. The specific implementation scheme can follow the cosine similarity calculation scheme mentioned above to calculate the average similarity between each candidate barrage text and the remaining candidate barrage texts. Since the average similarity is a decimal between 0 and 1, it can directly participate as the value-added score. Considering the impact of this item on the accuracy of the final answer, a weight system can be multiplied by the value-added score to adjust its importance.

[0111] In addition, when evaluating the barrage score of the corresponding second comment according to the barrage information, weights can be configured for each parameter included in the barrage information for weighted calculation to obtain the final score determined by the second comment based on the barrage information (that is, the score to be added to each second comment).

[0112] 2. When using the keyword offset position:

[0113] The earlier the position of any second keyword indicated by the keyword offset position in the corresponding second comment, the higher the score of the second keyword in the corresponding second comment (for example, the offset position scores that can be obtained corresponding to different positions can be pre-configured) to facilitate users to read the question answer more quickly. In this way, the score determined by a second comment based on the keyword offset position can be obtained by weighted summing the scores of each second keyword in the second comment, so as to adjust the importance of different second keywords. In addition, the offset position scores corresponding to different second comments can be normalized to obtain the final value-added scores corresponding to each second comment (that is, the scores to be added to each second comment).

[0114] Among them, the normalization process can be represented by the following formula:

[0115] weight*(1 - Sigmod(ln(mean(offset))))

[0116] In the formula, offset represents the score, mean() represents the average value function, ln() represents the natural logarithm function, Sigmod() represents the Sigmoid function (S function, that is, the activation function of the neural network), which is used to map any real number to between 0 and 1, and weight represents the preset weight.

[0117] 2.1. The specific example is as follows:

[0118] For ease of understanding, here it is taken as an example that the number of words in the target answer is less than the preset word count threshold.

[0119] Question: Who caught Person A in the target video?

[0120] Target answer: General Person C under Person B.

[0121] Each second comment: (1) Persons D and E also stood by and watched. Person A is about to be caught by Person B now. (2) Person A was too arrogant. Otherwise, Person B wouldn't have beheaded Person A after catching him. (3) Now he's going to be caught by Person C, who is under Person B.

[0122] Each second keyword obtained after segmenting the target answer: Person B, under, general, Person C.

[0123] Positions of each second keyword indicated by the keyword offset position in each second comment (the words in quotes are the second keywords): (1) Persons D and E also stood by and watched. Person A is about to be caught by "Person B" now. (2) Person A was too arrogant. Otherwise, "Person B" wouldn't have beheaded Person A after catching him. (3) Now he's going to be caught by "Person B"'s "under" "Person C".

[0124] It is not difficult to understand that in addition to more complete information, the second keyword in the second comment (3) appears earlier, enabling users to obtain the key information of the answer faster.

[0125] 3. When using the publication time corresponding to the second comment:

[0126] The later the publication time of the second comment, the higher the score determined based on the publication time for this second comment. In this embodiment, some viewpoints closer to the time when the question was raised can be pushed to the user to increase real-time performance.

[0127] Based on the above 1 - 3, the score of each second comment evaluated can be obtained by performing a weighted sum of the score determined based on the bullet screen information for each second comment, the score determined based on the publication time for each second comment, and / or the score determined based on the keyword offset position for each second comment.

[0128] Exemplarily, refer to Figure 3 , Figure 3 the bullet screens within the dashed box in, that is, the highlighted bullet screens determined by the target answer comment included in the target content in this embodiment (it should be understood that the highlighting can be processed in any color). It can be seen here that there doesn't necessarily need to be a significant difference between the style of this highlighted bullet screen and the general bullet screen style, and only some highlighting styles need to be given.

[0129] In addition, since the target answer comment can itself be selected from the bullet comments, when multiple people ask the same question, when other users see the highlighted bullet comment corresponding to this target answer comment, see Figure 3 , the question and the number of people who asked the question can also be displayed on this highlighted bullet comment, but the operation of guiding the evaluation of the target answer comment is usually only open to the user who raised this question this time.

[0130] In an optional implementation manner, the method further includes:

[0131] In the case where there is no such target answer comment, in the target video displayed on the user device that raised the question, the target answer is displayed in a highlighted bullet comment form and maintained for a preset duration, and stops displaying the target answer after reaching the preset duration.

[0132] After displaying the target answer or the target content, an evaluation interface is displayed, and a user operation acting on the evaluation interface is obtained, where the user operation indicates the evaluation score of the user for the displayed target answer or the displayed target content.

[0133] In this embodiment, when a suitable target answer comment cannot be found in the existing comments, the generated target answer can be displayed in the form of a bullet comment to inform the user of the answer to the question they raised, without the user having to search by themselves and interrupt the movie-watching process. In addition, the target answer only needs to be displayed in the target video displayed on the user device that raised the question, which is equivalent to not adding the target answer to the bullet comment list of the target video, which can prevent other users from seeing the target answer and thus affecting the real communication environment of other users, and also prevent spoiling for other users and thus affecting the movie-watching experience. Further, the display of the target answer will only last for a preset duration, so as to also prevent the user device that raised the question from seeing the generated target answer again when playing the target video later and affecting the real communication environment.

[0134] In an example, see Figure 3 , the displayed evaluation interface may be provided with a like button and a dislike button, where, when the user operation indicates pressing the like button, the evaluation score of the displayed target answer or the displayed target content is increased; when the user operation indicates pressing the dislike button, the evaluation score of the displayed target answer or the displayed target content is decreased.

[0135] Optionally, in combination with the above example, see Figure 4For the displayed target answer or the displayed target content whose evaluation score is higher than the first preset evaluation score threshold, it can be combined with the corresponding question to form a question-answer combination, which is used as model training corpus to update the language model, so as to achieve a virtuous cycle of improving the accuracy of machine learning question answering. On the other hand, the next time a similar question is encountered (for example, after collecting the question raised by the user, the similarity between the raised question and the questions in the question-answer combination is matched, and when the matching degree is higher than the matching degree threshold, it is determined as a similar question corresponding to the question in the question-answer combination), the answer with good evaluation effect (for example, the evaluation score is higher than the second preset evaluation score threshold) can be directly fed back to the user.

[0136] Second aspect, correspondingly, the embodiment of the present application further provides a question-answer interaction device, which can implement all the processes of the question-answer interaction method provided in the above embodiment.

[0137] See Figure 5 , which shows a schematic structural diagram of the question-answer interaction device provided by the embodiment of the present application, including:

[0138] A search module 503, configured to search whether there is a target answer comment in the existing comments of the target video based on a question related to the target video;

[0139] A target content display module 504, configured to display target content corresponding to the target answer comment when there is the target answer comment, where the target content includes a highlighted barrage of the target answer comment.

[0140] In an optional implementation manner, the device further includes:

[0141] A user instruction processing module 501, configured to obtain a user instruction and convert the user instruction into semantic text before searching whether there is a target answer comment in the existing comments of the target video;

[0142] A relevance judgment module 502, configured to judge whether the semantic text is relevant to the target video when it is determined that the semantic text is an interrogative sentence, and if so, use the semantic text as the question;

[0143] Wherein, the judgment of whether the semantic text is relevant to the target video includes any one of the following:

[0144] Identify the semantic text to obtain a first keyword, perform word matching based on the first keyword in a preset keyword library, and judge whether the semantic text is relevant to the target video according to the word matching result, where the keyword library is used to store keywords corresponding to the target video;

[0145] Use the pre-trained question filtering model to screen the semantic text, and determine whether the semantic text is relevant to the target video according to the screening result. Among them, the question filtering model is trained from a first pre-trained model based on the first text content of the target video and the text content of other videos, and the other videos are different from the target video.

[0146] In an optional implementation manner, finding whether there is a target answer comment in the existing comments of the target video based on the question related to the target video includes:

[0147] Input the question into the pre-trained language model for answer generation to obtain the target answer. Among them, the language model includes a first language model and / or a second language model. The first language model is trained from a second pre-trained model based on the second text content of the target video, and the second language model includes a large language model;

[0148] Perform text similarity matching between the existing comments and the target answer, and find whether there is the target answer comment in the existing comments according to the text similarity matching result.

[0149] In an optional implementation manner, performing text similarity matching between the existing comments and the target answer, and finding whether there is the target answer comment in the existing comments according to the text similarity matching result includes:

[0150] For any first comment in the existing comments whose publication time is within the target time range, if the text similarity between the first comment and the target answer meets the preset similarity condition, then determine the first comment as the second comment, where the target time range is determined by the time when the question is proposed;

[0151] When the number of the second comments is 1, use the second comment as the target answer comment;

[0152] When the number of the second comments is multiple, evaluate the score of each second comment based on the barrage information, publication time, and / or keyword offset position corresponding to each second comment, and find the target answer comment among the second comments based on the score;

[0153] Among them, the keyword offset position indicates the position of the second keyword in the corresponding second comment, and the second keyword is determined by the target answer.

[0154] In an alternative embodiment, the bullet screen information includes at least one of the following: the amount of bullet screens in the target video when the second comment is displayed, the number of bullet screen likes of the second comment, the number of bullet screen comments of the second comment, and the bullet screen similarity between the second comment and other simultaneously displayed bullet screens;

[0155] When the number of characters of the target answer is less than the preset character threshold, the second keyword is a word that meets the preset keyword type among the words obtained by performing word segmentation on the target answer;

[0156] When the number of characters of the target answer is not less than the preset character threshold, the second keyword is a word that meets the preset keyword type among the words obtained by performing word segmentation on the refined answer, and the refined answer is obtained by generating a text summary of the target answer.

[0157] In an alternative embodiment, the device further includes:

[0158] A target answer display module, configured to, in the absence of a comment on the target answer, display the target answer in a highlighted bullet screen form in the target video displayed on the user device that poses the question and maintain it for a preset duration, and stop displaying the target answer after reaching the preset duration.

[0159] An evaluation interface processing module, configured to display an evaluation interface after the target answer or the target content is displayed, and obtain a user operation on the evaluation interface, where the user operation indicates an evaluation score of the user for the displayed target answer or the displayed target content.

[0160] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the question-and-answer interaction method described in any one of the above are implemented.

[0161] In a fourth aspect, an embodiment of the present application provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the question-and-answer interaction method described in any one of the above are implemented.

[0162] In a fifth aspect, an embodiment of the present application provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, the steps of the question-and-answer interaction method described in any one of the above are implemented.

[0163] See Figure 6, the computer device of this embodiment includes: a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601, such as a question-and-answer interaction program. When the processor 601 executes the computer program, it implements the steps in each of the above-described embodiments of the question-and-answer interaction method, such as Figure 1 the steps S101 - S104 shown.

[0164] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 602 and executed by the processor 601 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device.

[0165] The computer device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor 601 and a memory 602. Those skilled in the art can understand that the schematic diagram is only an example of the computer device and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may further include input / output devices, network access devices, a bus, etc.

[0166] The processor 601 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor 601 may also be any conventional processor, etc. The processor 601 is the control center of the computer device and connects various parts of the entire computer device through various interfaces and lines.

[0167] The memory 602 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 602, and by invoking the data stored in the memory 602, the processor 601 realizes various functions of the computer device. The memory 602 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 602 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0168] Among them, if the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 601, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0169] In summary, the embodiments of this application have at least the following beneficial effects:

[0170] By adopting the embodiments of the present application, by searching for a target answer comment in the existing comments of the target video based on a problem related to the target video; in the case where the target answer comment exists, displaying the target content corresponding to the target answer comment, it is possible to directly search for whether there is a target answer comment in the existing comments of the target video after the user raises a question related to the target video, and when the search is successful, display it in the form of a highlighted bullet screen, thereby avoiding interrupting the user's viewing process while also answering the questions raised by the user during the viewing process to improve the user's viewing experience.

[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary hardware platform, and of course, it can also be implemented entirely by hardware. Based on such an understanding, all or part of the technical solution of the present application that contributes to the background technology can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0172] The above is the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.

Claims

1. A question-and-answer interaction method, characterized in that Including: Based on a problem related to the target video, check whether there is a target answer comment in the existing comments of the target video; When there is the target answer comment, display the target content corresponding to the target answer comment.

2. The Q&A interaction method according to claim 1, wherein Before checking whether there is a target answer comment in the existing comments of the target video, the method further includes: Obtain a user instruction and convert the user instruction into semantic text; When it is determined that the semantic text is an interrogative sentence, determine whether the semantic text is related to the target video. If it is related, use the semantic text as the problem; Among them, determining whether the semantic text is related to the target video includes any one of the following: Identify the semantic text to obtain a first keyword, perform word matching in a preset keyword library based on the first keyword, and determine whether the semantic text is related to the target video according to the word matching result. Among them, the keyword library is used to store keywords corresponding to the target video; Use a pre-trained question filtering model to screen the semantic text to determine whether the semantic text is related to the target video according to the screening result. Among them, the question filtering model is trained from a first pre-trained model based on the first text content of the target video and the text content of other videos, and the other videos are different from the target video.

3. The Q&A interaction method according to claim 1 or 2, characterized in that, Based on the problem related to the target video, checking whether there is a target answer comment in the existing comments of the target video includes: Input the problem into a pre-trained language model to generate an answer to obtain a target answer. Among them, the language model includes a first language model and / or a second language model. The first language model is trained from a second pre-trained model based on the second text content of the target video, and the second language model includes a large language model; Perform text similarity matching between the existing comments and the target answer, and check whether there is the target answer comment in the existing comments according to the text similarity matching result.

4. The Q&A interaction method according to claim 3, characterized in that, Performing text similarity matching between the existing comments and the target answer, and checking whether there is the target answer comment in the existing comments according to the text similarity matching result includes: For any first comment among the existing comments whose publication time is within the target time range, if the text similarity between the first comment and the target answer meets the preset similarity condition, determine the first comment as a second comment. Among them, the target time range is determined by the time when the problem is proposed; When the number of the second comments is 1, use the second comment as the target answer comment; When the number of the second comments is multiple, evaluate the score of each second comment based on the barrage information, publication time, and / or keyword offset position corresponding to each second comment, and find the target answer comment from each second comment based on the score; Among them, the keyword offset position indicates the position of a second keyword in the corresponding second comment, and the second keyword is determined by the target answer.

5. The Q&A interaction method according to claim 4, wherein: The bullet screen information includes at least one of the following: the amount of bullet screens in the target video when the second comment is displayed, the number of bullet screen likes for the second comment, the number of bullet screen comments for the second comment, the bullet screen similarity between the second comment and other bullet screens displayed simultaneously therewith; When the number of words of the target answer is less than a preset word count threshold, the second keyword is a word that meets a preset keyword type among the words obtained by performing word segmentation on the target answer; When the number of words of the target answer is not less than the preset word count threshold, the second keyword is a word that meets a preset keyword type among the words obtained by performing word segmentation on a refined answer, and the refined answer is obtained by generating a text summary of the target answer.

6. The Q&A interaction method according to claim 3, wherein The method further includes: In the case where there is no target answer comment, in the target video displayed on the user device that poses the question, the target answer is displayed in a highlighted bullet screen form and maintained for a preset duration, and the display of the target answer stops after the preset duration is reached. After the target answer or the target content is displayed, an evaluation interface is displayed, and a user operation acting on the evaluation interface is obtained, wherein the user operation indicates an evaluation score of the user for the displayed target answer or the displayed target content.

7. A question-and-answer interaction device, characterized in that, including: A search module, configured to search whether there is a target answer comment in the existing comments of the target video based on a question related to the target video; A target content display module, configured to display target content corresponding to the target answer comment in the case where there is the target answer comment.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the Q&A interaction method according to any one of claims 1-6.

9. A computer program product, comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, implement the Q&A interaction method according to any one of claims 1-6.

10. A computer device, characterized in that, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, the Q&A interaction method according to any one of claims 1-6 is implemented.

Citation Information

Cited By

  • System and method for short text matching

    US12585873B1