Information processing method and device, electronic equipment and readable storage medium
By generating questions and answers matching user interests during video playback, the problem that users cannot obtain information efficiently under long video playback is solved, and the user's viewing experience and information acquisition efficiency are improved.
Patent Information
- Application Number
- CN202510280021.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-18
AI Technical Summary
The video playback form of long videos makes it impossible for users to obtain information efficiently, reducing users' viewing experience. The existing technology such as video segmentation and video content thumbnails increase workload or requires frequent operations by users, making it difficult to meet users' needs to obtain information efficiently within fragmented time.
During the video playback process, by obtaining the user's user portrait, a question-and-answer question matching the video content and user interests is generated, candidate association questions are automatically generated, and target association questions are determined based on user selection until the question generation conditions are met, and a question-and-answer content that meets user needs is provided.
It improves users' viewing experience in long videos, helps users to understand the video content efficiently, meets the information acquisition needs in fragmented time, and reduces the frequency of users' operations.
Smart Images

Figure CN120336594A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to an information processing method, apparatus, electronic device, and readable storage medium. Background Art
[0002] The accelerating pace of life has fragmented people's spare time. To meet people's information acquisition needs during fragmented time, the number of short videos has increased rapidly. However, the video playback form of short videos has led to an increasing demand for information acquisition efficiency. Therefore, video playback forms such as long videos, from which users cannot quickly obtain information, will reduce the user's viewing experience. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide an information processing method, apparatus, electronic device, and readable storage medium to automatically generate questions and answers that meet the user's needs and have a high relevance to the currently played video during video playback, helping users understand the video content and thus enhancing the user's viewing experience.
[0004] In a first aspect, embodiments of the present invention provide an information processing method, the method comprising:
[0005] During video playback, obtain the user profile of the current user;
[0006] In each question-and-answer round, generate and output to the current user candidate associated questions of the target video in the current question-and-answer round according to the video content information of the target video, the user profile, and the answers to the target associated questions in the previous question-and-answer round, and determine the candidate associated question selected by the current user as the target associated question of the current question-and-answer round until the question generation condition is not met, where the target video is the video selected by the current user.
[0007] In a second aspect, embodiments of the present invention provide an information processing apparatus, the apparatus comprising:
[0008] A profile acquisition unit, configured to obtain the user profile of the current user during video playback;
[0009] An iteration unit, configured to generate and output to the current user candidate associated questions of the target video in the current question-and-answer round according to the video content information of the target video, the user profile, and the answers to the target associated questions in the previous question-and-answer round, and determine the candidate associated question selected by the current user as the target associated question of the current question-and-answer round until the question generation condition is not met, where the target video is the video selected by the current user.
[0010] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect.
[0011] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method as described in the first aspect is implemented.
[0012] In a fifth aspect, an embodiment of the present invention provides a computer program product, where the computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the method as described in the first aspect is implemented.
[0013] In the embodiment of the present invention, a user portrait of a user is obtained during the process of the user watching a video, and in each Q&A round, candidate associated questions for the current round are generated and output to the user according to the content information of the video selected by the user, the user portrait of the user, and the answers to the target associated questions related to the video selected by the user in the previous Q&A rounds, and then the candidate associated questions selected by the user are determined as the target associated questions for the current round until the question generation condition is not met. The embodiment of the present invention can automatically generate Q&A that meets the user's needs and has a high relevance to the currently played video during the video playback process, helping the user better understand the video content, and thus can improve the user's viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0015] Figure 1 is a flowchart of the information processing method according to an embodiment of the present invention;
[0016] Figure 2 is a flowchart of the information processing method according to an embodiment of the present invention;
[0017] Figure 3 is a data flowchart of the information processing method according to an embodiment of the present invention;
[0018] Figure 4 is a flowchart of the information processing method according to an embodiment of the present invention;
[0019] Figure 5 is a schematic diagram of the judgment process for generating candidate associated questions in an embodiment of the present invention;
[0020] Figure 6 is a schematic diagram of the information processing device according to an embodiment of the present invention;
[0021] Figure 7 It is a schematic diagram of the electronic device according to an embodiment of the present invention. Specific embodiments
[0022] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, components and circuits are not described in detail.
[0023] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0024] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document should be interpreted as having an inclusive meaning rather than an exclusive or exhaustive meaning; that is, it is the meaning of "including but not limited to".
[0025] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0026] For the solutions described in this specification and embodiments, if they involve personal information processing, they will all be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.
[0027] In daily life, people will use fragmented time to obtain information. Fragmented time is usually from a few minutes to dozens of minutes, and is mainly distributed during commuting, lunch breaks, queuing, etc., which makes people generally tend to be able to obtain information efficiently in a short time. To meet people's information acquisition needs, the number of short videos has increased rapidly. Compared with short videos, the content of long videos is more rich and complete, and can provide users with more abundant information. However, long videos have higher requirements for the playing duration, which causes some users to be unable to spend a long time playing long videos completely, resulting in users being unable to obtain all the information from long videos, thus reducing the user's viewing experience. At the same time, the user's willingness to watch long videos decreases, which will significantly reduce the play volume of long videos and is not conducive to the information dissemination of long videos.
[0028] To improve users' willingness to watch long videos, the prior art mainly adopts the methods of video segmentation and video content thumbnails to assist users in understanding the video content. Video segmentation mainly refers to splitting a long video into multiple segments and naming each segment, such as "Product 1 Evaluation", "Product 2 Evaluation", "Product 3 Evaluation", "Comprehensive Comparison", etc. During the process of watching a long video, users can click on the progress marker of any segment to jump to play the long video. However, this method will undoubtedly increase the workload of video producers, and the number of words in the segment names is usually small, and it is usually difficult to cover all the information in the segment. Therefore, users may miss the information they are interested in. Video content thumbnails mainly refer to displaying each frame of the long video through thumbnails. During the process of watching a long video, users can position to the specified frame by dragging the playback progress bar and preview the video content through the thumbnails. However, this method usually requires users to frequently drag the playback progress bar to locate the content they are more interested in, and the screen positioning effect presented by the long video will be negatively affected to varying degrees due to the change of the video playback progress, which is not conducive to users to obtain information.
[0029] To solve the above problems, embodiments of the present invention propose an information processing method, device, electronic device and readable storage medium to automatically generate questions and answers that meet the user's needs and have a high relevance to the currently played video during the video playback process, helping users understand the video content, thereby improving the user's viewing experience.
[0030] The following is illustrated through method embodiments. Figure 1 It is a flowchart of the information processing method of embodiments of the present invention. As Figure 1 shown, the method of this embodiment includes the following steps:
[0031] Step S101, during the video playback process, obtain the user profile of the current user.
[0032] In daily life, users can browse various existing video playback carriers such as video playback applications and video playback websites through terminals and select videos in the video playback carriers to play. The user profile of the current user can reflect the current user's preferences for different types of videos to a certain extent. Therefore, this embodiment can obtain the user profile of the current user during the video playback process.
[0033] The user profile of the current user mainly can include the personal attribute information, behavior information, social attribute information, etc. of the current user. Among them, the personal attribute information can include the age, gender, work status, main activity area, educational background, etc. of the current user. The behavior information can include the historical browsing records, historical query records, video sharing quantity (or frequency), video like quantity (or frequency), video comment quantity (or frequency), video collection frequency, etc. of the current user. The social attribute information can include the social behaviors of the current user, such as the video types shared, liked, commented on, collected in the video playback application, the types of followed accounts, etc., and the social relationships, such as the mutually followed accounts, bound accounts, types of social groups or social organizations joined, etc. According to the actual settings, the user profile can also include other information, and this embodiment does not limit this.
[0034] Step S102, in each question-and-answer round, generate and output to the current user the candidate associated questions of the target video in the current question-and-answer round according to the video content information of the target video, the user profile, and the answers to the target associated questions in the previous question-and-answer rounds, and determine the candidate associated question selected by the current user as the target associated question in the current question-and-answer round until the question generation condition is not satisfied.
[0035] This embodiment can determine the video selected by the user as the target video, generate multiple candidate associated questions according to the video content information of the target video, the user profile of the current user, and the answers to the target associated questions in the previous question-and-answer rounds, and output each generated candidate associated question to the current user. When the current user selects any one of the candidate associated questions, this embodiment can generate the answer to the candidate associated question. Furthermore, this embodiment can determine the candidate associated question selected by the user as the target associated question in the current question-and-answer round until the question generation condition is not satisfied.
[0036] According to the actual settings, this embodiment can output the candidate associated questions to the current user through various existing methods, such as displaying a list of associated questions of the target video on the terminal screen, where the list of associated questions includes the candidate associated questions, or can also broadcast each candidate associated question through the voice broadcast module, and this embodiment does not limit this.
[0037] Optionally, in order to further improve the matching degree between the candidate associated questions and the target video, this embodiment can also combine the target associated questions in the previous round when generating the candidate associated questions, that is, in each question-and-answer round, generate the candidate associated questions in the current round according to the video content information of the target video, the user profile of the current user, the target associated questions in the previous question-and-answer rounds, and the answers to the target associated questions in the previous question-and-answer rounds.
[0038] It is easy to understand that the previous Q&A round described in this embodiment can be the previous Q&A round of the current Q&A round, or multiple previous Q&A rounds of the current Q&A round. This embodiment does not limit this.
[0039] For example: in the first Q&A round, the user selects question Q1 as the target associated question. Then, this embodiment can generate the answer to question Q1, and in the second Q&A round, generate multiple candidate associated questions according to the video content information of the target video, the user profile of the current user, and the answer to question Q1, such as question Q2, question Q3, and question Q4. If in this Q&A round, the user selects question Q3 as the target associated question, then this embodiment can generate the answer to question Q3, and in the third Q&A round, generate multiple candidate associated questions according to the video content information of the target video, the user profile of the current user, the answer to question Q1, and the answer to question Q3.
[0040] Figure 2 is a flowchart of the information processing method according to an embodiment of the present invention. As Figure 2 shown, in an optional implementation manner, this embodiment can determine the video content information of the target video through the following steps:
[0041] Step S201: Perform speech recognition on the target video to determine the speech recognition information of the target video.
[0042] This embodiment can perform speech recognition on the target video through a speech recognition model, such as a Hidden Markov Model (HMM), a Deep Neural Network (DNN), etc., to determine the speech recognition information of the target video. Further, this embodiment can extract audio information from the target video and perform speech recognition on the extracted audio information to obtain the corresponding speech recognition information.
[0043] Step S202: Determine the video content information according to the speech recognition information.
[0044] The speech recognition information of the video is usually used to explain the video content and can effectively reflect the information that the video creator wants to convey to the viewer. Therefore, in this step, the video content information of each video can be determined according to the speech recognition information.
[0045] Optionally, the speech recognition information of the video can be directly determined as the video content information, or key information can be extracted from the speech recognition information through various existing methods as the video content information, and this embodiment does not limit this. For example, based on a preset statistical method, such as TF-IDF (Term Frequency–Inverse Document Frequency), bag of words, etc., the word frequency of each word in the speech recognition information can be statistically calculated, and the video content information can be determined according to the word frequency of each word; alternatively, the terminal or the server can also extract the video content information from the speech recognition information based on a natural language processing (NLP, Natural Language Processing) model, such as a model based on a neural network, a model based on an attention mechanism, a large language model, etc.
[0046] Optionally, if this embodiment extracts the video content information from the speech recognition information through a large language model, the prompt information corresponding to the large language model can be determined according to the speech recognition information and a pre-set question template, and this prompt information can be used as the input of the large language model to obtain the corresponding video content information.
[0047] For example, if the question template is "The following is the subtitle of a video. Please help me summarize in detail the content expressed by this video: {subtitle content}", "{subtitle content}" can be replaced with the speech recognition information of the target video to obtain the prompt information of the large language model, and then the corresponding video content information can be obtained.
[0048] According to actual requirements, this embodiment can also combine other information, such as the image recognition information obtained by performing image recognition on the target video, to obtain the video content information of the target video. Image recognition of the target video can be performed through various existing image recognition methods, such as models based on a convolutional neural network (CNN) architecture, models based on a recurrent neural network (RNN) architecture, etc., and this embodiment does not limit this.
[0049] Optionally, in step S102, in this embodiment, the initial associated questions of the target video may be generated according to the video content information of the target video and the user portrait of the current user and output to the current user. In the first Q&A round, the initially associated question selected by the current user is determined as the target associated question in the previous Q&A round. Then, in each Q&A round, according to the video content information of the target video, the user portrait of the current user, and the answers to the target associated questions in the previous Q&A round, the candidate associated questions of the target video in the current Q&A round are generated and output to the current user, and the candidate associated question selected by the current user is determined as the target associated question in the current Q&A round until the question generation condition is not met. This embodiment can generate corresponding candidate associated questions for the current user in a progressive manner to further match the user's viewing needs, so as to effectively improve the video viewing experience of the current user.
[0050] Among them, the question generation condition may be to detect the selection operation of the current user for the candidate associated question, and the total number of Q&A rounds does not exceed the preset maximum number of rounds. That is to say, if the current user does not select a candidate associated question as the target associated question, it means that the current user does not have the need to understand the target video by asking questions; if the total number of Q&A rounds exceeds the preset maximum number of rounds, it means that the target associated questions selected by the current user have basically covered the main content of the target video, and there is no need to generate candidate associated questions anymore. Therefore, if the question generation condition is not met, this embodiment may no longer generate new Q&A rounds, that is, generate new candidate associated questions.
[0051] Figure 3 is the data flow diagram of the information processing method of the embodiment of the present invention. As Figure 3 shown, in the initial stage, this embodiment may call the video service according to the video identifier 31 of the target video to obtain the video content information 33 of the target video, and call the user portrait service according to the user identifier 32 of the current user to obtain the user portrait 34 of the current user. Then, according to the video content information 33 of the target video and the user portrait 34 of the current user, the prompt information 35 of the large language model is generated, and thus the large language model service is called according to the prompt information 35 to obtain the candidate associated questions 36 of the target video.
[0052] Figure 4 is the flowchart of the information processing method of the embodiment of the present invention. As Figure 4 shown, in an alternative implementation, this embodiment may determine the answer to the target associated question through the following steps:
[0053] Step S401, perform intent recognition on the target associated question to determine the corresponding intent information.
[0054] In this embodiment, various existing methods can be used to identify the intent of the target associated problem, such as an intent recognition model, a large language model, etc., to determine the intent information of the target associated problem. Optionally, if this embodiment uses a large language model for intent recognition, the prompt information corresponding to the large language model can be determined according to the target associated problem and a pre-set question template, and this prompt information can be used as the input of the large language model to obtain the corresponding intent information.
[0055] For example, if the question template is "Please identify the intent of the following question. The intent options include: inquiry, consultation, appearance, interior, and evaluation {question content}", the "{question content}" can be replaced with the target associated problem, such as "How about the evaluation of the vehicle", to obtain the prompt information of the large language model as "Please identify the intent of the following question. The intent options include: inquiry, consultation, appearance, interior, and evaluation, question content: How about the evaluation of the vehicle", and then the corresponding intent information "evaluation" can be obtained.
[0056] Step S402, determine the video association information according to the intent information and the video content information.
[0057] After determining the intent information of the target associated problem, this embodiment can combine the video content information to determine the video association information of the target video. Specifically, if the target video is used to introduce the advantages and disadvantages of vehicles of different brands (or models), the target associated problem is "What is the on-road price of vehicle model A", and the intent information of the target associated problem is an inquiry, then the prices of vehicles of each brand (or model) involved in the target video can be queried as the video association information of the target video.
[0058] Step S403, based on the video association information, determine the answer based on a predetermined large language model.
[0059] After determining the video association information, this embodiment can determine the answer to the target associated problem based on a predetermined large language model. Specifically, this embodiment can determine the prompt information corresponding to the large language model according to the target associated problem, the video association information, the intent information of the target associated problem, and a pre-set question template, and use this prompt information as the input of the large language model to obtain the answer to the target associated problem.
[0060] For example, if the question template is "Please generate an answer to the following user question in combination with the video content description and the latest materials. Latest materials: {material content}", user question: {question content}, question intent: {intent information}, then this embodiment can replace "{material content}" with the video association information, replace {question content} with the target associated problem of the corresponding question and answer round, and replace {intent information} with the intent information of the target associated problem of the corresponding question and answer round to obtain the prompt information of the large language model, and then obtain the answer to the target associated problem of the corresponding question and answer round.
[0061] After determining the answer to the target associated question, in each Q&A round, this embodiment can determine the prompt information corresponding to the large language model based on the video content information of the target video, the user profile of the current user, the target associated questions in the previous Q&A rounds, the answers to the target associated questions in the previous Q&A rounds, and the preset question templates, and use this prompt information as the input of the large language model to obtain the candidate associated questions for the current Q&A round.
[0062] For example, the question template is "Please combine the video content description, user feature description, the user's recent several questions and the system answer content to output 5 questions that the user may be interested in. Video content description: {video content description}", user feature description: {user profile}, user question: {question content}, system answer to the question: {question answer}. Then this embodiment can replace "{video content description}" with the video content information, replace {user profile} with the user profile of the current user, replace {question content} with the target associated questions in the previous Q&A rounds, and replace {question answer} with the answers to the target associated questions in the previous Q&A rounds to obtain the prompt information of the large language model, thereby generating the candidate associated questions for the current Q&A round.
[0063] Figure 5 It is a schematic diagram of the judgment process for generating candidate associated questions in an embodiment of the present invention. As Figure 5 shown, in step S501, this embodiment can generate initial associated questions based on the video content information of the target video and the user profile of the current user. After the user selects any initial associated question, step S502 is executed to generate and output the answer to the target associated question to the current user, and then step S503 is executed to determine whether the question generation condition is met, that is, whether it is detected that the current user has selected an operation for the candidate associated question, and whether the total number of Q&A rounds does not exceed the preset maximum number of rounds. If so, step S504 is executed to generate candidate associated questions based on the video content information of the target video, the user profile of the current user, and the answers to the target associated questions in the previous Q&A rounds, and then after the user selects any initial associated question, return to execute step S502. If not, step S505 is executed to end the generation of candidate associated questions.
[0064] In the embodiment of the present invention, the user portrait of the user is obtained during the process of the user watching a video, and in each question-and-answer round, candidate associated questions for the current round are generated and output to the user according to the content information of the video selected by the user, the user portrait of the user, and the answers to the target associated questions related to the video selected by the user in the previous question-and-answer rounds, and then the candidate associated questions selected by the user are determined as the target associated questions for the current round until the question generation condition is not satisfied. The embodiment of the present invention can automatically generate questions and answers that meet the user's needs and have a high correlation with the currently played video during the video playback process, helping the user better understand the video content, and thus can improve the user's viewing experience.
[0065] Figure 6 It is a schematic diagram of the information processing device according to the embodiment of the present invention. As Figure 6 shown, the information processing device of this embodiment includes a portrait acquisition unit 601 and an iteration unit 602.
[0066] Among them, the portrait acquisition unit 601 is used to obtain the user portrait of the current user during the video playback process; the iteration unit 602 is used to generate and output candidate associated questions for the target video in the current question-and-answer round to the current user according to the video content information of the target video, the user portrait, and the answers to the target associated questions in the previous question-and-answer rounds, and determine the candidate associated questions selected by the current user as the target associated questions for the current question-and-answer round until the question generation condition is not satisfied, and the target video is the video selected by the current user.
[0067] In the embodiment of the present invention, the user portrait of the user is obtained during the process of the user watching a video, and in each question-and-answer round, candidate associated questions for the current round are generated and output to the user according to the content information of the video selected by the user, the user portrait of the user, and the answers to the target associated questions related to the video selected by the user in the previous question-and-answer rounds, and then the candidate associated questions selected by the user are determined as the target associated questions for the current round until the question generation condition is not satisfied. The embodiment of the present invention can automatically generate questions and answers that meet the user's needs and have a high correlation with the currently played video during the video playback process, helping the user better understand the video content, and thus can improve the user's viewing experience.
[0068] Figure 7 It is a schematic diagram of the electronic device according to the embodiment of the present invention. As Figure 7As shown, the electronic device 7 is a general-purpose data processing device, which includes a general computer hardware structure, and at least includes a processor 701 and a memory 702. The processor 701 and the memory 702 are connected through a bus 703. The memory 702 is adapted to store instructions or programs executable by the processor 701. The processor 701 can be an independent microprocessor or a set of one or more microprocessors. Thus, by executing the instructions stored in the memory 702, the processor 701 executes the method flow of the embodiment of the present invention as described above to implement data processing and control of other devices. The bus 703 connects the above-mentioned multiple components together and also connects the above-mentioned components to a display controller 704, a display device, and an input / output (I / O) device 705. The input / output (I / O) device 705 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a body-sensing input device, a printer, and other devices well-known in the art. Typically, the input / output (I / O) device 705 is connected to the system through an input / output (I / O) controller 706.
[0069] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device (equipment), or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0070] The present application is described with reference to the flowcharts of methods, devices (equipment), and computer program products according to the embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.
[0071] These computer program instructions can be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the process Figure 1 the functions specified in one process or multiple processes.
[0072] These computer program instructions can also be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for implementing the functions specified in Figure 1 one process or multiple processes.
[0073] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, which is used for a computer to execute the above-mentioned partial or all method embodiments.
[0074] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by specifying relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0075] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An information processing method, characterized in that, The method includes: During video playback, obtaining the user profile of the current user; In each Q&A round, generating and outputting to the current user the candidate associated questions of the target video in the current Q&A round according to the video content information of the target video, the user profile, and the answers to the target associated questions in the previous Q&A round, and determining the candidate associated question selected by the current user as the target associated question of the current Q&A round until the question generation condition is not met, where the target video is the video selected by the current user.
2. The method according to claim 1, wherein The video content information is determined in the following manner: Performing speech recognition on the target video to determine the speech recognition information of the target video; Determining the video content information according to the speech recognition information.
3. The method according to claim 1, wherein The answer is determined in the following manner: Performing intent recognition on the target associated question to determine the corresponding intent information; Determining video association information according to the intent information and the video content information; Determining the answer based on a predetermined large language model according to the video association information.
4. The method according to claim 1, wherein The generating and outputting to the current user the candidate associated questions of the target video in the current Q&A round according to the video content information of the target video, the user profile, and the answers to the target associated questions in the previous Q&A round until the question generation condition is not met includes: Generating and outputting to the current user the initial associated questions of the target video according to the video content information and the user profile; In the first Q&A round, determining the initial associated question selected by the current user as the target associated question in the previous Q&A round; In each Q&A round, generating and outputting to the current user the candidate associated questions of the target video in the current Q&A round according to the video content information, the user profile, and the answers in the previous Q&A round, and determining the candidate associated question selected by the current user as the target associated question of the current Q&A round until the question generation condition is not met.
5. The method according to claim 1, characterized in that, The outputting to the current user the candidate associated questions of the target video in the current Q&A round includes: Displaying a list of associated questions of the target video in the current Q&A round, where the list of associated questions includes the candidate associated questions.
6. The method according to claim 1, characterized in that, The question generation condition is detecting a selection operation of the current user for the candidate associated question and the total number of Q&A rounds not exceeding a preset maximum number of rounds.
7. An information processing apparatus, characterized in that, The device includes: A profile acquisition unit for obtaining the user profile of the current user during video playback; An iterative unit, configured to generate and output to the current user candidate associated questions of the target video in the current Q&A round according to the video content information of the target video, the user portrait, and the answers to the target associated questions in the previous Q&A rounds, and determine the candidate associated question selected by the current user as the target associated question in the current Q&A round until the question generation condition is not satisfied, where the target video is the video selected by the current user.
8. An electronic device, comprising a memory and a processor, characterized in that, The memory is configured to store one or more computer program instructions, where the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.
10. A computer program product, characterized in that, The computer program product includes a computer program / instructions, and when the computer program / instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.