Method, apparatus, device, and storage medium for displaying explanation information
By providing multiple explanation videos of different lengths and allowing users to switch, the problem of users losing interest when watching long videos is solved, and the explanation effect is improved.
Patent Information
- Application Number
- CN202410424341.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2044-04-09
AI Technical Summary
Users lose interest when watching long-term explanation videos, which affects the effect of item explanation.
A number of explanation videos are provided, wherein at least two videos have different lengths and contain some of the same explanation information. The user can switch between different videos through switching commands.
It improves users' viewing flexibility and item explanation effect, and avoids the limitation that users can only watch a complete long video.
Smart Images

Figure CN118354113B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technology, and in particular to a method, device, equipment, and storage medium for displaying explanation information. Background Art
[0002] With the development of Internet technology and the growing scale of Internet users, video resources have been widely disseminated, and recommending items in the form of videos has become a common method.
[0003] In order to facilitate the recommendation of items, a video is usually recorded while the explainer is explaining the item, thereby obtaining an explanation video of the item. The explanation video can then be published so that viewers who watch the explanation video can understand the item and then purchase the item.
[0004] However, if the explanation video is too long, users may not be interested in watching the entire explanation video, which will affect the explanation effect of the item. Summary of the Invention
[0005] The present disclosure provides a method, device, equipment, and storage medium for displaying explanation information, which prevents users from being restricted to viewing only one explanation video. If the user does not want to watch a complete explanation video, the user can switch to viewing other explanation videos, thereby increasing flexibility and further improving the explanation effect of the item. The technical solution of the present disclosure is as follows:
[0006] According to a first aspect of an embodiment of the present disclosure, a method for presenting explanation information is provided, comprising:
[0007] A first explanation video of the displayed item, wherein the first explanation video is used to record explanation information of the item;
[0008] In response to the explanation information switching instruction, switching the first explanation video of the item to the second explanation video for display;
[0009] The second explanation video is different in length from the first explanation video, the second explanation video is used to record explanation information of the item, and the explanation information contained in the second explanation video is at least partially the same as the explanation information contained in the first explanation video.
[0010] In some embodiments, the first explanation video and the second explanation video are generated based on the same historical explanation event;
[0011] The first explanation video is used to record the historical explanation event;
[0012] The second explanation video is used to record the essential information of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0013] In some embodiments, the second explanation video is a video clip extracted from the first explanation video; or
[0014] The second explanation video is a video generated again based on the first explanation video.
[0015] In some embodiments, the first explanation video of the displayed item includes:
[0016] An explanation entrance is displayed on the exhibition interface of the item, and the first explanation video is displayed in response to a triggering operation on the explanation entrance.
[0017] In some embodiments, the item display interface includes at least one of the following: a details interface of the item, an item list including the item, a search result interface including the item, and a recommendation result interface including the item.
[0018] In some embodiments, a plurality of explanation entrances are displayed on the exhibit interface of the item, the first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
[0019] In some embodiments, the explanation information switching instruction is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switching control.
[0020] In some embodiments, the method further comprises:
[0021] Converting the first voice explanation information in the first explanation video into text explanation information;
[0022] Based on the text explanation information, a second explanation video of the item is generated.
[0023] In some embodiments, generating a second explanation video of the item based on the text explanation information includes:
[0024] generating summary text explanation information corresponding to the text explanation information, wherein the length of the summary text explanation information is shorter than the length of the text explanation information;
[0025] Based on the summary text explanation information, a second explanation video of the item is generated, and the second explanation video is used to explain the item according to the summary text explanation information.
[0026] In some embodiments, generating a second explanation video of the item based on the summary text explanation information includes:
[0027] extracting a sound feature from the first voice explanation information, wherein the sound feature is used to represent the voice of the explainer;
[0028] generating second voice explanation information based on the summary text explanation information and the voice feature, wherein the summary text explanation information is spoken in the voice of the explainer;
[0029] Based on the second voice explanation information, the second explanation video is generated.
[0030] In some embodiments, generating the second voice explanation information based on the summary text explanation information and the sound features includes:
[0031] Segmenting the summary text explanation information into a plurality of summary text segments;
[0032] Based on the plurality of summary text segments and the voice features, a plurality of voice segments are generated respectively, wherein each of the summary text segments is spoken in the voice of the speaker;
[0033] The plurality of voice segments are concatenated to obtain the second voice explanation information.
[0034] In some embodiments, generating the second explanation video based on the second voice explanation information includes:
[0035] Based on the plurality of summary text segments, a plurality of video frames are generated, wherein the subtitles in each of the video frames constitute a summary text segment;
[0036] The second voice explanation information and the plurality of video frames are combined into the second explanation video, wherein the voice segment corresponding to the same summary text segment in the second explanation video is synchronized with the video frame.
[0037] In some embodiments, the method further comprises:
[0038] Segmenting the text explanation information into a plurality of text segments;
[0039] Generating text features corresponding to a plurality of text segments respectively;
[0040] The plurality of text segments and the corresponding plurality of text features are saved in a database of the object.
[0041] In some embodiments, the method further comprises:
[0042] Obtaining question information, wherein the question information is input through the display interface of the second explanation video;
[0043] Generating a problem feature corresponding to the problem information;
[0044] Querying the database for the text segment corresponding to the text feature associated with the question feature;
[0045] Determining the retrieved text segment as associated information of the question information;
[0046] Answer information is displayed based on the question information and the associated information.
[0047] According to a second aspect of an embodiment of the present disclosure, a method for presenting explanation information is provided, comprising:
[0048] Display the entry for explaining the item;
[0049] In response to a triggering operation on the explanation entrance, displaying at least one explanation video of the item, wherein the explanation video is used to record explanation information of the item;
[0050] The item has at least two explanation videos, the lengths of the at least two explanation videos are different, and the explanation information contained in the at least two explanation videos is at least partially the same.
[0051] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0052] The first explanation video and the second explanation video are generated based on the same historical explanation event;
[0053] The first explanation video is used to record the historical explanation event;
[0054] The second explanation video is used to record the essential information of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0055] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0056] The second explanation video includes at least one of the following:
[0057] The first explanation video contains the essential information;
[0058] Detailed information about the item in the first explanation video;
[0059] The first explanation video contains information related to the target topic.
[0060] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0061] The second explanation video is a video clip extracted from the first explanation video; or
[0062] The second explanation video is a video generated again based on the first explanation video.
[0063] In some embodiments, the at least one explanation video showing the item includes:
[0064] Displaying a first explanation video of the item, and in response to an explanation information switching instruction, switching the first explanation video of the item to a second explanation video for display; or
[0065] Display the second explanation video of the item, and in response to the explanation information switching instruction, switch the second explanation video of the item to the first explanation video for display.
[0066] In some embodiments, the explanation information switching instruction is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switching control.
[0067] In some embodiments, the display of the item explanation entrance includes:
[0068] Displaying the explanation entrance on the exhibit interface of the item;
[0069] The item display interface includes at least one of the following: an item details interface, an item list including the item, a search result interface including the item, and a recommendation result interface including the item.
[0070] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0071] The first explanation video includes textual explanation information of the item;
[0072] The second explanation video is generated based on the summary text explanation information corresponding to the text explanation information. The second explanation video is used to explain the item according to the summary text explanation information. The length of the summary text explanation information is shorter than the length of the text explanation information.
[0073] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0074] The first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
[0075] According to a third aspect of an embodiment of the present disclosure, there is provided a device for displaying explanation information, comprising:
[0076] a video display unit configured to display a first explanation video of an item, wherein the first explanation video is used to record explanation information of the item;
[0077] A video switching unit is configured to switch the first explanation video of the item to a second explanation video for display in response to an explanation information switching instruction;
[0078] The second explanation video is different in length from the first explanation video, the second explanation video is used to record explanation information of the item, and the explanation information contained in the second explanation video is at least partially the same as the explanation information contained in the first explanation video.
[0079] In some embodiments, the first explanation video and the second explanation video are generated based on the same historical explanation event;
[0080] The first explanation video is used to record the historical explanation event;
[0081] The second explanation video is used to record the essential information of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0082] In some embodiments, the second explanation video is a video clip extracted from the first explanation video; or
[0083] The second explanation video is a video generated again based on the first explanation video.
[0084] In some embodiments, the video display unit includes:
[0085] An entrance display subunit is configured to display an explanation entrance on the exhibit interface of the item;
[0086] The trigger display subunit is configured to display the first explanation video in response to a trigger operation on the explanation entrance.
[0087] In some embodiments, the item display interface includes at least one of the following: a details interface of the item, an item list including the item, a search result interface including the item, and a recommendation result interface including the item.
[0088] In some embodiments, a plurality of explanation entrances are displayed on the exhibit interface of the item, the first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
[0089] In some embodiments, the explanation information switching instruction is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switching control.
[0090] In some embodiments, the apparatus further comprises:
[0091] an information conversion unit configured to convert the first voice explanation information in the first explanation video into text explanation information;
[0092] The video generation unit is configured to generate a second explanation video of the item based on the text explanation information.
[0093] In some embodiments, the video generation unit is configured to generate summary text explanation information corresponding to the text explanation information, and the length of the summary text explanation information is less than the length of the text explanation information; based on the summary text explanation information, a second explanation video of the item is generated, and the second explanation video is used to explain the item according to the summary text explanation information.
[0094] In some embodiments, the video generation unit is configured to extract sound features from the first voice explanation information, wherein the sound features are used to represent the voice of the explainer; generate second voice explanation information based on the summary text explanation information and the sound features, in which the summary text explanation information is spoken in the explainer's voice; and generate the second explanation video based on the second voice explanation information.
[0095] In some embodiments, the video generation unit is configured to divide the summary text explanation information into multiple summary text segments; based on the multiple summary text segments and the sound features, generate multiple voice segments respectively, in which each summary text segment is spoken in the voice of the commentator; and splice the multiple voice segments to obtain the second voice explanation information.
[0096] In some embodiments, the video generation unit is configured to generate multiple video frames based on multiple summary text segments, and the subtitles in each video frame are one summary text segment; the second voice explanation information and multiple video frames are synthesized into the second explanation video, wherein the voice segment corresponding to the same summary text segment in the second explanation video is synchronized with the video frame.
[0097] In some embodiments, the apparatus further comprises:
[0098] a text segmentation unit configured to segment the text explanation information into a plurality of text segments;
[0099] a text feature generating unit, configured to respectively generate text features corresponding to the plurality of text segments;
[0100] The storage unit is configured to store the plurality of text segments and the corresponding plurality of text features in a database of the object.
[0101] In some embodiments, the apparatus further comprises:
[0102] a question acquisition unit configured to acquire question information, wherein the question information is input through the display interface of the second explanation video;
[0103] A feature generating unit, configured to generate a question feature corresponding to the question information;
[0104] a query unit configured to query the database for the text segment corresponding to the text feature associated with the question feature;
[0105] an association determination unit configured to determine the queried text segment as associated information of the question information;
[0106] The answer display unit is configured to display answer information based on the question information and the associated information.
[0107] According to a fourth aspect of an embodiment of the present disclosure, there is provided a device for displaying explanation information, comprising:
[0108] An entrance display unit is configured to display an explanation entrance of an item;
[0109] a video display unit configured to display at least one explanation video of the item in response to a triggering operation on the explanation entrance, wherein the explanation video is used to record explanation information of the item;
[0110] The item has at least two explanation videos, the lengths of the at least two explanation videos are different, and the explanation information contained in the at least two explanation videos is at least partially the same.
[0111] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0112] The first explanation video and the second explanation video are generated based on the same historical explanation event;
[0113] The first explanation video is used to record the historical explanation event;
[0114] The second explanation video is used to record the essential information of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0115] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0116] The second explanation video includes at least one of the following:
[0117] The first explanation video contains the essential information;
[0118] Detailed information about the item in the first explanation video;
[0119] The first explanation video contains information related to the target topic.
[0120] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0121] The second explanation video is a video clip extracted from the first explanation video; or
[0122] The second explanation video is a video generated again based on the first explanation video.
[0123] In some embodiments, the video display unit includes:
[0124] The first display subunit is configured to display a first explanation video of the item, and in response to an explanation information switching instruction, switch the first explanation video of the item to a second explanation video for display; or
[0125] The second display sub-unit is configured to display a second explanation video of the item, and in response to an explanation information switching instruction, switch the second explanation video of the item to the first explanation video for display.
[0126] In some embodiments, the explanation information switching instruction is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switching control.
[0127] In some embodiments, the entrance display unit is configured to display the explanation entrance on the exhibit interface of the item;
[0128] The item display interface includes at least one of the following: an item details interface, an item list including the item, a search result interface including the item, and a recommendation result interface including the item.
[0129] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0130] The first explanation video includes textual explanation information of the item;
[0131] The second explanation video is generated based on the summary text explanation information corresponding to the text explanation information. The second explanation video is used to explain the item according to the summary text explanation information. The length of the summary text explanation information is shorter than the length of the text explanation information.
[0132] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0133] The first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
[0134] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0135] processor;
[0136] a memory for storing instructions executable by the processor;
[0137] The processor is configured to execute the instructions to implement the method for displaying explanation information as described in the first aspect or the method for displaying explanation information as described in the second aspect.
[0138] According to the sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for displaying explanation information as described in the first aspect or the method for displaying explanation information as described in the second aspect.
[0139] According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method for displaying explanation information as described in the first aspect or the method for displaying explanation information as described in the second aspect.
[0140] In the embodiment of the present disclosure, the second explanation video and the first explanation video are used to record the explanation information of the item, and the second explanation video is different from the first explanation video in length, and the explanation information contained in the second explanation video is at least partially the same, which provides users with diverse choices. When displaying the first explanation video of the item, the first explanation video can be switched to the second explanation video for display through the explanation information switching instruction, avoiding the user having to watch only one explanation video. When the user does not want to watch a complete explanation video, the user can switch to watching other explanation videos, which improves flexibility and thereby improves the explanation effect of the item.
[0141] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0142] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0143] Figure 1 is a schematic diagram showing an implementation environment according to an exemplary embodiment;
[0144] Figure 2 is a flow chart showing a method for presenting explanation information according to an exemplary embodiment;
[0145] Figure 3 is a schematic diagram showing an item details interface according to an exemplary embodiment;
[0146] Figure 4 is a schematic diagram showing an item list in a live broadcast room according to an exemplary embodiment;
[0147] Figure 5 is a schematic diagram showing a search result interface according to an exemplary embodiment;
[0148] Figure 6 is a schematic diagram showing a recommendation result interface according to an exemplary embodiment;
[0149] Figure 7 is a flowchart showing a method for generating a second explanation video according to an exemplary embodiment;
[0150] Figure 8 is a flowchart showing another method for generating a second explanation video according to an exemplary embodiment;
[0151] Figure 9 is a schematic diagram showing a video screen according to an exemplary embodiment;
[0152] Figure 10 is a flow chart showing a question-answering method according to an exemplary embodiment;
[0153] Figure 11 is a schematic diagram illustrating an operational flow of a method for presenting explanation information according to an exemplary embodiment;
[0154] Figure 12 is a flow chart showing a method for presenting explanation information according to an exemplary embodiment;
[0155] Figure 13 This is a structural block diagram of a device for displaying explanation information according to an exemplary embodiment;
[0156] Figure 14 This is a structural block diagram of a device for displaying explanation information according to an exemplary embodiment;
[0157] Figure 15 is a structural block diagram of a terminal according to an exemplary embodiment;
[0158] Figure 16 The figure is a structural block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION
[0159] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0160] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0161] The user information involved in this disclosure may be information authorized by the user or fully authorized by all parties.
[0162] First, the concepts involved in this disclosure are explained as follows:
[0163] 1. ASR (Automatic Speech Recognition): Converts the vocabulary content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0164] Speaker recognition or speaker verification is used to identify or confirm the speaker who makes the speech, while ASR technology is used to confirm the vocabulary content contained in the speech.
[0165] 2. TTS (Text-To-Speech): This is a type of speech synthesis technology that converts text into natural speech output. It is the opposite of the ASR process.
[0166] 3. Large Language Model (LLM): An AI model designed to understand and generate human language. Trained on large amounts of text data, LLMs can perform a wide range of tasks, including summarization, translation, and sentiment analysis. LLMs are characterized by their massive size, containing billions of parameters, which helps them learn complex patterns in language data. LLMs are typically based on deep learning architectures, which contributes to their impressive performance on various NLP (Natural Language Processing) tasks.
[0167] Figure 1 is a schematic diagram showing an implementation environment according to an exemplary embodiment. Figure 1 , the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a wireless or wired network.
[0168] For example, the terminal 101 is installed with a target application provided by the server 102, and the terminal 101 can use the target application to perform functions such as video playback, live broadcast viewing, online shopping, etc. The target application can be a short video application, a shopping application, a social application, or other application, which is not limited in the embodiments of the present disclosure.
[0169] In the disclosed embodiment, server 102 obtains at least two explanation videos of an item, wherein the at least two explanation videos have different lengths and contain at least partially identical explanation information. Server 102 can then provide the at least one explanation video of the item to terminal 101, which can then display the video. Terminal 101 can also switch the currently displayed explanation video to another explanation video for display.
[0170] In addition, when the terminal 101 displays the explanation video, a question-and-answer function is also provided. When a user asks a question about the item, the terminal 101 can display the answer information corresponding to the question information to the user.
[0171] In some embodiments, the terminal 101 includes a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto.
[0172] In some embodiments, server 102 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of the present disclosure do not impose any restrictions on this.
[0173] The method for displaying explanation information provided by the embodiments of the present disclosure is applicable to various scenarios of displaying explanation information of an object.
[0174] For example, in a live broadcast explaining an item, if the host needs to recommend an item, they will explain the item during the live broadcast and enable the automatic recording function of the live broadcast room, thereby recording a first explanation video of the item. The first explanation video is used to record the explanation information of the item. However, during the live broadcast, the host's explanation content is relatively long, and they may also explain other content unrelated to the item or chat with the audience, resulting in a long explanation video that is not convenient for users to watch. In this way, a second explanation video is produced based on the first explanation video. The explanation information contained in the second explanation video is at least partially the same as the explanation information contained in the first explanation video, but the length is different. In this way, the audience can choose the explanation video to watch according to their needs.
[0175] Of course, the above-mentioned scenario of live broadcasting and explaining an object is only an example. The method provided in the embodiment of the present disclosure can also be applied to other scenarios, and the embodiment of the present disclosure does not limit this.
[0176] Figure 2 is a flow chart showing a method for presenting explanation information according to an exemplary embodiment. Figure 2 As shown, the method is performed by an electronic device, and the method includes the following steps:
[0177] In step 201, an electronic device displays a first explanation video of an item, where the first explanation video is used to record explanation information of the item.
[0178] Among them, the items may include any type of daily necessities, cosmetics, clothing, accessories, etc., and the embodiments of the present disclosure do not limit the types of items. The explanation information of the items may include the appearance, price, name, sales quantity, producer information and other aspects of the explanation information of the items, and the appearance may include style design, upper body effect, detail display, matching recommendation, etc. The price may include the original price, current price, price discount method, etc. The sales quantity may include the sales quantity in a recent period of time, the total sales quantity, the sales quantity during holidays, etc. The producer information includes the name or address of the producer, etc. The embodiments of the present disclosure do not limit the content of the explanation information. The explanation information of the items may include information displayed in the video screen, or information played in the voice, etc. The embodiments of the present disclosure do not limit the form of the explanation information.
[0179] In some embodiments, the host explains the item in the live broadcast room and turns on the automatic recording function of the live broadcast room to record a first explanation video.
[0180] In other embodiments, the explainer uses an electronic device to record a video while explaining the item to obtain a first explanation video.
[0181] In some embodiments, the electronic device displays the first explanation video, including: the electronic device displays video information of the first explanation video, the video information including a video cover, a video title, a video content summary, a play entry, etc. In this case, the electronic device has not yet played the first explanation video, and the user can first preview the video information of the first explanation video, and then trigger the video title or play entry of the first explanation video to play the first explanation video.
[0182] In other embodiments, the electronic device displays the first explanation video, including: the electronic device plays the first explanation video.
[0183] In step 202, the electronic device switches the first explanation video of the item to the second explanation video in response to the explanation information switching instruction.
[0184] The second explanation video is different in length from the first explanation video, and the second explanation video is also used to record explanation information about the item. The explanation information contained in the second explanation video is at least partially the same as the explanation information contained in the first explanation video. For example, the second explanation video is shorter than the first explanation video, or longer than the first explanation video.
[0185] The explanation information switching instruction is used to instruct switching of the explanation information of an item. When the electronic device receives the explanation information switching instruction while the first explanation video is already being displayed, the electronic device switches the currently displayed first explanation video to the second explanation video for display. In another embodiment, when the electronic device receives the explanation information switching instruction while the second explanation video is already being displayed, the electronic device switches the currently displayed second explanation video to the first explanation video for display.
[0186] In some embodiments, the number of the first explanation video is one, and the number of the second explanation video is one or more. The embodiment of the present disclosure does not limit the number of the second explanation videos.
[0187] In some embodiments, when there are multiple second explanation videos, the electronic device, while displaying the first explanation video, can, in response to the explanation information switching instruction, switch the currently displayed first explanation video to the second explanation video that is ranked first in the order of the multiple second explanation videos. Subsequently, upon receiving the explanation information switching instruction again, the currently displayed second explanation video can be switched to the second explanation video that is ranked next in the order of the multiple second explanation videos, and so on. In other embodiments, when the electronic device is displaying the second explanation video that is ranked last, in response to the explanation information switching instruction, switch the currently displayed second explanation video to the first explanation video for display.
[0188] In the method provided by the embodiment of the present disclosure, the second explanation video and the first explanation video are used to record the explanation information of the item, and the second explanation video is different from the first explanation video in length, and the explanation information contained in the second explanation video is at least partially the same, which provides users with diverse choices. When displaying the first explanation video of the item, the first explanation video can be switched to the second explanation video for display through the explanation information switching instruction, avoiding the user having to watch only one explanation video. When the user does not want to watch a complete explanation video, the user can switch to watching other explanation videos, which improves flexibility and thereby improves the explanation effect of the item.
[0189] Based on the above embodiments, the present disclosure further provides the following exemplary embodiments:
[0190] 1. In some embodiments, the first explanation video and the second explanation video are generated based on the same historical explanation event. The first explanation video is used to record the historical explanation event, and the second explanation video is used to record the essence of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0191] The historical explanation event is an event involving the explanation of the aforementioned item. During the explanation of the aforementioned item, explanation information is generated, and the first explanation video records the item explanation information. For example, if a person explains an item while recording a video, the first explanation video will include both the video footage of the person demonstrating the item and the audio of the person introducing the item. Both the video footage and the audio are considered item explanation information.
[0192] The difference between the second explanation video and the first explanation video is that the second explanation video will not record all the information of the historical explanation event, but will record the essence or derivative information of the historical explanation event to ensure that the length of the second explanation video is shorter than that of the first explanation video, thereby achieving the purpose of shortening the length of the explanation video.
[0193] The essence information is the information in the explanation information of the historical explanation event that meets the essence explanation conditions.
[0194] In some embodiments, a historical explanation event includes an opening phase, an explanation phase, and a closing phase, and the information that satisfies the essential explanation condition refers to the explanation information generated during the explanation phase of the historical explanation event. For example, each time a new phase of the historical explanation event begins, the explainer will speak a preset prompt text for that phase, and the spoken preset prompt text will be recorded in the explanation information. The electronic device can then use the preset prompt text in the explanation information to divide the explanation information into different phases. For example, the preset prompt text for the opening phase is "Welcome everyone," the preset prompt text for the explanation phase is "Beginning the introduction," and the preset prompt text for the closing phase is "End of introduction."
[0195] In other embodiments, the information that satisfies the "essential explanation" condition is explanation information from a time period in a historical explanation event where the number of views or popularity exceeds a threshold, or explanation information from a time period in a historical explanation event where the number of views or popularity is the highest. For example, the historical explanation event is a live broadcast event, and the first explanation video is a video recorded during the live broadcast. Each time period in the live broadcast has a corresponding number of views or popularity. The number of views is determined based on the number of viewers who enter the live broadcast room during this time period, and the popularity can be determined based on interaction data such as the number of viewers in the live broadcast room and the number of interactive operations that occur in the live broadcast room. Alternatively, the first explanation video is not a live broadcast video, but a video produced and released after recording the historical explanation event. After release and after one or more playbacks of the first explanation video, the number of views or popularity for each time period in the first explanation video can also be counted. The number of views is the number of times the video segment of the first explanation video is played during this time period, and the popularity indicates the popularity of the video segment of the first explanation video during this time period. The popularity can be determined based on interaction data generated during this time period, including the number of comments posted by viewers, the number of barrages posted by viewers, or the number of likes.
[0196] The embodiments of the present disclosure only provide the above two exemplary essence information. In other embodiments, the above essence explanation conditions may further include other conditions, and the above essence information may further include other information, which is not limited in the embodiments of the present disclosure.
[0197] Based on the same historical explanation event, a first explanation video and a second explanation video of different lengths are generated. The electronic device only records the essential information of the historical explanation event in the second explanation video instead of all the information of the historical explanation event. This can effectively reduce the amount of explanation information to obtain another explanation video of shorter length, providing users with diverse choices, and also retaining the key explanation information of the item, avoiding affecting the explanation effect of the item due to shortening the length of the explanation video.
[0198] Furthermore, derivative information refers to information derived from the explanatory information of a historical event. Derivative information includes different information from the explanatory information, thereby increasing the amount of new information. However, derivative information and the explanatory information can also include the same information. For example, if the explanatory information includes an image of an object, the derivative information replaces that image with a corresponding thumbnail, while retaining the rest of the original explanatory information. Alternatively, if the explanatory information includes a video of a speaker holding an object, the derivative information replaces that video with an image of the object, while retaining the rest of the original explanatory information.
[0199] The electronic device evolves based on the explanation information of the historical explanation event, and generates a second explanation video according to the derived information obtained by the evolution, thereby enriching the content of the explanation information and providing users with a variety of choices. Moreover, the second explanation video has the same explanation information as the first explanation video. When watching the second explanation video, the user can understand the general content of the first explanation video without having to watch the first explanation video repeatedly, avoiding the situation of watching both the first explanation video and the second explanation video, thereby saving the user's time as much as possible.
[0200] 2. In some embodiments, the second explanation video includes at least one of the following: essential information in the first explanation video; detailed information about the items in the first explanation video; information related to the target topic in the first explanation video.
[0201] First of all, the essential information is the same as the essential information in point 1 above, so I will not repeat it here.
[0202] Secondly, the item details include: item image, item name, item sales volume, item price, item reviews, etc. The item details can be extracted from the first explanation video or from the item information database. The embodiment of the present disclosure does not limit the source of the details.
[0203] Again, the first explanation video can include explanation information on one or more topics, with each topic explaining the item from different perspectives. For example, the multiple topics include style design, wearing effect, detail display, and matching recommendations. The target topic includes at least one of the one or more topics. The target topic can be determined by default by the electronic device, or by the person who created the second explanation video selecting from one or more topics, or by other methods.
[0204] 3. In some embodiments, the first explanation video and the second explanation video are independently generated videos, but because the first explanation video and the second explanation video contain explanation information of the same object, the electronic device switches between the first explanation video and the second explanation video for display.
[0205] In other embodiments, after the electronic device generates the first explanation video, it generates the second explanation video based on the first explanation video, or after generating the second explanation video, it generates the first explanation video based on the second explanation video.
[0206] The solution for generating a second explanation video based on a first explanation video includes the following two examples: the second explanation video is a video clip extracted from the first explanation video; or the second explanation video is a video generated again based on the first explanation video.
[0207] In some embodiments, the first explanation video includes multiple video segments, and the electronic device can extract one or more video segments from the multiple video segments to form the second explanation video. When extracting multiple video segments to form the second explanation video, the multiple video segments can be continuous or discontinuous video segments in the first explanation video.
[0208] Exemplarily, the electronic device automatically extracts one or more video clips from the first explanation video without the involvement of a technician. The extraction rules can be pre-set by the electronic device. For example, the electronic device automatically identifies and extracts video clips from the first explanation video that meet the essential explanation criteria. In other embodiments, after a technician determines the time period corresponding to the video clips to be extracted from the first explanation video, the electronic device extracts the corresponding video clips from the first explanation video according to the time period determined by the technician.
[0209] In some embodiments, after extracting a video clip from the first explanation video, the video clip can be edited and the edited video can be determined as the second explanation video. For example, the editing operation may include: adding special effects to the video screen, removing background voice, or adding a voice changer to the background voice, etc. The embodiment of the present disclosure does not limit the method of generating the second explanation video. In addition, the following Figure 7 and Figure 8 The illustrated embodiment also provides a detailed description of another method for generating a second explanation video based on the first explanation video, so it will not be repeated here.
[0210] Alternatively, a second explanation video can be generated based on the first explanation video. For example, a live broadcast operator can watch a live broadcast explaining an item, or refer to the first explanation video recorded during the live broadcast to learn about the item, and then create a second explanation video based on the learned information.
[0211] The scheme for generating the first explanation video based on the second explanation video includes the following two examples: the first explanation video is a video obtained by adding new video clips to the second explanation video, or the first explanation video is a video generated again based on the second explanation video.
[0212] In some embodiments, the electronic device selects information not included in the second explanation video from the detailed information of the item based on the generated second explanation video, generates a video clip based on the information, and splices the video clip with the second explanation video to obtain the first explanation video.
[0213] In other embodiments, the second explanation video is edited to obtain the first explanation video. The editing operation may include: adding special effects to the video screen, adding background voice, adding new screen elements, etc. The embodiment of the present disclosure does not limit the method of generating the first explanation video.
[0214] In another embodiment, the first explanation video is generated separately with reference to the second explanation video. For example, after the first explanation video is recorded during the live broadcast, the live broadcast operator watches the first explanation video, summarizes the first explanation video, and produces a second explanation video that is shorter than the first explanation video.
[0215] By generating a second explanation video based on the first explanation video, or generating the first explanation video based on the second explanation video, it is possible to refer to the explanation information in the existing explanation video, thereby reducing the difficulty of generating the explanation video and saving the time of generating the explanation video.
[0216] 4. In some embodiments, displaying a first explanation video of an item includes: displaying an explanation entrance on an exhibit interface of the item, and displaying the first explanation video in response to a triggering operation on the explanation entrance.
[0217] The exhibit interface displays one or more items. For example, the exhibit interface displays item information for each item, such as an item image, item name, item sales volume, and item price. The explanation portal is a control used to trigger the display of a first explanation video. The explanation portal can be a separately displayed button, or a display element within the item image, such as a decorative pendant. Displaying the explanation portal not only facilitates triggering the display of the first explanation video but also serves to beautify the item image. Alternatively, the explanation portal can be one item of item information. For example, an electronic device can associate the item name with the first explanation video, making the item name the explanation portal. Triggering the item name displays the associated first explanation video, eliminating the need to display the explanation portal separately, thereby saving display space.
[0218] In some embodiments, the triggering operation for the explanation entrance may include a click operation, a sliding operation, etc., which is not limited in the embodiments of the present disclosure.
[0219] In some embodiments, the item display interface includes at least one of the following: an item details interface, an item list containing the item, a search result interface containing the item, a recommendation result interface containing the item, etc. Other forms of interfaces may also be included, and the embodiments of the present disclosure do not limit the display interface.
[0220] Among them, the item details interface can be the interface used to display item information in the mall, or the interface used to display item information in the live broadcast room. The item list can be the list of items displayed in the live broadcast room, or the list of items displayed in the mall. The search results interface is the interface obtained after searching based on the search keyword, and the search results interface includes items that match the search keyword. The recommended results interface includes recommended items. The recommended results interface can be the main interface of the currently running application, the interface displayed after clicking the same city control, the interface displayed after clicking the group purchase control, etc.
[0221] For example, see Figure 3 In the item details interface, there is a "Replay" entry in the upper right corner of the item image. Triggering this entry will display the first explanation video of the item. Figure 4 In the list of items in the live broadcast room, for items that have been explained, a "Watch Explanation" entrance will be displayed below the image of the item. Triggering this entrance will replay the first explanation video of the item. Figure 5 In the search results interface, for the items in the search results, a "Replay" entry is displayed in the upper left corner of the item image. Triggering this entry will replay the first explanation video of the item. Figure 6 In the recommendation result interface, for the recommended items, a "Explanation Playback" entrance is displayed in the upper left corner of the image of the item. Triggering this entrance can replay the first explanation video of the item.
[0222] By displaying the tutorial entrance on the exhibit interface, users are provided with a short operation path to trigger the tutorial video, which simplifies user operation and improves operational efficiency. In addition, the exhibit interface includes a variety of situations, which enriches the display location of the tutorial entrance, thus avoiding the need for users to jump to multiple interfaces to find the tutorial entrance.
[0223] In some embodiments, multiple explanation entrances are displayed on the item's exhibit interface, and the first explanation video and the second explanation video correspond to different explanation entrances. The first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances. When a user triggers the explanation entrance corresponding to the first explanation video, the first explanation video is displayed, or when a user triggers the explanation entrance corresponding to the second explanation video, the second explanation video is displayed. Moreover, even if the first explanation video has already been displayed, the first explanation video can be switched to the second explanation video for display. Alternatively, even if the second explanation video has already been displayed, the second explanation video can be switched to the first explanation video for display.
[0224] By separately displaying the explanation entrances corresponding to the first explanation video and the second explanation video, users can easily trigger the explanation entrance corresponding to the explanation video they want to watch, without having to trigger the same explanation entrance first and then select the explanation entrance corresponding to the explanation video they want to watch from multiple explanation videos. This shortens the operation path and improves operation efficiency.
[0225] 5. In some embodiments, the instruction to switch the explanation information is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switch control.
[0226] Among them, the reserved area is located at any position in the current interface. For example, if the first explanation video is displayed in the current interface, the reserved area can be located in the playback area of the first explanation video, or in other areas outside the playback area of the first explanation video. The reserved gestures input in the reserved area include sliding gestures, dragging gestures, two-finger pinching gestures, two-finger stretching gestures, etc. The embodiment of the present disclosure does not limit the reserved area and the reserved gestures.
[0227] Among them, the switching control can be located on the right side of the playback area of the first explanation video, or in the control bar at the bottom of the playback area of the first explanation video. The predetermined gestures acting on the switching control include click gestures, sliding gestures, etc. The embodiment of the present disclosure does not limit the position of the switching control and the predetermined gestures acting on the switching control.
[0228] Based on a predetermined gesture input in a predetermined area and / or a predetermined gesture acting on a switching control, an explanation information switching instruction is triggered. The triggering operation is simple, fast and easy to implement.
[0229] It should be noted that the above-mentioned various exemplary embodiments can be combined in any manner to form new embodiments.
[0230] In the above Figure 2 Based on the illustrated embodiment, the disclosed embodiment also provides a method for generating a second explanation video based on a first explanation video, wherein the first explanation video includes video screen explanation information and first voice explanation information, wherein the first voice explanation information is provided by a narrator to explain the item, and the first voice explanation information in the first explanation video is converted into text explanation information, and a second explanation video of the item is generated based on the text explanation information.
[0231] Regarding the process of generating the second explanation video, the present disclosure provides an exemplary embodiment: Figure 7 is a flow chart showing a method for generating a second explanation video according to an exemplary embodiment. Figure 7 As shown, the method is performed by an electronic device, and the method includes the following steps:
[0232] In step 701, the electronic device converts first voice explanation information in a first explanation video into text explanation information.
[0233] In step 702, the electronic device generates summary text explanation information corresponding to the text explanation information, where the length of the summary text explanation information is shorter than the length of the text explanation information.
[0234] The text explanation information is used to explain the item, but if the content of the explanation is too much, the text explanation information will be too long. The summary text explanation information is a summary of the text explanation information, which can summarize the key content in the text explanation information to eliminate useless content, thereby shortening the text length.
[0235] In step 703, the electronic device generates a second explanation video of the item based on the summary text explanation information, where the second explanation video is used to explain the item according to the summary text explanation information.
[0236] Since the length of the summary text explanation information is shorter than the length of the text explanation information, the length of the second explanation video explaining the item according to the summary text explanation information will also be shorter than the length of the first explanation video, making it more convenient for users to watch the second explanation video.
[0237] After step 703, the method further includes: the electronic device publishing the second explanation video in the video playback record of the live broadcast room, so that users can watch the second explanation video in the video playback record of the live broadcast room, or the electronic device publishing the second explanation video in the item details interface as a supplement to the item details information. In the item details interface, users can not only view the item details but also watch the second explanation video, thereby gaining a full understanding of the item, thereby increasing the user's order rate. In addition, there is no need for the host to separately produce the second explanation video, which reduces the host's workload.
[0238] In addition, the electronic device can also publish the first explanation video, thereby providing users with more choices, and users can choose to watch the first explanation video or the second explanation video. In some embodiments, the electronic device embeds the playback entrance of the second explanation video into the first explanation video. When the terminal plays the first explanation video, the playback entrance of the second explanation video, such as "Simplified Explanation", will be displayed in the playback interface of the first explanation video. If the user finds that the first explanation video is too long, the playback entrance of the second explanation video can be triggered to play the second explanation video.
[0239] In some embodiments, if the text explanation information includes multiple topics, the summary text explanation information also includes multiple topics. The text explanation information for different topics explains the item from different perspectives. The electronic device can then directly generate a second explanation video based on the summary text explanation information. Alternatively, the text explanation information for the multiple topics can be separated, and a corresponding explanation video generated based on the text explanation information for each topic, thereby generating multiple second explanation videos. The second explanation videos corresponding to the multiple topics can be released separately. Alternatively, the multiple explanation videos generated can be combined into a single second explanation video, with each topic added to the second explanation video. Triggering any topic can jump to the corresponding segment of that topic. For example, the summary text explanation information may include summary text explanation information for the topic of "Product Details" and summary text explanation information for the topic of "Product Discounts," generating second explanation videos for each of the two topics. Users who have long-term interest in a product and already understand the product's details do not need to watch the second explanation video for the topic of "Product Details." Instead, they only need to watch the second explanation video for the topic of "Product Discounts" to learn about the current discounts available for purchasing the product. Users who are unfamiliar with the product can watch the second explanation video for the topic of "Product Details" to quickly learn about it.
[0240] If the explanation video is long, it will be inconvenient for users to watch, especially if it contains a lot of low-value content. Users will not be able to accurately locate the key content of the item in the explanation video, and users will not have the patience to watch the entire explanation video, resulting in a low completion rate of the explanation video.
[0241] In the embodiment of the present disclosure, the electronic device converts the first voice explanation information in the first explanation video into text explanation information, generates summary text explanation information corresponding to the text explanation information, and generates a second explanation video of the item based on the summary text explanation information. The second explanation video is used to explain the item according to the summary text explanation information. Since the length of the summary text explanation information is shorter than the length of the text explanation information, the length of the second explanation video is shorter than the length of the first explanation video, which shortens the length of the explanation video, makes it more convenient for users to watch, and thus improves the playback completion rate of the explanation video.
[0242] Moreover, since the summary text explanation information is a summary of the key content in the text explanation information and contains more key content, the generated second explanation video removes low-value content, highlights the key content, improves content dissemination efficiency, easily arouses user interest, and saves storage resources and playback resources.
[0243] Experiments have shown that explanation videos in related technologies are typically over 10 minutes long, with over 90% of them being played back for videos longer than one minute. However, the average viewing time is 40 milliseconds, and the completion rate is only 3%. The method provided in the embodiments of this disclosure can shorten explanation videos originally lasting over 10 minutes to under one minute, improving the completion rate while saving resources consumed by long videos.
[0244] above Figure 7 The embodiment shown is only a brief description of the method for generating the second explanation video. Figure 7 Based on the illustrated embodiment, another method for generating a second explanation video is provided. Figure 8 is a flowchart of another method for generating a second explanation video according to an exemplary embodiment. Figure 8 As shown, the method is performed by an electronic device, and the method includes the following steps:
[0245] In step 801, the electronic device converts first voice explanation information in a first explanation video into text explanation information.
[0246] In some embodiments, the electronic device extracts first voice explanation information from the first explanation video, and converts the first voice explanation information into text explanation information using ASR technology.
[0247] For example, the electronic device converts the first voice explanation information into the text explanation information using a voice recognition model. The voice recognition model may be a Whisper model or other model. The Whisper model is a deep learning-based voice recognition model that can be used for tasks such as voice recognition, voice translation, and language recognition. The Whisper model collects data from multiple data sources and is trained based on the collected data. The collected data includes multiple languages and data used to perform multiple tasks, which enables the Whisper model to be fully trained and its accuracy to be improved. The Whisper model has five levels: tiny (39M data); base (74M data); small (244M data); medium (769M data); and large (1150M data). The larger the Whisper model, the better the voice recognition effect. Any level of the Whisper model can be used in the embodiments of the present disclosure.
[0248] In some embodiments, after receiving the text explanation information, the electronic device preprocesses the text explanation information to remove useless information from the text explanation information, such as information unrelated to the item or repeated text. The electronic device can set preprocessing rules and preprocess the text explanation information in the first explanation video according to the rules.
[0249] In step 802, the electronic device generates summary text explanation information corresponding to the text explanation information, and the length of the summary text explanation information is shorter than the length of the text explanation information.
[0250] In some embodiments, the electronic device generates summary text explanation information corresponding to the text explanation information using a summary extraction model. The summary extraction model is configured to extract summary text corresponding to the text, thereby retaining useful information in the text and removing useless information from the text to shorten the text. The summary extraction model may be an LLM or other model.
[0251] In some embodiments, the process of training the summary extraction model includes: obtaining a large language model, which can be used for tasks such as speech translation, language recognition, and automatic question answering. To improve the accuracy of the large language model in extracting summary text, sample data is obtained, where the sample data includes sample text and sample summary text corresponding to the sample text, where the length of the sample summary text is less than the length of the sample text, and the sample summary text includes useful information in the sample text and removes useless information in the sample text. The large language model is trained based on the sample data until a training completion condition is met, thereby obtaining a trained summary extraction model.
[0252] In one possible implementation, the electronic device generates summary text explanation information corresponding to the text explanation information by using a summary extraction model, including the following steps 11-13:
[0253] 11. When the length of the text explanation information is greater than the target length of the summary extraction model, the text explanation information is divided into multiple text segments, and the length of each text segment is less than the target length.
[0254] The target length is the maximum length of text allowed by the summary extraction model. If the length of the text explanation exceeds the target length, it cannot be directly input into the summary extraction model. Instead, the text explanation is segmented into multiple text segments. Since the length of each text segment is less than the target length, the summary extraction model can process each text segment.
[0255] 12. Generate a summary text segment corresponding to each text segment through the summary extraction model.
[0256] In some embodiments, the electronic device processes each text segment separately, so the electronic device inputs one text segment into the summary extraction model each time, and generates a summary text segment corresponding to the text segment through the summary extraction model.
[0257] In other embodiments, the electronic device inputs the first text segment into a summary extraction model according to the order of multiple text segments, and generates a summary text segment corresponding to the first text segment through the summary extraction model. Then, the second text segment and the first summary text segment are input into the summary extraction model, and the summary extraction model generates a summary text segment corresponding to the second text segment, and so on, until a summary text segment corresponding to the last text segment is generated. The above-mentioned process of generating summary text segments takes into account the contextual relationship between adjacent text segments, thereby improving the accuracy of the generated summary text segments. In addition, adjacent summary text segments are continuous with each other, which can ensure that the summary text explanation information obtained by subsequently splicing multiple summary text segments is smooth and semantically accurate.
[0258] 13. Combine multiple summary text segments to obtain summary text explanation information.
[0259] After obtaining the summary text segments corresponding to the multiple text segments, the multiple summary text segments are spliced together in the order of the multiple text segments to obtain summary text explanation information.
[0260] In some embodiments, the electronic device may use a refine mode to combine multiple summary text segments, or use other methods to combine them, which are not limited in the embodiments of the present disclosure. The refine mode refers to using a large language model to optimize each summary text segment before combining them.
[0261] In another possible implementation, the electronic device generates summary text explanation information corresponding to the text explanation information by using a summary extraction model, including the following steps 21-22:
[0262] 21. Create a first prompt word, where the first prompt word includes text explanation information and a first task text, and the first task text is used to instruct the summary extraction model to generate summary text explanation information for the text explanation information.
[0263] 22. Input the first prompt word into the summary extraction model, and the electronic device generates summary text explanation information through the summary extraction model.
[0264] The summary extraction model can process the prompt words, which are used to issue tasks to the summary extraction model. Therefore, the electronic device creates a first prompt word including text explanation information and first task text. The first task text represents the task requirements for the summary extraction model. By issuing the first prompt word to the summary extraction model, the summary extraction model can understand that its task is to extract a summary for the text explanation information. Therefore, the summary extraction model can generate summary text explanation information corresponding to the text explanation information.
[0265] For example, the first prompt word is as follows:
[0266] Your job is to summarize the text in the video introducing the item in Chinese, using no more than 200 words. The text is as follows:
[0267] -------------------------
[0268] "Text explanation information"
[0269] -------------------------"
[0270] In some embodiments, the process of training the summary extraction model includes: obtaining a large language model, which can be used for tasks such as speech translation, language recognition, and automatic question answering. To improve the accuracy of the large language model in extracting summary text, sample data is obtained, the sample data including a sample prompt word and a sample summary text corresponding to the sample text in the sample prompt word, the length of the sample summary text being less than the length of the sample text, the sample prompt word including the sample text and a sample task text, the sample summary text including useful information in the sample text and removing useless information in the sample text. The large language model is trained based on the sample data until a training completion condition is met, thereby obtaining a trained summary extraction model.
[0271] It should be noted that steps 11-13 and steps 21-22 can be combined in any manner to form new embodiments of the present disclosure. For example, if the length of the text explanation information exceeds the target length of the summary extraction model, the electronic device segments the text explanation information into multiple text segments, each of which is less than the target length, creates multiple prompt words, each of which includes a text segment and the first task text, and inputs the multiple prompt words into the summary extraction model. The summary extraction model generates a summary text segment corresponding to each text segment, and the multiple summary text segments are concatenated to obtain the summary text explanation information.
[0272] After the electronic device generates the summary text explanation information, it can generate a second explanation video of the item based on the summary text explanation information. The second explanation video is used to explain the item according to the summary text explanation information. The process of generating the second explanation video is detailed in the following steps 803-805.
[0273] In step 803, the electronic device extracts sound features from the first voice explanation information, where the sound features are used to represent the voice of the explainer.
[0274] The first voice explanation information includes the voice of the explainer, and the sound features extracted from the first voice explanation information can represent the voice of the explainer. The sound features may include timbre features, voiceprint features, etc.
[0275] In some embodiments, the electronic device extracts sound features from the first voice explanation information through a feature extraction model, and the feature extraction model is an LSTM (Long Short-Term Memory) Speaker Encoder model or other models.
[0276] In step 804, the electronic device generates second voice explanation information based on the summary text explanation information and the voice feature, in which the summary text explanation information is spoken in the voice of the explainer.
[0277] The sound feature represents the voice of the speaker. Based on the summary text explanation information and the sound feature, the electronic device converts the summary text explanation information into audio according to the speaker's voice. The generated second voice explanation information then uses the speaker's voice to speak the summary text explanation information. For example, if the sound feature is a timbre feature, the generated second voice explanation information uses the speaker's timbre to speak the summary text explanation information.
[0278] In some embodiments, the electronic device uses TTS technology to generate the second voice explanation information based on the summary text explanation information and sound features. In other embodiments, the electronic device inputs the summary text explanation information and sound features into a speech synthesis model, and generates the second voice explanation information through the speech synthesis model. For example, the electronic device uses the Tacotron2 model (a speech synthesis model) to generate a spectrogram based on the summary text explanation information and sound features, and converts the spectrogram into a speech waveform using the WaveFlow model to obtain the playable second voice explanation information.
[0279] In some embodiments, the electronic device divides the summary text explanation information into multiple summary text segments, and generates multiple voice segments based on the multiple summary text segments and sound features. In each voice segment, each summary text segment is spoken in the voice of the commentator, and the multiple voice segments are spliced together to obtain second voice explanation information.
[0280] After the multiple summary text segments are segmented, the speech segments corresponding to each summary text segment may be generated sequentially, or in order to improve processing efficiency, the speech segments corresponding to each summary text segment may be generated in parallel.
[0281] In step 805, the electronic device generates a second explanation video based on the second voice explanation information.
[0282] The second explanation video includes second voice explanation information. Therefore, in the second explanation video, the summary text explanation information is spoken in the voice of the explainer for the user to listen to.
[0283] In addition, in addition to the second voice explanation information, the second explanation video also includes multiple video frames. The multiple video frames can be frames included in the first explanation video, and can retain the original frames of the explainer explaining the item in the first explanation video.
[0284] In some embodiments, the electronic device determines the required number of video frames according to the duration of the second voice explanation information, extracts the required number of video frames from the video frames of the first explanation video, and combines the second voice explanation information and the extracted multiple video frames into a second explanation video.
[0285] In some embodiments, in the first explanation video, multiple video frames correspond to the text explanation information in the first voice explanation information, and after the summary text explanation information is generated, the summary text explanation information will also include some text segments in the text explanation information. Then, the video frames corresponding to the text segments included in the summary text explanation information are obtained from the first explanation video, and the second voice explanation information and the multiple video frames are synthesized into a second explanation video.
[0286] Alternatively, the electronic device may use new video frames instead of the video frames in the first explanation video. That is, the video frames in the second explanation video may be separately generated video frames that include the object. In some embodiments, the method further includes: the electronic device acquiring multiple images of the object, wherein the object is displayed from different angles in any two images, and generating a three-dimensional model of the object based on the multiple images. Generating the second explanation video based on the second voice explanation information includes: generating multiple video frames based on the three-dimensional model, and synthesizing the second voice explanation information and the multiple video frames into the second explanation video. For example, the electronic device generates a three-dimensional model of the object based on the multiple images using a NeRF (Neural Radiance Field) model. Furthermore, the synthesis of the second explanation video may use FFmpeg (Fast Forward moving picture experts group) or other algorithms, which are not limited in the present embodiment.
[0287] Among them, each video frame is used to display the three-dimensional model of the object from a certain angle. By playing the multiple video frames continuously, an effect of rotating the three-dimensional model of the object can be formed, allowing users to view the three-dimensional model of the object from different angles, providing a full range of object display, and facilitating users to fully understand the appearance details of the object.
[0288] In addition, the electronic device can also add subtitles to the video screen based on the summary text explanation information. In some embodiments, the electronic device generates multiple video screens based on multiple summary text segments, and the subtitles in each video screen are a summary text segment. The second voice explanation information and the multiple video screens are combined into a second explanation video, wherein the voice segment corresponding to the same summary text segment in the second explanation video is synchronized with the video screen. Then, during the playback of the second explanation video, each time a video screen is played, a subtitle is displayed in the video screen, and the audio played at the same time also includes the content of the subtitle. The generated subtitles can be in WebVTT (Web Video Text Tracks) format or other formats.
[0289] For example, see Figure 9 The second explanation video includes a video screen showing a three-dimensional model of a mobile phone, and subtitles are included below the video screen. The subtitles are part of the text segment in the text explanation information.
[0290] In an embodiment of the present disclosure, the electronic device converts the first voice explanation information in the first explanation video into text explanation information, generates summary text explanation information corresponding to the text explanation information, and generates a second explanation video of the item based on the summary text explanation information. The second explanation video is used to explain the item according to the summary text explanation information. Since the length of the summary text explanation information is shorter than the length of the text explanation information, the length of the second explanation video is shorter than the length of the first explanation video, which shortens the length of the explanation video, facilitates user viewing, and thus improves the playback completion rate of the explanation video.
[0291] Furthermore, the second video uses the guide's voice to deliver the summary text, simulating the original guide's explanation of the item, avoiding the awkwardness of a change in guide, and improving the playback quality. Furthermore, if the guide spoke in a dialect or had unstandard pronunciation in the first video, some users might not be able to understand the text and, consequently, the meaning of the text explanation. While the second video retains the guide's voice, it can deliver the summary text in the official language, avoiding the presence of dialect or unstandard pronunciation, making it easier for users to understand and expanding the geographical scope of applicability.
[0292] Furthermore, the electronic device divides the summary text explanation information into multiple summary text segments and generates multiple voice segments based on the multiple summary text segments and voice characteristics. Each voice segment contains a summary text segment spoken by the speaker. The multiple voice segments are then concatenated to generate the second voice explanation information. This avoids the inaccurate audio explanation generated directly from the longer summary text explanation information, thereby improving accuracy. Furthermore, generating the voice segments corresponding to each summary text segment in parallel improves processing efficiency and reduces processing time.
[0293] Moreover, using each summary text segment as a subtitle in the video screen increases the amount of information in the video screen, making it easier for users to watch. Moreover, since the summary text explanation information has been divided into multiple summary text segments in the previous processing flow, the process of setting subtitles is relatively simple and does not waste a lot of processing time and processing resources.
[0294] In addition, a three-dimensional model of the object is generated based on multiple images of the object, and a video image is generated based on the three-dimensional model. This can create an effect of rotating the three-dimensional model of the object in the second explanation video, allowing users to view the three-dimensional model of the object from different angles, thereby improving the display effect.
[0295] Furthermore, using the summary extraction model to generate a corresponding summary text explanation for the text explanation information improves processing efficiency while ensuring summary extraction accuracy. Furthermore, for longer text explanations, the text explanation information can be divided into multiple segments, and the summary extraction model generates corresponding summary text segments for each segment, which are then concatenated to create the summary text explanation information. Therefore, this solution is applicable to text explanations of any length and has a wide range of applications.
[0296] In addition, by creating a first prompt word including text explanation information and the first task text, the first prompt word is sent to the summary extraction model so that the summary extraction model can fully understand the semantics of the first prompt word, thereby generating summary text explanation information corresponding to the text explanation information according to the requirements of the first prompt word, thereby improving the accuracy of the summary text explanation information.
[0297] The embodiments of the present disclosure are described above Figure 7 and Figure 8 Based on the illustrated embodiment, a question-answering method is also provided. Figure 10 is a flow chart of a question-answering method according to an exemplary embodiment. Figure 10 As shown, the method is performed by an electronic device, and the method includes the following steps:
[0298] In step 1001, the electronic device divides the text explanation information into multiple text segments and generates text features corresponding to the multiple text segments.
[0299] The text feature is used to represent the semantics of a text segment. The same text segment corresponds to the same text feature, and similar text segments correspond to the same or similar text features.
[0300] In some embodiments, the text explanation information is the text explanation information in the first explanation video in the above-described embodiment, or is text explanation information obtained by preprocessing the text explanation information in the first explanation video, where the preprocessing is used to remove useless information from the text explanation information, such as information unrelated to the item or repeated text. The electronic device can set preprocessing rules and preprocess the text explanation information in the first explanation video according to the rules.
[0301] In step 1002 , the electronic device saves a plurality of text segments and corresponding plurality of text features into a database of items.
[0302] In some embodiments, the text segment is used as an index and the text feature is used as data corresponding to the index, and the text segment and the text feature are saved in a database of items.
[0303] In some embodiments, the electronic device establishes databases for different items respectively, and associates the item identifiers corresponding to the items with the databases, thereby distinguishing the databases of different items.
[0304] In the disclosed embodiment, text explanation information is used to explain the item. The text explanation information includes a lot of information about the item. Therefore, the text explanation information is processed, and the obtained text segments can be used as the corpus resources of the item. Subsequently, the corpus resources can be queried to answer questions about the item for the user.
[0305] In step 1003, the electronic device obtains question information, which is input through the display interface of the second explanation video.
[0306] In some embodiments, the electronic device displays the second explanation video, and question information can be obtained through the display interface of the second explanation video.
[0307] In some embodiments, after the electronic device releases the second explanation video, one or more devices can display the second explanation video, and can obtain question information through the display interface of the second explanation video, thereby sending a question and answer request including the question information to the electronic device, and the electronic device obtains the question information.
[0308] The question information includes at least one of text or voice, and the question information is input by the user in the search bar in the display interface of the second explanation video, or the user selects from multiple preset question information provided by the electronic device, or input in other ways.
[0309] In step 1004, the electronic device generates a problem feature corresponding to the problem information.
[0310] Among them, question features are used to represent the semantics of question information.
[0311] In some embodiments, the electronic device may use the same feature extraction method as step 1001 above to extract question features corresponding to the question information, so as to ensure that the correlation between the question features and the text features can accurately reflect the correlation between the question information and the text explanation information.
[0312] In step 1005, the electronic device searches the database for a text segment corresponding to the text feature associated with the question feature, and determines the searched text segment as associated information of the question information.
[0313] The database includes text features corresponding to multiple text segments, and the text segment corresponding to the text feature associated with the question feature can be regarded as associated information of the question information, which is likely to contain the answer to the question information, or contain information associated with the answer to the question information.
[0314] For example, the text segment is "This dress comes in three sizes: S, M, and L," and the user asks the question "What sizes are available for this dress?" This text segment is associated with the question and can be considered when generating the answer to the question. This increases the amount of information and improves the accuracy of the answer.
[0315] In some embodiments, the electronic device determines the degree of association between each text feature in the database and the question feature, sorts the multiple text features according to the degree of association, and selects the top k text features with the highest degree of association with the question feature from the multiple text features as the text features associated with the question feature, where k is an integer greater than 1. The degree of association between any two features can be the cosine similarity between the two features or other data that can represent feature association, which is not limited in the embodiments of the present disclosure.
[0316] In some embodiments, the electronic device associates the second explanation video with an item identifier, determines the item identifier associated with the second explanation video, determines a database corresponding to the item identifier, and then queries the database for a text segment.
[0317] In step 1006 , the electronic device displays answer information based on the question information and the associated information.
[0318] The electronic device uses the associated information as supplementary material for the question information, and combines the question information and the associated information to generate answer information that can answer the question information. If the question information is input into the electronic device, the electronic device displays the answer information. Alternatively, if the question information is sent to the electronic device by another device, the electronic device can display the answer information after the electronic device sends the answer information to the device.
[0319] In some embodiments, displaying the answer information includes: displaying the question information and the answer information below the display interface of the explanation video, or displaying the question information and the answer information in a question-and-answer window, which is displayed by triggering a question-and-answer entry in the display interface of the explanation video, or displaying the answer information in other ways.
[0320] In some embodiments, the electronic device creates a second prompt word, the second prompt word includes question information, related information and second task text, the second task text is used to instruct the question-answering model to generate answer information for the question information and related information, the second prompt word is input into the question-answering model, and the answer information is generated by the question-answering model.
[0321] The question-answering model can process prompt words, and the prompt words are used to issue tasks to the question-answering model. Therefore, the electronic device creates a second prompt word including question information, related information and second task text. The second task text represents the task requirements for the question-answering model. By issuing the second prompt word to the question-answering model, the question-answering model can understand that its task is to answer the question information. Therefore, the question-answering model can generate answer information corresponding to the question information. In some embodiments, when creating the second prompt word, the electronic device can input the second prompt word into the question-answering model in a map_rerank manner. Among them, the map_rerank method refers to taking each piece of information in the related information as candidate answer information for the question information, and re-ranking them through the question-answering model to obtain the target answer information.
[0322] In some embodiments, the process of training the question-answering model includes: obtaining a large language model, which can be used for tasks such as speech translation, language recognition, and automatic question-answering. In order to improve the accuracy of the large language model in answering questions, sample data is obtained. The sample data includes sample prompt words and sample answer information corresponding to the sample question information in the sample prompt words. The sample prompt words include the sample question information, related information of the sample question information, and sample task text. The large language model is trained based on the sample data until the training completion conditions are met, and a trained question-answering model is obtained.
[0323] In other embodiments, if no relevant information for the question is found in the database, a default answer is displayed instead of using the question-answering model to generate an answer. The default answer may be, for example, "No relevant content found" to avoid inaccurate answer information being generated by the question-answering model.
[0324] In the related art, only the explanation video is played, and the user only watches the explanation video passively without being able to interact with the explainer, which lacks interactivity.
[0325] In the embodiment of the present disclosure, the electronic device saves multiple text segments and corresponding text features in the text explanation information into the database of the item as the corpus resource of the item. When the user asks a question about the item, the electronic device can combine the associated information related to the question information in the database to generate answer information corresponding to the question information, thereby realizing an intelligent question-and-answer function, expanding the function of the explanation video, and being able to use the text explanation information in the explanation video to quickly answer the user's questions about the item, thereby improving interactivity and user participation and reducing the cost of information access.
[0326] In addition, by creating a second prompt word including question information, related information and second task text, the second prompt word is sent to the question-answering model so that the question-answering model can fully understand the semantics of the second prompt word, and thus generate answer information corresponding to the question information according to the requirements of the second prompt word, thereby improving the accuracy of the answer information.
[0327] Based on the above embodiments, in some embodiments, the method further includes: determining a segment theme of at least one segment in the second explanation video, adding at least one segment theme in the second explanation video, associating each segment theme with a corresponding segment, and the association is used to jump to the segment corresponding to the segment theme when the segment theme is triggered.
[0328] The second explanation video can be divided into multiple segments of fixed duration, or divided into multiple segments through content recognition, or divided into segments using other methods. The segment theme is used to represent the theme of the segment, such as style design, wearing effect, detail display, matching recommendation, etc. Based on the segment theme, the general explanation content of the segment can be obtained.
[0329] After determining the segment theme of at least one segment in the second explanation video, at least one segment theme is added to the second explanation video, and the segment theme is associated with the corresponding segment. In this way, the at least one segment theme will be displayed when the second explanation video is displayed. After triggering any segment theme, you can jump to the segment corresponding to the segment theme without playing the entire second explanation video, which is convenient for users to locate the segment corresponding to the theme they want to watch in a timely manner.
[0330] Among them, there are many ways to add segment topics, such as adding at least one segment topic as a directory to the first video screen of the second explanation video, or adding at least one segment topic to the sidebar of the second explanation video, or adding at least one segment topic to the bottom of the video screen in the second explanation video, etc.
[0331] In other embodiments, the method further includes: the electronic device determining popular segments in the second explanation video, and based on the popular segments, deleting non-popular segments in the second explanation video to remove segments with low information content and retain key segments, thereby reducing the user's viewing cost. A popular segment is a segment that meets a target condition, such as a segment with a playback volume exceeding a target playback volume or a segment mentioned more than a target number of times in the comments of the second explanation video.
[0332] After deleting the non-hot segments, the electronic device determines the segment themes of the remaining segments in the second explanation video, adds at least one segment theme to the second explanation video, and associates each segment theme with a corresponding segment.
[0333] It should be noted that the above embodiment is only for adding segment themes or deleting non-hot segments for the second explanation video. The above processing can also be performed on the first explanation video, and the embodiment of the present disclosure does not limit this.
[0334] In the disclosed embodiment, an intelligent question-and-answer function is implemented. By inputting question information through the display interface of the explanation video, the answer information corresponding to the question information can be displayed, which expands the function of the explanation video. The explanation video can be used to quickly answer the user's questions about the item, thereby improving interactivity.
[0335] In addition, after determining the segment theme of at least one segment in the explanation video, at least one segment theme is added to the explanation video, and the segment theme is associated with the corresponding segment. In this way, the at least one segment theme will be displayed when the explanation video is displayed. After triggering any segment theme, you can jump to the segment corresponding to the segment theme without playing the entire explanation video, which is convenient for users to locate the segment corresponding to the theme they want to watch in a timely manner.
[0336] In summary, the method provided in the above embodiments, the embodiment of the present disclosure also provides an exemplary operation process, see Figure 11 , the operation process includes:
[0337] First, generate voice explanation information:
[0338] 1. The electronic device extracts voice explanation information from the first explanation video, converts the voice explanation information into text explanation information, and then extracts summary text explanation information corresponding to the text explanation information, thereby shortening the text length.
[0339] 2. Extract sound features from the voice explanation information and use the sound features to represent the voice of the speaker.
[0340] By combining the summary text explanation information and the sound characteristics, voice explanation information is produced, in which the summary text explanation information is spoken in the voice of the explainer.
[0341] Second, generate video screen:
[0342] A 3D model is generated based on multiple images of the object. The text segments are segmented based on the 3D model and textual explanation information. After audio and video processing, a video image is obtained, which contains the 3D model and subtitles.
[0343] Third, realize intelligent question and answer:
[0344] When displaying the explanation video, the user enters question information, and similar recall is performed in the database based on the question information, and the related information of the question information is queried, so that the answer information is displayed based on the question information and the related information.
[0345] Figure 12 is a flow chart showing a method for presenting explanation information according to an exemplary embodiment. Figure 12 As shown, the method is performed by an electronic device, and the method includes the following steps:
[0346] In step 1201, the electronic device displays an explanation entrance of an item.
[0347] The explanation entrance is a control used to trigger the display of the explanation video. The explanation entrance can be a separately displayed button, or the explanation entrance can be a display element in the object image, such as a decorative pendant. Displaying the explanation entrance can not only facilitate the triggering of the display of the first explanation video, but also beautify the object image. Alternatively, the explanation entrance is one of the object information of the object. For example, an electronic device associates the object name with the first explanation video, making the object name the explanation entrance. After triggering the object name, the explanation video of the object is displayed without the need to display the explanation entrance separately, thereby saving display space.
[0348] In some embodiments, in addition to the item explanation entrance, the current interface also displays other item information of the item, such as item image, item name, item sales volume, item price, etc.
[0349] In step 1202, the electronic device displays at least one explanation video of the item in response to a triggering operation on the explanation entrance, where the explanation video is used to record explanation information of the item.
[0350] Among them, the items may include any type of daily necessities, cosmetics, clothing, accessories, etc., and the embodiments of the present disclosure do not limit the types of items. The explanation information of the items may include the appearance, price, name, sales quantity, producer information and other aspects of the explanation information, and the appearance may include style design, upper body effect, detail display, matching recommendation, etc. The price may include the original price, current price, price discount method, etc. The sales quantity may include the sales quantity in a recent period of time, the total sales quantity, the sales quantity during holidays, etc. The producer information includes the name or address of the producer, etc. The embodiments of the present disclosure do not limit the content of the explanation information. The explanation information of the items may include information displayed in the video screen, or information played in the audio, etc. The embodiments of the present disclosure do not limit the form of the explanation information.
[0351] The item has at least two explanation videos, the lengths of the at least two explanation videos are different, and the explanation information contained in the at least two explanation videos is at least partially the same. Then, at least one explanation video can be displayed by triggering the explanation entrance.
[0352] In some embodiments, the electronic device responds to the triggering operation of the explanation entrance, plays one of the explanation videos of the item, and displays the video information of other explanation videos. The video information includes the video cover, video title, video content introduction, explanation entrance, etc., so that the user can understand the general content of the other explanation videos, and the user can trigger the explanation entrance corresponding to any other explanation video to switch to playing the explanation video.
[0353] In other embodiments, the electronic device responds to the triggering operation of the explanation entrance and displays the video information of each explanation video of the item. The video information includes the video cover, video title, video content introduction, explanation entrance, etc., so that the user can understand the general content of each explanation video and trigger the explanation entrance corresponding to any explanation video to play the explanation video.
[0354] In some embodiments, at least two explanation videos include a first explanation video and a second explanation video, the first explanation video and the second explanation video are generated based on the same historical explanation event, the first explanation video is used to record the historical explanation event, and the second explanation video is used to record the essence of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0355] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video, and the second explanation video includes at least one of the following: essential information in the first explanation video; detailed information of the items in the first explanation video; information belonging to the target topic in the first explanation video.
[0356] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video; the second explanation video is a video clip extracted from the first explanation video; or, the second explanation video is a video generated again based on the first explanation video.
[0357] In some embodiments, at least one explanation video of an item is displayed, including: displaying a first explanation video of the item, and switching the first explanation video of the item to a second explanation video for display in response to an explanation information switching instruction; or displaying a second explanation video of the item, and switching the second explanation video of the item to the first explanation video for display in response to an explanation information switching instruction.
[0358] In some embodiments, the instruction to switch the explanation information is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switch control.
[0359] In some embodiments, displaying the explanation entrance of an item includes: displaying the explanation entrance on the exhibit interface of the item; the exhibit interface of the item includes at least one of the following: an item details interface, an item list containing the item, a search result interface containing the item, and a recommendation result interface containing the item.
[0360] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video; the first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
[0361] It should be noted that the processes of the above embodiments are similar to those of the above Figure 2 The process of the embodiment shown is similar and will not be described again here.
[0362] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video; the first explanation video includes text explanation information of the item; the second explanation video is generated based on the summary text explanation information corresponding to the text explanation information, and the second explanation video is used to explain the item according to the summary text explanation information, and the length of the summary text explanation information is less than the length of the text explanation information. In some embodiments, the process of generating the second explanation video is the same as described above. Figure 7 and Figure 8 The same is true for the embodiment shown, which will not be described again here.
[0363] In the method provided by the embodiment of the present disclosure, the second explanation video and the first explanation video are used to record the explanation information of the item, and the second explanation video is different from the first explanation video in length, and the explanation information contained in the second explanation video is at least partially the same, which provides users with diverse choices. When displaying the first explanation video of the item, the first explanation video can be switched to the second explanation video for display through the explanation information switching instruction, avoiding the user having to watch only one explanation video. When the user does not want to watch a complete explanation video, the user can switch to watching other explanation videos, which improves flexibility and thereby improves the explanation effect of the item.
[0364] Figure 13 1 is a structural block diagram of a device for displaying explanation information according to an exemplary embodiment. Figure 13 , the device comprises:
[0365] The video display unit 1301 is configured to display a first explanation video of the item, where the first explanation video is used to record explanation information of the item;
[0366] The video switching unit 1302 is configured to switch the first explanation video of the item to the second explanation video for display in response to the explanation information switching instruction;
[0367] The second explanation video is different in length from the first explanation video. The second explanation video is used to record explanation information about the item. The explanation information contained in the second explanation video is at least partially the same as the explanation information contained in the first explanation video.
[0368] In some embodiments, the first explanation video and the second explanation video are generated based on the same historical explanation event;
[0369] The first explanation video is used to record historical explanation events;
[0370] The second explanation video is used to record the essential information of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0371] In some embodiments, the second explanation video is a video clip extracted from the first explanation video; or
[0372] The second explanation video is a video generated again based on the first explanation video.
[0373] In some embodiments, the video display unit 1301 includes:
[0374] The entrance display subunit is configured to display an explanation entrance on the exhibit interface of the item;
[0375] The trigger display subunit is configured to display the first explanation video in response to a trigger operation on the explanation entrance.
[0376] In some embodiments, the item display interface includes at least one of the following: an item details interface, an item list including the item, a search result interface including the item, and a recommendation result interface including the item.
[0377] In some embodiments, multiple explanation entrances are displayed on the exhibit interface of the item, the first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
[0378] In some embodiments, the instruction to switch the explanation information is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switch control.
[0379] In some embodiments, the apparatus further comprises:
[0380] an information conversion unit configured to convert the first voice explanation information in the first explanation video into text explanation information;
[0381] The video generation unit is configured to generate a second explanation video of the item based on the text explanation information.
[0382] In some embodiments, the video generation unit is configured to generate summary text explanation information corresponding to the text explanation information, and the length of the summary text explanation information is less than the length of the text explanation information; based on the summary text explanation information, a second explanation video of the item is generated, and the second explanation video is used to explain the item according to the summary text explanation information.
[0383] In some embodiments, the video generation unit is configured to extract sound features from the first voice explanation information, where the sound features are used to represent the voice of the explainer; generate second voice explanation information based on the summary text explanation information and the sound features, in which the summary text explanation information is spoken in the explainer's voice; and generate a second explanation video based on the second voice explanation information.
[0384] In some embodiments, the video generation unit is configured to divide the summary text explanation information into multiple summary text segments; based on the multiple summary text segments and sound features, generate multiple voice segments respectively, in which each summary text segment is spoken in the voice of the commentator; and splice the multiple voice segments to obtain second voice explanation information.
[0385] In some embodiments, the video generation unit is configured to generate multiple video frames based on multiple summary text segments, where the subtitles in each video frame are a summary text segment; and synthesize the second voice explanation information and the multiple video frames into a second explanation video, wherein the voice segment corresponding to the same summary text segment in the second explanation video is synchronized with the video frame.
[0386] In some embodiments, the apparatus further comprises:
[0387] A text segmentation unit configured to segment the text explanation information into a plurality of text segments;
[0388] a text feature generating unit configured to generate text features corresponding to a plurality of text segments respectively;
[0389] The storage unit is configured to store the plurality of text segments and the corresponding plurality of text features in a database of items.
[0390] In some embodiments, the apparatus further comprises:
[0391] a question acquisition unit configured to acquire question information, the question information being input through a display interface of the second explanation video;
[0392] A feature generating unit is configured to generate a question feature corresponding to the question information;
[0393] A query unit is configured to query a database for a text segment corresponding to a text feature associated with the question feature;
[0394] a correlation determination unit configured to determine the queried text segment as correlation information of the question information;
[0395] The answer display unit is configured to display answer information based on the question information and the associated information.
[0396] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0397] Figure 14 1 is a structural block diagram of a device for displaying explanation information according to an exemplary embodiment. Figure 14 , the device comprises:
[0398] The entrance display unit 1401 is configured to display the entrance of the item explanation;
[0399] The video display unit 1402 is configured to display at least one explanation video of the item in response to a trigger operation on the explanation entrance, where the explanation video is used to record explanation information of the item;
[0400] The item has at least two explanation videos, the lengths of the at least two explanation videos are different, and the explanation information contained in the at least two explanation videos is at least partially the same.
[0401] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0402] The first explanation video and the second explanation video are generated based on the same historical explanation event;
[0403] The first explanation video is used to record historical explanation events;
[0404] The second explanation video is used to record the essential information of the historical explanation event, or the second explanation video is used to record derivative information of the historical explanation event.
[0405] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0406] The second tutorial video includes at least one of the following:
[0407] First, explain the essential information in the video;
[0408] First, explain the details of the items in the video;
[0409] First, explain the information in the video that pertains to the target topic.
[0410] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0411] The second explanation video is a video clip extracted from the first explanation video; or
[0412] The second explanation video is a video generated again based on the first explanation video.
[0413] In some embodiments, the video display unit 1402 includes:
[0414] The first display subunit is configured to display a first explanation video of the item, and in response to an explanation information switching instruction, switches the first explanation video of the item to a second explanation video for display; or
[0415] The second display subunit is configured to display the second explanation video of the item, and in response to the explanation information switching instruction, switch the second explanation video of the item to the first explanation video for display.
[0416] In some embodiments, the instruction to switch the explanation information is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture applied to a switch control.
[0417] In some embodiments, the entrance display unit 1401 is configured to display the explanation entrance on the exhibit interface of the item;
[0418] The item display interface includes at least one of the following: an item details interface, an item list containing the item, a search result interface containing the item, and a recommendation result interface containing the item.
[0419] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0420] The first explanation video includes text explanation information of the item;
[0421] The second explanation video is generated based on the summary text explanation information corresponding to the text explanation information. The second explanation video is used to explain the item according to the summary text explanation information. The length of the summary text explanation information is shorter than the length of the text explanation information.
[0422] In some embodiments, the at least two explanation videos include a first explanation video and a second explanation video;
[0423] The first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
[0424] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0425] The embodiments of the present disclosure are executed by an electronic device, which is a terminal or a server.
[0426] Figure 15 This is a block diagram of a terminal structure according to an exemplary embodiment. In some embodiments, terminal 1500 includes a mobile phone, a desktop computer, a laptop computer, a tablet computer, a PDA, or other terminals. Terminal 1500 may also be referred to as user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other names.
[0427] Typically, the terminal 1500 includes a processor 1501 and a memory 1502 .
[0428] In some embodiments, processor 1501 includes one or more processing cores, such as a quad-core processor, an octal-core processor, etc.
[0429] In some embodiments, memory 1502 includes one or more non-transitory computer-readable storage media. In some embodiments, the non-transitory computer-readable storage media in memory 1502 is used to store executable instructions, which are executed by processor 1501 to implement the method for presenting explanation information provided in the method embodiments of the present disclosure.
[0430] In some embodiments, terminal 1500 may also optionally include a peripheral device interface 1503 and at least one peripheral device. In some embodiments, processor 1501, memory 1502, and peripheral device interface 1503 are connected via a bus or signal lines. In some embodiments, each peripheral device is connected to peripheral device interface 1503 via a bus, signal lines, or circuit boards. Specifically, the peripheral device includes at least one of a radio frequency circuit 1504, a display screen 1505, a camera assembly 1506, an audio circuit 1507, a positioning assembly 1508, and a power supply 1509.
[0431] In some embodiments, the terminal 1500 further includes one or more sensors 1510 , including but not limited to: an acceleration sensor 1511 , a gyroscope sensor 1512 , a pressure sensor 1513 , an optical sensor 1514 , and a proximity sensor 1515 .
[0432] Those skilled in the art will understand that Figure 15 The structure shown in the figure does not constitute a limitation on the terminal 1500, and the terminal 1500 can include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0433] Figure 16This is a schematic diagram of the structure of a server according to an exemplary embodiment. The server 1600 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1601 and one or more memories 1602, wherein the memory 1602 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 1601 to implement the methods provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.
[0434] In an exemplary embodiment, a computer-readable storage medium including instructions is further provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method for presenting explanatory information in the above-described embodiment. For example, the computer-readable storage medium may be a ROM (Read Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, or an optical data storage device.
[0435] In an exemplary embodiment, a computer program product is further provided, including a computer program, which implements the method for displaying explanation information in the above embodiment when the computer program is executed by a processor.
[0436] In some embodiments, the computer program involved in the embodiments of the present disclosure may be deployed and executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected through a communication network. Multiple electronic devices distributed at multiple locations and interconnected through a communication network may constitute a blockchain system.
[0437] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0438] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for presenting information, characterized in that: include: Converting the first voice explanation information in the first explanation video into text explanation information; If the length of the text explanation information is greater than the target length of the summary extraction model, the text explanation information is segmented into a plurality of text segments, each of which is less than the target length. The summary extraction model is used to extract summary text corresponding to the text to shorten the text length. The target length is the maximum length of text allowed by the summary extraction model. Creating a plurality of prompt words, each prompt word including a text segment and a first task word, wherein the first task word is used to instruct the summary extraction model to generate summary text explanation information for the text explanation information; inputting the plurality of prompt words corresponding to the plurality of text segments into the summary extraction model in the order of the plurality of text segments, and generating a summary text segment corresponding to each text segment by the summary extraction model; splicing the plurality of summary text segments in the order of the plurality of text segments to obtain summary text explanation information corresponding to the text explanation information, wherein the length of the summary text explanation information is shorter than the length of the text explanation information and the summary text explanation information is a summary of the text explanation information; extracting a sound feature from the first voice explanation information, wherein the sound feature is used to represent the voice of the explainer; generating second voice explanation information based on the summary text explanation information, the derivative information of the historical explanation event, and the voice feature, wherein the second voice explanation information is the summary text explanation information spoken in the voice of the explainer; Determining the number of required video frames based on the duration of the second voice explanation information, extracting the required number of video frames from the video frames of the first explanation video, and combining the second voice explanation information and the extracted plurality of video frames into a second explanation video for recording explanation information about the item; the first explanation video and the second explanation video are generated based on the same historical explanation event, the first explanation video is used to record the explanation information about the item targeted in the historical explanation event, and the second explanation video is used to record derivative information of the historical explanation event, the derivative information being information derived from the explanation information of the historical explanation event, the derivative information including different information from the explanation information, and the length of the second explanation video being shorter than the length of the first explanation video; Determining a segment theme of at least one segment in the second explanation video, adding at least one segment theme to the second explanation video, and associating each segment theme with a corresponding segment, wherein the association is used to jump to the segment corresponding to the segment theme when the segment theme is triggered; The first explanation video of the displayed item, wherein the playback interface of the first explanation video displays a playback entry for the second explanation video; In response to the explanation information switching instruction, the first explanation video of the item is switched to the second explanation video for display.
2. The method according to claim 1, characterized in that The second explanation video is used to record the essential information of the historical explanation event.
3. The method according to claim 1 or 2, characterized in that The second explanation video also includes a video clip extracted from the first explanation video; or The second explanation video also includes a video generated again based on the first explanation video.
4. The method according to claim 1, wherein The first explanation video of the displayed item includes: An explanation entrance is displayed on the exhibition interface of the item, and the first explanation video is displayed in response to a triggering operation on the explanation entrance.
5. The method according to claim 4, characterized in that The item display interface includes at least one of the following: an item details interface, an item list including the item, a search result interface including the item, and a recommendation result interface including the item.
6. The method according to claim 1 or 4, characterized in that A plurality of explanation entrances are displayed on the exhibit interface of the item, the first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
7. The method according to claim 1, characterized in that The explanation information switching instruction is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture acting on a switching control.
8. The method according to claim 1, characterized in that The method further comprises: Based on the plurality of summary text segments and the voice features, a plurality of voice segments are generated respectively, wherein each of the summary text segments is spoken in the voice of the speaker; The plurality of voice segments are concatenated to obtain the second voice explanation information.
9. The method according to claim 8, characterized in that The method further comprises: The second voice explanation information and the plurality of video frames are combined into the second explanation video, wherein the voice segment corresponding to the same summary text segment in the second explanation video is synchronized with the video frame.
10. A method for presenting information, characterized in that: The method comprises: Converting the first voice explanation information in the first explanation video into text explanation information; If the length of the text explanation information is greater than the target length of the summary extraction model, the text explanation information is segmented into a plurality of text segments, each of which is less than the target length. The summary extraction model is used to extract summary text corresponding to the text to shorten the text length. The target length is the maximum length of text allowed by the summary extraction model. Creating a plurality of prompt words, each prompt word including a text segment and a first task word, wherein the first task word is used to instruct the summary extraction model to generate summary text explanation information for the text explanation information; inputting the plurality of prompt words corresponding to the plurality of text segments into the summary extraction model in the order of the plurality of text segments, and generating a summary text segment corresponding to each text segment by the summary extraction model; splicing the plurality of summary text segments in the order of the plurality of text segments to obtain summary text explanation information corresponding to the text explanation information, wherein the length of the summary text explanation information is shorter than the length of the text explanation information and the summary text explanation information is a summary of the text explanation information; extracting a sound feature from the first voice explanation information, wherein the sound feature is used to represent the voice of the explainer; generating second voice explanation information based on the summary text explanation information, the derivative information of the historical explanation event, and the voice feature, wherein the second voice explanation information is the summary text explanation information spoken in the voice of the explainer; Determining the number of required video frames based on the duration of the second voice explanation information, extracting the required number of video frames from the video frames of the first explanation video, and combining the second voice explanation information and the extracted plurality of video frames into a second explanation video for recording explanation information about the item; the first explanation video and the second explanation video are generated based on the same historical explanation event, the first explanation video is used to record the explanation information about the item targeted in the historical explanation event, and the second explanation video is used to record derivative information of the historical explanation event, the derivative information being information derived from the explanation information of the historical explanation event, the derivative information including different information from the explanation information, and the length of the second explanation video being shorter than the length of the first explanation video; Determining a segment theme of at least one segment in the second explanation video, adding at least one segment theme to the second explanation video, and associating each segment theme with a corresponding segment, wherein the association is used to jump to the segment corresponding to the segment theme when the segment theme is triggered; Display the entry for explaining the item; In response to a triggering operation on the explanation entrance, displaying at least one explanation video of the item, the explanation video being used to record explanation information of the item; the at least one explanation video includes the first explanation video and the second explanation video; The at least one explanation video showing the item includes: The first explanation video is displayed, and in response to the explanation information switching instruction, the first explanation video is switched to the second explanation video for display, and the playback entrance of the second explanation video is displayed in the playback interface of the first explanation video.
11. The method according to claim 10, characterized in that The second explanation video is used to record the essential information of the historical explanation event.
12. The method according to claim 10, characterized in that The second explanation video further includes at least one of the following: The first explanation video contains the essential information; Detailed information about the item in the first explanation video; The first explanation video contains information related to the target topic.
13. The method according to claim 10, characterized in that The second explanation video also includes a video clip extracted from the first explanation video; or The second explanation video also includes a video generated again based on the first explanation video.
14. The method according to claim 10, characterized in that The at least one explanation video showing the item further includes: Display the second explanation video of the item, and in response to the explanation information switching instruction, switch the second explanation video of the item to the first explanation video for display.
15. The method according to claim 14, characterized in that The explanation information switching instruction is triggered based on the following operations: a predetermined gesture input in a predetermined area, and / or a predetermined gesture acting on a switching control.
16. The method according to claim 10, characterized in that The display of the explanation entrance of the item includes: Displaying the explanation entrance on the exhibit interface of the item; The item display interface includes at least one of the following: an item details interface, an item list including the item, a search result interface including the item, and a recommendation result interface including the item.
17. The method according to claim 10, wherein: The first explanation video and the second explanation video correspond to different explanation entrances, and the first explanation video and the second explanation video are triggered for display through their corresponding explanation entrances respectively.
18. A display device for explaining information, characterized in that: include: an information conversion unit configured to convert the first voice explanation information in the first explanation video into text explanation information; a text segmentation unit configured to, if the length of the text explanation information is greater than a target length of a summary extraction model, segment the text explanation information into a plurality of text segments, each of which is less than the target length, and wherein the summary extraction model is configured to extract summary text corresponding to the text to shorten the text length, wherein the target length is the maximum length of text allowed to be input by the summary extraction model; The video generation unit is configured to create a plurality of prompt words, each prompt word including a text segment and a first task text, the first task text being used to instruct the summary extraction model to generate summary text explanation information for the text explanation information; input a plurality of prompt words corresponding to the plurality of text segments into the summary extraction model in the order of the plurality of text segments, and generate a summary text segment corresponding to each text segment by the summary extraction model; concatenate the plurality of summary text segments in the order of the plurality of text segments to obtain summary text explanation information corresponding to the text explanation information, wherein the length of the summary text explanation information is less than the length of the text explanation information and the summary text explanation information is a summary of the text explanation information; extract sound features from the first voice explanation information, wherein the sound features are used to represent the voice of the speaker; generating second voice explanation information based on the summary text explanation information, derivative information of the historical explanation event, and the sound features, wherein the summary text explanation information is spoken in the second voice explanation information in the voice of the explainer; determining a required number of video frames according to the duration of the second voice explanation information, extracting the required number of video frames from the video frames of the first explanation video, and combining the second voice explanation information and the extracted multiple video frames into a second explanation video for recording explanation information of the item; the first explanation video and the second explanation video are generated based on the same historical explanation event, the first explanation video is used to record explanation information of the item targeted by the historical explanation event, and the second explanation video is used to record derivative information of the historical explanation event, the derivative information being information derived from the explanation information of the historical explanation event, the derivative information including different information from the explanation information, and the length of the second explanation video being shorter than the length of the first explanation video; determining a segment theme of at least one segment in the second explanation video, adding at least one segment theme to the second explanation video, and associating each segment theme with a corresponding segment, wherein the association is used to jump to the segment corresponding to the segment theme when a segment theme is triggered; a video display unit configured to display the first explanation video of the item, wherein a playback interface of the first explanation video displays a playback entry for the second explanation video; The video switching unit is configured to switch the first explanation video of the item to the second explanation video for display in response to the explanation information switching instruction.
19. A display device for explaining information, characterized in that: include: an information conversion unit configured to convert the first voice explanation information in the first explanation video into text explanation information; a text segmentation unit configured to, if the length of the text explanation information is greater than a target length of a summary extraction model, segment the text explanation information into a plurality of text segments, each of which is less than the target length, and wherein the summary extraction model is configured to extract summary text corresponding to the text to shorten the text length, wherein the target length is the maximum length of text allowed to be input by the summary extraction model; The video generation unit is configured to create a plurality of prompt words, each prompt word including a text segment and a first task text, the first task text being used to instruct the summary extraction model to generate summary text explanation information for the text explanation information; input a plurality of prompt words corresponding to the plurality of text segments into the summary extraction model in the order of the plurality of text segments, and generate a summary text segment corresponding to each text segment by the summary extraction model; concatenate the plurality of summary text segments in the order of the plurality of text segments to obtain summary text explanation information corresponding to the text explanation information, wherein the length of the summary text explanation information is less than the length of the text explanation information and the summary text explanation information is a summary of the text explanation information; extract sound features from the first voice explanation information, wherein the sound features are used to represent the voice of the speaker; generating second voice explanation information based on the summary text explanation information, derivative information of the historical explanation event, and the sound features, wherein the summary text explanation information is spoken in the second voice explanation information in the voice of the explainer; determining a required number of video frames according to the duration of the second voice explanation information, extracting the required number of video frames from the video frames of the first explanation video, and combining the second voice explanation information and the extracted multiple video frames into a second explanation video for recording explanation information of the item; the first explanation video and the second explanation video are generated based on the same historical explanation event, the first explanation video is used to record explanation information of the item targeted by the historical explanation event, and the second explanation video is used to record derivative information of the historical explanation event, the derivative information being information derived from the explanation information of the historical explanation event, the derivative information including different information from the explanation information, and the length of the second explanation video being shorter than the length of the first explanation video; determining a segment theme of at least one segment in the second explanation video, adding at least one segment theme to the second explanation video, and associating each segment theme with a corresponding segment, wherein the association is used to jump to the segment corresponding to the segment theme when a segment theme is triggered; an entrance display unit, configured to display an explanation entrance of the item; a video display unit configured to display at least one explanation video of the item in response to a triggering operation on the explanation entrance, wherein the explanation video is used to record explanation information of the item; The at least one explanation video includes the first explanation video and the second explanation video; The video display unit includes: The first display subunit is configured to display the first explanation video, and in response to the explanation information switching instruction, the first explanation video is switched to the second explanation video for display, and the playback entrance of the second explanation video is displayed in the playback interface of the first explanation video.
20. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method for displaying explanation information according to any one of claims 1 to 9, or to implement the method for displaying explanation information according to any one of claims 10 to 17.
21. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for displaying explanation information as described in any one of claims 1 to 9, or the method for displaying explanation information as described in any one of claims 10 to 17.
22. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for displaying explanation information according to any one of claims 1 to 9, or implements the method for displaying explanation information according to any one of claims 10 to 17.
Citation Information
Patent Citations
Live broadcast content processing method, device and system
CN109429074A
Question and answer information processing method and system, computer equipment and storage medium
CN110308947A
Method and device for controlling video display
CN113301356A
Video abstract generation method and device, electronic equipment and storage medium
CN113992973A
A method, apparatus, computer equipment, and storage medium for displaying product information.
CN114936896A