Information display methods, devices, electronic equipment and storage media
By identifying and displaying related information from audio clips in advertising videos, the problem of monotonous information display is solved, enabling the conversion between auditory and visual experiences, improving recommendation effectiveness, and reducing costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-12
- Publication Date
- 2026-03-31
AI Technical Summary
The information displayed for recommended items in existing advertising videos is too simplistic, resulting in poor recommendation effectiveness.
By acquiring the audio data of recommended videos, identifying target audio segments containing preset keywords, and displaying related information when the video is played on the client side, the system achieves a conversion between auditory and visual dimensions.
It displays diverse information, strengthens users' impression of the audio in recommended videos, improves recommendation effectiveness, expands the scope of application, and reduces production costs.
Smart Images

Figure CN115967822B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an information display method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, when playing advertising videos, the recommended item information can be displayed on the video playback page, such as an item image or purchase link, to facilitate viewing and purchasing the recommended item.
[0003] However, the product information displayed in existing advertising videos is relatively limited, resulting in poor recommendation effectiveness. Summary of the Invention
[0004] This disclosure provides an information display method, apparatus, electronic device, and storage medium to display diverse information and improve the recommendation effect of videos.
[0005] In a first aspect, embodiments of this disclosure provide an information display method, including:
[0006] Obtain the audio data of the recommended videos for the target audience;
[0007] Based on the audio data, the audio segment information of the target audio segment in the recommended video is determined, wherein the audio segment information includes identification information and association information, and the target audio segment contains preset keywords;
[0008] When an information retrieval request for the recommended video is received, the audio segment information is sent to the client so that the client displays the associated information when playing the target video segment corresponding to the target audio segment, wherein the information retrieval request is sent by the client.
[0009] Secondly, this disclosure also provides an information display method, including:
[0010] Send a request to the server to obtain information about recommended videos for a target object, and receive audio segment information returned by the server based on the information obtaining request. The audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords.
[0011] Play the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, display the associated information in the first display area of the video playback page.
[0012] Thirdly, embodiments of this disclosure also provide an information display device, including:
[0013] The data acquisition module is used to acquire the audio data of the recommended videos for the target object;
[0014] The information determination module is used to determine the audio segment information of the target audio segment in the recommended video based on the audio data, wherein the audio segment information includes identification information and association information, and the target audio segment contains preset keywords;
[0015] An information sending module is used to send the audio segment information to the client when it receives an information retrieval request for the recommended video, so that the client displays the associated information when playing the target video segment corresponding to the target audio segment, wherein the information retrieval request is sent by the client.
[0016] Fourthly, embodiments of this disclosure also provide an information display device, including:
[0017] The information receiving module is used to send an information acquisition request for recommended videos targeting a target object to the server, and to receive audio segment information returned by the server based on the information acquisition request. The audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords.
[0018] The video playback module is used to play the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, the associated information is displayed in the first display area of the video playback page.
[0019] Fifthly, embodiments of this disclosure also provide an electronic device, including:
[0020] One or more processors;
[0021] Memory, used to store one or more programs.
[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the information display method as described in the embodiments of this disclosure.
[0023] Sixthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the information display method as described in embodiments of this disclosure.
[0024] The information display method, apparatus, electronic device, and storage medium provided in this disclosure acquire audio data of a recommended video for a target object; determine the identification information and association information of a target audio segment containing preset keywords in the recommended video based on the audio data; and when receiving an information acquisition request for the recommended video from a client, send the identification information and association information of the target audio segment to the client, so that the client displays the association information when playing the target video segment corresponding to the target audio segment. By adopting the above technical solution, this disclosure controls the client to display the association information of the specific audio segment containing preset keywords on the video playback page, which not only displays diverse information but also achieves a conversion between auditory and visual dimensions, strengthening the user's impression of the audio broadcast in the recommended video, improving the recommendation effect of the recommended video, expanding the applicability of the recommended video, and reducing the production cost of the recommended video. Attached Figure Description
[0025] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0026] Figure 1 A flowchart illustrating an information display method provided in an embodiment of this disclosure;
[0027] Figure 2 A flowchart illustrating another information display method provided in this embodiment of the disclosure;
[0028] Figure 3 A flowchart illustrating yet another information display method provided in this disclosure embodiment;
[0029] Figure 4 This is a schematic diagram illustrating a display method for associated information provided in an embodiment of this disclosure;
[0030] Figure 5 This is a schematic diagram illustrating another method of displaying information associated with other information provided in this embodiment of the disclosure;
[0031] Figure 6 A flowchart illustrating the fourth information display method provided in this embodiment of the disclosure;
[0032] Figure 7 This is a schematic diagram illustrating a method for moving associated information according to an embodiment of this disclosure;
[0033] Figure 8 This is a schematic diagram illustrating another method of moving associated information provided in an embodiment of this disclosure;
[0034] Figure 9 This is a schematic diagram of a second display area provided in an embodiment of the present disclosure;
[0035] Figure 10 This is a schematic diagram of another second display area provided in an embodiment of the present disclosure;
[0036] Figure 11 A structural block diagram of an information display device provided in an embodiment of this disclosure;
[0037] Figure 12 A structural block diagram of another information display device provided in this disclosure embodiment;
[0038] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0039] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0040] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0041] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0042] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0043] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0044] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0045] Figure 1 This is a flowchart illustrating an information display method provided in an embodiment of this disclosure. The method can be executed by an information display device, which can be implemented in software and / or hardware and can be configured in an electronic device, typically in a server. The information display method provided in this disclosure is applicable to scenarios where information is displayed based on audio data from recommended videos. Figure 1 As shown, the information display method provided in this embodiment may include:
[0046] S101. Obtain the audio data of the recommended video for the target object.
[0047] The recommended video can be a video that recommends a specific object to the user, such as an advertising video or a product promotion video. Correspondingly, the target object can be any object recommended in the recommended video, such as an activity or item recommended in the video. This item can be a real item or a virtual item; this embodiment does not impose any restrictions on this.
[0048] In this embodiment, the server can pre-determine the audio segment information of the target audio segment in the recommended video of the target object, so that when the client plays the recommended video, it can determine the target audio segment based on the audio segment information and display the associated information.
[0049] Specifically, when the current conditions meet the preset conditions for determining the audio segment information of the target audio segment in the recommended video of the target object, such as when a recommended video is a newly uploaded recommended video, that is, when the recommended video is a recommended video whose audio segment information has not yet been determined, the server can obtain the audio data of the recommended video at the current moment and determine the audio segment information of the target audio segment in the recommended video based on the audio data.
[0050] S102. Determine the audio segment information of the target audio segment in the recommended video based on the audio data, wherein the audio segment information includes identification information and association information, and the target audio segment contains preset keywords.
[0051] The target audio segment can be an audio segment containing preset keywords in the audio data of the recommended video, that is, an audio segment whose corresponding text content contains preset keywords, such as the audio segment corresponding to the phrase (including sentence fragment) where the preset keyword is located. The preset keyword can be a preset type of keyword, such as keywords related to discount information (e.g., clearance, cost-effectiveness, price reduction, bargain, wholesale, instant reduction, discount, cheap, limited time, etc., or statements related to discount information), keywords related to descriptive information (e.g., words or statements related to size, color, shape, face value, etc.). The following explanation uses keywords related to discount information as an example. The identification information of the target audio segment is information that can be used to identify the target audio segment, such as the target audio segment flag, the target audio segment's ID, or the start and end time information of the target audio segment in the recommended video. The associated information of the target audio segment can be information related to the content played in the target audio segment, such as the content played by the target audio segment, the summary information or key information of the content played by the target audio segment, or information related to the object recommended by the target audio segment (i.e., the target object) (e.g., discount information), etc.
[0052] In this embodiment, when the server obtains the audio data of the recommended video, it can determine whether there is a target audio segment containing a preset keyword in the recommended video based on the audio data, and when it determines that there is a target audio segment, it further determines the identification information and association information of the target audio segment.
[0053] For example, after obtaining the audio data of the recommended video of the target object, the server can perform speech recognition on the audio data to obtain the speech recognition text corresponding to the audio data. It can then determine whether the speech recognition text contains a preset keyword. If so, it can determine that there is a target audio segment containing the preset keyword in the recommended video. Furthermore, it can determine the audio segment information of the target audio segment based on the position of the preset keyword. For example, the audio segment corresponding to the text statement (including text statement fragments) where the preset keyword is located can be determined as the target audio segment. Identification information can be set for the target audio segment, and further, the association information of the target audio segment can be obtained. For example, the summary information of the text statement / statement fragment where the preset keyword is located can be extracted as the association information of the target audio segment.
[0054] In addition, after obtaining the speech recognition text corresponding to the audio data of the recommended video of the target object, the server can first divide the speech recognition text into at least one text sentence according to the pre-set segmentation rules, such as according to grammatical and semantic rules and / or punctuation marks, and determine whether there is a text sentence containing a preset keyword in the at least one text sentence. If so, it is determined that there is a target audio segment containing the preset keyword in the recommended video, and the audio segment corresponding to the text sentence containing the preset keyword is identified as the target audio segment. Furthermore, the audio segment information of the target audio segment is determined, such as setting identification information for the target audio segment and extracting the summary information of the text sentence / sentence segment where the preset keyword is located as the association information of the target audio segment.
[0055] In this embodiment, the audio segment information of the target audio segment in the recommended video can be determined solely based on the speech recognition text of the recommended video; alternatively, the audio segment information of the target audio segment in the recommended video can be determined based on both the speech recognition text and the audio data features of the recommended video, thereby further improving the accuracy of the determined audio segment information. In this case, determining the audio segment information of the target audio segment in the recommended video based on the audio data may include: performing speech recognition on the audio data to obtain speech recognition text, and acquiring the Mel-frequency cepstral coefficient feature vector of the audio data; determining the time information corresponding to each word in the speech recognition text in the audio data based on the Mel-frequency cepstral coefficient feature vector; and determining the audio segment information of the target audio segment in the recommended video based on the speech recognition text and the time information.
[0056] For example, after obtaining the audio data of the recommended video for the target object, the server can first perform speech recognition on the audio data to obtain the speech recognition text of the recommended video; and then extract the Mel-scale Frequency Cepstral Coefficients (MFCC) feature vector of the audio data. For example, based on the audio data, the server obtains the audio waveform information of the recommended video, performs frame processing on the audio waveform information at fixed time intervals, and after the frame processing is completed, it uses a sliding window of fixed time length to divide the audio waveform information into multiple audio waveform segments, and performs Discrete Fourier Transform on each audio waveform segment to obtain the frequency domain information of the audio data. Furthermore, it performs Mel filtering on the frequency domain information to obtain the MFCC feature vector of the audio data. Then, the obtained MFCC feature vector is used to determine the time information corresponding to each word in the audio recognition text in the audio data (i.e., in the recommended video). For example, the server inputs the audio recognition text and the MFCC feature vector into a pre-trained detection model, and obtains the start and end times of each word in the audio recognition text in the audio data as the time information corresponding to each word in the audio data, as output by the detection model. Therefore, the target audio segment and its audio segment information can be determined based on the speech recognition text and the time information of each word in the speech recognition text.
[0057] S103. When an information retrieval request for the recommended video is received, the audio segment information is sent to the client so that the client displays the associated information when playing the target video segment corresponding to the target audio segment, wherein the information retrieval request is sent by the client.
[0058] The target video segment can be the video segment corresponding to the target audio segment in the recommended video, that is, the video segment in the recommended video located between the start and end time nodes corresponding to the target audio segment in the recommended video. The information retrieval request can be a request to obtain audio segment information of the target audio segment in the recommended video, and it can also further allow the user to obtain the video data of the recommended video, such as the information retrieval request being a request to obtain the video data of the recommended video.
[0059] For example, a client can request video data (including audio data) of a recommended video to be played from the server via a video data retrieval request, and retrieve audio segment information of a target audio segment within the recommended video via an information retrieval request. For instance, when a client wants to play a recommended video of a target object, it can generate an information retrieval request carrying the video identifier of the recommended video before, after, or simultaneously with generating the video data retrieval request, and send this information retrieval request to the server. Correspondingly, upon receiving an information retrieval request from a client, the server can parse the request to obtain the video identifier carried within it, retrieve the audio segment information of the target audio segment within the recommended video based on the video identifier, and return this audio segment information to the client. Thus, the client can determine the target audio segment within the recommended video based on the identifier information in the audio segment information, play the recommended video, and, when playing to the target video segment corresponding to the target audio segment, display the associated information of the target audio segment contained in the audio segment information on the video playback page.
[0060] Alternatively, the client can also request video data of the recommended video and audio segment information of the target audio segment within the recommended video from the server via an information retrieval request. For example, the information retrieval request can be a video data retrieval request. For instance, the client can generate a video data retrieval request carrying the video identifier of the recommended video and send it to the server. Correspondingly, upon receiving a video data retrieval request from a client, the server can parse the request to obtain the video identifier, retrieve the video data of the recommended video and the audio segment information of the target audio segment within the recommended video based on the video identifier, and return the video data and audio segment information to the client. Thus, the client can determine the target audio segment in the recommended video based on the identifier information in the audio segment information, play the recommended video based on the video data, and when playing to the target video segment corresponding to the target audio segment, display the associated information of the target audio segment contained in the audio segment information on the video playback page.
[0061] In this embodiment, the server can pre-determine target audio segments containing preset keywords and their associated information in each recommended video, and send them to the client in response to the client's information retrieval request. Thus, when the client plays a recommended video, it can display information based on the audio played in the video. For example, when the video plays audio related to the discount information of the target object, the discount information of the target object is automatically displayed on the video playback page; and / or, when the video plays audio describing the target object, the descriptive information of the target object is automatically displayed on the video playback page, etc. This enables the conversion between auditory and visual information in the video, further enhancing the user's understanding of the auditory information in the video and improving the recommendation effect of the target object. Furthermore, since information display does not need to be based on the currently playing video frame, it is not necessary for the target object to be a displayable actual item or to appear on the video frame. It also does not require the video producer to fully understand the recommended video and determine when to display information in advance. This expands the applicability of recommended videos, enabling the recommendation of actual items and even virtual items, reducing the production cost of recommended videos, and thus reducing the recommendation cost of the target.
[0062] In this embodiment, after the server initially determines the audio segment information of the target audio segment in the recommended video of the target object, it can also periodically or periodically detect whether the video content (including audio content) of the recommended video has changed. When a change is detected, the server re-identifies the target audio segment in the recommended video and determines the audio segment information of the target audio segment based on the audio data of the recommended video after the change. This is to avoid situations where the video sound and the displayed associated information are misaligned or mismatched when the client plays the recommended video, and to ensure the real-time performance and effectiveness of the displayed associated information. Preferably, after determining the audio segment information of the target audio segment in the recommended video based on the audio data, the server further includes: when a change in the content of the recommended video is detected, re-determining the audio segment information of the target audio segment in the recommended video based on the audio data of the changed recommended video.
[0063] The information display method provided in this embodiment acquires the audio data of a recommended video for a target object; determines the identification information and association information of a target audio segment containing preset keywords in the recommended video based on the audio data; and when a client sends an information acquisition request for the recommended video, the identification information and association information of the target audio segment are sent to the client so that the client displays the association information when playing the target video segment corresponding to the target audio segment. By adopting the above technical solution, this embodiment controls the client to display the association information of the specific audio segment containing preset keywords on the video playback page. This not only displays diverse information but also achieves a conversion between auditory and visual dimensions, strengthening the user's impression of the audio broadcast in the recommended video, improving the recommendation effect of the recommended video, expanding the applicability of the recommended video, and reducing the production cost of the recommended video.
[0064] Figure 2 This is a flowchart illustrating another information display method provided in an embodiment of this disclosure. The solution in this embodiment can be combined with one or more optional solutions in the above embodiments. Optionally, determining the audio segment information of the target audio segment in the recommended video based on the speech-recognized text and the time information includes: segmenting the speech-recognized text based on the time information to obtain at least one text statement; identifying text statements containing preset keywords in the at least one text statement as candidate text statements; filtering the candidate text statements based on preset filtering rules to obtain a target text statement; using the audio segment corresponding to the target text statement in the audio data as the target audio segment in the recommended video, and determining the audio segment information of the target audio segment.
[0065] Correspondingly, such as Figure 2 As shown, the information display method provided in this embodiment may include:
[0066] S201. Obtain the audio data of the recommended video for the target object.
[0067] S202. Perform speech recognition on the audio data to obtain the speech recognition text, and obtain the Mel frequency cepstral coefficient feature vector of the audio data.
[0068] S203. Determine the time information corresponding to each character in the audio recognition text in the audio data based on the Mel frequency cepstral coefficient feature vector.
[0069] S204. Based on the time information, the speech recognition text is segmented to obtain at least one text sentence.
[0070] In this embodiment, when segmenting the speech recognition text of the recommended video, the start and end times of each word in the speech recognition text during the audio data can be further considered to improve the accuracy of the segmented text sentences, thereby improving the rationality of the subsequently determined target audio segments and the audio segment information of the target audio segments.
[0071] For example, a Named Entity Recognition (NER) model for identifying target audio segments in recommended videos can be pre-trained based on multiple training samples. Then, after obtaining the speech recognition text of the recommended video and the time information of each character in the speech recognition text, the speech recognition text and the time information can be input into the NER model. Based on the time information of each character in the speech recognition text, and further combined with syntactic and semantic information and punctuation, the NER model can segment the speech recognition text into sentences, obtaining at least one text sentence.
[0072] S205. Identify text statements containing preset keywords in the at least one text statement as candidate text statements.
[0073] Specifically, after dividing the speech recognition text into at least one text statement, for each text statement, it can be determined whether the text statement contains a preset keyword. If it does, the text statement is determined as a candidate text statement; otherwise, the text statement is treated as a non-candidate text statement.
[0074] In this embodiment, when determining whether each text statement contains a preset keyword, the preset keyword corresponding to each text statement can be a keyword from the same keyword list or a keyword from different keyword lists. For example, keyword lists corresponding to different industries can be preset. When determining whether a text statement contains a preset keyword, all keywords in each keyword list can be used as the preset keyword corresponding to the text statement, and it can be determined whether the text statement contains at least one preset keyword. Alternatively, the industry to which the text statement belongs can be determined first, and the keywords in the keyword list corresponding to that industry can be used as the preset keyword corresponding to the text statement, and it can be determined whether the text statement contains at least one preset keyword.
[0075] Taking promotional information as an example, since promotional information is expressed differently in different industries—for instance, price discounts in the gaming industry can be expressed as "gift codes," while in the e-commerce industry they can be expressed as "buy X get X free" or "spend XXX and get XXX off"—this embodiment preferably uses keywords corresponding to the industry to which each text statement belongs as preset keywords for that text statement to be recognized. For example, after the NER model divides the speech recognition text into at least one text statement, it can classify each text statement according to industry. After classification, it determines whether each text statement contains the preset keywords of its industry and selects the text statements containing the preset keywords of their industry as candidate text statements. Furthermore, by combining the time information of each word in the candidate text statement in the audio data, the start time and duration of each candidate statement in the audio data are obtained, and each candidate text statement, along with its start time and duration, are output. The industry label of each candidate text statement can also be output. Thus, the server can store each candidate text statement, its start time and duration, and its industry label output by the NER model for subsequent operations.
[0076] S206. The candidate text statements are filtered based on preset filtering rules to obtain the target text statements.
[0077] In this embodiment, after obtaining candidate text statements, the audio segments corresponding to each candidate text statement in the audio data of the recommended video can be directly used as the target recommended video segments of the recommended video; alternatively, each candidate text statement can be further filtered based on preset filtering rules to further improve the display effect of the associated information of the final target audio segments.
[0078] In this embodiment, the preset filtering rules can be flexibly set as needed. For example, candidate text statements can be filtered based on preset time ranges, statement lengths, industry types, and / or text quality. For instance, candidate text statements in audio data whose start and end times are within preset time ranges, whose statement lengths are within preset statement length ranges, whose industry types match the industry type of the recommended video, and whose text quality is higher than a preset text quality threshold can be selected as target text statements. The text quality of each candidate text statement can be determined by a pre-trained text quality evaluation model.
[0079] Understandably, when multiple candidate text statements that meet the requirements are obtained based on preset filtering rules, a segment of the candidate text statements (such as the first candidate text statement or the candidate text statement with the highest text quality) can be used as the target text statement; alternatively, all of the multiple candidate text statements can be used as the target text statement.
[0080] S207. The audio segment corresponding to the target text statement in the audio data is taken as the target audio segment in the recommended video, and the audio segment information of the target audio segment is determined, wherein the audio segment information includes identification information and association information.
[0081] Specifically, the audio segment corresponding to the target text statement can be determined from the audio data based on the start time and duration of the target text statement, and used as the target audio segment for the recommended video. Furthermore, the identification information and association information of the target audio segment can be determined.
[0082] When determining the target audio segment, for example, when there is only one target text statement, the audio segment in the audio data that has the same start and end time as the target text statement can be used as the target audio segment; when there are multiple target text statements, the audio segment in the audio data that has the same start and end time as the target text statement can be used for each target text statement, thereby obtaining multiple target audio segments; alternatively, the multiple target text statements can be sorted according to the order in which they are played in the audio data, and the audio segment whose start time is the start time of the first target text statement and whose end time is the end time of the last text statement can be used as the target audio segment. This embodiment does not limit this.
[0083] When determining the audio segment information of a target audio segment, for example, the identification information of the target audio segment can be generated according to a preset generation rule, or the time node information (such as start and end time information, or start time information and duration information, etc.) corresponding to the target audio segment can be used as the identification information of the target audio segment; and / or, the target text statement corresponding to the target audio segment or the summary information of the target text statement corresponding to the target audio segment can be used as the association information of the target audio segment. In this embodiment, it is preferable to use the time node information corresponding to the target audio segment as the identification information of the target audio segment and the target text statement corresponding to the target audio segment as the association information of the target audio segment, so as to further simplify the computational workload required to determine the audio segment information and improve the playback effect of the target audio segment. In this case, determining the audio segment information of the target audio segment may include: using the time node information corresponding to the target audio segment in the recommended video as the identification information of the target audio segment and using the target text statement as the association information of the target audio segment.
[0084] S208. When an information retrieval request for the recommended video is received, the audio segment information is sent to the client so that the client displays the associated information when playing the target video segment corresponding to the target audio segment, wherein the information retrieval request is sent by the client.
[0085] The information display method provided in this embodiment performs text segmentation based on the time information of each character in the audio data of the speech recognition text, determines candidate text sentences in the segmented text sentences, and filters the determined candidate text sentences. This can further improve the accuracy of the determined target audio segments, improve the effect of displaying the associated information of the target audio segments, and avoid causing too much interference to the user's video viewing, thereby improving the user's viewing experience.
[0086] Figure 3 This is a flowchart illustrating an information display method provided in an embodiment of this disclosure. The method can be executed by an information display device, which can be implemented in software and / or hardware and can be configured in an electronic device, typically a mobile phone or tablet. The information display method provided in this disclosure is applicable to scenarios where information is displayed based on audio data from recommended videos. Figure 3 As shown, the information display method provided in this embodiment may include:
[0087] S301. Send a request to the server to obtain information about recommended videos for the target object, and receive audio segment information returned by the server based on the information obtaining request, wherein the audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords.
[0088] In this embodiment, before playing a recommended video, the client can obtain the association information of the target audio segment in the recommended video from the server, so as to display the association information to the user when playing the target video segment corresponding to the target audio segment.
[0089] For example, a user can watch a video through a video playback page. Correspondingly, the client can display the video playback page, switch the video played on the page based on the user's video switching operation, or automatically switch the video played on the page according to preset switching rules. When it is necessary to switch to a recommended video for a specific object (i.e., the target object), the client generates an information data retrieval request, or generates both a video data retrieval request and an information retrieval request, and sends them to the server. The server can then retrieve the video data and audio segment information of the recommended video based on the information retrieval request, or retrieve the video data and audio segment information of the recommended video based on the video data retrieval request, and return them to the client. Thus, the client can receive the video data of the recommended video and the audio segment information of the target audio segment contained in the recommended video sent by the server.
[0090] S302. Play the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, display the associated information in the first display area of the video playback page.
[0091] The first display area can be the area on the video playback page used to display the associated information. Its size and specific position can be set as needed. For example, it can be set to have the same width as the second display area, and / or it can be set to be above the second display area, in the middle of the video playback page, or on one side, etc.
[0092] Specifically, the client can determine the target audio segment based on its identifier information, play the recommended video on the video playback page based on the video frequency data of the recommended video, and display the associated information of the target audio segment in the first display area 40 of the video playback page when playing the target video segment corresponding to the target audio segment. Figure 4 and Figure 5 As shown. For example, after the recommended video scrolls onto the screen, the client can calculate the height and initialize the Hybrid framework, call JSB to load the display container for the associated information and set it to be invisible to the user, such as making it transparent; when the recommended video plays to the target video segment corresponding to the target audio segment, the display container is set to be visible to the user, and the associated information of the target audio segment is displayed in the display container.
[0093] In this embodiment, when the client plays a recommended video, information can be displayed based on the audio played in the video. For example, when the video plays audio related to the discount information of the target object, the discount information of the target object is automatically displayed on the video playback page; and / or, when the video plays audio describing the target object, the descriptive information of the target object is automatically displayed on the video playback page, etc. This enables the conversion between auditory and visual information in the video, further enhancing the user's understanding of the auditory information in the video and improving the recommendation effect of the target object. Furthermore, since information display does not need to be based on the currently playing video frame, it is not necessary for the target object to be a tangible item that can be displayed or to appear on the video frame. It also does not require the video producer to fully understand the recommended video in advance and determine when to display information. This expands the applicability of recommended videos, enabling the recommendation of both tangible and virtual items, reducing the production cost of recommended videos, and thus reducing the recommendation cost of the target object.
[0094] In one embodiment, displaying the associated information in the first display area of the video playback page includes: displaying a target text statement corresponding to the target audio segment in the first display area of the video playback page, wherein the target text statement is obtained by the server through speech recognition of the target audio segment.
[0095] In the above embodiments, the associated information of the target audio segment can be the target text statement corresponding to the target audio segment, that is, the text content corresponding to the target audio segment. Thus, by displaying the text statement corresponding to the audio segment on the video playback page when playing an audio segment containing preset keywords, such as synchronously displaying the currently broadcast promotional information in text form on the video playback page, the effect of "what you hear is what you see" is achieved, which strengthens the user's impression and understanding of the content broadcast in the audio segment and increases the probability of the user receiving recommendations for the target object.
[0096] Specifically, the client can play recommended videos of the target object on the video playback page. When the recommended video plays to the target video segment corresponding to the target audio segment containing preset keywords, the client can display the target text statement corresponding to that audio segment in the first display area of the video playback page. For example, when the recommended video plays to the video segment corresponding to the audio segment containing keywords related to discount information, the client can display the discount information broadcast in the video in the first display area of the video playback page. The display status of the target text statement, such as font, font size, and / or color, can be set as needed.
[0097] In the above embodiments, after the target text statement is displayed, the target text statement may include the text currently being played in the recommended video and the text not currently being played in the recommended video (including text that has been played and / or text that has not yet been played). Therefore, during the playback of the recommended video, the text currently being played and the text not currently being played in the recommended video can be displayed in the same display state; alternatively, different display states can be used to display the text currently being played and the text not currently being played in the recommended video, thereby increasing the user's attention to and impression of the text information (especially the text currently being played), and enhancing the interest in displaying the text information. Optionally, displaying the target text statement corresponding to the target audio segment in the first display area of the video playback page includes: displaying the currently played text in the target text statement in a first display state in the first display area of the video playback page, and displaying the non-currently played text in the target text statement in a second display state. The first display state and the second display state can be two display states with different display methods such as font, font size and / or color.
[0098] It should be noted that when the target text statement cannot be displayed simultaneously in the first display area, a portion of the target text statement, including the text currently being played in the recommended video, can be displayed in the first display area. The text displayed in the first display area should be updated as the recommended video plays, for example, using a scrolling update or a page-turning update method. Furthermore, when the target audio segment is not present in the recommended video, the associated information of the target audio segment may not be displayed during the playback of the recommended video.
[0099] In this embodiment, the display duration of the association information of the target audio segment can be set as needed. For example, the association information of the target audio segment can be set to display for a preset duration (such as 3s or 5s); the target audio segment can also be set to be displayed until the recommended video ends or the target video segment ends, and so on.
[0100] Considering the relevance of the displayed association information to the currently playing content, this embodiment preferably displays the association information of the target audio segment only during the playback of the target audio segment, i.e., only during the playback of the target video segment corresponding to the target audio segment. In this case, the information display method provided in this embodiment further includes: stopping the display of the association information in the first display area when the target video segment finishes playing, such as stopping the display of the association information of the target audio segment at any position on the video playback page, or moving the association information of the target audio segment to another display area on the video playback page other than the first display area for display, etc.
[0101] The information display method provided in this embodiment sends an information retrieval request for recommended videos targeting a specific object to a server, and receives identification information and association information of a target audio segment containing preset keywords returned by the server based on the information retrieval request; plays the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, displays the association information in the first display area of the video playback page. By adopting the above technical solution, this embodiment displays the association information of a specific voice containing preset keywords on the video playback page when the recommended video plays to that specific voice. This not only allows for the display of diverse information but also enables the conversion between auditory and visual dimensions, strengthening the user's impression of the voice broadcast in the recommended video, improving the recommendation effect of the recommended video, expanding the applicability of the recommended video, and reducing the production cost of the recommended video.
[0102] Figure 6 This is a flowchart illustrating another information display method provided in this embodiment. The solution in this embodiment can be combined with one or more optional solutions in the above embodiments. Optionally, the information display method provided in this embodiment further includes: when the target video segment finishes playing, moving the associated information from the first display area to the second display area of the video playback page, and displaying the associated information in a gradually shrinking manner during the movement, wherein the second display area is a simplified information display area of the target object or a detailed information display area of the target object.
[0103] Optionally, after moving the associated information from the first display area to the second display area of the video playback page, the method further includes: if the second display area is the simplified information display area, then displaying the associated information in the target display area; if the second display area is the detailed information display area, then displaying the detailed information corresponding to the associated information in the target display area, and canceling the display of the associated information.
[0104] Correspondingly, such as Figure 6As shown, the information display method provided in this embodiment may include:
[0105] S401. Send a request to the server to obtain information about recommended videos for the target object, and receive audio segment information returned by the server based on the information obtaining request, wherein the audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords.
[0106] S402. Play the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, display the associated information in the first display area of the video playback page.
[0107] S403. When the target video segment finishes playing, the associated information is moved from the first display area to the second display area of the video playback page, and during the movement, the associated information is displayed in a gradually shrinking manner. S404 or S405 is executed, wherein the second display area is a simplified information display area of the target object or a detailed information display area of the target object.
[0108] In this embodiment, when the target video segment finishes playing, the associated information displayed in the first display area can be moved to the second display area to guide the user to view the object information (such as brief information or detailed information) of the target object displayed in the second display area and interact with it in the second display area.
[0109] The second display area can be either a brief information display area or a detailed information display area for the target object. The brief information display area can be an area on the video playback page used to display brief information (such as a brief introduction) about the target object, while the detailed information display area can be an area on the video playback page used to display detailed information (such as a detailed description) about the target object. It can be a different area from the brief information display area, or it can be an extension of the established information display area. The brief information may include images, names, and promotional information about the target object; the detailed information may include images, names, prices, and promotional information about the target object.
[0110] For details, please continue to see Figure 4 and Figure 5When the client plays a recommended video of a target object on the video playback page, it can display brief information and / or detailed information of the target object in the second display area 50 of the video playback page, such as displaying both brief and detailed information simultaneously; or, displaying the brief and detailed information at different time points. For example, when the recommended video plays to the first time point, the brief information of the target object is displayed in the brief information display area of the video playback page, such as... Figure 4 As shown; when the recommended video plays to the second time point, the detailed information of the target object is displayed in the detailed information display area of the video playback page, and the display of the brief information is canceled, such as... Figure 5 As shown. Furthermore, when displaying detailed information about the target object, the user can also instruct the client to switch the currently displayed detailed information back to the simplified information by performing a toggle operation (such as turning off detailed information).
[0111] Therefore, when the recommended video plays to the target video segment corresponding to the target audio segment, the associated information of the target audio segment can be displayed in the first display area 40 of the video playback page, and the recommended video can continue playing; when the target video segment finishes playing, the associated information of the target audio segment displayed in the first display area 40 can be moved to the second display area 50, and during the movement, the display size of the associated information is gradually reduced, such as... Figure 7 and Figure 8 As shown, the display can be stopped, the related information can continue to be displayed in the second display area 50, or the detailed information of the related information can be displayed in the second display area 50. During the movement, the movement trajectory, movement speed, and shrinking speed of the related information can all be set as needed; this embodiment does not impose any limitations on these settings.
[0112] S404. If the second display area is the simplified information display area, then the associated information is displayed in the target display area, and the operation ends.
[0113] S405. If the second display area is the detailed information display area, then display the detailed information corresponding to the associated information in the target display area, and cancel the display of the associated information.
[0114] The first detailed information can be the detailed information of the associated information of the target audio segment. For example, when the target audio segment contains audio segments related to discount information, the detailed information can be the detailed discount information of the target object, such as the original price, current price and / or discount amount of the target object.
[0115] Specifically, when the target video segment corresponding to the target audio segment finishes playing, the client controls the associated information of the target audio segment to move from the first display area 40 to the second display area 50. Upon reaching the second display area 50, it further determines whether the second display area 50 is a simplified information display area or a detailed information display area; that is, it determines whether the information displayed in the second display area 50 is simplified information or detailed information about the target object. If the information displayed in the second display area 50 is simplified information, then the associated information continues to be displayed in the second display area 50. Figure 9 As shown; if the second display area 50 displays detailed information about the target object, then the detailed information of that related information can be displayed in the second display area 50, such as... Figure 10 As shown.
[0116] Taking the preset keywords as keywords related to discount information and the associated information of the target audio segment as the target text statement corresponding to the target audio segment as an example, when the second display area 50 is a small, simplified information display area, the target text statement can be controlled to move from the first display area 40 to the display position of the discount information of the target object in the second display area 50 (such as the subtitle position), and the discount information displayed at that position can be replaced with the associated information, such as... Figure 9 As shown; when the second display area 50 is a large detailed information display area, the target text statement can be moved from the first display area 40 to a certain position in the second display area 50 (such as the boundary position or the center position, etc.). The image and name of the target object originally displayed in the second display area 50 can be reduced in size and adjusted towards the boundary direction. The detailed information of the associated information can then be displayed in the center area of the second display area 50. Figure 10 As shown ( Figure 10 (Taking coupon information as an example).
[0117] In one implementation, a user can view detailed information about a target object by triggering the associated information of the displayed target audio segment. In this case, the information display method provided in this embodiment may further include: in response to a triggering operation on the associated information, switching the current page from the video playback page to the detailed page of the target object, and displaying the detailed information of the target object on the detailed page.
[0118] For example, when a client plays a recommended video of a target object on a video playback page, and the recommended video contains a target video segment corresponding to a target audio segment with preset keywords, the associated information of the target audio segment is displayed in the first display area of the video playback page. Thus, when a user wants to view detailed information about the target object recommended in the video, such as viewing detailed information about the associated information or purchasing the target object, they can trigger (e.g., click) the associated information. Correspondingly, when the client detects that the user has triggered the associated information displayed on the video playback page, it can switch the current page from the video playback page to the target object's details page, displaying the target object's detailed information for the user to view. This details page may reside in the same or a different application. When it resides in a different application, the application containing the details page can be instructed to display the target object's details page by calling the corresponding interface of the application containing the details page.
[0119] In addition, when the associated information of the target audio segment moves from the first display area to the second display area or is displayed in the second display area, if it is detected that the user triggers the associated information, or if it is detected that the user triggers any position in the second display area (including control positions or non-control positions in the second display area), the current display page can also be switched from the video playback page to the details page of the target object, so that the user can view the details of the target object.
[0120] The information display method provided in this embodiment, by displaying and moving the associated information of the target audio segment, and further displaying the associated information or the details of the associated information after the movement, can further guide users to interact and provide users with richer information display and interaction methods, thereby improving the user's viewing and interaction experience and the recommendation effect of recommended videos.
[0121] Figure 11 This is a structural block diagram of an information display device provided in an embodiment of this disclosure. The device can be implemented by software and / or hardware, and can be configured in an electronic device, typically in a server. It can control a client to display information based on audio data from recommended videos by executing an information display method. Figure 11 As shown, the information display device provided in this embodiment may include: a data acquisition module 1101, an information determination module 1102, and an information sending module 1103, wherein,
[0122] The data acquisition module 1101 is used to acquire the audio data of the recommended video of the target object;
[0123] The information determination module 1102 is used to determine the audio segment information of the target audio segment in the recommended video based on the audio data, wherein the audio segment information includes identification information and association information, and the target audio segment contains preset keywords;
[0124] The information sending module 1103 is used to send the audio segment information to the client when it receives an information acquisition request for the recommended video, so that the client displays the associated information when playing the target video segment corresponding to the target audio segment, wherein the information acquisition request is sent by the client.
[0125] The information display device provided in this embodiment acquires audio data of a recommended video for a target object through a data acquisition module; determines the identification information and association information of a target audio segment containing preset keywords in the recommended video based on the audio data through an information determination module; and sends the identification information and association information of the target audio segment to the client when the client receives an information acquisition request for the recommended video from the client, so that the client displays the association information when playing the target video segment corresponding to the target audio segment. By adopting the above technical solution, this embodiment controls the client to display the association information of the specific audio segment on the video playback page when playing a specific audio segment containing preset keywords. This not only displays diverse information but also achieves a conversion between auditory and visual dimensions, strengthening the user's impression of the audio broadcast in the recommended video, improving the recommendation effect of the recommended video, expanding the applicability of the recommended video, and reducing the production cost of the recommended video.
[0126] In the above scheme, the information determination module 1102 may include: a speech recognition unit, used to perform speech recognition on the audio data to obtain speech recognition text, and to obtain the Mel frequency cepstral coefficient feature vector of the audio data; a time determination unit, used to determine the time information corresponding to each word in the audio recognition text in the audio data based on the Mel frequency cepstral coefficient feature vector; and an information determination unit, used to determine the audio segment information of the target audio segment in the recommended video according to the speech recognition text and the time information.
[0127] In the above scheme, the information determination unit may include: a text segmentation subunit, used to segment the speech recognition text based on the time information to obtain at least one text statement; a statement determination subunit, used to identify text statements containing preset keywords in the at least one text statement as candidate text statements; a statement filtering subunit, used to filter the candidate text statements based on preset filtering rules to obtain target text statements; and an information determination subunit, used to take the audio segment corresponding to the target text statement in the audio data as the target audio segment in the recommended video, and determine the audio segment information of the target audio segment.
[0128] In the above scheme, the information determination subunit can be used to: use the time node information corresponding to the target audio segment in the recommended video as the identification information of the target audio segment, and use the target text statement as the association information of the target audio segment.
[0129] Furthermore, the information display device provided in this embodiment may also include: an information update module, used to, after determining the audio segment information of the target audio segment in the recommended video based on the audio data, when a change in the content of the recommended video is detected, re-determine the audio segment information of the target audio segment in the recommended video based on the audio data of the changed recommended video.
[0130] The information display device provided in this disclosure can execute the information display method executed by the server provided in this disclosure, and has the corresponding functional modules and beneficial effects for executing the information display method. Technical details not described in detail in this embodiment can be found in the information display method executed by the server provided in this disclosure.
[0131] Figure 12 This is a structural block diagram of an information display device provided in an embodiment of this disclosure. The device can be implemented by software and / or hardware, and can be configured in an electronic device, typically a mobile phone or tablet computer, to display information based on audio data of a recommended video through the execution of an information display method. Figure 12 As shown, the information display device provided in this embodiment may include: an information receiving module 1201 and a video playback module 1202, wherein,
[0132] The information receiving module 1201 is used to send an information acquisition request for recommended videos of a target object to the server, and to receive audio segment information returned by the server based on the information acquisition request, wherein the audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords;
[0133] The video playback module 1202 is used to play the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, it displays the associated information in the first display area of the video playback page.
[0134] The information display device provided in this embodiment sends an information acquisition request for recommended videos targeting a target object to a server through an information receiving module, and receives identification information and association information of a target audio segment containing preset keywords returned by the server based on the information acquisition request; the recommended video is played through a video playback module, and when the recommended video plays to the target video segment corresponding to the target audio segment, the association information is displayed in the first display area of the video playback page. By adopting the above technical solution, this embodiment displays the association information of a specific voice containing preset keywords on the video playback page when the recommended video plays to that specific voice, which not only displays diverse information but also achieves a conversion between auditory and visual dimensions, strengthening the user's impression of the voice broadcast in the recommended video, improving the recommendation effect of the recommended video, expanding the applicability of the recommended video, and reducing the production cost of the recommended video.
[0135] In the above scheme, the video playback module 1202 can be used to: display a target text statement corresponding to the target audio segment in the first display area of the video playback page, wherein the target text statement is obtained by the server through speech recognition of the target audio segment.
[0136] Furthermore, the information display device provided in this embodiment may also include: a moving module, used to move the associated information from the first display area to the second display area of the video playback page when the target video segment finishes playing, and to display the associated information in a gradually shrinking manner during the movement, wherein the second display area is a simplified information display area of the target object or a detailed information display area of the target object.
[0137] Furthermore, the information display device provided in this embodiment may further include: a first information display module, configured to, after the associated information is moved from the first display area to the second display area of the video playback page, if the second display area is the simplified information display area, display the associated information in the target display area; if the second display area is the detailed information display area, display the detailed information corresponding to the associated information in the target display area, and cancel the display of the associated information.
[0138] Furthermore, the information display device provided in this embodiment may further include: a second information display module, configured to display brief information of the target object in the brief information display area of the video playback page when the recommended video is played to a first time node; and to display detailed information of the target object in the detailed information display area of the video playback page when the recommended video is played to a second time node, and to cancel the display of the brief information.
[0139] The information display device provided in this disclosure can execute the information display method executed by the client provided in this disclosure, and has the corresponding functional modules and beneficial effects for executing the information display method. Technical details not described in detail in this embodiment can be found in the information display method executed by the client provided in this disclosure.
[0140] The following is for reference. Figure 13 The diagram illustrates a structural schematic of an electronic device (e.g., a server or terminal device) 1300 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 13 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0141] like Figure 13 As shown, electronic device 1300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1302 or a program loaded from storage device 1308 into random access memory (RAM) 1303. The RAM 1303 also stores various programs and data required for the operation of electronic device 1300. The processing device 1301, ROM 1302, and RAM 1303 are interconnected via bus 1304. Input / output (I / O) interface 1305 is also connected to bus 1304.
[0142] Typically, the following devices can be connected to I / O interface 1305: input devices 1306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1309. Communication device 1309 allows electronic device 1300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 13 An electronic device 1300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0143] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1309, or installed from storage device 1308, or installed from ROM 1302. When the computer program is executed by processing device 1301, it performs the functions defined in the methods of embodiments of this disclosure.
[0144] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0145] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0146] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0147] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0148] Obtain audio data of recommended videos for the target object; determine audio segment information of target audio segments in the recommended videos based on the audio data, wherein the audio segment information includes identification information and association information, and the target audio segment contains preset keywords; when an information retrieval request for the recommended video is received, send the audio segment information to the client, so that the client displays the association information when playing the target video segment corresponding to the target audio segment, wherein the information retrieval request is sent by the client. Alternatively,
[0149] A request to obtain information about recommended videos for a target object is sent to the server, and audio segment information returned by the server based on the request is received. The audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords. The recommended video is played, and when the recommended video plays to the target video segment corresponding to the target audio segment, the association information is displayed in the first display area of the video playback page.
[0150] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0152] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of modules do not, in some cases, constitute a limitation on the unit itself.
[0153] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0154] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0155] According to one or more embodiments of this disclosure, Example 1 provides an information display method, including:
[0156] Obtain the audio data of the recommended videos for the target audience;
[0157] Based on the audio data, the audio segment information of the target audio segment in the recommended video is determined, wherein the audio segment information includes identification information and association information, and the target audio segment contains preset keywords;
[0158] When an information retrieval request for the recommended video is received, the audio segment information is sent to the client so that the client displays the associated information when playing the target video segment corresponding to the target audio segment, wherein the information retrieval request is sent by the client.
[0159] According to one or more embodiments of this disclosure, Example 2 describes the method described in Example 1, wherein determining the audio segment information of the target audio segment in the recommended video based on the audio data includes:
[0160] The audio data is subjected to speech recognition to obtain the speech recognition text, and the Mel frequency cepstral coefficient feature vector of the audio data is obtained.
[0161] Based on the Mel frequency cepstral coefficient feature vector, determine the time information corresponding to each word in the audio recognition text in the audio data;
[0162] The audio segment information of the target audio segment in the recommended video is determined based on the speech recognition text and the time information.
[0163] According to one or more embodiments of this disclosure, Example 3 describes the method described in Example 2, wherein determining the audio segment information of the target audio segment in the recommended video based on the speech-recognized text and the time information includes:
[0164] The speech recognition text is segmented based on the time information to obtain at least one text sentence;
[0165] Identify text statements containing preset keywords in at least one text statement as candidate text statements;
[0166] The candidate text statements are filtered based on preset filtering rules to obtain the target text statements;
[0167] The audio segment corresponding to the target text statement in the audio data is used as the target audio segment in the recommended video, and the audio segment information of the target audio segment is determined.
[0168] According to one or more embodiments of this disclosure, Example 4, based on the method described in Example 3, includes determining the audio segment information of the target audio segment, including:
[0169] The time node information corresponding to the target audio segment in the recommended video is used as the identification information of the target audio segment, and the target text statement is used as the association information of the target audio segment.
[0170] According to one or more embodiments of this disclosure, Example 5, based on any one of Examples 1-4, further includes, after determining the audio segment information of the target audio segment in the recommended video based on the audio data:
[0171] When a change in the content of the recommended video is detected, the audio segment information of the target audio segment in the recommended video is re-determined based on the audio data of the changed recommended video.
[0172] According to one or more embodiments of this disclosure, Example 6 provides an information display method, including:
[0173] Send a request to the server to obtain information about recommended videos for a target object, and receive audio segment information returned by the server based on the information obtaining request. The audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords.
[0174] Play the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, display the associated information in the first display area of the video playback page.
[0175] According to one or more embodiments of this disclosure, Example 7, based on the method described in Example 6, includes displaying the associated information in a first display area of the video playback page, comprising:
[0176] The first display area of the video playback page displays the target text statement corresponding to the target audio segment, wherein the target text statement is obtained by the server through speech recognition of the target audio segment.
[0177] According to one or more embodiments of this disclosure, Example 8, based on the method described in Example 6 or 7, further includes:
[0178] When the target video segment finishes playing, the associated information is moved from the first display area to the second display area of the video playback page. During the movement, the associated information is displayed in a gradually shrinking manner. The second display area is either a simplified information display area for the target object or a detailed information display area for the target object.
[0179] According to one or more embodiments of this disclosure, Example 9, based on the method of Example 8, further includes, after moving the associated information from the first display area to the second display area of the video playback page:
[0180] If the second display area is the simplified information display area, then the associated information is displayed in the target display area;
[0181] If the second display area is the detailed information display area, then the detailed information corresponding to the associated information is displayed in the target display area, and the display of the associated information is canceled.
[0182] According to one or more embodiments of this disclosure, Example 10, based on the method described in Example 8, further includes:
[0183] When the recommended video reaches the first time point, brief information about the target object is displayed in the brief information display area of the video playback page;
[0184] When the recommended video reaches the second time point, the detailed information of the target object is displayed in the detailed information display area of the video playback page, and the display of the brief information is canceled.
[0185] According to one or more embodiments of this disclosure, Example 11 provides an information display device, including:
[0186] The data acquisition module is used to acquire the audio data of the recommended videos for the target object;
[0187] The information determination module is used to determine the audio segment information of the target audio segment in the recommended video based on the audio data, wherein the audio segment information includes identification information and association information, and the target audio segment contains preset keywords;
[0188] An information sending module is used to send the audio segment information to the client when it receives an information retrieval request for the recommended video, so that the client displays the associated information when playing the target video segment corresponding to the target audio segment, wherein the information retrieval request is sent by the client.
[0189] According to one or more embodiments of this disclosure, Example 12 provides an information display device, including:
[0190] The information receiving module is used to send an information acquisition request for recommended videos targeting a target object to the server, and to receive audio segment information returned by the server based on the information acquisition request. The audio segment information includes the identification information and association information of the target audio segment in the recommended video, and the target audio segment contains preset keywords.
[0191] The video playback module is used to play the recommended video, and when the recommended video plays to the target video segment corresponding to the target audio segment, the associated information is displayed in the first display area of the video playback page.
[0192] According to one or more embodiments of this disclosure, Example 13 provides an electronic device, including:
[0193] One or more processors;
[0194] Memory, used to store one or more programs.
[0195] When the one or more programs are executed by the one or more processors, the one or more processors implement the information display method as described in any of Examples 1-10.
[0196] According to one or more embodiments of the present disclosure, Example 14 provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the information display method as described in any of Examples 1-10.
[0197] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0198] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0199] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An information display method characterized by comprising: The method comprises: obtaining audio data of a recommended video of a target object; determining audio segment information of a target audio segment in the recommended video according to the audio data, wherein the audio segment information comprises identification information and association information of the target object, the association information comprises attribute information of a preset attribute of the target object, the preset attribute corresponds to a voice broadcast by the target audio segment, and the target audio segment comprises a preset keyword; when an information acquisition request for the recommended video is received, sending the audio segment information to a client, so that the client displays the association information when playing to a target video segment corresponding to the target audio segment, wherein the information acquisition request is sent by the client; wherein, after determining the audio segment information of the target audio segment in the recommended video according to the audio data, the method further comprises: when it is detected that the content of the recommended video is changed, re-determining the audio segment information of the target audio segment in the recommended video according to audio data of the changed recommended video.
2. The method of claim 1, wherein, The method of determining the audio segment information of the target audio segment in the recommended video according to the audio data comprises: performing voice recognition on the audio data to obtain voice recognition text and obtain a mel-frequency cepstrum coefficient feature vector of the audio data; determining time information corresponding to each character in the voice recognition text in the audio data based on the mel-frequency cepstrum coefficient feature vector; determining the audio segment information of the target audio segment in the recommended video according to the voice recognition text and the time information.
3. The method of claim 2, wherein, The method of determining the audio segment information of the target audio segment in the recommended video according to the voice recognition text and the time information comprises: segmenting the voice recognition text based on the time information to obtain at least one text sentence; identifying a text sentence containing a preset keyword in the at least one text sentence as a candidate text sentence; screening the candidate text sentence based on a preset screening rule to obtain a target text sentence; determining an audio segment corresponding to the target text sentence in the audio data as the target audio segment in the recommended video, and determining the audio segment information of the target audio segment.
4. The method of claim 3, wherein, The method of determining the audio segment information of the target audio segment comprises: determining time node information corresponding to the target audio segment in the recommended video as the identification information of the target audio segment, and determining the target text sentence as the association information of the target audio segment.
5. An information display method characterized by comprising: The method comprises: sending a recommendation video information acquisition request for a target object to a server, and receiving audio clip information returned by the server based on the information acquisition request, wherein the audio clip information comprises identification information of a target audio clip in the recommendation video and association information of the target object, the association information comprises attribute information of a preset attribute of the target object, the preset attribute corresponds to a voice broadcast by the target audio clip, and the target audio clip contains a preset keyword; the audio clip information supports being updated by the server according to audio data of a changed recommendation video when the server detects that the content of the recommendation video is changed; playing the recommendation video, and displaying the association information in a first display area of a video playing page when the recommendation video is played to a target video clip corresponding to the target audio clip.
6. The method of claim 5, wherein, The displaying of the association information in the first display area of the video playing page comprises: displaying a target text sentence corresponding to the target audio clip in the first display area of the video playing page, wherein the target text sentence is obtained by performing voice recognition on the target audio clip by the server.
7. The method according to claim 5 or 6, characterized in that, Further comprising: when the target video clip is played to the end, moving the association information from the first display area to a second display area of the video playing page, and displaying the association information in a manner of gradually reducing during the moving, wherein the second display area is a brief information display area of the target object or a detailed information display area of the target object.
8. The method of claim 7, wherein, After the moving of the association information from the first display area to the second display area of the video playing page, further comprising: if the second display area is the brief information display area, displaying the association information in a target display area; if the second display area is the detailed information display area, displaying detail information corresponding to the association information in the target display area, and canceling the display of the association information.
9. The method of claim 7, wherein, Further comprising: when the recommendation video is played to a first time node, displaying brief information of the target object in the brief information display area of the video playing page; when the recommendation video is played to a second time node, displaying detailed information of the target object in the detailed information display area of the video playing page, and canceling the display of the brief information.
10. An information display device, characterized by comprising: comprising: a data acquisition module configured to acquire audio data of a recommendation video of a target object; an information determination module configured to determine audio clip information of a target audio clip in the recommendation video according to the audio data, wherein the audio clip information comprises identification information and association information of the target object, the association information comprises attribute information of a preset attribute of the target object, the preset attribute corresponds to a voice broadcast by the target audio clip, and the target audio clip contains a preset keyword; The information sending module is configured to send the audio segment information to the client when receiving an information obtaining request for the recommended video, so that the client displays the associated information when playing to a target video segment corresponding to the target audio segment. The device further comprises: The information updating module is configured to, after determining the audio segment information of the target audio segment in the recommended video according to the audio data, re-determine the audio segment information of the target audio segment in the recommended video according to audio data of the recommended video after detecting a change in content of the recommended video.
11. An information display device, characterized by comprising: The information receiving module is configured to send an information obtaining request for a recommended video of a target object to a server, and receive audio segment information returned by the server based on the information obtaining request, wherein the audio segment information comprises identification information of a target audio segment in the recommended video and associated information of the target object, the associated information comprises attribute information of a preset attribute of the target object, the preset attribute corresponds to voice broadcast by the target audio segment, and the target audio segment comprises a preset keyword; and the audio segment information supports being updated by the server according to audio data of a recommended video after detecting a change in content of the recommended video; The video playing module is configured to play the recommended video, and display the associated information in a first display area of a video playing page when the recommended video plays to a target video segment corresponding to the target audio segment. One or more processors; 12. An electronic device, comprising: Memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the information display method according to any one of claims 1-9. The program is executed by the processor to implement the information display method according to any one of claims 1-9. 13. A computer readable storage medium having stored thereon a computer program, characterized in that,
Citation Information
Patent Citations
Method for displaying words and processing device and computer program product thereof
CN103474081A
Video and audio processing method in network live broadcast, computer equipment and medium
CN112839237A