A search information display method and device, computer equipment and storage medium

By filtering and processing the text information of multimedia content to generate search information, the problem of users having difficulty obtaining key information about trending events is solved, and efficient and timely information recommendation and search are achieved.

CN116521993BActive Publication Date: 2026-02-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310457953.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-25
Publication Date
2026-02-06
Estimated Expiration
2043-04-25

AI Technical Summary

Technical Problem

Users often struggle to accurately obtain key information about trending events after browsing multimedia content, and they are unable to access relevant information through appropriate channels if they are unfamiliar with these events.

Method used

By filtering multimedia content that meets preset conditions under specific attribute dimensions, extracting text information and generating search information, using character recognition and semantic processing to generate search information, and displaying it at an appropriate time.

Benefits of technology

It improves the efficiency of information recommendation and the timeliness of users' access to trending content. Users can directly initiate efficient searches and stay informed about trending topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116521993B_ABST
    Figure CN116521993B_ABST
Patent Text Reader

Abstract

The present disclosure provides a display method and device for search information, a computer device and a storage medium, wherein the method comprises: determining at least one target multimedia content; the target multimedia content has extractable text information, and the characteristics of the target multimedia content in at least one attribute dimension meet a preset screening condition; generating search information based on the text information extracted from each target multimedia content; and in response to meeting a recommendation display condition of the search information, recommending and displaying the search information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of information display, and in particular, to a search information display method and device, a computer device, and a storage medium. BACKGROUND

[0002] In life, hot event information of the moment can be expressed through multimedia content such as news and videos. After browsing the multimedia content, the user can obtain some information related to the hot event. Generally, a hot event will continue to ferment for a period of time, and a lot of associated multimedia content will be generated in this period of time.

[0003] When the user wants to understand a certain hot event, the user needs to know in advance what search keywords to use to initiate a search. If the search keywords used are inaccurate, some key information content of the hot event cannot be searched. Or, in the case where the user does not understand the current hot event, the user will not be able to obtain relevant information of the hot event through a suitable channel. SUMMARY

[0004] The present disclosure provides at least a search information display method and device, a computer device, and a storage medium.

[0005] In a first aspect, the present disclosure provides a search information display method, including: determining at least one target multimedia content; the target multimedia content has extractable text information, and a feature of the target multimedia content in at least one attribute dimension satisfies a preset screening condition; generating search information based on the text information extracted from each target multimedia content; and in response to a recommendation display condition of the search information being satisfied, recommending and displaying the search information.

[0006] In an optional implementation, in response to the target multimedia content including a video, the text information associated with the target multimedia content is determined in the following manner: at least one key frame in the video is determined; the key frame of the video includes at least one of the following: at least one video frame corresponding to a cover of the video, and at least one video frame selected from each video frame included in the video; and character recognition is performed on each key frame to obtain text information in each key frame as text information associated with the video.

[0007] In an alternative implementation, the search information is generated based on the text information, including: screening target text information from the text information according to a preset screening condition; the preset screening condition includes at least one of the following: the text information is in a text recognition area of a key frame corresponding to the region attribute information meeting a preset region attribute information requirement, and the semantic content corresponding to the text information indicates the theme information of the target multimedia content; wherein the region attribute information includes region area and / or region position; and the search information is determined based on the target text information.

[0008] In an alternative implementation, the search information is determined based on the target text information, including: performing semantic segmentation processing on the target text information to obtain a plurality of semantic words; screening at least one keyword expressing the theme of the target multimedia content from the plurality of semantic words to integrate as the search information.

[0009] In an alternative implementation, the search information is determined based on the target text information, including: determining an entity keyword and at least one limited keyword from the at least one keyword corresponding to the target text information; wherein the entity keyword indicates an entity object associated with the target multimedia content, and the limited keyword is used to indicate the information dimension of the entity object; and the search information corresponding to the target multimedia content is generated based on the entity keyword and the limited keyword.

[0010] In an alternative implementation, the search information corresponding to the target multimedia content is determined based on the keyword corresponding to the target multimedia content in the following manner: a text reconstruction model is trained based on each keyword sample and a sentence sample composed of a plurality of keyword samples to reconstruct the at least one keyword into a sentence containing the at least one keyword as the search information.

[0011] In an alternative implementation, it further includes: in response to the existence of a plurality of entity keywords corresponding to the target multimedia content being the same, determining the semantic similarity between the at least one limited keyword corresponding to each target multimedia content; and in response to the existence of at least two target multimedia contents corresponding to the semantic similarity exceeding a preset similarity threshold in a plurality of target multimedia contents, performing deduplication processing on the search information corresponding to the at least two target multimedia contents to obtain one search information corresponding to the at least two target multimedia contents.

[0012] In an alternative implementation, the at least one attribute dimension of the target multimedia content comprises at least one of a publishing time of the target multimedia content, a publishing channel of the target multimedia content, and a publishing frequency of the publishing channel in a preset time period.

[0013] In an alternative implementation, the method further comprises: in response to a search trigger operation on the search information, determining a plurality of multimedia contents associated with the search information; the plurality of multimedia contents comprise the target multimedia content in which the search information is determined; and in response to the plurality of multimedia contents comprising a video, if there is a target video frame containing a keyword in the search information in the video, replacing the target video frame with a cover of the video for display.

[0014] In a second aspect, the embodiments of the present disclosure further provide a display device for search information, comprising: a determination module configured to determine at least one target multimedia content; the target multimedia content has extractable text information, and a feature of the target multimedia content in at least one attribute dimension meets a preset filtering condition; a generation module configured to generate search information based on the text information extracted from each target multimedia content; and a display module configured to recommend and display the search information in response to a recommendation and display condition of the search information being met.

[0015] In an alternative implementation, in response to the target multimedia content comprising a video, the text information associated with the target multimedia content is determined in the following manner: at least one key frame in the video is determined; the key frame of the video comprises at least one of at least one video frame corresponding to a cover of the video, and at least one video frame selected from each video frame contained in the video; and character recognition is performed on each key frame to obtain text information in each key frame as the text information associated with the video.

[0016] In an alternative implementation, when generating the search information based on the text information, the generation module is configured to: from each text information, filter out target text information meeting a preset filtering condition; the preset filtering condition comprises at least one of the following: region attribute information corresponding to a text recognition region in the key frame of the text information meets a preset region attribute information requirement, and semantic content corresponding to the text information indicates theme information of the target multimedia content; wherein the region attribute information comprises region area and / or region position; and determine the search information based on the target text information.

[0017] In an alternative implementation, the generating module is configured to: perform semantic segmentation on the target text information to obtain a plurality of semantic words; and select at least one keyword expressing a theme of the target multimedia content from the plurality of semantic words, and integrate the at least one keyword to obtain the search information.

[0018] In an alternative implementation, the generating module is configured to: determine an entity keyword and at least one limit keyword from the at least one keyword corresponding to the target text information, wherein the entity keyword indicates an entity object associated with the target multimedia content, and the limit keyword is used to indicate an information dimension of the entity object; and generate the search information corresponding to the target multimedia content based on the entity keyword and the limit keyword.

[0019] In an alternative implementation, the search information corresponding to the target multimedia content is determined based on the keyword corresponding to the target multimedia content in the following manner: a text reconstruction model is trained based on each keyword sample and a sentence sample composed of a plurality of keyword samples, and the text reconstruction model is used to perform text reconstruction on the at least one keyword to obtain a sentence containing the at least one keyword as the search information.

[0020] In an alternative implementation, the generating module is further configured to: determine semantic similarity between the at least one limit keyword corresponding to each of the target multimedia contents in response to the existence of the same entity keyword corresponding to a plurality of target multimedia contents; and perform deduplication processing on the search information corresponding to at least two target multimedia contents in response to the existence of the at least two target multimedia contents with semantic similarity exceeding a preset similarity threshold in the plurality of target multimedia contents, to obtain one search information corresponding to the at least two target multimedia contents.

[0021] In an alternative implementation, the at least one attribute dimension of the target multimedia content includes at least one of the following: a publishing time of the target multimedia content, a publishing channel, and a publishing frequency of the publishing channel in a preset time period.

[0022] In an alternative implementation, the apparatus further includes a processing module configured to: determine a plurality of multimedia contents associated with the search information in response to a search trigger operation on the search information; and the plurality of multimedia contents include a target multimedia content used to determine the search information; and replace a target video frame containing a keyword in the search information with a cover of a video in response to the plurality of multimedia contents including the video.

[0023] In a third aspect, the present disclosure provides a computer device, a processor, and a memory storing machine readable instructions executable by the processor, wherein the processor is configured to execute the machine readable instructions stored in the memory, and the machine readable instructions, when executed by the processor, perform the steps of the first aspect or any possible implementation of the first aspect.

[0024] In a fourth aspect, the present disclosure provides a computer readable storage medium, which stores a computer program, and the computer program, when executed, performs the steps of the first aspect or any possible implementation of the first aspect.

[0025] The method and device for displaying search information provided by the embodiments of the present disclosure can filter target multimedia content (which can be multimedia content associated with a hot event) that meets a preset filtering condition in at least one attribute dimension and has extractable text information, generate search information by extracting relevant text information from the filtered target multimedia content, and display the search information when a recommendation display condition is met.

[0026] In this way, by extracting text information from target multimedia content that meets certain filtering conditions, search information can be generated for recommendation to users, and the search information can be displayed at appropriate times and locations to actively prompt users of currently searchable hot information. Thus, valuable search information can be efficiently and timely generated, and users can directly initiate a search based on the search information, thereby learning about hot related content in a timely manner, which improves the efficiency of search information recommendation and the timeliness of users obtaining hot content.

[0027] In order to make the above objectives, features and advantages of the present disclosure more obvious and easy to understand, a preferred embodiment is described below, and the accompanying drawings are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. The drawings herein are incorporated into the specification and form a part of the specification, which illustrate the embodiments consistent with the present disclosure, and are used to explain the technical solutions of the present disclosure together with the specification. It should be understood that the following drawings only show some embodiments of the present disclosure, and therefore should not be considered as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0029] Figure 1A flow chart of a search information display method is shown.

[0030] Figure 2 A schematic diagram of a video frame image is shown.

[0031] Figure 3 A schematic diagram of multiple display scenarios is shown.

[0032] Figure 4 A schematic diagram of a search information display device is shown.

[0033] Figure 5 A schematic diagram of a computer device is shown. DETAILED DESCRIPTION

[0034] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure and are not all the embodiments. The components of the embodiments of the present disclosure described and shown herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present disclosure.

[0035] It is found through research that some information related to a hot event can be obtained after a user browses multimedia content. When the user wants to understand a certain hot event, the user needs to know in advance what search keywords to use to initiate a search. If the search keywords used are inaccurate, some key information content of the hot event cannot be searched. Or, in the case where the user does not understand the current hot event, the user will not be able to obtain relevant information of the hot event through a suitable channel.

[0036] Based on the above research, the present disclosure provides a search information display method, which screens target multimedia content (which can be hot event associated multimedia content) that has a feature in at least one attribute dimension satisfying a preset screening condition and has extractable text information. For the screened target multimedia content, search information is generated by extracting relevant text information therefrom, and the search information is displayed if a recommendation display condition is met.

[0037] In this way, by extracting text information from target multimedia content meeting certain screening conditions, search information recommended for use by a user can be generated, and by displaying the search information at a suitable time and location, the user can be actively prompted about hot information that can be searched for at present. Thus, valuable search information can be efficiently and timely generated, and based on the search information, the user can directly initiate a search, thereby learning about hot content in a timely manner, which improves the efficiency of search information recommendation and the timeliness of obtaining hot content by the user.

[0038] The above-mentioned defects are the results of the inventors after practice and careful research, and thus the discovery process of the above-mentioned problems and the solutions proposed by the present disclosure to solve the above-mentioned problems should be the contributions of the inventors to the present disclosure.

[0039] It should be noted that similar reference numerals and letters refer to similar items in the following drawings, and thus, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings.

[0040] To facilitate understanding of the present embodiment, first, a search information display method disclosed by the present embodiment is described in detail. The execution subject of the search information display method provided by the present embodiment is generally a computer device with certain computing capability, which may, for example, include a terminal device or a server or other processing device. The terminal device may, for example, be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the search information display method can be implemented by a processor invoking computer readable instructions stored in a memory.

[0041] The search information display method provided by the present embodiment is described below. The search information display method provided by the present embodiment can be specifically applied to different application platforms such as a video platform and a web search platform. In these platforms, multimedia content including videos, news, etc. can be provided to a user, and multimedia content associated with information searched for by the user can be displayed according to the search demand of the user. The search information determined in the present embodiment can be displayed in a plurality of different pages, for example, arranged and displayed in the form of a hot search list, or displayed as a recommended search word in a search box, or associated with and displayed in specific displayed multimedia content. The actual situation can be determined, and the present embodiment does not make a specific limitation.

[0042] Referring to Figure 1 As shown in FIG. 1, a flowchart of a display method of search information provided by an embodiment of the present disclosure is shown, and the method comprises steps S101-S103, wherein:

[0043] S101: determining at least one target multimedia content; the target multimedia content has extractable text information, and a feature of the target multimedia content in at least one attribute dimension satisfies a preset screening condition;

[0044] S102: generating search information based on the text information extracted from each target multimedia content;

[0045] S103: in response to satisfying a recommendation display condition of the search information, recommending and displaying the search information.

[0046] The above S101-S103 will be described in detail below.

[0047] For the above S101, in the embodiment of the present disclosure, when determining the search information, the text information extracted by the determined at least one target multimedia content is specifically obtained. When determining the at least one target multimedia content, considering that the text information used to determine the search information can be directly extracted from the target multimedia content without relying on manual processing, and these target multimedia contents contain meaningful search information, when selecting the target multimedia content in the embodiment of the present disclosure, the target multimedia content with extractable text information and the feature in at least one attribute dimension satisfying the preset screening condition is selected. The screening methods of the two aspects will be described below.

[0048] First, for selecting the target multimedia content with extractable text information, since the search information associated with the target multimedia content is further determined based on the target multimedia content, such search information can search the target multimedia content and reflect the content essentially indicated by the target multimedia content, therefore, if the target multimedia content has extractable text information, these text information can reflect the essential content contained in the target multimedia content, and can also be used for text processing to obtain a short sentence suitable as search information. Here, for multimedia, whether it is a news type with direct text information, or a video, audio, etc. which can recognize text information through speech-to-text recognition, image-to-text recognition, or has user-added labels and summaries when publishing, it can be considered that the multimedia information has extractable text information.

[0049] Secondly, the features of at least one attribute dimension satisfying the preset filtering condition are described. Since the filtered target multimedia content is specifically used to determine the search information, and the search information is specifically used for user selection search, it is generally directed to search for hot events, current news or other information with search significance. However, not all multimedia content can have information containing specific search significance, such as a randomly taken landscape video, which may not be suitable as target multimedia content because the associated searchable content is less and is not the content that most users may be interested in. Therefore, when determining the target multimedia content from a large number of available multimedia content, the multimedia content published by the official media or the user with a large number of attention users in a relatively short time can be selected as the target multimedia content.

[0050] That is, the at least one attribute dimension of the target multimedia content includes at least one of the following: the publishing time of the target multimedia content, the publishing channel, and the publishing frequency of the publishing channel in a preset time period. In this way, when filtering the target multimedia content, the target multimedia content can be determined from the multimedia content by filtering conditions such as whether the publishing time is within the last two days, whether it is published by an official certified user, whether the number of attention users is greater than 10,000, whether the publishing account has published a video in the last day, etc. Under such conditions, the target multimedia content filtered is mostly multimedia content with an official publishing channel and a large number of users paying attention, which is more suitable for processing to obtain search information.

[0051] In this way, at least one target multimedia content can be determined by the above conditions.

[0052] For S102 described above, in the case of determining the target multimedia content, the search information can also be generated based on the text information extracted from each target multimedia content.

[0053] Firstly, the text information extracted from the target multimedia content is described below. Similar to the examples described above, if the target multimedia content itself is multimedia content containing text information, such as a news report, the text information in the news report can be directly used as the text information of the target multimedia content. For multimedia content in the form of video, audio, etc., in one possible case, there is associated text information such as tags, summaries, etc. at the time of publishing, which will be part of the text information that can be extracted from the target multimedia content. In another possible case, if there is identifiable text information in the multimedia content, such as a video with a picture that can recognize the text and / or speech that can translate the text, or audio with speech that can translate the text, these text information obtained by recognition can also be used as the text information that can be extracted from the target multimedia content.

[0054] In the embodiments of the present disclosure, the target multimedia content is taken as an example of a video, and in specific implementations, the text information associated with the target multimedia content is determined in the following manner: at least one key frame in the video is determined; the key frame of the video includes at least one of the following: at least one video frame corresponding to a cover of the video, and at least one video frame selected from each video frame included in the video; and character recognition is performed on each key frame to obtain text information in each key frame as the text information associated with the video.

[0055] In one possible case, for a video expressing a hot event or the like, the video usually has a cover expressing the content of the video, such as information displayed by text and pictures, so that a user can easily determine the main content of the video. The cover of the video can be a static video cover including one video frame, or a dynamic video cover including multiple video frames. That is, the cover of the video can include one video frame, or multiple video frames displayed in succession, or multiple video frames with changing content. In another possible case, the video can have a video frame containing specific information, such as a video frame displaying a report, a summary of an event, a result of an evaluation, or the like, which can also be used to determine search information for the target multimedia content, and thus can be selected as at least one key frame in the video.

[0056] For at least one key frame determined in the video, text information in each key frame can be obtained by performing character recognition on the key frame as the text information associated with the video. After the character recognition, the obtained information can be de-duplicated and combined with text information such as tags and summaries associated with the video to obtain the text information associated with the video.

[0057] On the basis of the text information determined for the target multimedia content, the specific manner of generating search information from the text information is described as follows. The text information determined can be long, and can also include inappropriate words, sentences irrelevant to the subject indicated by the target multimedia content, and other redundant words and sentences. Therefore, for the determined text information, target text information can be further selected by screening, and search information can be determined based on the target text information.

[0058] When filtering target text information that meets preset filtering conditions from text information, preset filtering conditions can be used for filtering. Specifically, the preset filtering conditions include at least one of the following: the regional attribute information corresponding to the text recognition region in the keyframe of the text information meets the preset regional attribute information requirements, and the semantic content corresponding to the text information indicates the theme information of the target multimedia content; wherein, the regional attribute information includes the region area and / or the region location.

[0059] Here, taking the target multimedia content, including video, as an example, see [link to relevant documentation]. Figure 2 The image shown is a schematic diagram of a video frame provided in an embodiment of this disclosure. Firstly, using... Figure 2 Taking the video frame image shown in (a) as an example, when filtering target text information, the region of the text information can be determined. In one possible scenario, the text suitable as target text information in a video frame usually has a corresponding text recognition region in several specific locations, such as near the top or bottom of the video frame, or perhaps vertically arranged on the left side of the video frame. These locations are also often where titles, summary paragraphs, and audio subtitles of the main character in the video are located, making it easier to extract search information. Figure 2 The middle (a) bottom section is the "XX Food Festival" selected with a dashed box above the video frame.

[0060] In another possible scenario, when filtering target text information, the area of ​​the text information can also be used to determine the target. Figure 2 For example, in the video frame image (b), there are multiple text information that can be identified, but these may include redundant information such as signage information and store trademark information. Figure 2 The sign below (b) reads "No Parking Here." However, in video frames, the area occupied by this text information is typically smaller than other separately displayed text recognition areas used to determine search information. For example, the text recognition area for "Precautions for Traveling During the Short Holiday" is larger than the text recognition area for the sign "No Parking Here." Therefore, "Precautions for Traveling During the Short Holiday" can be selected as the target text information, instead of "No Parking Here."

[0061] To obtain more concise search information when determining the corresponding target text information, the following method can be used: perform semantic segmentation on the target text information to obtain multiple semantic words; select at least one keyword from the multiple semantic words that expresses the theme corresponding to the target multimedia content, and integrate them into the search information.

[0062] For example, if for a certain target multimedia content, the target text information that can be determined includes "The first XX Food Festival was successfully held in urban area A today, and there were many tourists coming to visit and taste, and everyone is welcome to come and visit", that is, the target multimedia content is specifically around the theme of "XX Food Festival", and the target text information specifically expresses this theme, but there is obviously a problem of too long length and not suitable for being used as search information. Therefore, semantic segmentation processing can be performed first to obtain a plurality of semantic words, such as "first", "XX Food Festival", "today", "urban area A", "tourist", "visit", "taste", and the like. Among them, the keywords specifically around "XX Food Festival" can include "first", "XX Food Festival", "today", and "urban area A", and after integration, the search information "today the first XXX Food Festival opens in urban area A" can be obtained.

[0063] Based on the keywords corresponding to the target multimedia content described above, the search information corresponding to the target multimedia content is determined, that is, in the example, the keywords "first", "XX Food Festival", "today", and "urban area A" are integrated into "today the first XXX Food Festival opens in urban area A", and the following method can be used: based on the trained text reconstruction model, the at least one keyword is text reconstructed to obtain a sentence containing the at least one keyword as the search information; and the text reconstruction model is trained based on each keyword sample and a sentence sample composed of a plurality of keyword samples.

[0064] Here, the text reconstruction model can be a trained neural network, and when training the neural network, a sample training method can be used, that is, a plurality of training samples are provided, each training sample includes a plurality of keyword samples, and a sentence composed of a plurality of keywords is trained as a sentence sample to enable the neural network to learn the ability to combine keywords to construct search information. Alternatively, the text reconstruction model can also be other trained models, such as a model that can perform semantic reconstruction obtained by sample training, such as a model that can determine a template corresponding to search information according to keywords, such as the keywords in the above example, the model can determine the template "(XX time) (XX object) (XX event occurs)", and the keywords can be directly filled into the corresponding position or the expression way can be changed to fill into the corresponding position. The appropriate text reconstruction model can be selected according to the actual situation, and the specific limitation is not made here.

[0065] In another embodiment of the present disclosure, as there can be multiple topics around the same theme in association with the target multimedia content of a hot event or the like, such as around the topic of "Star A's wedding", some multimedia content expresses the essence of Star A's wedding photos, some multimedia content expresses the essence of Star A's wedding photos, and some multimedia content expresses the essence of the artists attending Star A's wedding. The obtained target text information includes, for example, "Star A's wedding has been publicly announced in the media by famous photographers", "Star A held a wedding in a hotel and posted wedding images on social media", and "Today, many friends attended Star A's wedding and artist B appeared at the wedding". In this case, multiple different search information can be generated, and the following methods can be used:

[0066] From the at least one keyword corresponding to the target text information, an entity keyword and at least one limited keyword are determined; wherein the entity keyword indicates an entity object associated with the target multimedia content, and the limited keyword is used to indicate the information dimension of the entity object; and based on the entity keyword and the limited keyword, search information corresponding to the target multimedia content is generated.

[0067] Taking the above examples as examples, the entity keyword is, for example, "Star A's wedding", that is, the entity object. And "wedding photos", "wedding photos", and "artist B" can all be used as limited keywords to indicate multiple information dimensions under "Star A's wedding". In this case, the search information corresponding to the target multimedia content can be composed of the entity keyword and the limited keyword, such as "Star A's wedding photos", "Star A's wedding photos", and "Star A's wedding artist B".

[0068] In this case, each limited keyword specifically indicates a dimension, but in multiple multimedia content related to the dimension, through the associated target text information, more than one limited keyword can be obtained. For example, the above determined limited keyword "wedding photos" indicates the dimension of photos taken by Star A during the wedding, and in other related multimedia content, through the associated target text information, the limited keyword obtained can be "wedding photos", "wedding photos", "wedding photos", and the like. Since these limited keywords are actually multiple different expressions around the dimension of "Star A's wedding photos" around the entity keyword "Star A's wedding", they are not suitable for display one by one as search information, but are more suitable for selection or re-summarization and display.

[0069] In a specific implementation, in the above case, the semantic similarity between the at least one limited keyword corresponding to each of the target multimedia content can be determined in response to the fact that the entity keywords corresponding to the plurality of target multimedia content are the same; and the search information corresponding to at least two target multimedia content, for which the semantic similarity exceeds a preset similarity threshold, can be de-duplicated to obtain one search information corresponding to the at least two target multimedia content.

[0070] Continuing with the above example, if five target multimedia contents a, b, c, d, and e are obtained, the entity keywords corresponding to the five target multimedia contents are all "Star A wedding", and the limited keywords corresponding to the five target multimedia contents are a: "wedding photos", b: "wedding photos", c: "wedding photos", d: "wedding photos", and e: "Artist B", respectively. The semantic similarity of the limited keywords corresponding to the target multimedia contents b, c, and d can be determined by semantic similarity, and thus the search information corresponding to the target multimedia contents b, c, and d can be de-duplicated. For example, if "wedding photos" is selected, the search information corresponding to the target multimedia contents b, c, and d is "Star A wedding wedding photos", and accordingly, the associated target multimedia contents b, c, and d can be found by searching using the search information "Star A wedding wedding photos".

[0071] In determining the semantic similarity between the limited keywords, a trained neural network can also be used for processing. In setting the similarity threshold, the similarity threshold can be determined according to the detection accuracy of the neural network and the judgment standard for determining whether the limited keywords can belong to the same semantic, and can be set to 70%, 75%, etc. The greater the similarity threshold, the more stringent the standard for determining whether the limited keywords are the same in semantic, and the number of search information obtained can also increase accordingly.

[0072] In this way, through the above embodiments, search information can be generated from the text information extracted from each target multimedia content.

[0073] For the generated search information, the search information can be recommended and displayed when the recommended display condition of the search information is met.

[0074] Here, the scenarios in which the search information can be displayed include multiple scenarios, and the corresponding recommended display conditions are different in different scenarios. Different scenarios will be described below. For example, as shown in FIG. 1, a plurality of display scenarios provided by an embodiment of the present disclosure are shown. Figure 3

[0075] For example, as shown in FIG. 1, a plurality of display scenarios provided by an embodiment of the present disclosure are shown. Figure 3 ​Taking the scenario shown in (a) as an example, this scenario is a search information recommendation scenario, specifically displaying multiple search information in the form of a list on the search recommendation page. In this scenario, if the user chooses to switch to displaying this search recommendation page, it is considered that the conditions for recommending and displaying search information are met, and the information is displayed accordingly in the form shown in the example, or in other optional forms.

[0076] by Figure 3 Taking the scenario shown in (b) as an example, this scenario involves displaying corresponding search information on the video playback page after consuming multimedia content. In this scenario, users can view search information associated with the multimedia content while consuming the video on the video playback page. The search information is not necessarily based on the multimedia content itself, but it reflects its essence and can serve as search terms for users to continue searching and expand to obtain more multimedia content. In this scenario, when a user selects to play a certain multimedia content, the recommended display conditions for the corresponding search information are met, and it is displayed accordingly in the form of being associated with video frames, as shown in the example.

[0077] The above only lists two possible scenarios. On other pages where search information can be displayed, the search information can also be displayed according to the actual situation. These will not be listed one by one here.

[0078] The displayed search information can be used to perform search functions, and users can obtain multimedia content related to the search information through simple triggering operations.

[0079] In specific implementation, the following approach can be adopted: in response to a search trigger operation on the search information, determine multiple multimedia contents associated with the search information; the multiple multimedia contents include the target multimedia content of the search information; in response to the multiple multimedia contents including video, if there is a target video frame in the video containing the keywords in the search information, replace the target video frame with the cover of the video for display.

[0080] Specifically, in response to a search-triggered operation on the search information, multiple associated multimedia content can be obtained. According to the description in the above embodiments, among the multiple multimedia content obtained, there may be target multimedia content that determines the search information, or there may be other related multimedia content.

[0081] In addition, in a possible case, taking a video in multimedia content as an example, since the video has multiple video frames, the video frame directly expressing the specific content of the search information may not be displayed at any position in the video. The video frame avoiding highlighting the search information is proposed in the middle or rear part of the video. For example, in the search information "wedding photos of star A", the multimedia content of 2 minutes in length obtained by association is displayed at 1:00, and the specific wedding photos are displayed. The target video frame in the video containing the keyword in the search information can be replaced with the cover of the video for display. In this way, the user does not need to spend time watching a long video to view the information that the user wants to view in the triggered search information intent. The user has a better experience, and in this display mode, the video cover after the change can also intuitively display the information that matches the search information better.

[0082] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0083] Based on the same inventive concept, the disclosure embodiments also provide a search information display device corresponding to the search information display method. Since the principle of solving problems of the device in the disclosure embodiments is similar to the above search information display method of the disclosure embodiments, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described.

[0084] Reference Figure 4 As shown in FIG. 1, a search information display device provided by the disclosure embodiments includes a determination module 41, a generation module 42, and a display module 43.

[0085] The determination module 41 is configured to determine at least one target multimedia content. The target multimedia content has extractable text information, and the feature of the target multimedia content in at least one attribute dimension satisfies a preset screening condition.

[0086] The generation module 42 is configured to generate search information based on the text information extracted from each target multimedia content.

[0087] The display module 43 is configured to recommend and display the search information in response to a recommendation display condition of the search information being satisfied.

[0088] In an alternative implementation, in response to the target multimedia content comprising a video, the text information associated with the target multimedia content is determined in the following manner: at least one key frame in the video is determined; the key frame of the video comprises at least one of the following: at least one video frame corresponding to a cover of the video, at least one video frame selected from each video frame included in the video; each of the key frames is subjected to character recognition to obtain text information in each of the key frames as the text information associated with the video.

[0089] In an alternative implementation, when the generation module 42 generates the search information based on the text information, it is configured to: from each of the text information, filter out target text information meeting a preset filtering condition; the preset filtering condition comprises at least one of the following: region attribute information corresponding to a text recognition region in the key frame of the text information meeting a preset region attribute information requirement, and semantic content corresponding to the text information indicating theme information of the target multimedia content; wherein the region attribute information comprises region area and / or region position; and determine the search information based on the target text information.

[0090] In an alternative implementation, when the generation module 42 determines the search information based on the target text information, it is configured to: perform semantic segmentation processing on the target text information to obtain a plurality of semantic words; filter out at least one keyword expressing a theme corresponding to the target multimedia content from the plurality of semantic words and integrate the keyword as the search information.

[0091] In an alternative implementation, when the generation module 42 determines the search information based on the target text information, it is configured to: determine an entity keyword and at least one limiting keyword from at least one keyword corresponding to the target text information; wherein the entity keyword indicates an entity object associated with the target multimedia content, and the limiting keyword is used to indicate an information dimension of the entity object; and generate the search information corresponding to the target multimedia content based on the entity keyword and the limiting keyword.

[0092] In an alternative implementation, the search information corresponding to the target multimedia content is determined based on the keyword corresponding to the target multimedia content in the following manner: the at least one keyword is subjected to text reconstruction based on a trained text reconstruction model to obtain a sentence containing the at least one keyword as the search information; and the text reconstruction model is trained based on each keyword sample and a sentence sample composed of a plurality of keyword samples.

[0093] In an alternative implementation, the generating module 42 is further configured to: in response to the same entity keyword corresponding to a plurality of target multimedia contents existing, determine semantic similarity between at least one limited keyword corresponding to each of the target multimedia contents; and in response to at least two target multimedia contents corresponding to the semantic similarity exceeding a preset similarity threshold existing in the plurality of target multimedia contents, perform deduplication processing on search information corresponding to the at least two target multimedia contents to obtain one search information corresponding to the at least two target multimedia contents.

[0094] In an alternative implementation, the at least one attribute dimension of the target multimedia content includes at least one of the following: a publishing time of the target multimedia content, a publishing channel, and a publishing frequency of the publishing channel in a preset time period.

[0095] In an alternative implementation, the apparatus further includes a processing module 44 configured to: in response to a search trigger operation on the search information, determine a plurality of multimedia contents associated with the search information; the plurality of multimedia contents including a target multimedia content of the search information; and in response to the plurality of multimedia contents including a video, if a target video frame containing a keyword in the search information exists in the video, replace the target video frame with a cover of the video for display.

[0096] The processing flow of each module in the apparatus and the interaction flow between the modules can refer to the related description in the above method embodiments, and will not be described in detail here.

[0097] The present disclosure also provides a computer device, as shown in Figure 5 The computer device structure schematic diagram provided by the present disclosure includes:

[0098] a processor 10 and a memory 20; the memory 20 stores machine readable instructions executable by the processor 10, and the processor 10 is configured to execute the machine readable instructions stored in the memory 20; when the machine readable instructions are executed by the processor 10, the processor 10 performs the following steps:

[0099] determine at least one target multimedia content; the target multimedia content has extractable text information, and a feature of the target multimedia content in at least one attribute dimension satisfies a preset screening condition; generate search information based on the text information extracted from each of the target multimedia contents; and in response to a recommendation display condition of the search information being satisfied, recommend and display the search information.

[0100] The memory 20 includes an internal memory 210 and an external memory 220; the internal memory 210 is also referred to as an internal storage, and is used to temporarily store operation data in the processor 10 and exchange data with the external memory 220 such as a hard disk; the processor 10 exchanges data with the external memory 220 through the internal memory 210.

[0101] The specific execution process of the above instructions can refer to the steps of the display method of search information described in the embodiments of the present disclosure, which will not be repeated here.

[0102] The embodiments of the present disclosure further provide a computer readable storage medium, which stores a computer program. When the computer program is run by a processor, the steps of the display method of search information described in the method embodiments are executed. The storage medium can be a volatile or non-volatile computer readable storage medium.

[0103] The embodiments of the present disclosure further provide a computer program product, which carries a program code. The instructions included in the program code can be used to execute the steps of the display method of search information described in the method embodiments. For details, refer to the method embodiments, which will not be repeated here.

[0104] The computer program product can be specifically implemented by hardware, software or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.

[0105] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system and device can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed system, device and method can be implemented by other ways. The device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can be another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.

[0106] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0107] In addition, each functional unit in various embodiments of the present disclosure can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0108] If the functions are realized in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present disclosure essentially or the part of the prior art or the part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various program code storage media.

[0109] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, and not to limit them, the protection scope of the present disclosure is not limited thereto, although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: any person skilled in the art in the technical range disclosed by the present disclosure, it can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and all should be covered in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for displaying search information, characterized in that, include: Identify at least one target multimedia content; The target multimedia content has extractable text information, and the features of the target multimedia content in at least one attribute dimension satisfy preset filtering conditions. Search information is generated based on the text information extracted from each of the target multimedia contents, wherein the search information is used by the user to select a search. In response to meeting the recommendation display criteria for search information, the search information is recommended and displayed. The step of generating search information based on the text information extracted from each of the target multimedia contents includes: From the various text information, select the target text information that meets the preset filtering conditions; From at least one keyword corresponding to the target text information, determine entity keywords and at least one qualifying keyword; wherein, the entity keyword indicates the entity object associated with the target multimedia content, and the qualifying keyword is used to indicate the information dimension of the entity object. In response to the existence of multiple target multimedia content corresponding to the same entity keywords, the semantic similarity between at least one limited keyword corresponding to each of the target multimedia content is determined; In response to the existence of at least two target multimedia contents with a semantic similarity exceeding a preset similarity threshold among multiple target multimedia contents, the search information corresponding to the at least two target multimedia contents is deduplicated to obtain a single search information that is commonly associated with the at least two target multimedia contents.

2. The method according to claim 1, characterized in that, In response to the target multimedia content including video, the text information associated with the target multimedia content is determined in the following manner: Determine at least one keyframe in the video; the keyframe of the video includes at least one of the following: at least one video frame corresponding to the cover of the video, or at least one video frame selected from the video frames contained in the video. Character recognition is performed on each of the keyframes to obtain the text information in each keyframe, which is used as the text information associated with the video.

3. The method according to claim 1 or 2, characterized in that, The preset filtering conditions include at least one of the following: the regional attribute information corresponding to the text recognition region in the keyframe of the text information satisfies the preset regional attribute information requirements, and the semantic content corresponding to the text information indicates the theme information of the target multimedia content; wherein, the regional attribute information includes the region area and / or the region location.

4. The method according to claim 1, characterized in that, Also includes: The target text information is semantically segmented to obtain multiple semantic words; At least one keyword expressing the theme corresponding to the target multimedia content is selected from the plurality of semantic words and integrated into the search information.

5. The method according to claim 1, characterized in that, Also includes: Based on the entity keywords and limiting keywords, search information corresponding to the target multimedia content is generated.

6. The method according to claim 4 or 5, characterized in that, The following method is used to determine the search information corresponding to the target multimedia content based on the keywords corresponding to the target multimedia content: Based on the trained text reconstruction model, the text is reconstructed for the at least one keyword to obtain a sentence containing the at least one keyword, which is used as the search information; the text reconstruction model is trained based on each keyword sample and a sentence sample composed of multiple keyword samples.

7. The method according to claim 1, characterized in that, At least one of the following attribute dimensions of the target multimedia content includes: The target multimedia content includes the release time, release channel, and the release frequency of the multimedia content released by the release channel within a preset time period.

8. The method according to claim 1, characterized in that, Also includes: In response to a search-triggered operation on the search information, multiple multimedia contents associated with the search information are determined; the multiple multimedia contents include the target multimedia content for which the search information is determined. In response to the multiple multimedia contents including videos, if there is a target video frame in the video that contains the keywords in the search information, the target video frame is replaced with the cover of the video for display.

9. A device for displaying search information, characterized in that, include: A determination module is used to determine at least one target multimedia content; The target multimedia content has extractable text information, and the features of the target multimedia content in at least one attribute dimension satisfy preset filtering conditions. A generation module is used to generate search information based on the text information extracted from each of the target multimedia contents, wherein the search information is used by the user to select a search. The display module is used to recommend and display search information in response to conditions that meet the search criteria. The step of generating search information based on the text information extracted from each of the target multimedia contents includes: From the various text information, select the target text information that meets the preset filtering conditions; From at least one keyword corresponding to the target text information, determine entity keywords and at least one qualifying keyword; wherein, the entity keyword indicates the entity object associated with the target multimedia content, and the qualifying keyword is used to indicate the information dimension of the entity object. In response to the existence of multiple target multimedia content corresponding to the same entity keywords, the semantic similarity between at least one limited keyword corresponding to each of the target multimedia content is determined; In response to the existence of at least two target multimedia contents with a semantic similarity exceeding a preset similarity threshold among multiple target multimedia contents, the search information corresponding to the at least two target multimedia contents is deduplicated to obtain a single search information that is commonly associated with the at least two target multimedia contents.

10. A computer device, characterized in that, include: A processor and a memory, the memory storing machine-readable instructions executable by the processor, the processor executing the machine-readable instructions stored in the memory, and when the machine-readable instructions are executed by the processor, the processor performing the steps of the method for displaying search information as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer device, performs the steps of the method for displaying search information as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for outputting information, equipment and storage medium

    CN111523019A

  • Natural language processing method and device and electronic equipment

    CN112800201A

  • Search recommendation information generation and display method and device, equipment and storage medium

    CN113486212A