Information display method and device, equipment and storage medium

By obtaining and displaying video frames, audio text and video file information related to the search information in the video editing interface, the time-consuming problem of selecting covers from multiple video frames is solved, and the efficiency of video editing and user experience are improved.

CN120658909APending Publication Date: 2025-09-16BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410302894.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

During the video editing process, selecting a suitable video frame from multiple video frames as a cover takes a long time, resulting in low editing efficiency.

Method used

By responding to the input of the material search box in the video editing interface, the correspondence between the video frame text, video frame information, audio text and video file information related to the search information is obtained, the first search result information is obtained from the preset database, and displayed in the interface so that users can quickly find suitable video materials.

Benefits of technology

It reduces the time users spend searching for video materials and improves the efficiency of video editing and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658909A_ABST
    Figure CN120658909A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an information display method and device, equipment and a storage medium, and the method comprises the steps: responding to search information inputted by a material search box of a video editing interface, and obtaining first search result information corresponding to the search information from a preset database, wherein the preset database comprises a corresponding relation between a video frame text and video frame information, a corresponding relation between an audio text and the video frame information, and a corresponding relation between video file information and the video frame information; displaying the first search result information in the video editing interface; wherein the first search result information comprises target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and an audio text corresponding to the target video frame information. The method and the device can reduce the time of looking for the materials by the user and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of video editing technology, and more particularly to an information display method, apparatus, device, and storage medium. Background Art

[0002] With the development of video editing technology, more and more users choose to add video covers to their videos through video editing software. Among them, a good video cover can not only show the content of the video, but also increase the attractiveness of the video.

[0003] In the related art, when adding a video cover to a video, the user can first watch the video through video preview or the like, and then select a suitable video frame from multiple video frames included in the video, and use the video frame as the video cover.

[0004] However, the inventors have discovered that the prior art has at least the following technical problems: when a video includes a large number of video frames, it takes a long time to select a suitable video frame from the multiple video frames, resulting in low efficiency in video editing. Summary of the Invention

[0005] The embodiments of the present disclosure provide an information display method, apparatus, device, and storage medium, which can reduce the time users spend looking for materials and improve the efficiency of video editing.

[0006] In a first aspect, an embodiment of the present disclosure provides an information display method, including:

[0007] In response to search information input into a material search box on a video editing interface, obtaining first search result information corresponding to the search information from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information;

[0008] The first search result information is displayed in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

[0009] In a second aspect, an embodiment of the present disclosure provides an information display device, comprising:

[0010] an acquisition module, configured to, in response to search information input into a material search box on a video editing interface, acquire first search result information corresponding to the search information from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information;

[0011] A display module is used to display the first search result information in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

[0012] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:

[0013] a processor, and a memory communicatively connected to the processor;

[0014] The memory stores computer-executable instructions;

[0015] The processor executes the computer-executable instructions stored in the memory to implement the information display method as described in the first aspect above.

[0016] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the information display method described in the first aspect above is implemented.

[0017] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the information display method described in the first aspect above.

[0018] The information display method, apparatus, device and storage medium provided in this embodiment include: in response to the search information input into the material search box of the video editing interface, obtaining first search result information corresponding to the search information from a preset database, wherein the preset database includes the correspondence between video frame text and video frame information, the correspondence between audio text and video frame information, and the correspondence between video file information and video frame information; displaying the first search result information in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information. In the embodiment of the present application, since video materials of multiple dimensions such as video file information, video frame information and audio text related to the search information can be searched through the search information, it is convenient for users to find suitable video materials for video editing, which reduces the time users spend looking for video materials, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 A schematic diagram of an application scenario of an information display method provided by an embodiment of the present disclosure;

[0021] Figure 2 A flowchart of an information display method provided in an embodiment of the present disclosure;

[0022] Figure 3 A schematic diagram of an information display method provided by an embodiment of the present disclosure;

[0023] Figure 4 A flowchart of another information display method provided by an embodiment of the present disclosure;

[0024] Figure 5 A flowchart of another information display method provided by an embodiment of the present disclosure;

[0025] Figure 6 A schematic diagram of an information display method provided by an embodiment of the present disclosure;

[0026] Figure 7 A schematic diagram of another information display method provided by an embodiment of the present disclosure;

[0027] Figure 8A flowchart of another information display method provided by an embodiment of the present disclosure;

[0028] Figure 9 A structural block diagram of an information display device provided in an embodiment of the present disclosure;

[0029] Figure 10 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] With the development of video editing technology, more and more users choose to add video covers to their videos through video editing software. Among them, a good video cover can not only show the content of the video, but also increase the attractiveness of the video.

[0033] In related art, when adding a video cover to a video, the user can first watch the video through video preview or other methods, then select a suitable video frame from the multiple video frames included in the video and use this video frame as the video cover. However, when the video includes a large number of video frames, it takes a long time to select the appropriate video frame from the multiple video frames, resulting in low video editing efficiency.

[0034] It can be seen that how to reduce the time users spend looking for materials to improve the efficiency of video editing is an urgent problem that needs to be solved.

[0035] In order to solve the above problems, this embodiment provides the following technical concept: users can search for video materials in multiple dimensions such as video files, video frame information, and video frame lines related to the search information through search terms, which makes it easier for users to find suitable video materials for video editing and reduces the time users spend looking for video materials.

[0036] The specific steps may include: first, in response to search information input into the material search box of the video editing interface, obtaining first search result information corresponding to the search information from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information; then, displaying the first search result information in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

[0037] Here, since the search information can be used to search for video materials of multiple dimensions such as video file information, video frame information and audio text related to the search information, it is convenient for users to find suitable video materials for video editing, reducing the time users spend looking for video materials, thereby improving user experience.

[0038] The following explains the application scenarios of the embodiments of the present disclosure:

[0039] The information display method provided by the embodiments of the present disclosure can be applied to various video editing scenarios. For example, the information display method provided by the embodiments of the present disclosure can provide users with video materials corresponding to video A. Users can create a cover for video A using the video materials corresponding to video A. For another example, the information display method provided by the embodiments of the present disclosure can provide users with multiple video materials corresponding to multiple videos. Users can classify the same type of materials in multiple video materials and then complete later video editing using the classified material videos.

[0040] Figure 1 This is a schematic diagram of an application scenario of an information display method provided by an embodiment of the present disclosure. Figure 1 As shown, the video editing interface features an "Import Material" button. Users can use this button to upload Video A. Furthermore, users can enter search information in the "Material Search Box" to search for relevant video materials across multiple dimensions, including video file information, video frame information, and audio text. This facilitates user search for suitable video materials for editing, reduces the time spent searching for video materials, and thus improves the user experience.

[0041] The information display method provided by the embodiment of the present disclosure is described in detail below using a detailed embodiment.

[0042] Figure 2 This is a flow chart of an information display method provided by an embodiment of the present disclosure. The execution subject of the information display method may be an electronic device. Figure 2As shown, the method includes:

[0043] S201. In response to search information input into a material search box in a video editing interface, first search result information corresponding to the search information is obtained from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information.

[0044] In an embodiment of the present disclosure, a video file may include an audio file and multiple video frame information. The video frame information may include a video frame image, and the video frame text may be the text within the video frame image. The audio text may be the dialogue text corresponding to the video file. The video file information may include any information related to the video file. Optionally, the video file information may include the video file, the video file name, and the video file duration.

[0045] In some embodiments, obtaining first search result information corresponding to the search information from a preset database may include: determining the target video frame information corresponding to at least one target video frame text including the search information from the correspondence between the video frame text and the video frame information included in the preset database; and, for each target video frame information, determining the video file information corresponding to the target video frame information from the correspondence between the video file information and the video frame information included in the preset database; and, for each target video frame information, determining the audio text corresponding to the target video frame information from the correspondence between the audio text and the video frame information included in the preset database.

[0046] For example, Figure 3 As shown, the search information is "dumplings." The target video frame information corresponding to at least one target video frame text containing "dumplings" is determined to be: video frame 1 and video frame 2. The video file information corresponding to video frame 1 is determined to be: video 1. The video file information corresponding to video frame 2 is determined to be: video 2. The audio text corresponding to video frame 1 is determined to be: the lines corresponding to video frame 1.

[0047] S202. Display first search result information in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

[0048] For example, Figure 3As shown, the video file information corresponding to the video frame information is displayed in the video editing interface: Video 1 and Video 2. At least one video frame information: Video Frame 1 and Video Frame 2. The audio text corresponding to the video frame information: the lines corresponding to Video Frame 1.

[0049] In an embodiment of the present disclosure, in response to search information input into the material search box of the video editing interface, before obtaining the first search result information corresponding to the search information from the preset database, the correspondence between the video frame text and the video frame information, the correspondence between the audio text and the video frame information, and the correspondence between the video file information and the video frame information can be stored in the preset database.

[0050] Accordingly, if Figure 4 As shown, the specific steps of storing the corresponding relationship may include:

[0051] S401: In response to an import operation of a video file, obtain video file information of the video file, an audio file of the video file, and multiple video frame information of the video file, wherein the video frame information at least includes a video frame image.

[0052] 402. For each video frame information, perform text recognition processing on the video frame image in the video frame information to obtain the video frame text corresponding to the video frame information, establish a correspondence between the video frame text and the video frame information, and store the correspondence between the video frame text and the video frame information in a preset database; and, perform speech recognition processing on the audio file to obtain the audio text corresponding to each of the multiple video frame information of the video file, establish a correspondence between the audio text and the video frame information, and store the correspondence between the audio text and the video frame information in a preset database; and establish a correspondence between the video file information and the video frame information, and store the correspondence between the video file information and the video frame information in a preset database.

[0053] In some embodiments, the video frame information also includes first time information corresponding to the video frame image; accordingly, the audio file is subjected to speech recognition processing to obtain audio texts corresponding to multiple video frame information of the video file, including: performing speech recognition processing on the audio file to obtain multiple sub-text information and second time information corresponding to each sub-text information; for each video frame information, determining the sub-text information corresponding to the video frame information based on the first time information in the video frame information and the second time information corresponding to each sub-text information, and obtaining audio texts corresponding to multiple video frame information of the video file.

[0054] Optionally, performing speech recognition processing on the audio file to obtain multiple sub-text information and second time information corresponding to each sub-text information includes: performing speech recognition processing on the audio file to obtain initial text information of the audio file, translating the initial text information into text information in a preset language through a preset translation model, dividing the text information in the preset language according to preset time intervals, and obtaining multiple sub-text information and second time information corresponding to each sub-text information.

[0055] For example, if the initial text information is in English and the preset language is Chinese, the initial text information is translated into Chinese text information using a preset translation model, and the Chinese text information is divided into preset time intervals to obtain multiple sub-text information and second time information corresponding to each sub-text information.

[0056] In some embodiments, the video frame information also includes first time information corresponding to the video frame image; accordingly, establishing a correspondence between the audio text and the video frame information includes: determining that the audio text includes multiple sub-text information and second time information corresponding to each sub-text information; for each video frame information, determining the sub-text information corresponding to the video frame information based on the first time information in the video frame information and the second time information corresponding to each sub-text information; establishing a correspondence between the sub-text information and the video frame information.

[0057] In the embodiment of the present disclosure, since the correspondence between video frame text and video frame information, the correspondence between audio text and video frame information, and the correspondence between video file information and video frame information are pre-stored in a preset database, it is convenient to determine the first search result information corresponding to the search information, thereby improving the search efficiency of video materials.

[0058] The disclosed embodiment provides an information display method: in response to search information input into a material search box of a video editing interface, first search result information corresponding to the search information is obtained from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information; the first search result information is displayed in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information. In the embodiment of the present application, since video materials of multiple dimensions such as video file information, video frame information, and audio text related to the search information can be searched through the search information, it is convenient for users to find suitable video materials for video editing, which reduces the time users spend looking for video materials, thereby improving user experience.

[0059] Figure 5This is a flow chart of an information display method provided by an embodiment of the present disclosure. This application can search for video materials by cluster identification. Accordingly, if Figure 5 As shown, the method includes:

[0060] S501: In response to a preset operation on a material search box in a video editing interface, display at least one clustering identifier in a preset area of ​​the video editing interface.

[0061] In the embodiment of the present disclosure, the preset operation may be any trigger operation for the material search box. Optionally, the preset operation may be a click operation or a selection operation.

[0062] Optionally, the preset area can be located at any position in the video editing interface. For example, Figure 3 As shown, the preset area is located below the material search box.

[0063] Optionally, the cluster identifier can be search information recommended to the user. In some embodiments, the cluster identifier can be a character image, which is used to recommend to the user to search for video materials related to a certain character. In other embodiments, the cluster identifier can also be a virtual character image, which is used to recommend to the user to search for video materials related to a certain virtual character. For example, Figure 6 As shown, at least one cluster identifier includes images of 8 people.

[0064] S502. In response to a triggering operation on a target cluster identifier in at least one cluster identifier, second search result information corresponding to the target cluster identifier is displayed in the video editing interface; wherein the second search result information includes video file information and / or at least one video frame information, and the second search result information is determined based on the correspondence between the pre-stored cluster identifier, the video file information and the at least one video frame information.

[0065] In an embodiment of the present disclosure, this step may include: in response to a triggering operation on a target cluster identifier in at least one cluster identifier, determining the video file information and / or at least one video frame information corresponding to the target cluster identifier from the correspondence between the pre-stored cluster identifier, video file information and at least one video frame information, and displaying the video file information and / or at least one video frame information in the video editing interface.

[0066] In some embodiments, the video file information may include any information related to the video file. Optionally, the video file information may include the video file, the video file name, and the video file duration. The video file may include the cluster identifier. Optionally, the cluster identifier may be a person image, and the video file may include the person image.

[0067] It should be noted that the second search result information is video material related to the target cluster identifier.

[0068] In some embodiments, the video frame information may include any information related to the video frame. Optionally, the video frame information may include the video frame, the name of the video file to which the video frame belongs, and the playback time of the video frame in the video file. The video frame may include the cluster identifier. Optionally, the cluster identifier may include a person image, and the video file may include the video file of the person image.

[0069] In some embodiments, the video file information and the video frame information may be displayed in the form of an image. The target cluster identifier may be person A. Optionally, the second search result information corresponding to person A includes: video frames (pictures) including person A and video files including person A.

[0070] For example, Figure 7 As shown, the second search result information is video 1 and video 2 including person A, and video frame 1 and video frame 2 including person A.

[0071] It should be noted that the method may further include: in response to a trigger operation on a cluster identifier, determining a preset object corresponding to the cluster identifier; and displaying graphic information including the cluster identifier and the preset object in the material search box. Figure 7 As shown, the cluster identifier can be an image of person A, and the preset object corresponding to the cluster identifier is: person. The graphic information displayed in the material search box is "person and image of person A".

[0072] In an embodiment of the present application, since video materials can be recommended to users in the form of image clustering, users can quickly find suitable materials for video editing based on cluster identification, which reduces the time users spend looking for materials, improves the efficiency of video editing, and thus improves user experience.

[0073] It should be noted that, in response to a preset operation on the material search box in the video editing interface, before displaying at least one cluster identifier in the preset area of ​​the video editing interface, a correspondence between the cluster identifier, the video file information and the at least one video frame information may be established first. Figure 8 As shown, the method further includes:

[0074] S801: In response to an import operation of a video file, obtain video file information corresponding to the video file, and determine multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0075] In the embodiment of the present disclosure, the video frame information corresponding to the cluster identifier can be obtained by using a preset recognition model. Accordingly, determining multiple cluster identifiers corresponding to a video file and at least one video frame information corresponding to each cluster identifier can include the following steps (1) to (3):

[0076] (1) Extract multiple video frame information from the video file.

[0077] In some embodiments, multiple video frames can be extracted from a video file at preset time intervals. For example, if the playback duration of video file A is 32 seconds, one video frame can be extracted every 1 second, resulting in 32 video frames.

[0078] In other embodiments, multiple video frame information may be extracted from a video file at a preset video frame interval. For example, if video file A includes 100 video frame information, one video frame information may be extracted every five video frames, resulting in 20 video frame information.

[0079] (2) For each video frame information, if image recognition is performed on the video frame information through a preset recognition model to obtain image information of a preset object image and an image feature vector of the preset object image, the video frame information is determined to be video frame information to be classified.

[0080] In the disclosed embodiments, the preset object image can be any type of object image. Optionally, the preset object image can be a facial image. Accordingly, if image recognition is performed on the video frame information using a preset recognition model to obtain image information of the facial image and an image feature vector of the facial image, the video frame information is determined to be video frame information to be classified.

[0081] Optionally, the image information may include the image area and image score of the image recognition. Exemplarily, the preset recognition model may be a face recognition model. The image information may include the face area and the face image score. The higher the score of the recognized image, the higher the quality of the face image (e.g., face clarity, facial expression, etc.). Optionally, if the score of the recognized image is greater than a preset value, the recognized image is determined to be a face image; if the score of the recognized image is less than the preset value, the recognized image is determined not to be a face image.

[0082] In some embodiments, the resolution of the video frame information can be adjusted before performing image recognition on the video frame information using a preset recognition model. Optionally, the resolution of the video frame information can be reduced to a preset resolution. The present disclosure does not specifically limit the value of the preset resolution and can be set and modified as needed. For example, the preset resolution can be: 320×240, 480×320, 640×480, etc.

[0083] Here, by adjusting the resolution of the video frame information, the size of the video frame information can be reduced, the resources occupied by the video frame information can be reduced, and the efficiency of image recognition can be improved.

[0084] (3) Dividing the plurality of video frame information to be classified according to the image information of the preset object image and the image feature vector of the preset object image corresponding to each of the plurality of video frame information to be classified to obtain a plurality of cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0085] In some embodiments, the plurality of video frames to be classified are divided by a preset clustering model. Accordingly, this step may include the following steps (a) to (c):

[0086] (a) Clustering the plurality of video frames to be classified using a preset clustering model according to image feature vectors of preset object images corresponding to the plurality of video frames to be classified, thereby obtaining a plurality of clustering result information, wherein the clustering result information includes at least one video frame information having a vector similarity greater than a preset threshold.

[0087] The present disclosure does not specifically limit the value of the preset threshold, and may be set and modified as needed. Optionally, the preset object may be a person. The preset object image may be a person image. The preset clustering model may be a portrait clustering model.

[0088] (b) For each clustering result information, determine a preset object image corresponding to each video frame information in the clustering result information to obtain at least one preset object image.

[0089] (c) Based on the image information corresponding to each of the at least one preset object images, a cluster identifier corresponding to the clustering result information is selected from the at least one preset object image to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0090] Optionally, the image information may include an image area and an image score for image recognition. Accordingly, selecting a cluster identifier corresponding to the clustering result information from the at least one preset object image based on the image information corresponding to each of the at least one preset object images includes: selecting the preset object image with the highest image score from the at least one preset object image based on the image score corresponding to each of the at least one preset object images, and using the preset object image with the highest image score as the cluster identifier corresponding to the clustering result information.

[0091] S802: For each cluster identifier, establish a corresponding relationship between the cluster identifier, video file information, and at least one video frame information, and store the corresponding relationship in a preset database.

[0092] In the embodiment of the present disclosure, since a correspondence between the cluster identifier, video file information and at least one video frame information is first established before displaying at least one cluster identifier in a preset area of ​​the video editing interface, it is convenient to determine the video file information and / or at least one video frame information corresponding to the target cluster identifier from the correspondence between the cluster identifier, video file information and at least one video frame information pre-stored in a preset database, thereby improving the search efficiency of video materials.

[0093] Figure 9 This is a structural block diagram of an information display device provided by an embodiment of the present disclosure. Figure 9 , the device includes: an acquisition module 901 and a display module 902;

[0094] An acquisition module 901 is configured to, in response to search information input into a material search box on a video editing interface, acquire first search result information corresponding to the search information from a preset database, wherein the preset database includes correspondences between video frame text and video frame information, between audio text and video frame information, and between video file information and video frame information;

[0095] Display module 902 is used to display the first search result information in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

[0096] According to one or more embodiments of the present disclosure, the device also includes: a storage module; the storage module is used to obtain the video file information of the video file, the audio file of the video file and multiple video frame information of the video file in response to the import operation of the video file, wherein the video frame information at least includes a video frame image; for each video frame information, perform text recognition processing on the video frame image in the video frame information to obtain the video frame text corresponding to the video frame information, establish a correspondence between the video frame text and the video frame information, and store the correspondence between the video frame text and the video frame information in a preset database; and perform speech recognition processing on the audio file to obtain the audio text corresponding to each of the multiple video frame information of the video file, establish a correspondence between the audio text and the video frame information, and store the correspondence between the audio text and the video frame information in the preset database; and establish a correspondence between the video file information and the video frame information, and store the correspondence between the video file information and the video frame information in the preset database.

[0097] According to one or more embodiments of the present disclosure, the video frame information also includes first time information corresponding to the video frame image; accordingly, the storage module performs speech recognition processing on the audio file to obtain audio texts corresponding to multiple video frame information of the video file, specifically including: performing speech recognition processing on the audio file to obtain multiple sub-text information and second time information corresponding to each sub-text information; for each video frame information, determining the sub-text information corresponding to the video frame information based on the first time information in the video frame information and the second time information corresponding to each sub-text information, to obtain audio texts corresponding to multiple video frame information of the video file.

[0098] According to one or more embodiments of the present disclosure, the acquisition module 901 acquires the first search result information corresponding to the search information from a preset database, specifically including: determining the target video frame information corresponding to each of at least one target video frame texts including the search information from the correspondence between the video frame text and the video frame information included in the preset database; and, for each target video frame information, determining the video file information corresponding to the target video frame information from the correspondence between the video file information and the video frame information included in the preset database; and, for each target video frame information, determining the audio text corresponding to the target video frame information from the correspondence between the audio text and the video frame information included in the preset database.

[0099] According to one or more embodiments of the present disclosure, the display module 902 is also used to display at least one cluster identifier in a preset area of ​​the video editing interface in response to a preset operation on the material search box in the video editing interface; in response to a trigger operation on a target cluster identifier among the at least one cluster identifier, display second search result information corresponding to the target cluster identifier in the video editing interface; wherein the second search result information includes video file information and / or at least one video frame information, and the second search result information is determined based on the correspondence between the pre-stored cluster identifier, the video file information and the at least one video frame information.

[0100] According to one or more embodiments of the present disclosure, the storage module is also used to obtain video file information corresponding to the video file in response to an import operation of the video file, and to determine multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier; for each cluster identifier, establish a correspondence between the cluster identifier, the video file information and at least one video frame information, and store the correspondence in a preset database.

[0101] According to one or more embodiments of the present disclosure, the storage module determines multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier, specifically including: extracting multiple video frame information from the video file; for each video frame information, if image recognition is performed on the video frame information through a preset recognition model to obtain image information of a preset object image and an image feature vector of the preset object image, then determining that the video frame information is video frame information to be classified; dividing the multiple video frame information to be classified according to the image information of the preset object image and the image feature vector of the preset object image corresponding to each of the multiple video frame information to be classified to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0102] According to one or more embodiments of the present disclosure, the storage module divides the multiple video frame information to be classified according to the image information of the preset object images corresponding to each of the multiple video frame information to be classified and the image feature vectors of the preset object images, and obtains multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier, specifically including: clustering the multiple video frame information to be classified according to the image feature vectors of the preset object images corresponding to each of the multiple video frame information to be classified through a preset clustering model to obtain multiple clustering result information, wherein the clustering result information includes at least one video frame information whose vector similarity is greater than a preset threshold; for each clustering result information, determining the preset object image corresponding to each video frame information in the clustering result information to obtain at least one preset object image; according to the image information corresponding to each of the at least one preset object image, selecting the cluster identifier corresponding to the clustering result information from the at least one preset object image to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0103] According to one or more embodiments of the present disclosure, the display module 902 is further configured to determine a preset object corresponding to the cluster identifier in response to a trigger operation on the cluster identifier; and display graphic information including the cluster identifier and the preset object in the material search box.

[0104] The first display module 901 and the second display module 902 are connected in sequence. The information display device provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.

[0105] Figure 10 This is a hardware structure diagram of an electronic device provided by an embodiment of the present disclosure. Figure 10The electronic device 1000 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), and vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0106] like Figure 10 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0107] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 10 The electronic device 1000 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0108] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0109] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0110] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0111] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0112] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0114] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0115] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0116] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an information presentation method, comprising:

[0118] In response to search information input into a material search box on a video editing interface, obtaining first search result information corresponding to the search information from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information;

[0119] The first search result information is displayed in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

[0120] According to one or more embodiments of the present disclosure, the response to the search information input in the material search box of the video editing interface, before obtaining the first search result information corresponding to the search information from the preset database, also includes: in response to the import operation of the video file, obtaining the video file information of the video file, the audio file of the video file and multiple video frame information of the video file, wherein the video frame information at least includes a video frame image; for each video frame information, performing text recognition processing on the video frame image in the video frame information to obtain the video frame text corresponding to the video frame information, establishing a correspondence between the video frame text and the video frame information, and storing the correspondence between the video frame text and the video frame information in the preset database; and, performing speech recognition processing on the audio file to obtain the audio text corresponding to each of the multiple video frame information of the video file, establishing a correspondence between the audio text and the video frame information, and storing the correspondence between the audio text and the video frame information in the preset database; and, establishing a correspondence between the video file information and the video frame information, and storing the correspondence between the video file information and the video frame information in the preset database.

[0121] According to one or more embodiments of the present disclosure, the video frame information also includes first time information corresponding to the video frame image; accordingly, the audio file is subjected to speech recognition processing to obtain audio texts corresponding to each of the multiple video frame information of the video file, including: performing speech recognition processing on the audio file to obtain multiple sub-text information and second time information corresponding to each sub-text information; for each video frame information, determining the sub-text information corresponding to the video frame information based on the first time information in the video frame information and the second time information corresponding to each sub-text information, to obtain audio texts corresponding to each of the multiple video frame information of the video file.

[0122] According to one or more embodiments of the present disclosure, obtaining first search result information corresponding to the search information from a preset database includes: determining target video frame information corresponding to each of at least one target video frame texts including the search information from the correspondence between the video frame text and the video frame information included in the preset database; and, for each target video frame information, determining the video file information corresponding to the target video frame information from the correspondence between the video file information and the video frame information included in the preset database; and, for each target video frame information, determining the audio text corresponding to the target video frame information from the correspondence between the audio text and the video frame information included in the preset database.

[0123] According to one or more embodiments of the present disclosure, it also includes: in response to a preset operation on a material search box in a video editing interface, displaying at least one cluster identifier in a preset area of ​​the video editing interface; in response to a trigger operation on a target cluster identifier among the at least one cluster identifier, displaying second search result information corresponding to the target cluster identifier in the video editing interface; wherein the second search result information includes video file information and / or at least one video frame information, and the second search result information is determined based on the correspondence between the pre-stored cluster identifier, video file information and at least one video frame information.

[0124] According to one or more embodiments of the present disclosure, the response to the preset operation of the material search box in the video editing interface, before displaying at least one cluster identifier in the preset area of ​​the video editing interface, also includes: in response to the import operation of the video file, obtaining the video file information corresponding to the video file, and determining the multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier; for each cluster identifier, establishing a correspondence between the cluster identifier, the video file information and the at least one video frame information, and storing the correspondence in a preset database.

[0125] According to one or more embodiments of the present disclosure, determining the multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier includes: extracting multiple video frame information from the video file; for each video frame information, if image recognition is performed on the video frame information through a preset recognition model to obtain image information of a preset object image and an image feature vector of the preset object image, then determining that the video frame information is video frame information to be classified; dividing the multiple video frame information to be classified according to the image information of the preset object image and the image feature vector of the preset object image corresponding to each of the multiple video frame information to be classified to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0126] According to one or more embodiments of the present disclosure, the multiple video frame information to be classified is divided according to the image information of the preset object images corresponding to each of the multiple video frame information to be classified and the image feature vectors of the preset object images, to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier, including: clustering the multiple video frame information to be classified according to the image feature vectors of the preset object images corresponding to each of the multiple video frame information to be classified through a preset clustering model to obtain multiple clustering result information, wherein the clustering result information includes at least one video frame information whose vector similarity is greater than a preset threshold; for each clustering result information, determining the preset object image corresponding to each video frame information in the clustering result information to obtain at least one preset object image; according to the image information corresponding to each of the at least one preset object image, selecting the cluster identifier corresponding to the clustering result information from the at least one preset object image to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0127] According to one or more embodiments of the present disclosure, the method further includes: determining a preset object corresponding to the cluster identifier in response to a trigger operation on the cluster identifier; and displaying graphic and text information including the cluster identifier and the preset object in the material search box.

[0128] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an information display device, comprising:

[0129] an acquisition module, configured to, in response to search information input into a material search box on a video editing interface, acquire first search result information corresponding to the search information from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information;

[0130] A display module is used to display the first search result information in the video editing interface; wherein the first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

[0131] According to one or more embodiments of the present disclosure, the device also includes: a storage module; the storage module is used to obtain the video file information of the video file, the audio file of the video file and multiple video frame information of the video file in response to the import operation of the video file, wherein the video frame information at least includes a video frame image; for each video frame information, perform text recognition processing on the video frame image in the video frame information to obtain the video frame text corresponding to the video frame information, establish a correspondence between the video frame text and the video frame information, and store the correspondence between the video frame text and the video frame information in a preset database; and perform speech recognition processing on the audio file to obtain the audio text corresponding to each of the multiple video frame information of the video file, establish a correspondence between the audio text and the video frame information, and store the correspondence between the audio text and the video frame information in the preset database; and establish a correspondence between the video file information and the video frame information, and store the correspondence between the video file information and the video frame information in the preset database.

[0132] According to one or more embodiments of the present disclosure, the video frame information also includes first time information corresponding to the video frame image; accordingly, the storage module performs speech recognition processing on the audio file to obtain audio texts corresponding to multiple video frame information of the video file, specifically including: performing speech recognition processing on the audio file to obtain multiple sub-text information and second time information corresponding to each sub-text information; for each video frame information, determining the sub-text information corresponding to the video frame information based on the first time information in the video frame information and the second time information corresponding to each sub-text information, to obtain audio texts corresponding to multiple video frame information of the video file.

[0133] According to one or more embodiments of the present disclosure, the acquisition module obtains the first search result information corresponding to the search information from a preset database, specifically including: determining the target video frame information corresponding to each of the at least one target video frame texts including the search information from the correspondence between the video frame text and the video frame information included in the preset database; and, for each target video frame information, determining the video file information corresponding to the target video frame information from the correspondence between the video file information and the video frame information included in the preset database; and, for each target video frame information, determining the audio text corresponding to the target video frame information from the correspondence between the audio text and the video frame information included in the preset database.

[0134] According to one or more embodiments of the present disclosure, the display module is also used to display at least one cluster identifier in a preset area of ​​the video editing interface in response to a preset operation on the material search box in the video editing interface; and to display second search result information corresponding to the target cluster identifier in the video editing interface in response to a trigger operation on the target cluster identifier among the at least one cluster identifier; wherein the second search result information includes video file information and / or at least one video frame information, and the second search result information is determined based on the correspondence between the pre-stored cluster identifier, the video file information and the at least one video frame information.

[0135] According to one or more embodiments of the present disclosure, the storage module is also used to obtain video file information corresponding to the video file in response to an import operation of the video file, and to determine multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier; for each cluster identifier, establish a correspondence between the cluster identifier, the video file information and at least one video frame information, and store the correspondence in a preset database.

[0136] According to one or more embodiments of the present disclosure, the storage module determines multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier, specifically including: extracting multiple video frame information from the video file; for each video frame information, if image recognition is performed on the video frame information through a preset recognition model to obtain image information of a preset object image and an image feature vector of the preset object image, then determining that the video frame information is video frame information to be classified; dividing the multiple video frame information to be classified according to the image information of the preset object image and the image feature vector of the preset object image corresponding to each of the multiple video frame information to be classified to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0137] According to one or more embodiments of the present disclosure, the storage module divides the multiple video frame information to be classified according to the image information of the preset object images corresponding to each of the multiple video frame information to be classified and the image feature vectors of the preset object images, and obtains multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier, specifically including: clustering the multiple video frame information to be classified according to the image feature vectors of the preset object images corresponding to each of the multiple video frame information to be classified through a preset clustering model to obtain multiple clustering result information, wherein the clustering result information includes at least one video frame information whose vector similarity is greater than a preset threshold; for each clustering result information, determining the preset object image corresponding to each video frame information in the clustering result information to obtain at least one preset object image; according to the image information corresponding to each of the at least one preset object image, selecting the cluster identifier corresponding to the clustering result information from the at least one preset object image to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

[0138] According to one or more embodiments of the present disclosure, the display module is further used to determine the preset object corresponding to the cluster identifier in response to a trigger operation on the cluster identifier; and display graphic information including the cluster identifier and the preset object in the material search box.

[0139] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0140] The memory stores computer-executable instructions;

[0141] The processor executes the computer-executable instructions stored in the memory to implement the information display method described in the first aspect and various possible designs of the first aspect.

[0142] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the information display method described in the first aspect and various possible designs of the first aspect is implemented.

[0143] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the information display method described in the first aspect and various possible designs of the first aspect.

[0144] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0145] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0146] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An information display method, characterized in that: include: In response to search information input into a material search box on a video editing interface, obtaining first search result information corresponding to the search information from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information; Displaying the first search result information in the video editing interface; The first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

2. The method according to claim 1, characterized in that Before obtaining first search result information corresponding to the search information from the preset database in response to the search information input into the material search box of the video editing interface, the method further includes: In response to an import operation of a video file, obtaining video file information of the video file, an audio file of the video file, and a plurality of video frame information of the video file, wherein the video frame information includes at least a video frame image; For each video frame information, text recognition processing is performed on the video frame image in the video frame information to obtain the video frame text corresponding to the video frame information, a correspondence between the video frame text and the video frame information is established, and the correspondence between the video frame text and the video frame information is stored in a preset database; and, speech recognition processing is performed on the audio file to obtain the audio text corresponding to each of the multiple video frame information of the video file, a correspondence between the audio text and the video frame information is established, and the correspondence between the audio text and the video frame information is stored in the preset database; and, a correspondence between the video file information and the video frame information is established, and the correspondence between the video file information and the video frame information is stored in a preset database.

3. The method according to claim 2, characterized in that The video frame information also includes first time information corresponding to the video frame image; Accordingly, the audio file is subjected to speech recognition processing to obtain audio texts corresponding to the plurality of video frame information of the video file, including: Performing speech recognition processing on the audio file to obtain a plurality of sub-text information and second time information corresponding to each sub-text information; For each video frame information, the subtext information corresponding to the video frame information is determined according to the first time information in the video frame information and the second time information corresponding to each subtext information, and the audio text corresponding to each of the multiple video frame information of the video file is obtained.

4. The method according to claim 1, wherein Acquiring first search result information corresponding to the search information from a preset database includes: From the correspondence between the video frame text and the video frame information included in the preset database, determine the target video frame information corresponding to each of the at least one target video frame texts including the search information; and, for each target video frame information, determine the video file information corresponding to the target video frame information from the correspondence between the video file information and the video frame information included in the preset database; and, for each target video frame information, determine the audio text corresponding to the target video frame information from the correspondence between the audio text and the video frame information included in the preset database.

5. The method according to claim 1, wherein Also includes: In response to a preset operation on a material search box in a video editing interface, displaying at least one cluster identifier in a preset area of ​​the video editing interface; In response to a triggering operation on a target cluster identifier among the at least one cluster identifier, displaying second search result information corresponding to the target cluster identifier in the video editing interface; The second search result information includes video file information and / or at least one video frame information, and the second search result information is determined based on a pre-stored correspondence between a cluster identifier, video file information and at least one video frame information.

6. The method according to claim 5, characterized in that In response to a preset operation on a material search box in a video editing interface, before displaying at least one cluster identifier in a preset area of ​​the video editing interface, the method further includes: In response to an import operation of a video file, obtaining video file information corresponding to the video file, and determining a plurality of cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier; For each cluster identifier, a corresponding relationship among the cluster identifier, the video file information and at least one video frame information is established, and the corresponding relationship is stored in a preset database.

7. The method according to claim 6, characterized in that The determining of the plurality of cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier includes: extracting a plurality of video frame information from the video file; For each video frame information, if image recognition is performed on the video frame information using a preset recognition model to obtain image information of a preset object image and an image feature vector of the preset object image, then the video frame information is determined to be video frame information to be classified; According to the image information of the preset object image and the image feature vector of the preset object image corresponding to each of the multiple video frame information to be classified, the multiple video frame information to be classified are divided to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

8. The method according to claim 7, characterized in that The method further comprises dividing the plurality of video frames to be classified according to the image information of the preset object images and the image feature vectors of the preset object images corresponding to the plurality of video frames to be classified, and obtaining a plurality of cluster identifiers corresponding to the video files and at least one video frame information corresponding to each cluster identifier, including: Clustering the plurality of video frames to be classified using a preset clustering model according to image feature vectors of preset object images corresponding to the plurality of video frames to be classified, to obtain a plurality of clustering result information, wherein the clustering result information includes at least one video frame information having a vector similarity greater than a preset threshold; For each clustering result information, determining a preset object image corresponding to each video frame information in the clustering result information to obtain at least one preset object image; According to the image information corresponding to each of the at least one preset object images, a cluster identifier corresponding to the clustering result information is selected from the at least one preset object image to obtain multiple cluster identifiers corresponding to the video file and at least one video frame information corresponding to each cluster identifier.

9. The method according to claim 5, characterized in that Also includes: In response to a triggering operation on the cluster identifier, determining a preset object corresponding to the cluster identifier; Graphic and text information including the cluster identifier and the preset object is displayed in the material search box.

10. An information display device, characterized in that: include: an acquisition module, configured to, in response to search information input into a material search box on a video editing interface, acquire first search result information corresponding to the search information from a preset database, wherein the preset database includes a correspondence between video frame text and video frame information, a correspondence between audio text and video frame information, and a correspondence between video file information and video frame information; A display module, configured to display the first search result information in the video editing interface; The first search result information includes: target video frame information corresponding to at least one target video frame text of the search information, video file information corresponding to the target video frame information, and audio text corresponding to the target video frame information.

11. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the information display method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the information display method according to any one of claims 1 to 9 is implemented.

13. A computer program product, characterized in that The invention comprises a computer program, which implements the information display method according to any one of claims 1 to 9 when executed by a processor.