Multimedia search method and device, electronic equipment and storage medium

By employing a dual search method using keywords and keyframe data in multimedia, the problem of pre-set tags failing to accurately match user interests is solved, resulting in higher search accuracy and a better user experience.

CN115658938BActive Publication Date: 2026-05-08BEIJING XINTANG SICHUANG EDUCATIONAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XINTANG SICHUANG EDUCATIONAL TECH CO LTD
Filing Date
2022-11-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

As video length and content increase, preset tags may fail to accurately match user interests, leading to decreased accuracy in multimedia searches and a poor user experience.

Method used

A dual search is performed using keyword data and keyframe data from multimedia frames. First and second search results are obtained from a multimedia database and a preset search engine, respectively, and the two are compared to determine the target search result.

Benefits of technology

It improved the accuracy of search recommendations on multimedia platforms and enhanced the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115658938B_ABST
    Figure CN115658938B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multimedia search method and device, electronic equipment and storage medium, and belongs to the field of Internet search. The method comprises the following steps: determining a multimedia frame to be searched in a target multimedia, and determining keyword data and key frame data of the multimedia frame; searching a first search result in a multimedia database based on the keyword data of the multimedia frame; searching a second search result through a preset search engine based on the key frame data of the multimedia frame; and comparing the first search result and the second search result to determine a target search result according to a comparison result. The present disclosure can improve the accuracy of search recommendation of a multimedia platform to a user and improve user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet search, and more particularly to a multimedia search method, apparatus, electronic device, and storage medium. Background Technology

[0002] When users browse internet content, internet platforms can recommend search results to them, thereby improving their search efficiency.

[0003] With the rise of short video platforms, these platforms pre-load tags onto each video before uploading it. For example, a sports video might be pre-loaded with keywords such as "sports," "celebrity," and "event." When users watch the video or read the text content, if they click to search for a tag, the platform can then perform a search based on that tag.

[0004] However, as video and content lengths continue to increase, pre-set tags may no longer accurately match user interests, thus necessitating a new search method. Summary of the Invention

[0005] In view of this, the present disclosure provides a multimedia search method, apparatus, electronic device, and storage medium, which can improve the accuracy of multimedia platforms in making search recommendations to users and enhance user experience.

[0006] According to one aspect of this disclosure, a multimedia search method is provided, the method comprising:

[0007] Identify the multimedia frames to be searched in the target multimedia, and determine the keyword data and keyframe data of the multimedia frames;

[0008] Based on the keyword data of the multimedia frame, a first search result is obtained by searching the multimedia database;

[0009] Based on the keyframe data of the multimedia frame, a second search result is obtained by searching through a preset search engine;

[0010] Compare the first search result and the second search result, and determine the target search result based on the comparison result.

[0011] According to another aspect of this disclosure, a multimedia search device is provided, the device comprising:

[0012] The first determining module is used to determine the multimedia frame to be searched in the target multimedia, and to determine the keyword data and keyframe data of the multimedia frame.

[0013] The search module is used to search the multimedia database based on the keyword data of the multimedia frame to obtain a first search result; and to search the keyframe data of the multimedia frame through a preset search engine to obtain a second search result.

[0014] The second determining module is used to compare the first search result and the second search result, and determine the target search result based on the comparison result.

[0015] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0016] Processor; and

[0017] Stored program memory,

[0018] The program includes instructions that, when executed by the processor, cause the processor to perform the multimedia search method.

[0019] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided that stores computer instructions, wherein the computer instructions are used to cause a computer to perform the multimedia search method described above.

[0020] In this disclosure, after identifying the multimedia frame to be searched within the target multimedia, a first search result can be obtained based on keyword data of the multimedia frame, and a second search result can be obtained based on keyframe data of the multimedia frame. The first and second search results are then compared to determine the target search result. By employing both text and image searches, the accuracy of search recommendations made by the multimedia platform to users can be improved, thus enhancing the user experience. Attached Figure Description

[0021] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0022] Figure 1 A flowchart of a multimedia search method provided according to an exemplary embodiment of the present disclosure is shown;

[0023] Figure 2 A schematic block diagram of a multimedia search device provided according to an exemplary embodiment of the present disclosure is shown;

[0024] Figure 3 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0027] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0030] This disclosure provides a multimedia search method that can perform both text and image searches, further improving search accuracy. This method can be performed by a terminal, server, and / or other devices with processing capabilities. The method provided in the embodiments of this disclosure can be performed by any of the aforementioned devices, or by multiple devices working together; this disclosure does not limit this.

[0031] The following will refer to Figure 1 The flowchart of the multimedia search method is shown, and the method is described below. The method includes the following steps 101-104.

[0032] Step 101: Identify the multimedia frames to be searched in the target multimedia, and determine the keyword data and keyframe data of the multimedia frames.

[0033] Multimedia can refer to a combination of various media, generally including text, sound, and images. In this embodiment, multimedia includes at least sound and images, such as video.

[0034] Optionally, the multimedia frame to be searched can be determined within the target multimedia after playback has ended.

[0035] In one possible implementation, a user can watch a video on a terminal (such as a smartphone, computer, etc.) and can switch out of the video when playback is complete or during playback. When a user's switching operation is detected, the currently playing video frame can be used as the multimedia frame to be searched.

[0036] Of course, the multimedia frame to be searched can be arbitrarily selected by the user. For example, when watching video content of interest, the user can pause the video playback and click the search option, thus using the currently playing video frame as the multimedia frame to be searched. This embodiment does not limit the specific form of selecting the multimedia frame.

[0037] The process of determining keyword data and keyframe data for multimedia frames will be described below.

[0038] Optionally, the method for determining keyword data in a multimedia frame can be as follows:

[0039] Obtain the time point of the multimedia frame in the target multimedia;

[0040] Based on the time nodes and the pre-stored mapping relationship between playback time and keywords, determine at least one sound keyword and at least one image keyword;

[0041] The keyword data of a multimedia frame is obtained by combining at least one sound keyword and at least one image keyword.

[0042] The mapping relationship between playback time and keywords stores keywords for multiple playback time periods. The keywords for each playback time period can be extracted based on the audio and image data of that playback time period in the target multimedia.

[0043] Optionally, the mapping relationship between playback time and keywords can be constructed as follows:

[0044] The target multimedia is analyzed to generate first text information corresponding to the target multimedia. The first text information includes first text content corresponding to the sound data and second text content corresponding to the image data.

[0045] The target multimedia is divided into multiple multimedia segments, and each multimedia segment is analyzed to generate second text information corresponding to each multimedia segment. The second text information includes third text content corresponding to the sound data and fourth text content corresponding to the image data.

[0046] For each multimedia segment, the first and third text contents corresponding to the sound data are combined to generate the text description information of the sound of the multimedia segment; the second and fourth text contents corresponding to the image data are combined to generate the text description information of the image of the multimedia segment.

[0047] Extract at least one audio keyword from the text description information of the audio in each multimedia segment, and extract at least one image keyword from the text description information of the image in each multimedia segment. Construct and store the mapping relationship between the playback time period of the multimedia segment and its keywords.

[0048] In one possible implementation, the audio and image data of the entire video can be analyzed using the Python language to generate the overall text content of the video (i.e., the first text information). This text content can include text content corresponding to the audio data and text content corresponding to the image data, respectively.

[0049] Since the speaking speed of each video may not be uniform, to improve adaptability, the video can be segmented into multiple video segments, each with its corresponding playback time period. Then, Python can be used to analyze the audio and / or image data of each video segment to generate text content (i.e., second text information) for each segment. Similarly, this text content can include text corresponding to the audio data and text corresponding to the image data separately.

[0050] For each video segment, the text content of the entire video and the text content of that specific video segment can be combined to generate a text description for that segment. Optionally, the text description for each video segment can be modified by the user. Specifically, the text content corresponding to the audio data of the entire video and the video segment can be combined to generate a text description for the audio data of that video segment; similarly, the text content corresponding to the image data of the entire video and the video segment can be combined to generate a text description for the image data of that video segment.

[0051] Furthermore, the text description information of each video segment can be processed separately, such as word segmentation and deletion of interjections, to extract keywords. This embodiment does not limit the specific processing for keyword extraction. Specifically, keywords can be extracted from the text description information of the audio in a video segment to obtain the audio keywords of that video segment; keywords can be extracted from the text description information of the images in a video segment to obtain the image keywords of that video segment.

[0052] The playback time period of each video segment is mapped to its keywords, thereby obtaining the mapping relationship between the playback time of the video (i.e., the target multimedia) and the keywords, and stored in the memory-based persistent key-value database Redis.

[0053] Subsequently, whenever a video frame to be searched is determined, the time node of that video frame in the entire video can be obtained. Thus, in the above mapping relationship between playback time and keywords, the playback time period in which that time node falls can be determined and the corresponding keywords, including sound keywords and image keywords, can be obtained.

[0054] Furthermore, the acquired keywords can be processed, such as combining sound keywords and image keywords, to obtain the keyword data for that video frame.

[0055] Optionally, to further improve search accuracy, keywords can be extracted from the context of video frames for searching. The corresponding processing can be as follows:

[0056] The first playback time range is determined by moving the time node backward by a first preset time length and moving the time node forward by a first preset time length.

[0057] In the pre-stored mapping relationship between playback time and keywords, at least one sound keyword and at least one image keyword corresponding to the first playback time range are determined.

[0058] In one possible implementation, the current time node ± N seconds (i.e., the first preset time length, where N is greater than 0) can be used as the first playback time range. In the above mapping relationship between playback time and keywords, the playback time period that intersects with the first playback time range can be determined and the corresponding keywords can be obtained. In this way, a corresponding keyword matrix can be constructed, and search recommendations can be made based on the keyword matrix.

[0059] Optionally, the selected keyword data can be recommended to users based on current keyword popularity, using a popularity-weighted approach. Furthermore, when recommending keywords to users, current trending keywords can be retrieved from a trending keyword database, and the selected keyword data and trending keywords can be recommended to the user.

[0060] The process of constructing keyframe data for multimedia frames will be described below.

[0061] Video image data can be composed of multiple keyframes in the order of playback. Whenever a video frame to be searched is determined, the keyframe corresponding to that video frame can be obtained so that it can be used as the object of image search.

[0062] Optionally, to further improve the accuracy of the search, keyframes in the context of the video frame can be extracted for the search. The corresponding processing can be as follows:

[0063] Obtain the time point of the multimedia frame in the target multimedia;

[0064] The second playback time range is determined by moving the time node backward by a second preset time length and moving the time node forward by a second preset time length.

[0065] In the image data of the target multimedia, multiple keyframes within the second playback time range are acquired;

[0066] These multiple keyframes are combined to obtain the keyframe data of the multimedia frame.

[0067] In one possible implementation, the current time node ± M seconds (i.e., the second preset time length, M is greater than 0) can be used as the second playback time range, and multiple keyframes within the second playback time range can be obtained in the video to construct keyframe data for searching.

[0068] Optionally, after determining the keyword data or keyframe data, a search can be recommended to the user. For example, the user can be shown the determined keywords or keyframes and select one or more keywords to search in step 102, or select one or more keyframes to search in step 103. This disclosure performs dual searches on text and images. Therefore, if the user selects keywords but not keyframes, a search can be performed based on the user-selected keywords and the determined keyframes; if the user selects keyframes but not keywords, a search can be performed based on the user-selected keyframes and the determined keywords.

[0069] Of course, you can also choose not to show users specific keywords or keyframes, and instead perform the search based on the keyword data and keyframe data mentioned above.

[0070] Step 102: Based on the keyword data of the multimedia frame, search the multimedia database to obtain the first search result.

[0071] In one possible implementation, videos corresponding to keywords can be searched for in the database of the video platform based on keyword data.

[0072] Step 103: Based on the keyframe data of the multimedia frame, a second search result is obtained by searching through a preset search engine.

[0073] In one possible implementation, images containing keyframes or content with high similarity to keyframes can be searched for on the Internet using a preset search engine based on keyframe data. This embodiment does not limit the specific search engine.

[0074] Optionally, the processing in step 103 above can be as follows:

[0075] The keyframe data of the multimedia frame is searched through a preset search engine to obtain multiple image results;

[0076] Access the source of each image result and obtain the corresponding multimedia content for each image result;

[0077] A second search result is constructed based on the multimedia content corresponding to each image result.

[0078] In one possible implementation, the keyframe image is searched using a preset search engine. The search results are also images, carrying source information such as the URL. The source of the image result is accessed, and its text, video, or other content is crawled to construct a second search result that can be compared with the first search result. For example, if an image similar to the keyframe image is found on the internet, the URL of that image can be looked up, and the text and other content within that URL can be crawled to construct a corresponding second search result.

[0079] Step 104: Compare the first search result and the second search result, and determine the target search result based on the comparison results.

[0080] In one possible implementation, the search result with higher confidence can be selected by comparing the first search result and the second search result. For example, the search result that exists in both has higher confidence and can be used as one of the target search results.

[0081] Optionally, the processing in step 104 above can be as follows:

[0082] Compare the first and second search results to determine the overlap ratio;

[0083] If the overlap ratio is greater than or equal to the preset overlap threshold, then the first search result and the second search result will be used as the target search result.

[0084] If the overlap ratio is less than the preset overlap threshold, the first search result will be used as the target search result.

[0085] Taking a preset overlap threshold of 80% as an example, if the overlap ratio between the first and second search results is determined to be greater than or equal to 80%, meaning that greater than or equal to 80% of the search results contain the same content, then the first search result based on keywords and the second search result based on keyframes are considered to correspond to the same content, with the same confidence level. Therefore, the first and second search results can be used as target search results, and both can be displayed to the user simultaneously. Optionally, a mixed layout can be used to display the first and second search results, sorted by popularity or relevance. If the popularity or relevance is the same, the first search result is prioritized. Simultaneously, for identical search results, only one is retained.

[0086] If the overlap between the first and second search results is less than 80%, meaning there's a significant discrepancy between the keyword-based first search result and the keyframe-based second search result, then the keyword-based first search result is considered to have higher confidence and can be used as the target search result for displaying to the user. Optionally, a toggle option can be added to the search results display page, allowing users to switch the displayed search result from the first to the second. In cases where there may be mismatches between text and images on the internet, users might be more interested in the image portion. The toggle option allows users to change the content of the search results, quickly displaying the keyframe-based second search result.

[0087] In this embodiment, after identifying the multimedia frame to be searched within the target multimedia content, a first search result can be obtained based on the keyword data of the multimedia frame, and a second search result can be obtained based on the keyframe data of the multimedia frame. The first and second search results are then compared to determine the target search result. By employing both text and image searches, the accuracy of the multimedia platform's search recommendations to users can be improved, thus enhancing the user experience.

[0088] This disclosure provides a multimedia search device for implementing the multimedia search method described above. Figure 2 The schematic block diagram shown indicates that the multimedia search device 200 includes: a first determining module 201, a search module 202, and a second determining module 203.

[0089] The first determining module 201 is used to determine the multimedia frame to be searched in the target multimedia, and to determine the keyword data and keyframe data of the multimedia frame.

[0090] The search module 202 is used to search for a first search result in a multimedia database based on the keyword data of the multimedia frame; and to search for a second search result through a preset search engine based on the keyframe data of the multimedia frame.

[0091] The second determining module 203 is used to compare the first search result and the second search result, and determine the target search result based on the comparison result.

[0092] Optionally, the first determining module 201 is used to:

[0093] Obtain the time node of the multimedia frame in the target multimedia;

[0094] Based on the time nodes and the pre-stored mapping relationship between playback time and keywords, at least one sound keyword and at least one image keyword are determined;

[0095] The keyword data of the multimedia frame is obtained by combining the at least one sound keyword and the at least one image keyword.

[0096] Optionally, the first determining module 201 is used to:

[0097] The first playback time range is determined by moving the time node backward by a first preset time length and moving the time node forward by the first preset time length.

[0098] In the pre-stored mapping relationship between playback time and keywords, at least one sound keyword and at least one image keyword corresponding to the first playback time range are determined.

[0099] Optionally, the first determining module 201 is further configured to:

[0100] The target multimedia is analyzed to generate first text information corresponding to the target multimedia. The first text information includes first text content corresponding to the sound data and second text content corresponding to the image data.

[0101] The target multimedia is divided into multiple multimedia segments, and each multimedia segment is analyzed to generate second text information corresponding to each multimedia segment. The second text information includes third text content corresponding to the sound data and fourth text content corresponding to the image data.

[0102] For each multimedia segment, the first and third text contents corresponding to the sound data are combined to generate text description information of the sound of the multimedia segment; the second and fourth text contents corresponding to the image data are combined to generate text description information of the image of the multimedia segment.

[0103] Extract at least one audio keyword from the text description information of the audio in each multimedia segment, and extract at least one image keyword from the text description information of the image in each multimedia segment. Construct and store the mapping relationship between the playback time period of the multimedia segment and its keywords.

[0104] Optionally, the first determining module 201 is used to:

[0105] Obtain the time node of the multimedia frame in the target multimedia;

[0106] The second playback time range is determined by moving the time node backward by a second preset time length and moving the time node forward by a second preset time length.

[0107] In the image data of the target multimedia, multiple keyframes within the second playback time range are obtained;

[0108] The multiple keyframes are combined to obtain the keyframe data of the multimedia frame.

[0109] Optionally, the search module 202 is used for:

[0110] The keyframe data of the multimedia frame is searched through a preset search engine to obtain multiple image results;

[0111] Access the source of each image result and obtain the corresponding multimedia content for each image result;

[0112] A second search result is constructed based on the multimedia content corresponding to each image result.

[0113] Optionally, the second determining module 203 is used for:

[0114] Compare the first search result and the second search result to determine the overlap ratio;

[0115] If the overlap ratio is greater than or equal to a preset overlap threshold, then the first search result and the second search result are used as the target search result.

[0116] If the overlap ratio is less than a preset overlap threshold, then the first search result is taken as the target search result.

[0117] In this embodiment, after identifying the multimedia frame to be searched within the target multimedia content, a first search result can be obtained based on the keyword data of the multimedia frame, and a second search result can be obtained based on the keyframe data of the multimedia frame. The first and second search results are then compared to determine the target search result. By employing both text and image searches, the accuracy of the multimedia platform's search recommendations to users can be improved, thus enhancing the user experience.

[0118] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.

[0119] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.

[0120] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.

[0121] refer to Figure 3 The present invention describes a structural block diagram of an electronic device 300 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0122] like Figure 3 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of the device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0123] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, output unit 307, storage unit 308, and communication unit 309. Input unit 306 can be any type of device capable of inputting information to electronic device 300. Input unit 306 can receive input digital or text information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 307 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 308 may include, but is not limited to, disks and optical discs. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0124] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above. For example, in some embodiments, the multimedia search method described above can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. In some embodiments, the computing unit 301 can be configured to perform the multimedia search method described above by any other suitable means (e.g., by means of firmware).

[0125] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0126] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0127] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0129] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0130] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

Claims

1. A multimedia search method, characterized in that, The method includes: Identify the multimedia frames to be searched in the target multimedia, and determine the keyword data and keyframe data of the multimedia frames; Based on the keyword data of the multimedia frame, a first search result is obtained by searching the multimedia database; Based on the keyframe data of the multimedia frame, a second search result is obtained by searching through a preset search engine; Compare the first search result and the second search result, and determine the target search result based on the comparison result; The keyword data for determining the multimedia frame includes: Obtain the time node of the multimedia frame in the target multimedia; Based on the time nodes and the pre-stored mapping relationship between playback time and keywords, at least one sound keyword and at least one image keyword are determined; The keyword data of the multimedia frame is obtained by combining the at least one sound keyword and the at least one image keyword.

2. The method according to claim 1, characterized in that, The step of determining at least one sound keyword and at least one image keyword based on the time node and the pre-stored mapping relationship between playback time and keywords includes: The first playback time range is determined by moving the time node backward by a first preset time length and moving the time node forward by the first preset time length. In the pre-stored mapping relationship between playback time and keywords, at least one sound keyword and at least one image keyword corresponding to the first playback time range are determined.

3. The method according to claim 2, characterized in that, The mapping relationship between playback time and keywords is constructed based on the following method: The target multimedia is analyzed to generate first text information corresponding to the target multimedia. The first text information includes first text content corresponding to the sound data and second text content corresponding to the image data. The target multimedia is divided into multiple multimedia segments, and each multimedia segment is analyzed to generate second text information corresponding to each multimedia segment. The second text information includes third text content corresponding to the sound data and fourth text content corresponding to the image data. For each multimedia segment, the first and third text contents corresponding to the sound data are combined to generate text description information of the sound of the multimedia segment; the second and fourth text contents corresponding to the image data are combined to generate text description information of the image of the multimedia segment. Extract at least one audio keyword from the text description information of the audio in each multimedia segment, and extract at least one image keyword from the text description information of the image in each multimedia segment. Construct and store the mapping relationship between the playback time period of the multimedia segment and its keywords.

4. The method according to any one of claims 1-3, characterized in that, Determining the keyframe data of the multimedia frame includes: Obtain the time node of the multimedia frame in the target multimedia; The second playback time range is determined by moving the time node backward by a second preset time length and moving the time node forward by a second preset time length. In the image data of the target multimedia, multiple keyframes within the second playback time range are obtained; The multiple keyframes are combined to obtain the keyframe data of the multimedia frame.

5. The method according to any one of claims 1-3, characterized in that, The keyframe data based on the multimedia frame is used to obtain a second search result through a preset search engine, including: The keyframe data of the multimedia frame is searched through a preset search engine to obtain multiple image results; Access the source of each image result and obtain the corresponding multimedia content for each image result; A second search result is constructed based on the multimedia content corresponding to each image result.

6. The method according to any one of claims 1-3, characterized in that, The step of comparing the first search result and the second search result, and determining the target search result based on the comparison result, includes: Compare the first search result and the second search result to determine the overlap ratio; If the overlap ratio is greater than or equal to a preset overlap threshold, then the first search result and the second search result are used as the target search result; If the overlap ratio is less than a preset overlap threshold, then the first search result is taken as the target search result.

7. A multimedia search device, characterized in that, The device includes: The first determining module is used to determine the multimedia frame to be searched in the target multimedia, and to determine the keyword data and keyframe data of the multimedia frame; wherein, determining the keyword data of the multimedia frame includes: obtaining the time node of the multimedia frame in the target multimedia; determining at least one sound keyword and at least one image keyword according to the time node and the pre-stored mapping relationship between playback time and keywords; and combining the at least one sound keyword and at least one image keyword to obtain the keyword data of the multimedia frame; The search module is used to search the multimedia database based on the keyword data of the multimedia frame to obtain a first search result; and to search the keyframe data of the multimedia frame through a preset search engine to obtain a second search result. The second determining module is used to compare the first search result and the second search result, and determine the target search result based on the comparison result.

8. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Video content searching method and device

    CN110362714A

  • Multimedia data search method and device, computer equipment and medium

    CN113190695A

  • Data search method and system based on big data

    CN114357212A