Multimedia resource search method, device, equipment and medium
By constructing an inverted index table and using the first keyword and classification tag for searching, the problem of low efficiency and low accuracy in searching multiple different types of text fragments in the prior art is solved, and efficient and accurate multimedia resource search is achieved.
Patent Information
- Application Number
- CN202210855628.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-07-20
AI Technical Summary
In the prior art, finding a variety of different types of text fragments is inefficient and low accuracy.
By extracting text content from multiple different types of text, obtaining text fragments and storing them in a preset database, word segmentation obtains the first keyword, constructing an inverted index table for word search, storing the classification tag of the text fragments into the inverted index table, receiving user query requests, searching for text fragments corresponding to the first keyword associated with the inverted index table and the second keyword, and then outputting them to the user after rating and sorting.
It realizes searching under a unified indexing architecture for multiple different types of text content, reducing the cost and search time of building multiple content libraries, improving the accuracy and efficiency of searches, and reducing user manual operations.
Smart Images

Figure CN115203445B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimedia resource search method, device, equipment and medium. Background Art
[0002] With the rapid development of the Internet, multimedia resource search has become an important topic. Usually, multimedia resource search will build multiple content libraries for different types of text (web page text, PDF text, image text, video text). Later, when the user enters a keyword to search for content, the background searches multiple content libraries separately and returns all different types of text associated with the keyword to the user. The user needs to switch back and forth between types in the display interface, and the user needs to spend time to identify the text fragments he wants in each text, which not only wastes the user's time, but also may result in low accuracy of the found text fragments due to the user's manual operation. Summary of the invention
[0003] In view of the above, the present invention provides a multimedia resource search method, device, equipment and medium, which aims to solve the technical problems of low efficiency and low accuracy in searching for multiple different types of text fragments in the prior art.
[0004] To achieve the above object, the present invention provides a multimedia resource search method, the method comprising:
[0005] Extract text content from multiple different types of texts to obtain one or more text segments and store them in a preset database, and segment each text segment to obtain the first keyword of each text segment;
[0006] constructing an inverted index table for word search according to the first keyword, and storing the classification labels of the text segments in the inverted index table to construct a multimedia library;
[0007] Receive a query request sent by a user terminal, extract a second keyword from the query request, search the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and read a corresponding text segment from the preset database according to the retrieved classification label;
[0008] The similarities between the multiple text segments are scored, the obtained score values are sorted according to a preset sorting order, and a preset number of text segments are selected according to the sorting order, rendered into corresponding texts and output to the user terminal.
[0009] Preferably, the multiple different types of texts include web page texts, PDF texts, image texts, and video texts. The steps of extracting text content from the multiple different types of texts to obtain one or more text segments and storing them in a preset database include:
[0010] Each type of text is divided into a format part and a text content part, and the text content part is segmented to obtain one or more text segments and store them in a preset database.
[0011] Preferably, the step of segmenting each text segment to obtain the first keyword of each text segment includes:
[0012] According to a preset word segmentation algorithm, the long text sentence of each text fragment is divided into multiple phrases;
[0013] The similarity values between adjacent phrases are calculated, and the phrases with similarity values less than a preset threshold are used as the first keywords.
[0014] Preferably, after constructing an inverted index table for word search according to the first keyword, the method further comprises:
[0015] Counting the frequency of occurrence of the first keyword in the corresponding text segment;
[0016] Compare the word frequency value with a preset word frequency value, and if the word frequency value is greater than or equal to the preset word frequency value, fill the first keyword into the high-frequency word queue in the inverted index table;
[0017] If the word frequency value is less than the preset word frequency value, the first keyword is added to the low-frequency word queue in the inverted index table.
[0018] Preferably, before storing the classification labels of the text segments in the inverted index table to construct the multimedia library, the method further comprises:
[0019] Read the text sequence of the first keyword of each text segment, input the text sequence into a preset classification model for tag embedding, and obtain word vector features;
[0020] According to the word vector feature, the classification label of the text segment is matched from the label module of the preset classification model, and a mapping relationship is established between the classification label and the first keyword of the text segment.
[0021] Preferably, extracting the second keyword from the query request includes:
[0022] Segment the information in the query request to obtain multiple segmented words;
[0023] A dictionary tree is generated according to a pre-constructed dictionary word list, and the multiple word segments are input into the dictionary tree for traversal to obtain the second keyword.
[0024] Preferably, searching the multimedia library for a classification label of a text segment corresponding to the first keyword associated with the second keyword according to the inverted index table and the second keyword, and reading the corresponding text segment from the preset database according to the retrieved classification label includes:
[0025] Inputting the second keyword into a search engine of the inverted index table;
[0026] Traversing the first keyword in the inverted index table according to the search engine to obtain a first keyword associated with the second keyword;
[0027] The classification label of the associated first keyword is read according to the mapping relationship, and the corresponding text segment is read from the preset database according to the retrieved classification label.
[0028] To achieve the above object, the present invention further provides a multimedia resource search device, the device comprising:
[0029] Extraction module: used to extract text content from multiple different types of texts, obtain one or more text fragments and store them in a preset database, and segment each text fragment to obtain the first keyword of each text fragment;
[0030] Storage module: used for constructing an inverted index table for word search according to the first keyword, and storing the classification labels of each text segment in the inverted index table to construct a multimedia library;
[0031] Query module: used for receiving a query request sent by a user terminal, extracting a second keyword from the query request, searching the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and reading a corresponding text segment from the preset database according to the retrieved classification label;
[0032] Output module: used to score the similarities between the multiple text fragments, sort the obtained score values according to a preset sorting order, select a preset number of text fragments according to the sorting order, render them into corresponding texts and output them to the user end.
[0033] To achieve the above object, the present invention further provides an electronic device, the electronic device comprising:
[0034] at least one processor; and,
[0035] a memory communicatively connected to the at least one processor; wherein,
[0036] The memory stores a program executable by the at least one processor, and the program is executed by the at least one processor so that the at least one processor can execute the multimedia resource search method according to any one of claims 1 to 7.
[0037] To achieve the above-mentioned purpose, the present invention also provides a computer-readable medium, wherein the computer-readable medium stores multimedia resources, and when the multimedia resources are executed by a processor, the steps of the multimedia resource search method as described in any one of claims 1 to 7 are implemented.
[0038] The present invention extracts the first keywords and text fragments of various different types of texts, constructs an inverted index table for word search based on all the first keywords, stores the classification labels of all text fragments in the inverted index table to construct a multimedia library, and realizes searching the contents of various different types of texts under a unified index architecture, thereby reducing the cost of constructing multiple content libraries and the search time.
[0039] According to the inverted index table and the second keyword of the user's query, the multimedia library is searched to obtain multiple text fragments of the first keyword related to the second keyword. After the similarities of the multiple text fragments are scored and sorted, the previously sorted text fragments are selected and rendered into corresponding texts and output to the user end, thereby realizing the use of text fragments as search results and mixed display of multiple different types of texts in the display interface, reducing user manual operations and improving the accuracy and efficiency of search. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A flowchart diagram of a preferred embodiment of the multimedia resource search method of the present invention;
[0041] Figure 2 A schematic diagram of modules of a preferred embodiment of a multimedia resource search device of the present invention;
[0042] Figure 3 A schematic diagram of a preferred embodiment of the electronic device of the present invention;
[0043] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0045] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0046] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] The present invention provides a multimedia resource search method. Figure 1 FIG. 1 is a flow chart of a method of searching for multimedia resources according to an embodiment of the present invention. The method can be executed by an electronic device, and the electronic device can be implemented by software and / or hardware. The method of searching for multimedia resources includes the following steps S10-S40:
[0048] Step S10: extracting text content from a plurality of different types of texts respectively to obtain one or more text segments and storing them in a preset database, and performing word segmentation on each text segment to obtain the first keyword of each text segment.
[0049] In this embodiment, the various types of texts include but are not limited to web page texts, PDF texts, image texts, and video texts. The methods for extracting text content from various types of texts are different. The length of the extracted text content is relatively long. The text content is divided into at least one text segment according to punctuation marks (periods, exclamation marks, semicolons) or paragraphs in the text content. A text segment refers to a pause in the text caused by transitions, emphasis, pauses, etc. when expressing the ideological content of the text, which is usually called a "natural paragraph". Each text segment is segmented, and the words representing the important meanings and semantics of the text segment are used as the first keyword of the text segment. The first keyword is one of the main search index methods of the multimedia library, and is also the specific name of the company's products, services, etc. that users want to know.
[0050] In one embodiment, the multiple different types of texts include web page texts, PDF texts, image texts, and video texts. The steps of extracting text content from the multiple different types of texts to obtain one or more text segments and storing them in a preset database include:
[0051] Each type of text is divided into a format part and a text content part, and the text content part is segmented to obtain one or more text segments and store them in a preset database.
[0052] In one embodiment, the formats of the multiple different types of text include: using the HTML code of the web page text, the coordinate information of the text content of the PDF text, the coordinate information of the text content of the image text, and the starting time period of the video text on the playback timeline as the text format.
[0053] Dividing the format part and text content part of various types of text is the basic condition for searching text fragments in the text. It is also a prerequisite for displaying various different types of texts in a mixed format on the user interface and reducing the time users need to spend on identifying the text fragments they want in each text.
[0054] Each type of text is divided into a format part and a text content part, including:
[0055] Web page text: Separate the HTML code part and text content part of the web page text. For example, when you open a web page text, click the "Show Web Page Source Code" button with the right mouse button, the current web page text will display the HTML code and text content together, for example, " <title> Come and see: the final complete form of China's space station! | Manned spacecraft | Astronauts | Shenzhou | Rockets_NetEase Subscription< / title> ......"Read HTML code (format)" <title>< / title> ” and the text content “Come and see: the final complete form of China’s space station! | Manned spacecraft | Astronauts | Shenzhou | Rocket_NetEase Subscription” are separated, and the text content is divided into a set containing at least one text fragment according to the title and paragraph of the text content, and the collection of HTML codes (formats) and text fragments are respectively stored in the preset database.
[0056] PDF text: The text content part and the coordinate information part of the text content are extracted through the OCR (text recognition) algorithm, and the text content is divided into a set containing at least one text fragment according to the title and paragraph of the text content, and the coordinate information of the text content and the set of text fragments are stored in the preset database respectively. The coordinate information of the text content refers to the coordinate information of a line of text, and the coordinate information includes the coordinate information of the x-axis and y-axis of the vertices of the rectangular box of the line of text, as well as four elements such as the length and width of the rectangular box. OCR uses the preset text recognition model to train and determine which area in the PDF text may contain text, and then perform text recognition on the area. For example, if there is a PDF text, the text recognition model first generates some candidate rectangular boxes, determines the possibility of these boxes containing text, and then recognizes the text in the box.
[0057] Image text: The text content part and the coordinate information part of the image text are extracted through the OCR (character recognition) algorithm, and the text content is divided into a set containing at least one text fragment according to the title and paragraph of the text content, and the coordinate information of the text content and the set of text fragments are stored in the preset database respectively.
[0058] Video text: The subtitles and voice in the video text are identified and extracted through the ASR (automatic speech recognition) algorithm to obtain the text content part. According to the similarity of the subtitle keywords and / or the pauses of the voice, the text content is divided into a set containing at least one text segment. The starting time period of each text segment on the playback timeline is read and stored in the preset database together with the set of text segments.
[0059] In one embodiment, segmenting each text segment to obtain the first keyword of each text segment includes:
[0060] According to a preset word segmentation algorithm, the long text sentence of each text fragment is divided into multiple phrases;
[0061] The similarity values between adjacent phrases are calculated, and the phrases with similarity values less than a preset threshold are used as the first keywords.
[0062] The preset word segmentation algorithms include but are not limited to the greedy algorithm and the stuttering algorithm. The long text sentences of each text segment are divided to obtain a word sequence vector, wherein the word sequence vector includes multiple phrases obtained by segmenting the text segment, and the similarity values between adjacent phrases are calculated. It is determined whether the similarity value is less than a preset threshold (for example, the preset threshold is 1), and the phrase less than the preset threshold is used as the first keyword.
[0063] The present invention extracts the first keyword of each text fragment, because the first keyword represents the central theme and core idea of each text fragment. The corresponding text fragment can be screened out through the first keyword. If the correlation of the extracted first keyword is greater, the search efficiency and accuracy will be accelerated.
[0064] Step S20: constructing an inverted index table for word search according to the first keyword, and storing the classification labels of the text segments in the inverted index table to construct a multimedia library.
[0065] In this embodiment, the inverted index table is used to record a list of first keywords contained in the text fragments. The classification labels of all text fragments are stored in the queue of the first keywords corresponding to the inverted index table to construct a multimedia library. The multimedia library uses the inverted index table to search the contents of multiple different types of texts under a unified index architecture, reducing the cost and search time of building multiple content libraries.
[0066] In a collection of text fragments, there will be many text fragments containing the same first keyword. Each text fragment records the information of each first keyword in the inverted index table with a document number (Doc ID) (for example, the arrangement number of the first keyword in the inverted index table, the number of times it is shared), and also records the number of times the first keyword appears in this text fragment (IDF) and the positions where the first keyword appears in the text fragment. The information related to a text fragment is used as an inverted index item (Posting), and a series of inverted index items containing all the first keywords form the structure of the inverted index table.
[0067] The core of the inverted index table contains the contents of two parts (the word dictionary and the inverted list):
[0068] 1. Dictionary word list: records all the first keywords to form a list. The split granularity of the first keyword can be implemented according to specific needs. Dictionary word lists are generally large and can be implemented through B+ trees or hash lists to meet high-performance insertion and query and custom editing (for example, deletion, addition, and modification of the first keyword).
[0069] 2. Inverted index list: mainly records the relationship between the first keyword and the corresponding text fragment. The attributes in the relationship between them are called inverted index items. The inverted index items include the Doc ID of the text fragment, the word frequency (the word frequency refers to the number of times the first keyword appears in the text fragment, which can be used to calculate the relevance), and the position of the first keyword in the text fragment (the position refers to the starting position and the ending position).
[0070] In one embodiment, after constructing an inverted index table for word search according to the first keyword, the method further includes:
[0071] Counting the frequency of occurrence of the first keyword in the corresponding text segment;
[0072] Compare the word frequency value with a preset word frequency value, and if the word frequency value is greater than or equal to the preset word frequency value, fill the first keyword into the high-frequency word queue in the inverted index table;
[0073] If the word frequency value is less than the preset word frequency value, the first keyword is added to the low-frequency word queue in the inverted index table.
[0074] Through programming models such as MapReduce, the frequency statistics of each first keyword obtained can be performed. According to the preset frequency value (for example, the preset frequency value is 3), the first keyword greater than or equal to the preset frequency value is used as a high-frequency word, and the first keyword less than the preset frequency value is used as a low-frequency word, so as to fill it into the queue of high-frequency words or low-frequency words in the inverted index table. The high-frequency word queue and the low-frequency word queue are respectively generated into respective inverted indexes. By generating respective inverted indexes, the accuracy and speed of searching for the first keyword can be improved, and the resources of the search engine can also be reduced. For example, the high-frequency word queue is generated as a reverse index and the low-frequency word queue is generated as a forward index, or the high-frequency word queue is generated as a forward index and the low-frequency word queue is generated as a reverse index. It is also possible to generate the high-frequency word queue and the low-frequency word queue as a reverse index at the same time. The setting is based on the actual business scenario and is not limited here.
[0075] In one embodiment, before storing the classification labels of the text segments in the inverted index table to construct the multimedia library, the method further includes:
[0076] Read the text sequence of the first keyword of each text segment, input the text sequence into a preset classification model for tag embedding, and obtain word vector features;
[0077] According to the word vector feature, the classification label of the text segment is matched from the label module of the preset classification model, and a mapping relationship is established between the classification label and the first keyword of the text segment.
[0078] The preset classification model refers to a sample set of text fragments containing different keywords that are collected and manually annotated, and the sample set is trained through a preset model (BERT modeling) to obtain a classification model.
[0079] For example, the text sequence of each first keyword of text segment A is read and input into the preset classification model for tag embedding. The text sequence is represented by a matrix through the encoder, and the word vector features of each first keyword are output. The word vector features are matched for similarity through the feature representation fusion layer and the fully connected layer, so that the label module of the classification model outputs the label with the greatest similarity to the word vector feature as the classification label of text segment A, and a mapping relationship is established between the classification label and each first keyword of text segment A. By establishing a mapping relationship between the classification label and the first keyword of the text segment, when searching, the corresponding text segment can be found through the classification label as long as the first keyword is determined. There is no need to search the text segment for any keywords. The text segment only needs to be stored in the preset database, which improves the running speed of the inverted index table.
[0080] Step S30: Receive a query request sent by the user terminal, extract a second keyword from the query request, search the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword based on the inverted index table and the second keyword, and read the corresponding text segment from the preset database based on the retrieved classification label.
[0081] In this embodiment, the user inputs the content of the query request on the interface of the search engine (search engine of the inverted index table) of the multimedia library on the user side, and after clicking the "search" button, the search engine program processes the content, such as Chinese-specific word segmentation, removing stop words, judging whether it is necessary to start integrated search, judging whether there are spelling errors or typos, etc. The query request can be analyzed and the second keyword can be extracted. After obtaining the second keyword, the search engine program starts matching work, finds all the first keywords with the same or close semantics as the second keyword from the inverted index table, and then searches to obtain the classification label associated with the first keyword according to the mapping relationship, and reads multiple text fragments of the first keyword from the preset database according to the classification label.
[0082] In one embodiment, extracting the second keyword from the query request includes:
[0083] Segment the information in the query request to obtain multiple segmented words;
[0084] A dictionary tree is generated according to a pre-constructed dictionary word list, and the multiple word segments are input into the dictionary tree for traversal to obtain the second keyword.
[0085] Based on a preset word segmentation algorithm (for example, the textrank word segmentation algorithm), the content of the query request is segmented, relevant words in the query request are extracted and stop words are removed, a word association matrix is constructed based on the relevant words, spelling errors or typos in the content of the query request are corrected, the importance value of each word is obtained through the word segmentation algorithm formula, and a preset number of words ranked first are selected as word segmentations according to the order of importance values from large to small.
[0086] Based on the pre-recorded first keywords, the word list of the first keyword and the word maximum prefix are matched to obtain a dictionary word list, and the key value and character string of each first keyword in the dictionary word list are used as nodes to generate a tree-structured dictionary tree. According to the pre-statistics of the word frequency of the segmented words in the historical query requests of all users, the character prefix features of multiple segmented words are read, and the traversal matching starts along the root node of the dictionary tree. The word with the same character prefix feature of the node of the dictionary tree and the character prefix feature of the segmented word is used as the second keyword. According to the dictionary word list, the content of the query request is matched with the user's historical search behavior to obtain the second keyword, which solves the technical problems of typos, grammatical errors, and unclear expressions in the content input by the user.
[0087] In one embodiment, searching the multimedia library for a classification label of a text segment corresponding to the first keyword associated with the second keyword according to the inverted index table and the second keyword, and reading the corresponding text segment from the preset database according to the retrieved classification label includes:
[0088] Inputting the second keyword into a search engine of the inverted index table;
[0089] Traversing the first keyword in the inverted index table according to the search engine to obtain a first keyword associated with the second keyword;
[0090] The classification label of the associated first keyword is read according to the mapping relationship, and the corresponding text segment is read from the preset database according to the retrieved classification label.
[0091] Input the second keyword into the search engine of the inverted index table, according to the different characteristics of the high-frequency word queue and the low-frequency word queue of the inverted index table in data reading, for example, the search engine uses the reverse index method to traverse the high-frequency word queue of the inverted index table, and uses the forward index method to traverse the low-frequency word queue of the inverted index table, and obtains the first keyword associated with the second keyword. The associated first keyword refers to the first keyword with the same or similar semantics as the second keyword. As for which index method to use, it is set according to the actual business scenario and is not limited here. According to the mapping relationship established in step S20, the classification label of the associated first keyword is read to obtain multiple text fragments of the first keyword. The use of different indexing methods can solve the technical problems that only a single index method in the prior art occupies more physical space of the search engine, and when adding, deleting and modifying the data in the inverted index table, the index must also be dynamically maintained, which reduces the speed of data maintenance, effectively saving the occupied physical space and improving the convenience of data maintenance.
[0092] Step S40: scoring the similarities between the multiple text segments, sorting the obtained score values according to a preset sorting order, selecting a preset number of text segments according to the sorting order, rendering them into corresponding texts and outputting them to the user terminal.
[0093] In this embodiment, after obtaining multiple text fragments of the first keyword, these text fragments may include various types of text fragments such as web page text, PDF text, image text, video text, etc., similarity calculation is performed on these text fragments, and the similarity of these text fragments is scored according to a preset scoring algorithm. The obtained scoring values are sorted according to a preset sorting order (for example, the scoring values are sorted in order from high to low), and a preset number (for example, 10) of text fragments with high rankings are selected according to the sorting order, and the formats of these 10 text fragments are read from a preset database to render them into corresponding texts and output them to the user end.
[0094] For example, if the first keywords are "Shenzhou" and "rocket", and the selected text segment is a web page text, all text segments related to the two keywords are returned based on the first keywords "Shenzhou" and "rocket". After calculation, the top 10 text segments are selected and the corresponding HTML codes are obtained. The code is rendered into the original web page text on the user side and displayed to the user. If the selected text segment is a PDF text or an image text, the corresponding text segment and coordinate information are read for rendering. If the selected text segment is a video text, the corresponding text segment and the starting time period of the playback timeline are read for rendering.
[0095] The preset scoring algorithm includes:
[0096]
[0097] in, is the score of any text fragment, is the number of the first keyword of the text segment, The first The first keyword, The first The IDF value of the first keyword, The second keyword for the user's query, The first public word, is the similarity between the text segment and the second keyword of the user query.
[0098] By scoring the similarities between text fragments and selecting the text fragments and the formats of the text fragments with higher scores in the user's query request for rendering, the text fragments that the user wants to view can be automatically and quickly obtained without the user having to spend time switching between types in the display interface, nor does the user have to spend time identifying the text fragments they want from each text. This reduces the time for obtaining text and the time for rendering, and achieves the effect of using text fragments as search results and displaying multiple different types of texts in a mixed manner in the display interface.
[0099] Reference Figure 2 FIG. 1 is a schematic diagram of functional modules of the multimedia resource search device 100 of the present invention.
[0100] The multimedia resource search device 100 of the present invention can be installed in an electronic device. According to the functions to be implemented, the multimedia resource search device 100 may include an extraction module 110, an extraction module 20, a query module 130 and an output module 140. The module described in the present invention may also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0101] In this embodiment, the functions of each module / unit are as follows:
[0102] Extraction module 110: used to extract text content from multiple different types of texts, obtain one or more text segments and store them in a preset database, and perform word segmentation on each text segment to obtain the first keyword of each text segment;
[0103] Storage module 120: used to construct an inverted index table for word search according to the first keyword, and store the classification label of each text segment in the inverted index table to construct a multimedia library;
[0104] Query module 130: used to receive a query request sent by a user terminal, extract a second keyword from the query request, search the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and read a corresponding text segment from the preset database according to the retrieved classification label;
[0105] Output module 140: used to score the similarities between the multiple text segments, sort the obtained score values according to a preset sorting order, select a preset number of text segments according to the sorting order, render them into corresponding texts and output them to the user terminal.
[0106] In one embodiment, the multiple different types of texts include web page texts, PDF texts, image texts, and video texts. The steps of extracting text content from the multiple different types of texts to obtain one or more text segments and storing them in a preset database include:
[0107] Each type of text is divided into a format part and a text content part, and the text content part is segmented to obtain one or more text segments and store them in a preset database.
[0108] In one embodiment, segmenting each text segment to obtain the first keyword of each text segment includes:
[0109] According to a preset word segmentation algorithm, the long text sentence of each text fragment is divided into multiple phrases;
[0110] The similarity values between adjacent phrases are calculated, and the phrases with similarity values less than a preset threshold are used as the first keywords.
[0111] In one embodiment, after constructing an inverted index table for word search according to the first keyword, the method further includes:
[0112] Counting the frequency of occurrence of the first keyword in the corresponding text segment;
[0113] Compare the word frequency value with a preset word frequency value, and if the word frequency value is greater than or equal to the preset word frequency value, fill the first keyword into the high-frequency word queue in the inverted index table;
[0114] If the word frequency value is less than the preset word frequency value, the first keyword is added to the low-frequency word queue in the inverted index table.
[0115] In one embodiment, before storing the classification labels of the text segments in the inverted index table to construct the multimedia library, the method further includes:
[0116] Read the text sequence of the first keyword of each text segment, input the text sequence into a preset classification model for tag embedding, and obtain word vector features;
[0117] According to the word vector feature, the classification label of the text segment is matched from the label module of the preset classification model, and a mapping relationship is established between the classification label and the first keyword of the text segment.
[0118] In one embodiment, extracting the second keyword from the query request includes:
[0119] Segment the information in the query request to obtain multiple segmented words;
[0120] A dictionary tree is generated according to a pre-constructed dictionary word list, and the multiple word segments are input into the dictionary tree for traversal to obtain the second keyword.
[0121] In one embodiment, searching the multimedia library for a classification label of a text segment corresponding to the first keyword associated with the second keyword according to the inverted index table and the second keyword, and reading the corresponding text segment from the preset database according to the retrieved classification label includes:
[0122] Inputting the second keyword into a search engine of the inverted index table;
[0123] Traversing the first keyword in the inverted index table according to the search engine to obtain a first keyword associated with the second keyword;
[0124] The classification label of the associated first keyword is read according to the mapping relationship, and the corresponding text segment is read from the preset database according to the retrieved classification label.
[0125] Reference Figure 3 FIG. 1 is a schematic diagram of a preferred embodiment of an electronic device 1 of the present invention.
[0126] The electronic device 1 includes but is not limited to: a memory 11, a processor 12, a display 13 and a network interface 14. The electronic device 1 is connected to a network through the network interface 14 to obtain raw data. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobilecommunication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, or a call network.
[0127] Among them, the memory 11 includes at least one type of readable medium, and the readable medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a hard disk or memory of the electronic device 1. In other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped with the electronic device 1. Of course, the memory 11 can also include both the internal storage unit of the electronic device 1 and its external storage device. In this embodiment, the memory 11 is generally used to store the operating system and various application software installed on the electronic device 1, such as the program code of the multimedia resource search 10, etc. In addition, the memory 11 can also be used to temporarily store various types of data that have been output or are to be output.
[0128] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 12 is generally used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication. In this embodiment, the processor 12 is used to run the program code stored in the memory 11 or process data, such as running the program code of the multimedia resource search 10.
[0129] The display 13 may be referred to as a display screen or a display unit. In some embodiments, the display 13 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an organic light-emitting diode (OLED) touch device, etc. The display 13 is used to display information processed in the electronic device 1 and to display a visual working interface, such as displaying the results of data statistics.
[0130] The network interface 14 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface). The network interface 14 is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0131] Figure 3Only the electronic device 1 having components 11 - 14 and the multimedia resource search 10 is shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0132] Optionally, the electronic device 1 may further include a user interface, which may include a display (Display), an input unit such as a keyboard (Key board), and the optional user interface may also include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an organic light-emitting diode (Organic Light-Emitting Diode, OLED) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.
[0133] The electronic device 1 may further include a radio frequency (RF) circuit, a sensor, an audio circuit, etc., which will not be described in detail here.
[0134] In the above embodiment, the processor 12 may implement the following steps when executing the multimedia resource search 10 stored in the memory 11:
[0135] Extract text content from multiple different types of texts to obtain one or more text segments and store them in a preset database, and segment each text segment to obtain the first keyword of each text segment;
[0136] constructing an inverted index table for word search according to the first keyword, and storing the classification labels of the text segments in the inverted index table to construct a multimedia library;
[0137] Receive a query request sent by a user terminal, extract a second keyword from the query request, search the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and read a corresponding text segment from the preset database according to the retrieved classification label;
[0138] The similarities between the multiple text segments are scored, the obtained score values are sorted according to a preset sorting order, and a preset number of text segments are selected according to the sorting order, rendered into corresponding texts and output to the user terminal.
[0139] The storage device may be the memory 11 of the electronic device 1 , or may be another storage device that is communicatively connected to the electronic device 1 .
[0140] For a detailed description of the above steps, please refer to the above Figure 2 Functional module diagram of the embodiment of the multimedia resource search device 100 and Figure 1 Description of a flowchart of an embodiment of a multimedia resource search method.
[0141] In addition, an embodiment of the present invention further proposes a computer-readable medium, which may be non-volatile or volatile. The computer-readable medium may be any one or any combination of a hard disk, a multimedia card, an SD card, a flash memory card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, etc. The computer-readable medium includes a data storage area and a program storage area, the data storage area stores data created according to the use of the blockchain node, and the program storage area stores multimedia resources 10. When the multimedia resource search 10 is executed by the processor, the following operations are implemented:
[0142] Extract text content from multiple different types of texts to obtain one or more text segments and store them in a preset database, and segment each text segment to obtain the first keyword of each text segment;
[0143] constructing an inverted index table for word search according to the first keyword, and storing the classification labels of the text segments in the inverted index table to construct a multimedia library;
[0144] Receive a query request sent by a user terminal, extract a second keyword from the query request, search the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and read a corresponding text segment from the preset database according to the retrieved classification label;
[0145] The similarities between the multiple text segments are scored, the obtained score values are sorted according to a preset sorting order, and a preset number of text segments are selected according to the sorting order, rendered into corresponding texts and output to the user terminal.
[0146] The specific implementation of the computer-readable medium of the present invention is substantially the same as the specific implementation of the multimedia resource search method described above, and will not be described in detail herein.
[0147] In another embodiment, the multimedia resource search method provided by the present invention can further ensure the privacy and security of all the above data, and all the above data can also be stored in a blockchain node. For example, the first keyword and the second keyword can be stored in the blockchain node.
[0148] It should be noted that the blockchain referred to in the present invention is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, platform product service layer, and application service layer.
[0149] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that the process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or the method also includes elements inherent to such process, device, article or method. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, device, article or method including the element.
[0150] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product is stored in a medium (such as ROM / RAM, magnetic disk, optical disk) as described above, including a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, an electronic device, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0151] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A multimedia resource search method, characterized in that: The method comprises: Extract text content from multiple different types of texts to obtain one or more text segments and store them in a preset database, and segment each text segment to obtain the first keyword of each text segment; constructing an inverted index table for word search according to the first keyword, and storing the classification labels of the text segments in the inverted index table to construct a multimedia library; Receive a query request sent by a user terminal, extract a second keyword from the query request, search the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and read a corresponding text segment from the preset database according to the retrieved classification label; Scoring the similarities between the multiple text segments, sorting the obtained score values according to a preset sorting order, selecting a preset number of text segments according to the sorting order, rendering them into corresponding texts and outputting them to the user terminal; Before storing the classification labels of the text segments in the inverted index table to construct the multimedia library, the method further includes: reading a text sequence of the first keyword of each text segment, inputting the text sequence into a preset classification model for tag embedding, and obtaining a word vector feature; matching the classification label of the text segment from the label module of the preset classification model according to the word vector feature, and establishing a mapping relationship between the classification label and the first keyword of the text segment; The extracting the second keyword from the query request includes: segmenting the information of the query request to obtain a plurality of segmented words; generating a dictionary tree according to a pre-constructed dictionary word list, inputting the plurality of segmented words into the dictionary tree for traversal, and obtaining the second keyword; The method of searching the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and reading a corresponding text segment from the preset database according to the retrieved classification label, comprises: inputting the second keyword into a search engine of the inverted index table; traversing the first keyword in the inverted index table according to the search engine to obtain a first keyword associated with the second keyword; reading the classification label of the associated first keyword according to a mapping relationship, and reading a corresponding text segment from the preset database according to the retrieved classification label.
2. The multimedia resource search method according to claim 1, characterized in that: The multiple different types of texts include web page texts, PDF texts, image texts, and video texts. The steps of extracting text content from the multiple different types of texts to obtain one or more text segments and storing them in a preset database include: Each type of text is divided into a format part and a text content part, and the text content part is segmented to obtain one or more text segments and store them in a preset database.
3. The multimedia resource search method according to claim 1, characterized in that: The step of segmenting each text segment to obtain the first keyword of each text segment includes: According to a preset word segmentation algorithm, the long text sentence of each text fragment is divided into multiple phrases; The similarity values between adjacent phrases are calculated, and the phrases with similarity values less than a preset threshold are used as the first keywords.
4. The multimedia resource search method according to claim 1, characterized in that: After constructing an inverted index table for word search according to the first keyword, the method further includes: Counting the frequency of occurrence of the first keyword in the corresponding text segment; Compare the word frequency value with a preset word frequency value, and if the word frequency value is greater than or equal to the preset word frequency value, fill the first keyword into the high-frequency word queue in the inverted index table; If the word frequency value is less than the preset word frequency value, the first keyword is added to the low-frequency word queue in the inverted index table.
5. A multimedia resource search device, used to implement the multimedia resource search method according to any one of claims 1 to 4, characterized in that: The device comprises: Extraction module: used to extract text content from multiple different types of texts, obtain one or more text fragments and store them in a preset database, and segment each text fragment to obtain the first keyword of each text fragment; Storage module: used for constructing an inverted index table for word search according to the first keyword, and storing the classification labels of each text segment in the inverted index table to construct a multimedia library; Query module: used for receiving a query request sent by a user terminal, extracting a second keyword from the query request, searching the multimedia library for a classification label of a text segment corresponding to a first keyword associated with the second keyword according to the inverted index table and the second keyword, and reading a corresponding text segment from the preset database according to the retrieved classification label; Output module: used to score the similarities between the multiple text fragments, sort the obtained score values according to a preset sorting order, select a preset number of text fragments according to the sorting order, render them into corresponding texts and output them to the user end.
6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a program executable by the at least one processor, and the program is executed by the at least one processor so that the at least one processor can execute the multimedia resource search method according to any one of claims 1 to 4.
7. A computer-readable medium, characterized in that The computer-readable medium stores multimedia resources, and when the multimedia resources are executed by a processor, the multimedia resource search method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Method and device for searching multimedia file
CN102867042A
Multimedia conceptual search system and associated search method
US20070130112A1