Content search methods and related devices

By recognizing the user's intent to answer video questions and using multi-dimensional similarity calculations, the system obtains and filters video search results from preset video question-and-answer pairs, solving the problem of inaccurate search results in existing technologies and achieving more efficient video search result returns.

CN115269961BActive Publication Date: 2025-10-28TENCENT TECH (CHENGDU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210912133.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-10-28
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

Existing search engines fail to identify users' actual needs, resulting in inaccurate search results that cannot meet users' video Q&A requirements.

Method used

By recognizing the user's intent to answer video questions in the input content, preset video question-answer pairs are obtained, and video search results are recalled and filtered based on multi-dimensional similarity, including similarity calculations in the dimensions of optical character recognition, speech recognition, image recognition, video title and summary, to determine the most relevant video search results.

Benefits of technology

It improves the accuracy of search results, returning intuitive and concise video search results to meet users' video Q&A needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115269961B_ABST
    Figure CN115269961B_ABST
Patent Text Reader

Abstract

This application discloses a content search method and related equipment. Related embodiments can be applied to various scenarios such as cloud technology, artificial intelligence, and smart transportation. It can identify video question-and-answer intent information of target content. When video question-and-answer intent information is identified, at least one preset video question-and-answer pair is obtained. Based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, video search results are recalled. Based on the recall frequency information corresponding to the recalled video search results, the target video search result corresponding to the target content is determined from the recalled video search results. This application can identify video question-and-answer intent in the user-input search content. If video question-and-answer intent is present, more intuitive and concise video search results can be returned, which helps improve the accuracy of search results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a content search method and related equipment. Background Technology

[0002] With the development of internet technology, online information is growing rapidly, and the internet is filled with a large amount of redundant information. Users need to rely on search engines to find the information they need online. A search engine is a software system used on the internet. It uses certain strategies to collect and discover information online, processes it, and then provides users with information search services. Search engines typically provide a web interface that allows users to submit their search queries. The search application then retrieves search results that match the user's input and returns these results to the user.

[0003] However, in current related technologies, the search results returned by search applications are generally complex text search results, without identifying the specific needs based on the user's actual search content. This can easily lead to search results that are not satisfactory to the user, resulting in insufficient accuracy of the search results. Summary of the Invention

[0004] This application provides a content search method and related equipment. The related equipment may include a content search device, an electronic device, a computer-readable storage medium, and a computer program product, which can return more intuitive and concise video search results, thereby improving the accuracy of search results.

[0005] This application provides a content search method, including:

[0006] Obtain the target content to be searched and identify the video question-and-answer intent information of the target content;

[0007] When the target content is identified to contain video question-and-answer intent information, at least one preset video question-and-answer pair is obtained, the preset video question-and-answer pair including the search content and the video search results corresponding to the search content;

[0008] Based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, the video search results in the preset video question-and-answer pair are recalled. The at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension.

[0009] Based on the recall frequency information corresponding to the recalled video search results, the target video search results corresponding to the target content are determined from the recalled video search results.

[0010] Accordingly, embodiments of this application provide a content search device, including:

[0011] An intent recognition unit is used to acquire the target content to be searched and to identify the video question-and-answer intent information of the target content.

[0012] The acquisition unit is used to acquire at least one preset video question-and-answer pair when it is identified that the target content contains video question-and-answer intent information. The preset video question-and-answer pair includes the search content and the video search results corresponding to the search content.

[0013] The recall unit is used to recall video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the content information of the target content and the video search results in the preset video question-and-answer pair in at least one dimension. The at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension and summary dimension.

[0014] The determining unit is used to determine the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results.

[0015] Optionally, in some embodiments of this application, the intent recognition unit may include a feature extraction subunit and an intent recognition subunit, as follows:

[0016] The feature extraction subunit is used to extract temporal features from each text unit in the target content to obtain the content temporal feature information of the target content.

[0017] The intent recognition subunit is used to identify video question-and-answer intent information of the target content based on the content temporal feature information.

[0018] Optionally, in some embodiments of this application, the feature extraction subunit may be used to extract features from each text unit in the target content to obtain word-level feature information corresponding to each text unit; process the word-level feature information of each text unit based on the word-level feature information of the text units in the context corresponding to each text unit; and fuse the processed word-level feature information of each text unit to obtain the content temporal feature information of the target content.

[0019] Optionally, in some embodiments of this application, the recall unit may include an index graph acquisition subunit, a node search subunit, and a search result recall subunit, as follows:

[0020] The index graph acquisition subunit is used to acquire the content index graph corresponding to the search content in the preset video question and answer pair, and the content index graph corresponding to the content information of the video search results in the preset video question and answer pair under at least one dimension. The content index graph includes various index layers arranged from top to bottom with the number of nodes increasing sequentially. Each index layer includes at least one node, and the node content corresponding to each node is the content information of a search content or video search result under at least one dimension.

[0021] The node search subunit is used to perform node search on each index layer in the content index graph in a top-to-bottom order, based on the similarity between the target content and the node content corresponding to the node, for each content index graph, so as to find similar nodes corresponding to the target content in the nodes of the target index layer.

[0022] The search result recall subunit is used to recall video search results in the preset video question-and-answer pair based on the node content corresponding to the similar nodes, and obtain the recall result corresponding to the content index graph.

[0023] Optionally, in some embodiments of this application, the recall unit may include a first recall subunit, an extraction subunit, and a second recall subunit, as follows:

[0024] The first recall subunit is used to recall video search results in the preset video question-and-answer pair based on the similarity between the content feature vector of the target content and the content feature vector of the search content in the preset video question-and-answer pair, and obtain a first recall result.

[0025] Extraction subunits are used to vectorize the content information of video search results in the preset video question-and-answer pair in at least one dimension to obtain the content feature vector in the at least one dimension.

[0026] The second recall subunit is used to recall video search results in the preset video question-and-answer pair based on the similarity between the content feature vector of the target content and the content feature vector under the at least one dimension, and obtain the second recall result under the at least one dimension.

[0027] Optionally, in some embodiments of this application, the extraction subunit may specifically be used to vectorize the optical character recognition (OCR) text of the video search results in the preset video question-and-answer pair to obtain a content feature vector in the OCR dimension, wherein the OCR text is the content information of the video search results in the OCR dimension; to vectorize the speech recognition information of the video search results to obtain a content feature vector in the speech recognition dimension, wherein the speech recognition information is the content information of the video search results in the speech recognition dimension; to vectorize the video frame image sequence of the video search results to obtain a content feature vector in the image dimension, wherein the video frame image sequence is the content information of the video search results in the image dimension; to vectorize the video title of the video search results to obtain a content feature vector in the video title dimension, wherein the video title is the content information of the video search results in the video title dimension; and to perform summary extraction processing on the video search results based on the OCR text and the speech recognition information to obtain a content feature vector of the video search results in the summary dimension.

[0028] Optionally, in some embodiments of this application, the at least one dimension further includes cross-dimensionality; the extraction subunit can specifically be used to obtain the optical character recognition text, speech recognition information, and video frame image sequence in the preset video question-and-answer pair under the optical character recognition dimension, the speech recognition information under the speech recognition dimension, and the video frame image sequence under the image dimension of the video search results; and to perform feature vector interaction processing on the optical character recognition text, the speech recognition information, the video frame image sequence, and the search content in the preset video question-and-answer pair to obtain the content feature vector under the cross-dimensionality.

[0029] Optionally, in some embodiments of this application, the second recall result under at least one dimension includes the second recall result under each dimension; the determining unit may include a statistical subunit and a result determining subunit, as follows:

[0030] The statistical subunit is used to perform aggregate statistical processing on the first recall result and the second recall result under each dimension to obtain the recall frequency information corresponding to each recalled video search result.

[0031] The result determination subunit is used to determine the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information.

[0032] Optionally, in some embodiments of this application, the determining unit may include an acquiring subunit and a determining subunit, as follows:

[0033] The acquisition subunit is used to acquire quality information of the recalled video search results in at least one dimension.

[0034] A determining subunit is configured to determine the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results and the quality information in the at least one dimension.

[0035] Optionally, in some embodiments of this application, the determining subunit may be specifically used to fuse the recall frequency information and quality information in the at least one dimension corresponding to the recalled video search results to obtain fused feature information; based on the fused feature information, predict the probability that the recalled video search results meet the preset quality conditions; and based on the probability, determine the target video search results corresponding to the target content from the recalled video search results.

[0036] Optionally, in some embodiments of this application, the intent recognition unit may be specifically used to identify video question-and-answer intent information of the target content through an intent recognition model.

[0037] Optionally, in some embodiments of this application, the content search device may further include a training unit for training an intent recognition model. Specifically, the training unit may acquire training data, including sample content and the expected probability that the sample content contains video question-and-answer intent information. Using the intent recognition model, temporal features are extracted from each text unit in the sample content to obtain the content temporal feature information of the sample content. Based on the content temporal feature information, the actual probability that the sample content contains video question-and-answer intent information is predicted. Based on the expected probability and the actual probability, the parameters of the intent recognition model are adjusted to obtain the trained intent recognition model.

[0038] An electronic device provided in this application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the content search method provided in this application.

[0039] This application also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps in the content search method provided in this application.

[0040] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the content search method provided in embodiments of this application.

[0041] This application provides a content search method and related equipment, which can acquire the target content to be searched and identify video question-and-answer intent information of the target content; when video question-and-answer intent information is identified in the target content, at least one preset video question-and-answer pair is acquired, the preset video question-and-answer pair including search content and video search results corresponding to the search content; based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, the video search results in the preset video question-and-answer pair are recalled, the at least one dimension including optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension; according to the recall frequency information corresponding to the recalled video search results, the target video search results corresponding to the target content are determined from the recalled video search results. This application can identify video question-and-answer intent of the user-input search content, and if video question-and-answer intent exists, it can return more intuitive and concise video search results, which is beneficial to improving the accuracy of search results. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1a This is a schematic diagram of a scenario for the content search method provided in an embodiment of this application;

[0044] Figure 1b This is a flowchart of the content search method provided in the embodiments of this application;

[0045] Figure 1c This is a model structure diagram of the content search method provided in the embodiments of this application;

[0046] Figure 1d This is another model structure diagram of the content search method provided in the embodiments of this application;

[0047] Figure 1e This is another model structure diagram of the content search method provided in the embodiments of this application;

[0048] Figure 1f This is another flowchart of the content search method provided in the embodiments of this application;

[0049] Figure 1g This is a page illustration of the content search method provided in an embodiment of this application;

[0050] Figure 1h This is another page illustration of the content search method provided in the embodiments of this application;

[0051] Figure 2 This is another flowchart of the content search method provided in the embodiments of this application;

[0052] Figure 3 This is a schematic diagram of the structure of the content search device provided in the embodiments of this application;

[0053] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] This application provides a content search method and related equipment. The related equipment may include a content search device, an electronic device, a computer-readable storage medium, and a computer program product. Specifically, the content search device may be integrated into an electronic device, which may be a terminal or a server, etc.

[0056] It is understood that the content search method in this embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as limiting this application.

[0057] like Figure 1a As shown, the content search method is implemented jointly by a terminal and a server as an example. The content search system provided in this application includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected via a network, such as a wired or wireless network, etc., wherein the content search device can be integrated into the server.

[0058] Terminal 10 can be used to: acquire the target content currently to be searched in the target application, send the target content to server 11 to trigger the server to search for the target content, and obtain the target video search result corresponding to the target content; terminal 10 can also receive the target video search result sent by server 11 and display the target video search result on the corresponding search result page. Terminal 10 may include mobile phones, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, tablet computers, laptops, or personal computers (PCs), etc. A client can also be set on terminal 10, which can be an application client or a browser client, etc.

[0059] Server 11 can be used to: receive the target content to be searched sent by terminal 10, and identify video question-and-answer intent information of the target content; when video question-and-answer intent information is identified in the target content, obtain at least one preset video question-and-answer pair, the preset video question-and-answer pair including search content and video search results corresponding to the search content; recall the video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension; and determine the target video search result corresponding to the target content from the recalled video search results according to the recall frequency information corresponding to the recalled video search results. Server 11 can be a single server, a server cluster composed of multiple servers, or a cloud server. In the content search method or apparatus disclosed in this application, multiple servers can form a blockchain, and the server is a node on the blockchain.

[0060] The steps such as content search performed on the aforementioned server 11 can also be executed by the terminal 10.

[0061] The content search method provided in this application relates to computer vision technology, speech technology, and natural language processing in the field of artificial intelligence.

[0062] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. AI software technology mainly includes computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.

[0063] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes for target recognition and measurement, and further processes images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0064] Key technologies in speech technology include automatic speech recognition, speech synthesis, and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech technology emerging as one of the most promising methods.

[0065] Natural Language Processing (NLP) is an important area within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close connection with linguistic research. NLP technologies typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0066] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0067] This embodiment will be described from the perspective of a content search device, which can be integrated into an electronic device, such as a server or a terminal.

[0068] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0069] The content search method of this application embodiment can be applied to scenarios such as browser search. This embodiment can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0070] like Figure 1b As shown, the specific process of this content search method can be as follows:

[0071] 101. Obtain the target content to be searched and identify the video question-and-answer intent information of the target content.

[0072] The target content refers to the content to be searched, and its type is not limited. For example, the target content can be text, audio, or an image. Specifically, the target content is what the user is querying, and can be represented by a query. Specifically, if the target content is audio, it can be converted into text through speech recognition before the content search is performed.

[0073] Specifically, it can obtain the target content to be searched in the target application. The target application can be a content search platform, and users can search for information through the search entry and built-in search engine provided by the target application. For example, the target application can be a browser.

[0074] In a specific scenario, the target application is a browser. The content currently entered by the user in the browser's search input box can be considered the target content. When the user performs a search operation on the target content in the search input box, the target content can be used as the current search term, and the corresponding search results can be obtained through a search engine. In some embodiments, this search operation can be a press-the-Enter operation performed on the search input box; in other embodiments, the search operation can also be a trigger operation on the search control in the browser's corresponding application page. This trigger operation can be a click operation or a swipe operation. In response to the trigger operation on the search control, the browser can use the content in the current search input box as the target content to be searched, perform relevant searches based on the target content, and thus return the search results to the user.

[0075] In current technologies, search applications typically return complex text results without identifying the user's actual search needs. This can lead to search results that are not satisfactory to the user, resulting in insufficient accuracy.

[0076] The content search method provided in this application can identify the video question-and-answer intent of the user's input search content. If the video question-and-answer intent exists, it can return more intuitive and concise video search results, which helps to improve the accuracy of search results.

[0077] Optionally, in this embodiment, the step of "identifying the video question-and-answer intent information of the target content" may include:

[0078] Temporal features are extracted from each text unit in the target content to obtain the content temporal feature information of the target content;

[0079] Based on the temporal feature information of the content, the intent information of video question answering is identified in the target content.

[0080] The process involves first segmenting the target content into words to obtain individual text units, and then extracting temporal features from each text unit. A text unit can be a word or a character; this embodiment does not impose any restrictions on this.

[0081] Specifically, based on the temporal features of the content, it can be predicted whether the target content contains video question-and-answer intent information. This can be achieved using a classifier. The classifier can be a Support Vector Machine (SVM), a Recurrent Neural Network (RNN), a Fully Connected Deep Neural Network (DNN), etc., and this embodiment does not impose any limitations on this.

[0082] If a search result contains video question-and-answer intent information, it means that the user needs to obtain answers for that search result, that is, the user has a question-and-answer requirement and needs search results for the video type of that search result.

[0083] Optionally, in this embodiment, the step "extracting temporal features from each text unit in the target content to obtain the content temporal feature information of the target content" may include:

[0084] Feature extraction is performed on each text unit in the target content to obtain word-level feature information corresponding to each text unit;

[0085] Based on the word-level feature information of each text unit corresponding to its context, the word-level feature information of each text unit is processed.

[0086] The word-level feature information of each processed text unit is fused to obtain the content temporal feature information of the target content.

[0087] The word-level feature information of a text unit can specifically be the word vector of the text unit, or the feature information obtained by fusing the content vector, type vector and position vector of the text unit.

[0088] Specifically, the content vector corresponding to a text unit can be the word vector of the text unit, the type vector can represent the information type to which the text unit belongs, and the position vector can represent the position of the text unit in the target content, which can be the beginning of a sentence, the end of a sentence, etc.

[0089] There are various ways to fuse content vectors, type vectors, and position vectors, and this embodiment does not limit this. For example, the fusion method can be concatenation, and the concatenation order is not limited. For instance, it can be concatenated in the order of content vector, type vector, and position vector, or in the reverse order, that is, in the order of position vector, type vector, and content vector. The fusion method can also be weighted fusion, where the weights of the content vector, type vector, and position vector are first determined, and then fused according to the weights.

[0090] Specifically, the text units in the context corresponding to a text unit can be other text units in the target content besides the text unit itself. This embodiment can fuse the word-level feature information of the text units in the various contexts corresponding to a text unit to obtain the context feature information corresponding to that text unit, and then process the word-level feature information of that text unit based on the context feature information. There are various ways to fuse the processed word-level feature information of the various text units, such as weighted summation, and this embodiment does not limit this approach.

[0091] Optionally, in this embodiment, the step of "identifying the video question-and-answer intent information of the target content" may include:

[0092] The intent recognition model is used to identify the intent information of the target content in video question-and-answer sessions.

[0093] The intent recognition model can be a temporal model. This temporal model can include Long Short-Term Memory (LSTM) networks, Bidirectional Encoder Representations from Transformers (BERT), and so on.

[0094] LSTM is a type of recurrent neural network, specifically a recurrent neural network (RNN). LSTM is well-suited for extracting semantic features from time-series data and is frequently used in natural language processing tasks to extract semantic features from contextual information. LSTM uses a three-gate structure (input gate, forget gate, and output gate) to selectively forget some historical data, add some current input data, and ultimately integrate it into the current state to generate the output state.

[0095] BERT is an open-source temporal model based on the Transformer architecture. BERT consists of multiple layers of bidirectional Transformers, typically 12 or 24 layers. BERT can be obtained through pre-training and fine-tuning. BERT training mainly includes two tasks: first, randomly removing words from the training corpus and replacing them with masks, then having the model predict the removed words; second, each training data point consists of a sentence and its preceding and following sentences, where some sentences are truly related, while others are unrelated, requiring the model to determine the relationship between the sentences. The model is optimized based on the loss values ​​from these two tasks. BERT's training process can fully utilize contextual information, giving the model stronger expressive power. After pre-training, the model can be fine-tuned for specific tasks. Fine-tuning is a common transfer learning technique in deep learning, allowing the model to be better adapted to language knowledge in specific scenarios.

[0096] It should be noted that the intent recognition model can be trained by other devices and then provided to the content search device, or it can be trained by the content search device itself.

[0097] If the content search device is trained independently, the content search method may further include the following steps before the step "identifying video question-and-answer intent information of the target content through the intent recognition model":

[0098] Acquire training data, which includes sample content and the expected probability that the sample content contains video question-answering intent information;

[0099] By using an intent recognition model, temporal features are extracted from each text unit in the sample content to obtain the content temporal feature information of the sample content.

[0100] Based on the temporal feature information of the content, predict the actual probability that the sample content contains video question-and-answer intent information;

[0101] Based on the expected probability and the actual probability, the parameters of the intent recognition model are adjusted to obtain the trained intent recognition model.

[0102] If the expected probability is 1, it indicates that the sample content contains video question-answering intent information; if the expected probability is 0, it indicates that the sample content does not contain video question-answering intent information. Here, "sample content" refers to the sample search content.

[0103] The training process can involve first calculating the actual probability that the sample content contains video question-and-answer intent information. Then, using a backpropagation algorithm, the parameters of the intent recognition model are adjusted. Based on the actual and expected probabilities of the sample content containing video question-and-answer intent information, the parameters of the intent recognition model are optimized so that the actual probability of the sample content containing video question-and-answer intent information approaches the expected probability, resulting in a trained intent recognition model. Specifically, the loss value between the calculated actual probability and the expected probability can be made less than a preset value, which can be set according to the actual situation.

[0104] In one specific embodiment, a video intent operator can also be provided through a video intent model, which can identify whether the user has a video requirement for the query, i.e., a need for video-type search results; and a question-and-answer intent operator can be provided through a question-and-answer intent model, which can identify whether the user has a question-and-answer intent for the query. By combining the video intent operator and the question-and-answer intent operator, it can be determined whether the target content (query) to be searched contains video question-and-answer intent information.

[0105] Among them, the video intent model and the question-answering intent model can be time-series models, which can be LSTM models or BERT models.

[0106] For example, a video intent model can employ a BERT-based binary classification model and be trained using manually labeled training data. The model structure diagram for the video intent model can be found by referring to... Figure 1c The left side; where '[CLS]' can be regarded as a sequence of position labels, Tok1, ..., TokN-1, TokN, TokN+1, and TokM represent the text units in the query content (specifically, the sample search content), and '[SEP]' is the separator for each text unit. The BERT model can extract features from each text unit in the query content based on the CLS marker, generating a set of feature vectors T1, ..., T... N-1 T N T N+1 …T M The model is then fine-tuned using a fully connected layer, which can be a CRF model. CRF stands for Conditional Random Fields. A CRF model can be viewed as a task-related layer of a BERT model; specifically, this task-related layer can be used for video intent prediction tasks.

[0107] Table 1

[0108]

[0109] The training data can carry labels indicating whether there is a video requirement. Specifically, the search content of a sample can be used to determine whether it has a video requirement based on the content type and keywords it contains. The content type can include news events, practical life skills, and operation tutorials, as shown in Table 1.

[0110] During the training of the video intent model, the parameters of the video intent model can be adjusted based on the label of the search content for each sample (i.e., the expected probability of the existence of video intent) and the predicted actual probability of the existence of video intent, so as to obtain the trained video intent model.

[0111] It should be noted that the question-answering intent model can be similar to the video intent model, and the question-answering intent model will not be elaborated on here.

[0112] 102. When it is identified that the target content contains video question-and-answer intent information, at least one preset video question-and-answer pair is obtained, wherein the preset video question-and-answer pair includes the search content and the video search results corresponding to the search content.

[0113] In one embodiment, at least one preset video question-and-answer pair can be obtained from a video question-and-answer library. The video question-and-answer library can store a certain number of video question-and-answer pairs. Each preset video question-and-answer pair includes a search content and a video search result corresponding to the search content. The search content is the question content query, and the video search result is the answer corresponding to the question content. The video search result is specifically a video type search result.

[0114] Specifically, the video question-and-answer database can also store relevant information for each preset video question-and-answer pair. This relevant information can be divided into basic information, extended information, and quality information. Basic information may include video identification information (id, Identity document), video URL (Universal Resource Locator, also known as web address), video title, video cover, video author, video duration, and video release date. Extended information may include video OCR information, video ASR information, answer summary, and answer frame-by-frame summary.

[0115] Specifically, the video OCR information (i.e., optical character recognition text) can be text information obtained by performing optical character recognition (OCR) on the video frame corresponding to the video search result. More specifically, the video OCR information can be the video subtitle text information corresponding to the video search result. In some embodiments, video frames corresponding to the video search result can be extracted, perhaps every 5 frames. Subtitle extraction processing is then performed on the extracted video frames, and the extracted subtitles are deduplicated. The deduplicated subtitles can be used as video OCR information. Additionally, the start and end time points corresponding to the video OCR information in the video can be obtained, i.e., the start and end time points of the corresponding subtitles appearing in the video. These start and end time points can be used as positioning points during video playback, allowing users to locate the video segment corresponding to the video OCR information.

[0116] The video ASR information (i.e., speech recognition information) can be text information obtained by performing automatic speech recognition (ASR) on the audio information corresponding to the video search results. In some embodiments, the audio information of the video can be converted into text using a speech recognition model, which can be a fusion model of RNNT (Recurrent Neural Network Transducer) and LAS (listen, attention, and spell) models. In addition, the start and end time points corresponding to the audio information in the video can be obtained as the start and end time points of the video ASR information. These start and end time points can be used as positioning points during video playback, allowing users to locate the video segment corresponding to the video ASR information.

[0117] The answer summary and frame-by-frame summary information can be obtained by summarizing the text information acquired through OCR / ASR. In one embodiment, the text information acquired through OCR / ASR can be fed into a trained MRC (Machine Reading Comprehension) long answer model (such as Multi-passageBERT) to obtain the long answer. After obtaining the long answer, it is fed into a summary generation model trained based on T5-pegasus (a Chinese generation model) to obtain the answer summary corresponding to the full text. This model has the functions of simplification, completion, and error correction. Then, a trained binary classification model can be used to determine whether the text information obtained by OCR / ASR needs to be framed. If the output probability is below 0.5, it indicates that the answer summary is a continuous answer and does not need to be framed. If the output probability is above 0.5, it indicates that there are multiple methods or multiple operation categories in the text information and the answer summary can be framed. That is, the text information obtained by OCR / ASR is divided into multiple segments using the answer summary, and the start and end time points of each segment in the video are determined. These start and end time points can be used as positioning points when the video is played, so that users can locate and jump to the video segment corresponding to a certain method or operation category in the video for viewing.

[0118] The first step in frame segmentation can be aligning the extracted answer summary with the text information obtained from OCR / ASR. Specifically, the text information obtained from OCR / ASR can be segmented into sentences based on periods. Then, an optimized edit distance matching method is used to align the answer summary with each text sentence in the OCR / ASR sequentially. If no matching text sentence is found, considering that the text information corresponding to the OCR / ASR may contain typos, it can be converted into pinyin before matching. If a matching text sentence still cannot be found, the video is considered unsuitable for frame segmentation.

[0119] 103. Based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, the video search results in the preset video question-and-answer pair are recalled. The at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension.

[0120] In some embodiments, the similarity between the target content and the search content in the preset video question-and-answer pair, as well as the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, can be fused to determine the target similarity between the target content and the preset video question-and-answer pair. Then, the video search results in the preset video question-and-answer pair can be recalled based on the target similarity. Specifically, the video search results of preset video question-and-answer pairs with a target similarity greater than a preset value can be recalled, or the preset video question-and-answer pairs can be sorted from largest to smallest based on the target similarity, and the video search results in the top n sorted preset video question-and-answer pairs can be recalled. There are various methods for similarity fusion, and this embodiment does not limit this; for example, it can be a weighted summation.

[0121] In other embodiments, the video search results in the preset video question-and-answer pair can be retrieved in multiple ways based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension.

[0122] Optionally, in this embodiment, the step of "recalling video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension of content information" may include:

[0123] Based on the similarity between the content feature vector of the target content and the content feature vector of the search content in the preset video question and answer pair, the video search results in the preset video question and answer pair are recalled to obtain the first recall result.

[0124] The content information of the video search results in the preset video question-and-answer pair is vectorized in at least one dimension to obtain the content feature vector in the at least one dimension.

[0125] Based on the similarity between the content feature vector of the target content and the content feature vector under the at least one dimension, the video search results in the preset video question-and-answer pair are recalled to obtain the second recall result under the at least one dimension.

[0126] Specifically, the similarity between the target content and the search content in the preset video question-and-answer pair can be determined based on the vector distance between the content feature vectors of the target content and the search content. The larger the vector distance, the lower the similarity; conversely, the smaller the vector distance, the higher the similarity. The vector distance can be calculated using Euclidean distance, cosine distance, etc., and this embodiment does not limit the method used.

[0127] In this process, a semantic recognition model can be used to obtain the feature vectors of the target content and the search content. This semantic recognition model can be an LSTM model, a BERT model, or something similar.

[0128] In one embodiment, the semantic recognition model can provide a query vectorization operator that can represent the user's search query (target content or search terms) as a 256-dimensional semantic vector. For example, this semantic recognition model also employs a BERT-based model, such as... Figure 1c The right side shows the model structure diagram of this semantic recognition model, and... Figure 1c The difference between the model on the left and the model on the right is that the output layer of this semantic recognition model can be normalized using L2_normalize (L2 norm), and the final output is a 256-dimensional representation vector.

[0129] After calculating the similarity between the target content and the search content in the preset video question-and-answer pairs, the video search results in the preset video question-and-answer pairs with a similarity greater than a preset value can be recalled. Alternatively, the preset video question-and-answer pairs can be sorted according to similarity, such as sorting from largest to smallest, and the video search results in the top n preset video question-and-answer pairs after sorting can be recalled.

[0130] In this embodiment, the content information of video search results in at least one dimension may include content information in various dimensions such as optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension. This embodiment does not impose any limitations on this.

[0131] Specifically, for content information in the optical character recognition dimension, text in video frame images can be converted into text information using OCR (Optical Character Recognition); for content information in the speech recognition dimension, speech information can be converted into text information using ASR (Automated Speech Recognition) technology.

[0132] Optionally, in this embodiment, the step "vectorizing the content information of the video search results in the preset video question-and-answer pair in at least one dimension to obtain the content feature vector in the at least one dimension" may include:

[0133] The optical character recognition text of the video search results in the preset video question-and-answer pair is vectorized to obtain the content feature vector under the optical character recognition dimension. The optical character recognition text is the content information of the video search results under the optical character recognition dimension.

[0134] The speech recognition information of the video search results is vectorized to obtain the content feature vector under the speech recognition dimension, and the speech recognition information is the content information of the video search results under the speech recognition dimension;

[0135] The video frame image sequence of the video search result is vectorized to obtain the content feature vector in the image dimension, and the video frame image sequence is the content information of the video search result in the image dimension.

[0136] The video title of the video search result is vectorized to obtain a content feature vector in the dimension of the video title, where the video title is the content information of the video search result in the dimension of the video title.

[0137] Based on the optical character recognition text and the speech recognition information, the video search results are processed to extract a summary, resulting in a content feature vector of the video search results under the summary dimension.

[0138] In some embodiments, the content information of video search results in at least one dimension may also include content information in a fusion dimension. This fusion dimension may also be referred to as cross-dimensional content information. Cross-dimensional content information may include search content (query), video information corresponding to the video search results (i.e., the video itself, which may be a sequence of video frame images), video title, optical character recognition text, and speech recognition information, etc. By performing feature interaction on these content information, cross-dimensional content feature information, i.e., cross-dimensional vector, can be obtained.

[0139] For example, it can be done through, as Figure 1d The cross-dimensional vectorization model shown is used to extract cross-dimensional vectors. Specifically, this model can be a BERT model. The model can simultaneously input the query of the video question-answering pair, the video frame image sequence, the video title, the optical character recognition text, and the speech recognition information, so that information from multiple dimensions can interact in the model to obtain a richer feature vector representation.

[0140] Optionally, in this embodiment, the at least one dimension further includes cross-dimensionality; the step "vectorizing the content information of the video search results in the preset video question-and-answer pair under at least one dimension to obtain the content feature vector under the at least one dimension" may include:

[0141] The video search results in the preset video question-and-answer pair are obtained as follows: optical character recognition text in the optical character recognition dimension, speech recognition information in the speech recognition dimension, and video frame image sequence in the image dimension.

[0142] The optical character recognition text, the speech recognition information, the video frame image sequence, and the search content in the preset video question-and-answer pair are subjected to interactive feature vector processing to obtain the content feature vector in the cross-dimensional context.

[0143] Among them, the content feature vector in the cross-dimensional context is the cross-dimensional vector in the above embodiment.

[0144] Specifically, in some embodiments, video search results in a preset video question-and-answer pair can be retrieved through multiple channels based on the similarity between the target content and the content information of video search results under each dimension, thereby obtaining retrieval results corresponding to each dimension. In other embodiments, the similarity between the target content and the content information of video search results under each dimension can be fused, and video search results in a preset video question-and-answer pair can be retrieved based on the fused similarity.

[0145] In one specific embodiment, for each preset video question-and-answer pair, the video itself, video title, video OCR information, video ASR information, answer summary, and query content corresponding to its video search result can be obtained. The video frame image sequence, video title, video OCR information, video ASR information, answer summary, and query are then vectorized to obtain video vector, video title vector, OCR vector, ASR vector, answer summary vector, and query vector, respectively. The query vector model can use a BERT model, and the video vector model, video title vector model, OCR vector model, ASR vector model, and answer summary vector model can all adopt a structure similar to the query vector model, see [link to relevant documentation]. Figure 1e The input to the model can be any of the five types of data mentioned above, the difference being that the input information, features, and training data are different.

[0146] Based on the aforementioned video vectors, video title vectors, OCR vectors, ASR vectors, answer summary vectors, query vectors, and cross-dimensional vectors, the target content can be used to calculate the similarity with each of these seven vectors. Then, based on the similarity, the video search results for the preset video question-and-answer pairs are recalled to obtain seven recall results.

[0147] Specifically, the HNSW algorithm can be used to recall video search results. HNSW stands for Hierarchical Navigable Small World graphs, a graph-based algorithm in the field of neural network search. The HNSW algorithm constructs an approximate small-world network from a vector set in a certain way, and then, for the query vector, randomly selects an initial point for fast retrieval.

[0148] In this embodiment, the HNSW algorithm can be used to construct seven index libraries for the seven vectors corresponding to each preset video question-answer pair in the video question-answer library. The target content to be searched is then compared with each vector in the seven index libraries to perform vector similarity recall processing. Candidate answers for the target content are then retrieved from the index libraries through seven-way recall.

[0149] The following example uses query vectors to illustrate the process of building an index library for query vectors, and how to retrieve video search results (i.e., candidate answers for the target content) from the query vector index library based on the target content to be searched. The construction of other index libraries can refer to this process, and this embodiment will not elaborate on it further.

[0150] First, the query vectors of each preset video question-answer pair in the video question-answering database can be obtained. The centroids of m partitions of these query vectors can be calculated, and these queries can be divided into m partitions such that the distance between the query in each partition and the centroid of that partition is less than a preset value. Then, n query vectors are uniformly sampled from each partition as sample data. The distribution information of the query vectors in the partition is represented by the sample data, and the partition label of the partition in which the sample data belongs is given. Thus, a global index is built using the sample data, and the index library of query vectors is completed.

[0151] After completing the query index construction, the p sample data points with the closest vector distance to the target content can be found in the global index to obtain preliminary result vectors. Then, based on the partition labels corresponding to these p sample data points, the number of preliminary result vectors contained in each partition is counted. According to the number, each partition is sorted in descending order, and the top s partitions with a non-zero number of preliminary result vectors are selected as the query partitions. For each query partition, the k query vectors with the closest vector distance to the target content in that query partition are obtained as the partition result vectors. According to the similarity between the vectors corresponding to the target content, the partition result vectors of each query partition are sorted, such as sorted in descending order. The top l query vectors are then used as the target query vectors, and the video search results corresponding to the target query vectors are used as candidate answers for the target content.

[0152] Optionally, in this embodiment, the step of "recalling video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension of content information" may include:

[0153] Obtain the content index map corresponding to the search content in the preset video question and answer pair, and the content index map corresponding to the content information of the video search results in the preset video question and answer pair under at least one dimension. The content index map includes various index layers arranged from top to bottom with the number of nodes increasing sequentially. Each index layer includes at least one node, and the node content corresponding to each node is the content information of a search content or video search result under at least one dimension.

[0154] For each content index graph, based on the similarity between the target content and the node content corresponding to the node, a node search is performed on each index layer in the content index graph in a top-to-bottom order to find similar nodes corresponding to the target content in the nodes of the target index layer.

[0155] Based on the node content corresponding to the similar nodes, the video search results in the preset video question-and-answer pairs are recalled to obtain the recall results corresponding to the content index graph.

[0156] The content index graph can be constructed using the HNSW algorithm. For example, for the content index graph corresponding to search content, the search content can be clustered and partitioned based on the similarity between the search content in each preset video question-and-answer pair, resulting in at least one content partition. The distance between the search content in each content partition and the cluster center of that content partition is less than a preset value. Then, each content partition undergoes multiple sampling processes, where each sampling process is performed on the sampling result of the previous sampling of the corresponding content partition (i.e., the search content obtained from the previous sampling). This yields multiple sampling results for each content partition, and the search content in the same sampling result is aggregated into the same index layer. The top layer of the content index graph (i.e., the first index layer) includes the search content from the last sampling result of each content partition. It should be noted that the construction process of the content index graph corresponding to the content information of video search results in at least one dimension can refer to the construction process of the content index graph corresponding to search content. Specifically, the content index graph corresponding to the content information of video search results in at least one dimension can include content index graphs corresponding to the content information in each dimension.

[0157] Specifically, the node search for each index layer can include using the start node of the current index layer as the initial current node, searching from the current node and its connected neighboring nodes to find the node closest to the feature vector of the target content as the updated current node, and determining the current node that reaches the end of the search as the first node, through which the search proceeds to the next index layer; and this first node is used as the start node of the next index layer. The start node can be any node chosen.

[0158] The target index layer can be the last index layer in the content index graph, that is, the bottom layer. It should be noted that the last index layer of the content index graph corresponding to the search content includes all nodes corresponding to the search content.

[0159] 104. Based on the recall frequency information corresponding to the recalled video search results, determine the target video search results corresponding to the target content from the recalled video search results.

[0160] In some embodiments, video search results can be retrieved through multiple channels, and the retrieved video search results can be regarded as candidate video answers for the target content. Since a certain video search result may be retrieved repeatedly in multiple channels, this embodiment can aggregate and statistically analyze the retrieval results from each channel to determine the retrieval frequency information of the retrieved video search results.

[0161] Optionally, in this embodiment, the second recall result under at least one dimension includes the second recall result under each dimension; the step "determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information" may include:

[0162] The first recall result and the second recall results under each dimension are aggregated and statistically processed to obtain the recall frequency information corresponding to each recalled video search result.

[0163] Based on the recall frequency information, the target video search results corresponding to the target content are determined from the recalled video search results.

[0164] Optionally, in some embodiments, video search results with recall frequency information exceeding a preset number can be determined as target video search results corresponding to the target content. Alternatively, the recalled video search results can be sorted according to the recall frequency information, such as sorting from largest to smallest, and the top n video search results in the sorted results can be determined as target video search results corresponding to the target content.

[0165] Optionally, in this embodiment, the step "determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results" may include:

[0166] Obtain quality information of the recalled video search results in at least one dimension;

[0167] Based on the recall frequency information corresponding to the recalled video search results and the quality information in the at least one dimension, the target video search results corresponding to the target content are determined from the recalled video search results.

[0168] Specifically, the quality information of video search results in at least one dimension may include: overall video score f1, video clarity f2, video cover score f3, publisher's overall rating f4, publisher's popularity and influence f5, publisher's credibility f6, publisher's overall quality f7, publisher's domain focus f8, consistency between video content and publisher's domain f9, and explicit domain distribution of account posts f1. 10 Does the publisher have a predetermined number of followers? 11 The probability f that the question and text do not match 12 Video views f 13 Video likes f 14 Video comment count f 15 Video sharing volume f 16 Are the videos in the video Q&A pairs relevant to the search content? 17Video access control 18 wait.

[0169] Optionally, in this embodiment, the step "determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results and the quality information in the at least one dimension" may include:

[0170] The recall frequency information and quality information in at least one dimension corresponding to the recalled video search results are fused to obtain fused feature information.

[0171] Based on the fused feature information, the probability that the recalled video search results meet the preset quality conditions is predicted.

[0172] Based on the probability, the target video search results corresponding to the target content are determined from the recalled video search results.

[0173] There are various ways to fuse recall frequency information and quality information in at least one dimension. This embodiment does not limit this method. For example, the fusion method can be splicing.

[0174] In this process, a logarithmic linear operation can be performed on the fused feature information to obtain the probability that the recalled video search results meet the preset quality conditions. The preset quality conditions can be set according to the actual situation, and this embodiment does not limit them. For example, the preset quality condition can be that the video search results are video answers that match the target content to be searched.

[0175] In some embodiments, video search results with a probability greater than a preset value can be identified as target video search results corresponding to the target content; in other embodiments, the recalled video search results can be sorted according to probability, such as sorting from largest to smallest, and the top n video search results in the sorted video search results can be identified as target video search results corresponding to the target content.

[0176] In one embodiment, a logistic regression model can be used to determine the target video search results corresponding to the target content from the recalled video search results; specifically, the quality information may include 18 dimensions, and the quality information in each dimension is denoted as f1, f2, ..., f 18 The recall frequency information can be denoted as f. 19 Based on these 18 quality information and recall frequency information, a 19-dimensional feature representation vector x can be constructed. This feature representation vector x serves as the input to the logistic regression model. Specifically, the feature representation vector x can be represented as shown in equation (1):

[0177] x = concat(f1, f2, ..., f19 (1)

[0178] Here, `concat` is a function that combines text from multiple strings.

[0179] Logistic regression is a classification method in statistical learning. It can be a binary log-linear model that can predict whether a video answer is a good answer. Its predicted conditional probability distribution can be shown in equations (1) and (2):

[0180]

[0181]

[0182] Here, w and b are model parameters, whose actual values ​​can be obtained through training. P(Y=1|x) represents the probability that the video answer is a good video answer, and P(Y=0|x) represents the probability that the video answer is a relatively bad video answer.

[0183] After obtaining the target video search results corresponding to the target content, the target video search results can be displayed on the corresponding search results page.

[0184] In a specific scenario, such as Figure 1f The diagram shown is a flowchart of a content search device based on this application. The content search device may mainly include a search control module, a QU module, and a video question-and-answer backend module.

[0185] The search control module can receive the target content (i.e., query content) entered by the user in the content search platform, and send a request instruction to the QU module to extract the content feature vector of the target content. It can also receive the content feature vector of the target content sent by the QU module and then pass the content feature vector of the target content to the video Q&A backend module.

[0186] The QU (Query Understanding) module provides natural language processing capabilities such as video intent operators, question-and-answer intent operators, and query vectorization operators. Its main function is to deeply understand the target content, extract the content feature vector of the target content, and identify whether the target content has video question-and-answer intent information.

[0187] The video Q&A backend module can retrieve high-quality video answers from the video Q&A library based on the content feature vector of the target content. The video Q&A backend module can be divided into two sub-modules: the multi-path retrieval sub-module and the ranking sub-module.

[0188] The multi-path recall submodule can be used to recall candidate video answers from a video question-answering database. Specifically, more relevant candidate video answers can be obtained through multi-path recall. In one embodiment, the multi-path recall submodule may include recall sub-dimensions corresponding to seven recall paths: query recall submodule, title recall submodule, video recall submodule, optical character recognition recall submodule, speech recognition recall submodule, summary recall submodule, and cross-dimensional recall submodule.

[0189] The query recall submodule retrieves video search results from the video question-and-answer pair based on the similarity between the target content and the search query, resulting in a recall result r1. The video recall submodule retrieves video search results from the video question-and-answer pair based on the similarity between the target content and the video content itself (i.e., the video frame sequence), resulting in a recall result r2. The title recall submodule retrieves video search results from the video question-and-answer pair based on the similarity between the target content and the video title, resulting in a recall result r3. The optical character recognition (OCR) recall submodule retrieves video search results from the video question-and-answer pair based on the similarity between the target content and the video OCR information (i.e., optical character recognition text), resulting in a recall result r4. The speech recognition recall submodule retrieves video search results from the video question-and-answer pair based on the similarity between the target content and the video ASR information (i.e., speech recognition information), resulting in a recall result r5. The summary recall submodule can be used to recall video search results in video question-and-answer pairs based on the similarity between the target content and the answer summary of the video question-and-answer pair, resulting in recall result r6. The cross-dimensional recall submodule can be used to recall video search results in video question-and-answer pairs based on the similarity between the target content and the cross-dimensional vector of the video question-and-answer pair, resulting in recall result r7.

[0190] Each recall pathway retrieves video search results from the top 10 video question-and-answer pairs most relevant to the target content. The candidate video answers retrieved from each pathway are then aggregated and sent to the ranking submodule. The top 10 is an empirical value that can be adjusted flexibly according to actual needs. By recalling candidate video answers through seven pathways, with each pathway retrieving the top 10, a total of 70 candidate video answers can be output to the ranking submodule.

[0191] The ranking submodule can be used to sort the candidate video answers retrieved from multiple channels and select the best video answer to display on the search results page. Specifically, this ranking submodule can comprehensively consider various quality information of each candidate video answer, score all candidate video answers, and finally select the answer with the highest overall video quality for online display.

[0192] First, the 70 candidate video answers retrieved from multiple channels can be aggregated to obtain recall frequency information. This involves accumulating the recall counts for video answers that are retrieved repeatedly. For example, if a candidate video answer exists in the recall results r2 output by the video recall submodule, r4 output by the optical character recognition recall submodule, r5 output by the speech recognition recall submodule, and r7 output by the cross-dimensional recall submodule, then the recall frequency information for that candidate video answer is 4.

[0193] The ranking strategy can employ a logistic regression model. Based on the quality information and recall frequency information of candidate video answers in various dimensions, a feature representation vector is constructed as the input to the logistic regression model. The logistic regression model can predict the probability that the candidate video answer belongs to a high-quality video answer. Then, based on the probability, all 70 candidate video answers are ranked, and the candidate video answer with the highest probability is taken as the optimal video answer within the target, thereby displaying the video answer on the search results page.

[0194] This solution can retrieve video answers to queries from a video question-and-answer database. The information presented in the video answers is more intuitive, which can improve user satisfaction.

[0195] refer to Figure 1g Pages a and b in the text, and Figure 1h Pages c and d in this document provide examples of online multi-channel video answer retrieval results displayed using the content search method provided in this application. When a user enters text or voice queries such as "how to parallel park," "how to edit and modify content in a PDF," "how to complete real-name authentication for an application," or "how to change the password for a mobile Wi-Fi network" into the browser's search input box, video answer-type search results can be provided. Since video answers are more intuitive and clear, this better meets user needs. Furthermore, while providing video answers, related text answers can also be displayed on the search results page, further satisfying the question-and-answer needs of different users and improving the user experience.

[0196] PDF stands for Portable Document Format; Wi-Fi is a wireless network communication technology.

[0197] This application proposes a video question-answering recall method that can comprehensively utilize the feature vectors of the current search content and video question-answer pairs in multiple dimensions to perform multi-path recall of video answers. By introducing mutual verification of data from various paths and dimensions in the recall process, the accuracy and precision of video search result recall are greatly improved, the relevance of search results to the user's search content is enhanced, and the user's search experience is significantly improved.

[0198] As can be seen from the above, this embodiment can obtain the target content to be searched and identify video question-and-answer intent information of the target content; when video question-and-answer intent information is identified in the target content, at least one preset video question-and-answer pair is obtained, the preset video question-and-answer pair includes the search content and the video search results corresponding to the search content; based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, the video search results in the preset video question-and-answer pair are recalled, the at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension; according to the recall frequency information corresponding to the recalled video search results, the target video search results corresponding to the target content are determined from the recalled video search results. This application can identify video question-and-answer intent of the user-input search content, and if video question-and-answer intent exists, more intuitive and concise video search results can be returned, which is beneficial to improving the accuracy of search results.

[0199] Based on the method described in the preceding embodiments, the following will provide a more detailed explanation by taking the specific integration of the content search device into a server as an example.

[0200] This application provides a content search method, such as... Figure 2 As shown, the specific process of this content search method can be as follows:

[0201] 201. The server obtains the target content to be searched and identifies the video question-and-answer intent information of the target content.

[0202] The target content refers to the content to be searched, and its type is not limited. For example, the target content can be text, audio, or an image. Specifically, the target content is what the user is querying, and can be represented by a query. Specifically, if the target content is audio, it can be converted into text through speech recognition before the content search is performed.

[0203] Optionally, in this embodiment, the step of "identifying the video question-and-answer intent information of the target content" may include:

[0204] Temporal features are extracted from each text unit in the target content to obtain the content temporal feature information of the target content;

[0205] Based on the temporal feature information of the content, the intent information of video question answering is identified in the target content.

[0206] The process involves first segmenting the target content into words to obtain individual text units, and then extracting temporal features from each text unit. A text unit can be a word or a character; this embodiment does not impose any restrictions on this.

[0207] Specifically, based on the temporal features of the content, it is possible to predict whether the target content contains video question-and-answer intent information. This can be achieved using a classifier.

[0208] If a search result contains video question-and-answer intent information, it means that the user needs to obtain answers for that search result, that is, the user has a question-and-answer requirement and needs search results for the video type of that search result.

[0209] Optionally, in this embodiment, the step "extracting temporal features from each text unit in the target content to obtain the content temporal feature information of the target content" may include:

[0210] Feature extraction is performed on each text unit in the target content to obtain word-level feature information corresponding to each text unit;

[0211] Based on the word-level feature information of each text unit corresponding to its context, the word-level feature information of each text unit is processed.

[0212] The word-level feature information of each processed text unit is fused to obtain the content temporal feature information of the target content.

[0213] The word-level feature information of a text unit can specifically be the word vector of the text unit, or the feature information obtained by fusing the content vector, type vector and position vector of the text unit.

[0214] Specifically, the text units in the context corresponding to a text unit can be other text units in the target content besides the text unit itself. This embodiment can fuse the word-level feature information of the text units in the various contexts corresponding to a text unit to obtain the context feature information corresponding to that text unit, and then process the word-level feature information of that text unit based on the context feature information. There are various ways to fuse the processed word-level feature information of the various text units, such as weighted summation, and this embodiment does not limit this approach.

[0215] 202. When the target content is identified to contain video question-and-answer intent information, the server obtains at least one preset video question-and-answer pair, which includes the search content and the video search results corresponding to the search content.

[0216] In one embodiment, at least one preset video question-and-answer pair can be obtained from a video question-and-answer library. The video question-and-answer library can store a certain number of video question-and-answer pairs. Each preset video question-and-answer pair includes a search content and a video search result corresponding to the search content. The search content is the question content query, and the video search result is the answer corresponding to the question content. The video search result is specifically a video type search result.

[0217] 203. The server recalls video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension. The at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension.

[0218] In some embodiments, the similarity between the target content and the search content in the preset video question-and-answer pair, as well as the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, can be fused to determine the target similarity between the target content and the preset video question-and-answer pair. Then, the video search results in the preset video question-and-answer pair can be recalled based on the target similarity. Specifically, the video search results of preset video question-and-answer pairs with a target similarity greater than a preset value can be recalled, or the preset video question-and-answer pairs can be sorted from largest to smallest based on the target similarity, and the video search results in the top n sorted preset video question-and-answer pairs can be recalled. There are various methods for similarity fusion, and this embodiment does not limit this; for example, it can be a weighted summation.

[0219] In other embodiments, the video search results in the preset video question-and-answer pair can be retrieved in multiple ways based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension.

[0220] Optionally, in this embodiment, the step of "recalling video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension of content information" may include:

[0221] Based on the similarity between the content feature vector of the target content and the content feature vector of the search content in the preset video question and answer pair, the video search results in the preset video question and answer pair are recalled to obtain the first recall result.

[0222] The content information of the video search results in the preset video question-and-answer pair is vectorized in at least one dimension to obtain the content feature vector in the at least one dimension.

[0223] Based on the similarity between the content feature vector of the target content and the content feature vector under the at least one dimension, the video search results in the preset video question-and-answer pair are recalled to obtain the second recall result under the at least one dimension.

[0224] Specifically, the similarity between the target content and the search content in the preset video question-and-answer pair can be determined based on the vector distance between the content feature vectors of the target content and the search content. The larger the vector distance, the lower the similarity; conversely, the smaller the vector distance, the higher the similarity. The vector distance can be calculated using Euclidean distance, cosine distance, etc., and this embodiment does not limit the method used.

[0225] After calculating the similarity between the target content and the search content in the preset video question-and-answer pairs, the video search results in the preset video question-and-answer pairs with a similarity greater than a preset value can be recalled. Alternatively, the preset video question-and-answer pairs can be sorted according to similarity, such as sorting from largest to smallest, and the video search results in the top n preset video question-and-answer pairs after sorting can be recalled.

[0226] In this embodiment, the content information of video search results in at least one dimension may include content information in various dimensions such as optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension. This embodiment does not impose any limitations on this.

[0227] Specifically, for content information in the optical character recognition dimension, text in video frame images can be converted into text information using OCR (Optical Character Recognition); for content information in the speech recognition dimension, speech information can be converted into text information using ASR (Automated Speech Recognition) technology.

[0228] Optionally, in this embodiment, the step "vectorizing the content information of the video search results in the preset video question-and-answer pair in at least one dimension to obtain the content feature vector in the at least one dimension" may include:

[0229] The optical character recognition text of the video search results in the preset video question-and-answer pair is vectorized to obtain the content feature vector under the optical character recognition dimension. The optical character recognition text is the content information of the video search results under the optical character recognition dimension.

[0230] The speech recognition information of the video search results is vectorized to obtain the content feature vector under the speech recognition dimension, and the speech recognition information is the content information of the video search results under the speech recognition dimension;

[0231] The video frame image sequence of the video search result is vectorized to obtain the content feature vector in the image dimension, and the video frame image sequence is the content information of the video search result in the image dimension.

[0232] The video title of the video search result is vectorized to obtain a content feature vector in the dimension of the video title, where the video title is the content information of the video search result in the dimension of the video title.

[0233] Based on the optical character recognition text and the speech recognition information, the video search results are processed to extract a summary, resulting in a content feature vector of the video search results under the summary dimension.

[0234] Optionally, in this embodiment, the at least one dimension further includes cross-dimensionality; the step "vectorizing the content information of the video search results in the preset video question-and-answer pair under at least one dimension to obtain the content feature vector under the at least one dimension" may include:

[0235] The video search results in the preset video question-and-answer pair are obtained as follows: optical character recognition text in the optical character recognition dimension, speech recognition information in the speech recognition dimension, and video frame image sequence in the image dimension.

[0236] The optical character recognition text, the speech recognition information, the video frame image sequence, and the search content in the preset video question-and-answer pair are subjected to interactive feature vector processing to obtain the content feature vector in the cross-dimensional context.

[0237] Among them, the content feature vector in the cross-dimensional context is the cross-dimensional vector in the above embodiment.

[0238] Specifically, in some embodiments, video search results in a preset video question-and-answer pair can be retrieved through multiple channels based on the similarity between the target content and the content information of video search results under each dimension, thereby obtaining retrieval results corresponding to each dimension. In other embodiments, the similarity between the target content and the content information of video search results under each dimension can be fused, and video search results in a preset video question-and-answer pair can be retrieved based on the fused similarity.

[0239] Optionally, in this embodiment, the step of "recalling video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension of content information" may include:

[0240] Obtain the content index map corresponding to the search content in the preset video question and answer pair, and the content index map corresponding to the content information of the video search results in the preset video question and answer pair under at least one dimension. The content index map includes various index layers arranged from top to bottom with the number of nodes increasing sequentially. Each index layer includes at least one node, and the node content corresponding to each node is the content information of a search content or video search result under at least one dimension.

[0241] For each content index graph, based on the similarity between the target content and the node content corresponding to the node, a node search is performed on each index layer in the content index graph in a top-to-bottom order to find similar nodes corresponding to the target content in the nodes of the target index layer.

[0242] Based on the node content corresponding to the similar nodes, the video search results in the preset video question-and-answer pairs are recalled to obtain the recall results corresponding to the content index graph.

[0243] The content index graph can be constructed using the HNSW algorithm. For example, for the content index graph corresponding to the search content, the search content can be clustered and partitioned according to the similarity between the search content in each preset video question-and-answer pair to obtain at least one content partition. The distance between the search content in each content partition and the cluster center of that content partition is less than a preset value. Then, multiple sampling processes are performed on each content partition. Each sampling process is performed on the sampling result of the previous sampling of the corresponding content partition (i.e., the search content obtained from the previous sampling). This can obtain multiple sampling results for each content partition, and the search content in the same sampling result is aggregated into the same index layer. The top layer of the content index graph (i.e., the first index layer) includes the search content in the last sampling result of each content partition.

[0244] Specifically, the node search for each index layer can include using the start node of the current index layer as the initial current node, searching from the current node and its connected neighboring nodes to find the node closest to the feature vector of the target content as the updated current node, and determining the current node that reaches the end of the search as the first node, through which the search proceeds to the next index layer; and this first node is used as the start node of the next index layer. The start node can be any node chosen.

[0245] The target index layer can be the last index layer in the content index graph, that is, the bottom layer. It should be noted that the last index layer of the content index graph corresponding to the search content includes all nodes corresponding to the search content.

[0246] 204. The server determines the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results.

[0247] In some embodiments, video search results can be retrieved through multiple channels, and the retrieved video search results can be regarded as candidate video answers for the target content. Since a certain video search result may be retrieved repeatedly in multiple channels, this embodiment can aggregate and statistically analyze the retrieval results from each channel to determine the retrieval frequency information of the retrieved video search results.

[0248] Optionally, in this embodiment, the second recall result under at least one dimension includes the second recall result under each dimension; the step "determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information" may include:

[0249] The first recall result and the second recall results under each dimension are aggregated and statistically processed to obtain the recall frequency information corresponding to each recalled video search result.

[0250] Based on the recall frequency information, the target video search results corresponding to the target content are determined from the recalled video search results.

[0251] Optionally, in some embodiments, video search results with recall frequency information exceeding a preset number can be determined as target video search results corresponding to the target content. Alternatively, the recalled video search results can be sorted according to the recall frequency information, such as sorting from largest to smallest, and the top n video search results in the sorted results can be determined as target video search results corresponding to the target content.

[0252] Optionally, in this embodiment, the step "determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results" may include:

[0253] Obtain quality information of the recalled video search results in at least one dimension;

[0254] Based on the recall frequency information corresponding to the recalled video search results and the quality information in the at least one dimension, the target video search results corresponding to the target content are determined from the recalled video search results.

[0255] Optionally, in this embodiment, the step "determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results and the quality information in the at least one dimension" may include:

[0256] The recall frequency information and quality information in at least one dimension corresponding to the recalled video search results are fused to obtain fused feature information.

[0257] Based on the fused feature information, the probability that the recalled video search results meet the preset quality conditions is predicted.

[0258] Based on the probability, the target video search results corresponding to the target content are determined from the recalled video search results.

[0259] There are various ways to fuse recall frequency information and quality information in at least one dimension. This embodiment does not limit this method. For example, the fusion method can be splicing.

[0260] In this process, a logarithmic linear operation can be performed on the fused feature information to obtain the probability that the recalled video search results meet the preset quality conditions. The preset quality conditions can be set according to the actual situation, and this embodiment does not limit them. For example, the preset quality condition can be that the video search results are video answers that match the target content to be searched.

[0261] In some embodiments, video search results with a probability greater than a preset value can be identified as target video search results corresponding to the target content; in other embodiments, the recalled video search results can be sorted according to probability, such as sorting from largest to smallest, and the top n video search results in the sorted video search results can be identified as target video search results corresponding to the target content.

[0262] This application proposes a video question-answering recall method that can comprehensively utilize the feature information of the current search content and video question-answer pairs under multiple dimensions to perform multi-path recall of video answers. By introducing mutual verification of data from various paths and dimensions in the recall process, the accuracy and precision of video search result recall are greatly improved, the relevance of search results to the user's search content is enhanced, and the user's search experience is significantly improved.

[0263] As can be seen from the above, this embodiment can obtain the target content to be searched through a server and identify the video question-and-answer intent information of the target content; when the target content is identified to have video question-and-answer intent information, at least one preset video question-and-answer pair is obtained, the preset video question-and-answer pair includes the search content and the video search result corresponding to the search content; based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search result in the preset video question-and-answer pair in at least one dimension, the video search result in the preset video question-and-answer pair is recalled, the at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension; according to the recall frequency information corresponding to the recalled video search result, the target video search result corresponding to the target content is determined from the recalled video search result. This application can identify the video question-and-answer intent of the user-input search content. If video question-and-answer intent exists, more intuitive and concise video search results can be returned, which is beneficial to improving the accuracy of search results.

[0264] To better implement the above methods, embodiments of this application also provide a content search device, such as... Figure 3 As shown, the content search device may include an intent recognition unit 301, an acquisition unit 302, a recall unit 303, and a determination unit 304, as follows:

[0265] (1) Intent recognition unit 301;

[0266] The intent recognition unit is used to acquire the target content to be searched and to identify the video question-and-answer intent information of the target content.

[0267] Optionally, in some embodiments of this application, the intent recognition unit may include a feature extraction subunit and an intent recognition subunit, as follows:

[0268] The feature extraction subunit is used to extract temporal features from each text unit in the target content to obtain the content temporal feature information of the target content.

[0269] The intent recognition subunit is used to identify video question-and-answer intent information of the target content based on the content temporal feature information.

[0270] Optionally, in some embodiments of this application, the feature extraction subunit may be used to extract features from each text unit in the target content to obtain word-level feature information corresponding to each text unit; process the word-level feature information of each text unit based on the word-level feature information of the text units in the context corresponding to each text unit; and fuse the processed word-level feature information of each text unit to obtain the content temporal feature information of the target content.

[0271] Optionally, in some embodiments of this application, the intent recognition unit may be specifically used to identify video question-and-answer intent information of the target content through an intent recognition model.

[0272] Optionally, in some embodiments of this application, the content search device may further include a training unit for training an intent recognition model. Specifically, the training unit may acquire training data, including sample content and the expected probability that the sample content contains video question-and-answer intent information. Using the intent recognition model, temporal features are extracted from each text unit in the sample content to obtain the content temporal feature information of the sample content. Based on the content temporal feature information, the actual probability that the sample content contains video question-and-answer intent information is predicted. Based on the expected probability and the actual probability, the parameters of the intent recognition model are adjusted to obtain the trained intent recognition model.

[0273] (2) Obtain unit 302;

[0274] The acquisition unit is used to acquire at least one preset video question-and-answer pair when it is identified that the target content contains video question-and-answer intent information. The preset video question-and-answer pair includes the search content and the video search results corresponding to the search content.

[0275] (3) Recall Unit 303;

[0276] The recall unit is used to recall video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension. The at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension.

[0277] Optionally, in some embodiments of this application, the recall unit may include an index graph acquisition subunit, a node search subunit, and a search result recall subunit, as follows:

[0278] The index graph acquisition subunit is used to acquire the content index graph corresponding to the search content in the preset video question and answer pair, and the content index graph corresponding to the content information of the video search results in the preset video question and answer pair under at least one dimension. The content index graph includes various index layers arranged from top to bottom with the number of nodes increasing sequentially. Each index layer includes at least one node, and the node content corresponding to each node is the content information of a search content or video search result under at least one dimension.

[0279] The node search subunit is used to perform node search on each index layer in the content index graph in a top-to-bottom order, based on the similarity between the target content and the node content corresponding to the node, for each content index graph, so as to find similar nodes corresponding to the target content in the nodes of the target index layer.

[0280] The search result recall subunit is used to recall video search results in the preset video question-and-answer pair based on the node content corresponding to the similar nodes, and obtain the recall result corresponding to the content index graph.

[0281] Optionally, in some embodiments of this application, the recall unit may include a first recall subunit, an extraction subunit, and a second recall subunit, as follows:

[0282] The first recall subunit is used to recall video search results in the preset video question-and-answer pair based on the similarity between the content feature vector of the target content and the content feature vector of the search content in the preset video question-and-answer pair, and obtain a first recall result.

[0283] Extraction subunits are used to vectorize the content information of video search results in the preset video question-and-answer pair in at least one dimension to obtain the content feature vector in the at least one dimension.

[0284] The second recall subunit is used to recall video search results in the preset video question-and-answer pair based on the similarity between the content feature vector of the target content and the content feature vector under the at least one dimension, and obtain the second recall result under the at least one dimension.

[0285] Optionally, in some embodiments of this application, the extraction subunit may specifically be used to vectorize the optical character recognition (OCR) text of the video search results in the preset video question-and-answer pair to obtain a content feature vector in the OCR dimension, wherein the OCR text is the content information of the video search results in the OCR dimension; to vectorize the speech recognition information of the video search results to obtain a content feature vector in the speech recognition dimension, wherein the speech recognition information is the content information of the video search results in the speech recognition dimension; to vectorize the video frame image sequence of the video search results to obtain a content feature vector in the image dimension, wherein the video frame image sequence is the content information of the video search results in the image dimension; to vectorize the video title of the video search results to obtain a content feature vector in the video title dimension, wherein the video title is the content information of the video search results in the video title dimension; and to perform summary extraction processing on the video search results based on the OCR text and the speech recognition information to obtain a content feature vector of the video search results in the summary dimension.

[0286] Optionally, in some embodiments of this application, the at least one dimension further includes cross-dimensionality; the extraction subunit can specifically be used to obtain the optical character recognition text, speech recognition information, and video frame image sequence in the preset video question-and-answer pair under the optical character recognition dimension, the speech recognition information under the speech recognition dimension, and the video frame image sequence under the image dimension of the video search results; and to perform feature vector interaction processing on the optical character recognition text, the speech recognition information, the video frame image sequence, and the search content in the preset video question-and-answer pair to obtain the content feature vector under the cross-dimensionality.

[0287] (4) Determine unit 304;

[0288] The determining unit is used to determine the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results.

[0289] Optionally, in some embodiments of this application, the second recall result under at least one dimension includes the second recall result under each dimension; the determining unit may include a statistical subunit and a result determining subunit, as follows:

[0290] The statistical subunit is used to perform aggregate statistical processing on the first recall result and the second recall result under each dimension to obtain the recall frequency information corresponding to each recalled video search result.

[0291] The result determination subunit is used to determine the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information.

[0292] Optionally, in some embodiments of this application, the determining unit may include an acquiring subunit and a determining subunit, as follows:

[0293] The acquisition subunit is used to acquire quality information of the recalled video search results in at least one dimension.

[0294] A determining subunit is configured to determine the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results and the quality information in the at least one dimension.

[0295] Optionally, in some embodiments of this application, the determining subunit may be specifically used to fuse the recall frequency information and quality information in the at least one dimension corresponding to the recalled video search results to obtain fused feature information; based on the fused feature information, predict the probability that the recalled video search results meet the preset quality conditions; and based on the probability, determine the target video search results corresponding to the target content from the recalled video search results.

[0296] As can be seen from the above, this embodiment can obtain the target content to be searched through the intent recognition unit 301 and identify the video question-and-answer intent information of the target content; when the target content is identified to have video question-and-answer intent information, the acquisition unit 302 obtains at least one preset video question-and-answer pair, the preset video question-and-answer pair including the search content and the video search result corresponding to the search content; the recall unit 303 recalls the video search result in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search result in the preset video question-and-answer pair in at least one dimension, the at least one dimension including optical character recognition dimension, speech recognition dimension, image dimension, video title dimension and summary dimension; the determination unit 304 determines the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results. This application can perform video question-and-answer intent recognition on the user-input search content, and if video question-and-answer intent exists, it can return more intuitive and concise video search results, which is beneficial to improving the accuracy of search results.

[0297] This application also provides an electronic device, such as... Figure 4The diagram shows a schematic representation of the structure of an electronic device according to an embodiment of this application. This electronic device can be a terminal or a server, specifically:

[0298] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0299] The processor 401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 402, and calls data stored in the memory 402, to perform various functions and process data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0300] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0301] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0302] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0303] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0304] The system acquires the target content to be searched and identifies video question-and-answer intent information within the target content. When video question-and-answer intent information is identified in the target content, at least one preset video question-and-answer pair is acquired. The preset video question-and-answer pair includes the search content and the corresponding video search result. Based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search result in the preset video question-and-answer pair in at least one dimension, the system recalls the video search results in the preset video question-and-answer pair. The at least one dimension includes optical character recognition, speech recognition, image, video title, and summary dimensions. Based on the recall frequency information corresponding to the recalled video search results, the system determines the target video search result corresponding to the target content from the recalled video search results.

[0305] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0306] As can be seen from the above, this embodiment can obtain the target content to be searched and identify video question-and-answer intent information of the target content; when video question-and-answer intent information is identified in the target content, at least one preset video question-and-answer pair is obtained, the preset video question-and-answer pair includes the search content and the video search results corresponding to the search content; based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, the video search results in the preset video question-and-answer pair are recalled, the at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension; according to the recall frequency information corresponding to the recalled video search results, the target video search results corresponding to the target content are determined from the recalled video search results. This application can identify video question-and-answer intent of the user-input search content, and if video question-and-answer intent exists, more intuitive and concise video search results can be returned, which is beneficial to improving the accuracy of search results.

[0307] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0308] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the content search methods provided in embodiments of this application. For example, the instructions can execute the following steps:

[0309] The system acquires the target content to be searched and identifies video question-and-answer intent information within the target content. When video question-and-answer intent information is identified in the target content, at least one preset video question-and-answer pair is acquired. The preset video question-and-answer pair includes the search content and the corresponding video search result. Based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search result in the preset video question-and-answer pair in at least one dimension, the system recalls the video search results in the preset video question-and-answer pair. The at least one dimension includes optical character recognition, speech recognition, image, video title, and summary dimensions. Based on the recall frequency information corresponding to the recalled video search results, the system determines the target video search result corresponding to the target content from the recalled video search results.

[0310] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0311] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0312] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the content search methods provided in the embodiments of this application, the beneficial effects that any of the content search methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0313] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations of the above-described content search aspects.

[0314] The above provides a detailed description of a content search method and related devices provided by the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A content search method, characterized in that, include: Obtain the target content to be searched and identify the video question-and-answer intent information of the target content; When the target content is identified to contain video question-and-answer intent information, at least one preset video question-and-answer pair is obtained, the preset video question-and-answer pair including the search content and the video search results corresponding to the search content; Based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, the video search results in the preset video question-and-answer pair are recalled. The at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension, and summary dimension. Based on the recall frequency information corresponding to the recalled video search results, the target video search results corresponding to the target content are determined from the recalled video search results.

2. The method according to claim 1, characterized in that, The identification of video question-and-answer intent information for the target content includes: Temporal features are extracted from each text unit in the target content to obtain the content temporal feature information of the target content; Based on the temporal feature information of the content, the intent information of video question answering is identified in the target content.

3. The method according to claim 2, characterized in that, The step of extracting temporal features from each text unit in the target content to obtain the content temporal feature information of the target content includes: Feature extraction is performed on each text unit in the target content to obtain word-level feature information corresponding to each text unit; Based on the word-level feature information of each text unit corresponding to its context, the word-level feature information of each text unit is processed. The word-level feature information of each processed text unit is fused to obtain the content temporal feature information of the target content.

4. The method according to claim 1, characterized in that, The step of recalling video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, includes: Obtain the content index map corresponding to the search content in the preset video question and answer pair, and the content index map corresponding to the content information of the video search results in the preset video question and answer pair under at least one dimension. The content index map includes various index layers arranged from top to bottom with the number of nodes increasing sequentially. Each index layer includes at least one node, and the node content corresponding to each node is the content information of a search content or video search result under at least one dimension. For each content index graph, based on the similarity between the target content and the node content corresponding to the node, a node search is performed on each index layer in the content index graph in a top-to-bottom order to find similar nodes corresponding to the target content in the nodes of the target index layer. Based on the node content corresponding to the similar nodes, the video search results in the preset video question-and-answer pairs are recalled to obtain the recall results corresponding to the content index graph.

5. The method according to claim 1, wherein The step of recalling video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the target content and the video search results in the preset video question-and-answer pair in at least one dimension, includes: Based on the similarity between the content feature vector of the target content and the content feature vector of the search content in the preset video question and answer pair, the video search results in the preset video question and answer pair are recalled to obtain the first recall result. The content information of the video search results in the preset video question-and-answer pair is vectorized in at least one dimension to obtain the content feature vector in the at least one dimension. Based on the similarity between the content feature vector of the target content and the content feature vector under the at least one dimension, the video search results in the preset video question-and-answer pair are recalled to obtain the second recall result under the at least one dimension.

6. The method according to claim 5, characterized in that, The step of vectorizing the content information of the video search results in the preset video question-and-answer pair in at least one dimension to obtain the content feature vector in the at least one dimension includes: The optical character recognition text of the video search results in the preset video question-and-answer pair is vectorized to obtain the content feature vector under the optical character recognition dimension. The optical character recognition text is the content information of the video search results under the optical character recognition dimension. The speech recognition information of the video search results is vectorized to obtain the content feature vector under the speech recognition dimension, and the speech recognition information is the content information of the video search results under the speech recognition dimension; The video frame image sequence of the video search result is vectorized to obtain the content feature vector in the image dimension, and the video frame image sequence is the content information of the video search result in the image dimension. The video title of the video search result is vectorized to obtain a content feature vector in the dimension of the video title, where the video title is the content information of the video search result in the dimension of the video title. Based on the optical character recognition text and the speech recognition information, the video search results are processed to extract a summary, resulting in a content feature vector of the video search results under the summary dimension.

7. The method according to claim 6, characterized in that, The at least one dimension also includes cross-dimensional processing; the vectorization of the content information of the video search results in the preset video question-and-answer pair under at least one dimension to obtain the content feature vector under the at least one dimension includes: The video search results in the preset video question-and-answer pair are obtained as follows: optical character recognition text in the optical character recognition dimension, speech recognition information in the speech recognition dimension, and video frame image sequence in the image dimension. The optical character recognition text, the speech recognition information, the video frame image sequence, and the search content in the preset video question-and-answer pair are subjected to interactive feature vector processing to obtain the content feature vector in the cross-dimensional context.

8. The method according to claim 5, characterized in that, The second recall result under at least one dimension includes the second recall result under each dimension; determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information includes: The first recall result and the second recall results under each dimension are aggregated and statistically processed to obtain the recall frequency information corresponding to each recalled video search result. Based on the recall frequency information, the target video search results corresponding to the target content are determined from the recalled video search results.

9. The method according to claim 1, characterized in that, The step of determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information includes: Obtain quality information of the recalled video search results in at least one dimension; Based on the recall frequency information corresponding to the recalled video search results and the quality information in the at least one dimension, the target video search results corresponding to the target content are determined from the recalled video search results.

10. The method according to claim 9, characterized in that, The step of determining the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information and quality information in at least one dimension includes: The recall frequency information and quality information in at least one dimension corresponding to the recalled video search results are fused to obtain fused feature information. Based on the fused feature information, the probability that the recalled video search results meet the preset quality conditions is predicted. Based on the probability, the target video search results corresponding to the target content are determined from the recalled video search results.

11. The method according to claim 1, characterized in that, The identification of video question-and-answer intent information for the target content includes: The intent recognition model is used to identify the intent information of the target content in video question-and-answer sessions.

12. The method according to claim 11, characterized in that, Before identifying the video question-and-answer intent information of the target content using the intent recognition model, the method further includes: Acquire training data, which includes sample content and the expected probability that the sample content contains video question-answering intent information; By using an intent recognition model, temporal features are extracted from each text unit in the sample content to obtain the content temporal feature information of the sample content. Based on the temporal feature information of the content, predict the actual probability that the sample content contains video question-and-answer intent information; Based on the expected probability and the actual probability, the parameters of the intent recognition model are adjusted to obtain the trained intent recognition model.

13. A content search device, characterized in that, include: An intent recognition unit is used to acquire the target content to be searched and to identify the video question-and-answer intent information of the target content. The acquisition unit is used to acquire at least one preset video question-and-answer pair when it is identified that the target content contains video question-and-answer intent information. The preset video question-and-answer pair includes the search content and the video search results corresponding to the search content. The recall unit is used to recall video search results in the preset video question-and-answer pair based on the similarity between the target content and the search content in the preset video question-and-answer pair, and the similarity between the content information of the target content and the video search results in the preset video question-and-answer pair in at least one dimension. The at least one dimension includes optical character recognition dimension, speech recognition dimension, image dimension, video title dimension and summary dimension. The determining unit is used to determine the target video search result corresponding to the target content from the recalled video search results based on the recall frequency information corresponding to the recalled video search results.

14. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor runs the application program within the memory to perform the operations in the content search method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the content search method according to any one of claims 1 to 12.

16. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the content search method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Searching method and device, electronic equipment and storage medium

    CN113204697A

  • Similar video detection method and device

    CN113469152A