Video searching method and device
By obtaining the query text in video search and determining the query type, combining video meta information and text content, and using ASR and LLM for sparse and intensive search, the problems of low correlation, low efficiency and high cost in existing video search solutions are solved, and efficient and accurate video search results are achieved.
Patent Information
- Application Number
- CN202510413005.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-25
AI Technical Summary
The existing video search solutions have problems such as low correlation with search requests, difficulty in adapting to diverse search requests, high calculation costs and low search efficiency.
By obtaining the query text and determining the query type, video search is performed based on the meta information and text content of the video in the query text and preset video library, automatic speech recognition (ASR) technology is used to improve content understanding, and sparse and intensive searches are combined with the large language model (LLM), and highly correlated search results are generated.
It improves the relevance and efficiency of video search, reduces calculation costs, can quickly provide accurate search results, and improves user experience.
Smart Images

Figure CN120372050A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of Internet technologies, and in particular, to a video search method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of Internet technologies, more and more information is shared in the form of videos, resulting in a substantial increase in the demand for video search. However, existing video search solutions still have problems such as low relevance between search results and search requests, difficulty in adapting to diverse search requests, high computational costs, and low search efficiency, thus affecting the search experience.
[0003] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention
[0004] Embodiments of the present application provide a video search method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above-mentioned technical problems.
[0005] One aspect of embodiments of the present application provides a video search method, the method including: Obtaining a query text and determining a query type based on the query text, the query type including: information query or question query; When the query type is the information query, determining a plurality of first reference videos based on the query text and meta-information associated with each of a plurality of videos in a preset video library, and generating a first search result, the first search result including the plurality of first reference videos; When the query type is the question query, determining a plurality of second reference videos based on the query text, meta-information associated with each of the plurality of videos, and text content associated with each of the plurality of videos; generating a second search result based on the query text and the plurality of second reference videos, the second search result including a response text for the query text; wherein the text content associated with the video is obtained by performing speech recognition on the video.
[0006] Optionally, before determining the plurality of second reference videos based on the query text, meta-information associated with each of the plurality of videos, and text content associated with each of the plurality of videos, the video search method further includes: Determining the similarity between the query text and the meta-information associated with each of the plurality of videos; Determining the videos with the similarity greater than a first preset threshold as the first reference videos.
[0007] Optionally, the text content associated with the video is obtained by performing speech recognition on the video, including: For each video, obtain the associated meta-information; Input the meta-information into a pre-supervised and fine-tuned first large language model to extract keywords from the meta-information through the first large language model and correct the keywords; Provide the corrected keywords to a pre-configured speech recognition system as bias words; Perform speech recognition on the video through the speech recognition system based on the bias words to obtain the text content.
[0008] Optionally, determining multiple second reference videos based on the query text and the meta-information and text content associated with each of the multiple videos includes: Enhance the query text through a pre-trained second large language model to obtain an enhanced query text; Process the text content of the multiple videos through a pre-trained third large language model to generate corresponding summary content for each of the multiple videos; Perform sparse retrieval based on the enhanced query text and the corresponding summary content of each of the multiple videos to obtain a first retrieval result, where the first retrieval result includes videos with word matches with the enhanced query text; Perform dense retrieval based on the enhanced query text and the corresponding summary content of each of the multiple videos to obtain a second retrieval result, where the second retrieval result includes videos with semantic matches with the enhanced query text; Determine the multiple second reference videos based on the first retrieval result and the second retrieval result.
[0009] Optionally, enhancing the query text through a pre-trained second large language model to obtain an enhanced query text includes: Obtain historical query texts and a list of historical viewed videos; Input the historical query texts, the historical viewing list, and the query text into the second large language model to generate extended query words related to the query text through the second large language model; Fuse the extended query words with the query text to obtain an extended query text; Perform word segmentation on the extended query text to obtain multiple sub-words and assign corresponding weights to each sub-word; Determine the enhanced query text based on the multiple sub-words and their corresponding weights.
[0010] Optionally, perform sparse retrieval based on the enhanced query text and the respective summary content of the multiple videos to obtain a first retrieval result, including: Extract multiple target elements from the respective summary content of the multiple videos through a pre-trained fourth large language model, and construct an inverted index based on the multiple target elements and the respective summary content of the multiple videos, where the inverted index is used to represent the summary content corresponding to each target element; Perform sparse retrieval based on the inverted index and the enhanced query text to obtain the first retrieval result.
[0011] Optionally, perform dense retrieval based on the enhanced query text and the respective summary content of the multiple videos to obtain a second retrieval result, including: Obtain the meta-information associated with each of the multiple videos; Fuse the meta-information and summary content of each video to construct a video index set; Input the video index set and the enhanced query text into a pre-trained semantic recall model to obtain the second retrieval result through the semantic recall model.
[0012] Optionally, generate a second search result based on the query text and the multiple second reference videos, including: Obtain the text content associated with each of the multiple second reference videos, and perform chunking processing on each text content to obtain multiple text content chunks; where each text content chunk is associated with video information, and the video information includes the meta-information of the second reference video corresponding to the text content chunk and the time information of the text content chunk in the corresponding second reference video; Input the multiple text content chunks and the query text into a pre-trained vector retrieval model to determine a preset number of text content chunks with the highest matching degree with the query text through the vector retrieval model; Input the query text, the preset number of text content chunks, and their respective associated video information into a fifth large language model that has been pre-supervised and fine-tuned to obtain the second search result through the fifth large language model.
[0013] Another aspect of the embodiments of the present application provides a video search device, where the device includes: An acquisition module, configured to acquire a query text, and determine a query type based on the query text, where the query type includes: information query or question query; A first search module, configured to, when the query type is the information query, determine a plurality of first reference videos based on the query text and meta-information associated with each of a plurality of videos in a preset video library, and generate a first search result, where the first search result includes the plurality of first reference videos; A second search module, configured to, when the query type is the question query, determine a plurality of second reference videos based on the query text, meta-information associated with each of the plurality of videos, and text content associated with each of the plurality of videos; and generate a second search result based on the query text and the plurality of second reference videos, where the second search result includes a reply text for the query text; Wherein, the text content associated with the video is obtained by performing speech recognition on the video.
[0014] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0015] Another aspect of the embodiments of the present application provides a computer-readable storage medium, where computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.
[0016] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0017] The embodiments of the present application adopting the above technical solutions may include the following advantages: Obtain the query text and determine its corresponding query type. In the case where the query type is an information query, determine multiple first reference videos based on the query text and the meta-information associated with each of the multiple videos in the preset video library, and generate a first search result. The first search result includes multiple first reference videos. In the case where the query type is a question query: determine multiple second reference videos based on the query text, the meta-information associated with each of the multiple videos in the preset video library, and the text content (obtained through automatic speech recognition ASR); generate a second search result based on the query text and the multiple second reference videos, where the second search result includes a response text for the query text. It can be seen that the embodiments of the present application customize video search solutions for different query types: quickly provide the first reference videos as the first search result based on the meta-information during information query; introduce ASR and combine with the meta-information to enhance the understanding of the video content during question query, which can improve the relevance between the second reference videos and the query text while controlling the computing cost, and generate a response text for the query text based on the second reference videos, which can provide a second search result that is more intuitive and accurate than the first search result, effectively optimizing the search experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings exemplarily show embodiments and form part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0019] Figure 1 Schematically shows an operating environment diagram of the video search method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of the video search method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 a sub-step flowchart of step S204; Figure 4 Schematically shows the first training data according to Embodiment 1 of the present application; Figure 5 Schematically shows the second search result according to Embodiment 1 of the present application; Figure 6 Schematically shows Figure 2 a sub-step flowchart of step S204; Figure 7 Schematically shows Figure 6 a sub-step flowchart of step S600; Figure 8 Schematically shows the first search result according to Embodiment 1 of the present application; Figure 9 Schematically shows Figure 6 The sub-step flowchart of step S606 in Figure 10 Schematically shows Figure 6 The sub-step flowchart of step S606 in Figure 11 Schematically shows the structure of the semantic recall model according to Embodiment 1 of the present application; Figure 12 Schematically shows Figure 2 The sub-step flowchart of step S204 in Figure 13 Schematically shows the flowchart of the video search method according to Embodiment 1 of the present application; Figure 14 Schematically shows the block diagram of the video search device according to Embodiment 2 of the present application; and Figure 15 Schematically shows the schematic diagram of the hardware architecture of the computer device according to Embodiment 3 of the present application. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0021] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.
[0023] First, provide the term explanations involved in the present application: ASR (Automatic Speech Recognition): Automatic Speech Recognition is a technology that converts audio into text.
[0024] LLM (Large Language Model): A large language model is a deep learning model trained with large-scale data that can understand and generate natural language text.
[0025] OCR (Optical Character Recognition): Optical Character Recognition is a technology that recognizes printed or handwritten text from images and converts it into editable text.
[0026] Frame Embedding: Converts video frames into low-dimensional vectors that are easy to process and compare.
[0027] ANN (Approximate Nearest Neighbor): Approximate Nearest Neighbor accelerates the search process through an approximate method to find the data point closest to the target point.
[0028] DPO (Direct Preference Optimization): Direct Preference Optimization.
[0029] SFT (Supervised Fine-Tuning): Supervised Fine-Tuning refers to further training on a pre-trained model using a specific labeled dataset to make the model better adapt to a specific task or application scenario.
[0030] Secondly, to facilitate the understanding of the technical solutions provided by the embodiments of the present application by those skilled in the art, the related technologies are described below: With the development of Internet technology, more and more information is shared in the form of videos, resulting in a substantial increase in the demand for video search. However, the applicant has learned that the related video search solutions still have the following defects: (1) In the text-based search solution, a large amount of video content is not fully indexed, resulting in a low correlation between the search results and the search request, affecting the search experience.
[0031] (2) The image-based search solution has problems of high computational cost and poor scalability. For example, mapping the query text and the video to the same space for similarity matching to achieve end-to-end inference from text to video. Although this method performs well in short videos, as the video duration increases, the accuracy decreases and the computational cost rises. And this method requires more time to control and interpret the search results and cannot ensure a response can be returned within a few hundred milliseconds, thus limiting large-scale deployment.
[0032] (3) Difficult to adapt to diverse search requests and low search efficiency.
[0033] For this reason, the embodiments of the present application provide a technical solution for video search. In this technical solution: (1) Customize video search solutions for different query types, quickly provide videos or responses as search results, better adapt to diverse queries, save computing costs, and effectively improve the search experience; (2) In question queries, integrate ASR and LLM for video search, where ASR provides a cost-effective way to understand videos, enabling the video search solution to be applied on a large scale. LLM can correct ASR and optimize the relevance judgment in video search, and quickly provide answers based on the retrieved videos, which can significantly improve the search experience. See the following for details.
[0034] Finally, for ease of understanding, an exemplary operating environment is provided below.
[0035] As Figure 1 shown, the operating environment diagram includes: a service platform 2, and clients (4A, 4B,..., 4N).
[0036] The service platform 2 can connect to the clients (4A, 4B,..., 4N) via a network.
[0037] The service platform 2 can be a single server, a server cluster, or a cloud computing service center.
[0038] The service platform 2 can provide Video storage services, video streaming services, video search services, etc.
[0039] The service platform 2 can be located in a data center such as a single location, or distributed in different geographical locations (for example, in multiple locations). The service platform 2 can provide services via a network. The network includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and combinations thereof, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.
[0040] Clients (4A, 4B, …, 4N) can be configured to access the content and services of service platform 2. The clients (4A, 4B, …, 4N) can include electronic devices with or external to a display panel, such as mobile devices, tablet devices, laptop computers, workstations, virtual reality devices, gaming devices, digital streaming devices, vehicle terminals, smart TVs, set-top boxes, etc., and can also include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load the virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices.
[0041] Clients (4A, 4B, …, 4N) can be associated with one or more users. A single user can also use one or more of the clients (4A, 4B, …, 4N) to access service platform 2. The clients (4A, 4B, …, 4N) can travel to various locations and use different networks to access service platform 2.
[0042] Clients (4A, 4B, …, 4N) can include an interface. The interface can include a touchpad, a touch screen, a mouse, a keyboard, or other sensing elements. For example, the input element can be configured to receive user instructions, and the user instructions can cause the clients (4A, 4B, …, 4N) to perform various operations, such as video search, video viewing, browsing search results (such as reply text), etc.
[0043] Note that the above devices are exemplary, and the number and types of devices can be adjusted in different scenarios or according to different requirements.
[0044] The following takes service platform 2 as the execution entity and introduces the technical solutions of this application through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.
[0045] Embodiment 1 Figure 2 A flowchart of a video search method according to Embodiment 1 of the present application is schematically shown.
[0046] As Figure 2 and Figure 13 shown, the video search method can include steps S200 to S204, where: Step S200, obtain a query text, and determine a query type based on the query text, where the query type includes: information query, or question query.
[0047] Step S202, when the query type is the information query, determine multiple first reference videos based on the query text and the meta-information associated with each of the multiple videos in the preset video library, and generate a first search result, where the first search result includes the multiple first reference videos.
[0048] Step S204, when the query type is the question query, determine multiple second reference videos based on the query text, the meta-information associated with each video among the multiple videos, and the text content associated with each video among the multiple videos; generate a second search result based on the query text and the multiple second reference videos, where the second search result includes a response text for the query text. Among them, the text content associated with the video is obtained by performing speech recognition on the video.
[0049] The video search method provided in this embodiment obtains a query text and determines its corresponding query type. When the query type is the information query, multiple first reference videos are determined based on the query text and the meta-information associated with each of the multiple videos in the preset video library, and a first search result is generated. Among them, the first search result includes multiple first reference videos. When the query type is the question query: multiple second reference videos are determined based on the query text, the meta-information associated with each of the multiple videos in the preset video library, and the text content (obtained by automatic speech recognition ASR); a second search result is generated based on the query text and the multiple second reference videos. Among them, the second search result includes a response text for the query text. It can be seen that the embodiment of the present application customizes a video search solution for different query types: quickly provides the first reference videos as the first search result based on the meta-information during the information query; introduces ASR and combines the meta-information to improve the understanding of the video content during the question query, which can improve the relevance between the second reference videos and the query text while controlling the computing cost, and generates a response text for the query text based on the second reference videos, which can provide a second search result that is more intuitive and accurate than the first search result, effectively optimizing the search experience.
[0050] The following combines Figure 2 to elaborate in detail on each step in steps S200 - S204 and other optional steps.
[0051] Step S200 Obtain a query text, and determine a query type based on the query text, where the query type includes: information query, or question query.
[0052] The client can receive the query text input by the user through interfaces (such as touchscreens, keyboards, etc.), for example: "iPhone 16", "The differences between iPhone 16 and iPhone 16 pro", etc. Based on the query text, a search request can be generated and sent to the service platform 2. After receiving the search request, the service platform 2 can obtain the query text by parsing the search request and determine the corresponding query type based on the query text. The query type can include information query and question query. Among them, the query input of the information query can be a keyword or a phrase, and the query target can be relevant documents, web pages, etc. Users tend to obtain information by browsing relevant documents and web pages. The query input of the question query can be a complete question or natural language, and the query target can be the answer corresponding to the question. Users tend to obtain direct answers rather than further reading a large number of documents. For example, query texts such as "iPhone 16", "Hongshan Zoo" are more likely that users are looking for content related to iPhone 16 and Hongshan Zoo (such as videos). Therefore, the type of such texts can be determined as an information query, and the relevant videos searched can be directly pushed to the client accordingly. However, a query text like "The differences between iPhone 16 and iPhone 16 pro" is more likely that users want to directly obtain the answer to this specific question. In this case, displaying the answer to the query text can better meet the actual needs and improve the search experience compared to displaying relevant videos. Therefore, the embodiments of this application propose to accurately identify the query type from the query text and customize corresponding video search schemes for different query types to adapt to diverse queries, thereby improving search satisfaction. Exemplarily, a pre-developed four-layer lightweight BERT model can be used to accurately identify the query type. In actual deployment, the FastText model can also be used to distill the results of the BERT model to achieve faster response and reduce the deployment resources required. In practical applications, the recognition accuracy of this model can reach more than 90%, and the computing resources occupied are relatively small.
[0053] Step S202 In the case where the query type is the information query, based on the query text and the meta-information associated with each of multiple videos in a preset video library, multiple first reference videos are determined, and a first search result is generated, where the first search result includes the multiple first reference videos.
[0054] The preset video library can be a database pre-maintained in the service platform for storing videos, which can include multiple videos. Each video can be associated with meta information. The meta information of the video can include, but is not limited to, video title, video description, video tags, release date, and / or publisher, etc. In the case where the query type is information query, the similarity between the query text and the meta information of each video in the preset video library can be calculated, and multiple videos with higher similarity are obtained from the multiple videos as the first reference videos, and a first search result is generated. The service platform can send the first search result to the client and display it through the client. As Figure 8 shown, the first search result can be a video list, and the video list corresponds to multiple first reference videos.
[0055] In this embodiment, based on the similarity between the query text and the video meta information, relevant videos can be quickly obtained and returned, which can improve the information search efficiency.
[0056] In an alternative embodiment, step S202 may include: determining the similarity between the query text and the meta information associated with each of the multiple videos; determining the videos with the similarity greater than a first preset threshold as the first reference videos.
[0057] Exemplarily, the video title, publisher identifier (ID), video tags, etc. can be obtained from the video meta information, the similarity between the query text and the video title, publisher identifier, and video tags is calculated, and the videos with the similarity higher than the first preset threshold are determined as the first reference videos.
[0058] In this embodiment, by calculating the similarity between the query text and the meta information, videos with higher relevance can be accurately screened out as search results, improving the search experience.
[0059] Step S204 , in the case where the query type is the question query, multiple second reference videos are determined based on the query text, the meta information associated with each video in the multiple videos, and the text content associated with each video in the multiple videos; a second search result is generated based on the query text and the multiple second reference videos, and the second search result includes a reply text for the query text; wherein, the text content associated with the video is obtained by performing speech recognition on the video.
[0060] The text content associated with the video can be obtained by performing automatic speech recognition (ASR) on the video. Specifically, a pre-configured speech recognition system can be used to extract the speech from the audio track of the video and convert it into text content that can be used for analysis, thereby deepening the understanding of the video content and making the video content more easily indexed, searched, and analyzed. In practical applications, the text content generated by ASR may contain recognition errors such as homophones (e.g., Yurao, Yurao), near-homophones (e.g., Jinhaishi, Jinghaishi), and entity objects (e.g., personal names, place names, brand names, etc.). The existence of these errors may lead to misunderstandings of the video content when applying the text content generated by ASR to video search, directly affecting the accuracy of the search. Therefore, to improve the search accuracy, the speech recognition process can be optimized to obtain high-quality text content. Exemplary solutions are provided below.
[0061] In an alternative embodiment, as Figure 3 shown, step S204 may include: Step S300, for each video, obtain the associated meta-information.
[0062] Step S302, input the meta-information into a first large language model that has been pre-supervised and fine-tuned, so as to extract keywords from the meta-information through the first large language model and correct the keywords.
[0063] Step S304, provide the corrected keywords to a pre-configured speech recognition system as bias words.
[0064] Step S306, perform speech recognition on the video by the speech recognition system based on the bias words to obtain the text content.
[0065] Exemplarily, the meta-information associated with the video can be obtained and input into a first large language model that has been pre-supervised and fine-tuned, so as to extract keywords from the meta-information through the first large language model, such as video titles, video tags, video descriptions, etc., identify errors in the keywords, and correct them. Among them, the errors may include homophones, near-homophones, entity objects, etc. The correction may include: replacing the incorrect homophones and near-homophones and entity objects based on a pre-configured standard word library and standard entity object library. It should be noted that the terms "first", "second", etc. in the embodiments of the present application are only used for identification, not for limitation. For example: the first large language model, the second large language model, and the third large language model may be the same large language model, or different large language models, or different versions of the same large language model obtained through different supervised fine-tuning.
[0066] After extracting and correcting keywords through the LLM, the corrected keywords can be provided to a pre-configured speech recognition system as bias words. Among them, the bias words can be used to guide or optimize the output of the speech recognition system and improve the recognition accuracy. By performing speech recognition on the video under the guidance of the bias words through the speech recognition system, high-quality text content can be obtained. For example, the bias word can be "iPhone". The speech recognition system may misrecognize "iPhone" in the audio as "eye phone". The bias word can make the speech recognition system more inclined to select the correct "iPhone" for output. For example, the speech recognition system can implement context bias by combining the WFST beam search algorithm to improve the accuracy of the text content. The algorithm can be represented as follows: ; where x is the acoustic feature (extracted from the audio track of the video), y is the text content recognized by ASR, logPC(y) is the context score from the bias word, logP(y|x) is obtained by the speech recognition system, λ is an adjustable hyperparameter, and y* represents the final score of the text content generated by ASR. The higher the final score, the higher the accuracy of the text content.
[0067] In this embodiment, the LLM that has been pre-supervised and fine-tuned is used to extract and correct keywords, which are provided to the speech recognition system as bias words. The accuracy of speech recognition is improved by the method of context bias to generate high-quality text content. Especially when recognizing homophones, near-homophones, and entity objects, a high accuracy rate can be guaranteed.
[0068] The first large language model can be supervised and fine-tuned based on the first training data. The first training data can be as Figure 4 shown. Exemplarily, training data pairs can be constructed based on the interactive data of the video (such as comments, bullet screens, and replies). The input Input--text can come from the interactive data (such as "The unreasonable part is, did Yu Rao go back and forth to Ganlu Temple so quickly?"). The output can be the possible keywords, such as: "Yu Rao", "Ganlu Temple". These data pairs are reviewed by machine or manually to ensure accuracy, and it is verified whether the output has no error characters. In the case of having error characters, the incorrect keywords are replaced with the correct near-homophones, homophones, and entity objects according to possible pinyin variants, entity objects, etc. For example, in Figure 4 it, Yu Rao yurao is replaced with Yu Rao yurao. The large language model LLM is supervised and fine-tuned with the constructed first training data pair, so that the supervised and fine-tuned LLM can mine and correct keywords from comments, bullet screens, replies, and the text content generated by ASR.
[0069] In some embodiments, the LLM can also be supervised fine-tuned (SFT) with the first training data first, and the audio of the video is converted into the original text content through ASR. The original text content may include errors such as homophones and entity objects. The LLM after supervised fine-tuning is directly used to correct the original text content to obtain higher-quality text content. In practical applications, this solution performs worse than the previous solution in terms of recall rate, precision, generality, etc. of keywords, that is, first use the LLM to extract keywords and then perform context biasing. Therefore, the previous solution can be preferably used to improve the quality of text content and further optimize the search precision.
[0070] The above-mentioned multiple embodiments exemplarily introduce how to obtain high-quality text content. Based on the query text, combined with the meta-information and text content of the video, video search can be performed on the premise of deeply understanding the video to obtain multiple second reference videos with high relevance. Based on the query text and multiple second reference videos, a second search result can be generated. The second search result may include a response text for the query text, and may also include relevant information of the second reference videos (such as: video title, video description, etc.), as Figure 5 shown.
[0071] In this embodiment, by introducing the ASR technology, the accuracy and relevance of video search in content understanding can be enhanced. The resources and computing requirements of ASR in the video are less than 10% of those required by OCR or frame embedding, and it can be widely applied, especially for science and education videos, which include a large amount of text content. Integrating ASR into video search can improve the relevance of search results and enhance search satisfaction. Based on the query text, combined with meta-information and text content for video search, second reference videos with improved relevance can be obtained, thereby providing an effective answer to the query text.
[0072] The following will further introduce step S204 in combination with multiple exemplary solutions.
[0073] In an alternative embodiment, as Figure 6 shown, step S204 may include: Step S600, enhancing the query text through a pre-trained second large language model to obtain an enhanced query text.
[0074] Step S602, processing the text content of the multiple videos through a pre-trained third large language model to generate respective summary contents for the multiple videos.
[0075] Step S604, performing sparse retrieval based on the enhanced query text and the respective summary contents of the multiple videos to obtain a first retrieval result, where the first retrieval result includes videos with word matches with the enhanced query text.
[0076] Step S606: Based on the enhanced query text and the summary content corresponding to each of the multiple videos, perform dense retrieval to obtain a second retrieval result, where the second retrieval result includes videos that have semantic matching with the enhanced query text.
[0077] Step S608: Based on the first retrieval result and the second retrieval result, determine the multiple second reference videos.
[0078] Exemplarily, the query text (query) can be enhanced by a pre-trained second large language model (LLM) to obtain an enhanced query text. Then, based on the enhanced query text and the video, a match is made to retrieve more relevant and high-quality second reference videos, improving the search results and providing high-quality references for subsequent answer generation. The following provides an exemplary solution.
[0079] In an alternative embodiment, as Figure 7 shown, step S600 may include: Step S700: Obtain the historical query text and the historical viewing video list.
[0080] Step S702: Input the historical query text, the historical viewing list, and the query text into the second large language model to generate extended query terms related to the query text through the second large language model.
[0081] Step S704: Fuse the extended query terms with the query text to obtain an extended query text.
[0082] Step S706: Perform word segmentation on the extended query text to obtain multiple sub-words, and assign corresponding weights to each sub-word.
[0083] Step S708: Based on the multiple sub-words and their respective weights, determine the enhanced query text.
[0084] Exemplarily, the historical query text of the user and the list of historical watched videos can be obtained. It should be noted that all the data involved in the embodiments of this application are obtained on the premise of compliance and with the consent of the user, and have been desensitized. The historical query text and the historical watch list can be used as prompt words to improve the understanding ability of the large language model. The prompt words and the query text are input into the second large language model together. Optionally, manually written examples can also be input into the second large language model to guide the LLM to generate extended query words related to the query text. For example, the query text can be "Beijing flowers", and the extended query words generated by the LLM can be "capital", "flowers", etc. Fusing the extended query words with the query text can obtain an extended query text. It should be noted that the fusion involved in the embodiments of this application can be simple splicing or generating new text. Through multiple tokenizers, such as Jieba, HanLP, Spacy, etc., the extended query text is tokenized respectively to obtain multiple subwords and assign corresponding weights to each subword. Based on the multiple subwords and their respective weights, an enhanced query text can be generated. Due to the performance differences of the tokenizers, for query texts with large tokenization differences among multiple tokenizers, the LLM can be used for evaluation and correction. For example: some tokenizers do not include professional terms in specific fields, resulting in splitting professional terms into individual characters, or not performing tokenization, resulting in conflicts in the tokenization results with other tokenizers. Optimizing such query texts through the LLM can increase the accuracy from 87.5% to 95.%, and this data is obtained by evaluating 160 million randomly selected query texts, while the term weight increases from 89.75% to 92.75%.
[0085] In this embodiment, by expanding the query text through the LLM and performing tokenization, the comprehensiveness of the query text can be improved, thereby enhancing the relevance and accuracy of video search.
[0086] In some embodiments, the extended query text can also be directly used as the enhanced query text for video retrieval to balance search efficiency and search precision.
[0087] LLMs can be used to enhance the understanding of query texts and also to enhance the text content generated by ASR. Since the text content generated by ASR is expressed in a relatively informal manner, with low information density and high noise, it is necessary to generate effective summary content to improve the search experience. Displaying the summary content of videos in search results can increase search trust and interactivity. Exemplarily, the video topic can be obtained from the meta-information and interaction data of the video corresponding to the text content, such as: video title, high-frequency comment words, etc. Inputting the video topic and the text content into a pre-trained LLM together can generate high-quality summary content. Further, for longer text content, a two-stage generation scheme can be adopted, for example: segmenting the text content and extracting summary content for each segment through the LLM. Integrating these summary contents together and re-inputting them into the LLM can generate the final summary content. In this embodiment, high-quality summary content can be obtained through the LLM for video retrieval, which can reduce the interference of noise on video search and improve search accuracy.
[0088] After obtaining the enhanced query text and the summary content of each video through the LLM, in the search engine, two methods of sparse retrieval and dense retrieval can be combined to recall relevant videos (second reference videos). Among them, sparse retrieval can ensure the recall accuracy, while dense retrieval can explore more semantically related videos. Video retrieval based on high-quality summary content can achieve real-time and accurate recall. The following provides an exemplary solution.
[0089] In an alternative embodiment, as Figure 9 shown, step S606 may include: Step S900, extracting multiple target elements from the summary content corresponding to each of the multiple videos through a pre-trained fourth large language model, and constructing an inverted index based on the multiple target elements and the summary content corresponding to each of the multiple videos. The inverted index is used to represent the summary content corresponding to each of the target elements.
[0090] Step S902, performing sparse retrieval based on the inverted index and the enhanced query text to obtain the first retrieval result.
[0091] Exemplarily, the target elements (core elements) such as the theme, time, etc. can be extracted from the abstract content through BERT-based UIE. By recording in which abstract contents each target element appears, an inverted index can be constructed. Based on the inverted index and the enhanced query text, sparse retrieval can be performed to obtain the first retrieval result. The first retrieval result may include a second reference video that has a word match with the enhanced query text. Different from ANN, this method does not require index reconstruction when facing new videos, thus ensuring that newly released videos can also be quickly recalled and providing stronger interpretability for rare retrievals.
[0092] In this embodiment, the real-time performance and accuracy of video retrieval are improved by combining LLM and sparse retrieval.
[0093] In an alternative embodiment, as Figure 10 shown, step S606 may include: Step S1000, obtaining the meta-information associated with each of the multiple videos.
[0094] Step S1002, fusing the meta-information and the abstract content of each video to construct a video index set.
[0095] Step S1004, inputting the video index set and the enhanced query text into a pre-trained semantic recall model to obtain the second retrieval result through the semantic recall model.
[0096] Exemplarily, the meta-information of a video can be obtained, such as: video title, video tags, etc. By fusing the abstract content with the video title and video tags, a video index set can be constructed. Inputting the video index set into a pre-trained semantic recall model can obtain the second retrieval result. Among them, the second retrieval result may include a second reference video that has a semantic match with the query text.
[0097] In this embodiment, a sufficient video index is constructed by combining the abstract content, video title, and video tags of the video. Based on the enhanced query text, comprehensive semantic recall can be performed to improve the search accuracy and relevance.
[0098] Among them, the semantic recall model can be trained using a contrastive learning framework on a pre-trained BGE model, and the contrastive learning loss formula can be expressed as follows: ; where hi and hi + respectively represent the vector of the enhanced query text and the vector representation of the video index, which are output by the BGE model. N represents the training batch size. The structure of the model is as Figure 11 shown.
[0099] In some embodiments, the set of search queries with the most video interactions can be mined into the model to improve model performance.
[0100] In some embodiments, the LLM can be trained to score and rank the relevance of each second reference video to the query text. For example, the user search interaction logs can be mined first to collect samples. The samples are the reference videos corresponding to the historical query texts. The semantic understanding ability of the LLM can also be utilized to filter the samples through prompts. Different from the number of user clicks and viewing duration, the selection is based on whether the reference video can contribute to generating the response text. In practical applications, there is a user behavior bias. The viewing duration and the number of clicks are usually dominated by popular or attractive videos, while videos based on knowledge and experience may perform poorly in terms of viewing duration and the number of clicks. Therefore, prompts can be constructed based on the contribution degree to the response text to guide the LLM to select more appropriate samples. Model training is performed based on the samples to obtain the basic model in the first stage. Then, the basic model can be distilled to optimize the model performance. For example, the MiniiRBT Chinese pre-trained model and FastTokenizer can be used to optimize the performance of the basic model to obtain the final model. In the relevance evaluation, the time for evaluating 300 video segments can be reduced from 90 milliseconds to 23 milliseconds.
[0101] Based on the first retrieval result and the second retrieval result, a set of second reference videos that match words and semantics can be obtained to get multiple second reference videos. Based on the text content of the multiple second reference videos and the query text, the required answer can be generated through the LLM. The following provides an exemplary solution.
[0102] In an alternative embodiment, as Figure 12 shown, step S204 may include: Step S1200, obtaining the text content associated with each of the multiple second reference videos, and performing chunking processing on each text content to obtain multiple text content chunks. Among them, each text content chunk is associated with video information, and the video information includes the meta information of the second reference video corresponding to the text content chunk and the time information of the text content chunk in the corresponding second reference video.
[0103] Step S1202, inputting the multiple text content chunks and the query text into a pre-trained vector retrieval model to determine a preset number of text content chunks with the highest matching degree to the query text through the vector retrieval model.
[0104] Step S1204: Input the query text, the preset number of text content chunks, and their respective associated video information into the fifth large language model that has been pre-supervised and fine-tuned, so as to obtain the second search result through the fifth large language model.
[0105] Exemplarily, after obtaining multiple second reference videos returned by the search engine, according to the relevance scores of each second reference video provided by the search engine, the top K videos most relevant to the query text can be obtained (where K can be any value, such as 5 to 10, specifically depending on the threshold of the relevance score), and the text content of these videos can be retrieved from the storage. Then, each text content is segmented into multiple text content chunks, and each text content chunk can include S Chinese characters, where S ∈ {64, 128, 256, 512, 1024}. Each text content chunk can be associated with video information, and the video information can include the meta information of the corresponding second reference video (such as the video title), and the time information of the text content chunk in the corresponding second reference video (such as which interval of the video it is located in). Through a pre-trained vector retrieval model, N text content chunks most relevant to the query text can be identified, ensuring that N * S ≤ the context window length of the LLM (such as 3200). Through experiments, it is found that using S = 256 and N = 12 in the LLM prompt can achieve the best coherence and diversity. Finally, combine the query text, these N text content chunks, and their respective associated video information to form a specific prompt, and input it into the LLM that has been pre-supervised and fine-tuned according to general dialogue data and search-based question guidance, and the final answer can be generated. In practical semantics, multiple prompts with similar meanings but different formats can be used to enhance the diversity of the answers.
[0106] In this embodiment, introducing the meta information of the video to provide more structured information for the fused text content chunks can enable the LLM to better understand the search results.
[0107] Among them, the vector retrieval model can be obtained by performing supervised fine-tuning on the BGE-base-zh-v1.5 model using a pre-configured data pair of high-quality query text and text content chunks. To generate the data pair, the query text and each text content chunk can be input into the LLM together, processed through the same prompt, and an automatic scoring system can be used to evaluate the quality of the LLM's answer. A higher answer quality can indicate a stronger correlation between the query text and the text content chunk.
[0108] Finally, the service platform can send the second search result to the client. The client can display the second search result through the interface and provide interaction options. For example: liking or disliking the answer. Collecting these user feedback data for the response text can be used for the supervised fine-tuning of the LLM or applying the DPO algorithm to learn user preferences. Some exemplary solutions are provided below.
[0109] In some embodiments, to reduce the inference cost of the LLM, a prediction model can be established based on the like and dislike data for the answers. When the LLM generates an answer, the prediction model can be used to: sort the generated answer with the cached answers and finally output the answer that the user is most likely to like.
[0110] In some embodiments, for the answers provided by the search engine, the user can like or dislike them by selecting the control. These data are collected to represent the user's satisfaction with the answer. For answers whose like or dislike ratio exceeds the set threshold, the answer along with the like or dislike data can be re-input into the LLM as a prompt to guide the LLM to generate a better answer. These data can also be added to the supervised fine-tuning training data of the LLM for data augmentation.
[0111] In some embodiments, good answers often rely on the text content they reference. Therefore, the corresponding text content can be added to the prompt of the vector retrieval model to optimize the training of the vector retrieval model.
[0112] In some embodiments, to improve the model's ability to follow instructions and align with human preferences, like and dislike data with high confidence can be selected as positive and negative samples, and the DPO algorithm can be used to maximize the likelihood difference between the positive and negative samples. The algorithm can be expressed as follows: ; where, π ref is the fine-tuned LLM model, π θ is the optimized model, x is the question, y l is the liked answer (liked), y d is the disliked answer, D is the DPO training set, and β is the hyperparameter.
[0113] Embodiment 2 Figure 14 Schematically shows a block diagram of a video search device according to Embodiment 2 of the present application. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 14As shown, the device 1400 may include: an acquisition module 1410, a first search module 1420, and a second search module 1430, where: The acquisition module 1410 is configured to acquire a query text and determine a query type based on the query text, where the query type includes: information query or question query; The first search module 1420 is configured to, when the query type is the information query, determine a plurality of first reference videos based on the query text and the meta-information associated with each of the plurality of videos in a preset video library, and generate a first search result, where the first search result includes the plurality of first reference videos; The second search module 1430 is configured to, when the query type is the question query, determine a plurality of second reference videos based on the query text, the meta-information associated with each of the plurality of videos, and the text content associated with each of the plurality of videos; generate a second search result based on the query text and the plurality of second reference videos, where the second search result includes a reply text for the query text; Wherein, the text content associated with the video is obtained by performing speech recognition on the video.
[0114] As an optional embodiment, before determining the plurality of second reference videos based on the query text, the meta-information associated with each of the plurality of videos, and the text content associated with each of the plurality of videos, the device 1400 is further configured to: Determine the similarity between the query text and the meta-information associated with each of the plurality of videos; Determine the videos with the similarity greater than a first preset threshold as the first reference videos.
[0115] As an optional embodiment, that the text content associated with the video is obtained by performing speech recognition on the video includes: Acquire the meta-information associated with the video; Input the meta-information into a pre-supervised and fine-tuned first large language model to extract keywords from the meta-information through the first large language model and correct the keywords; Provide the corrected keywords to a pre-configured speech recognition system as bias words; Perform speech recognition on the video through the speech recognition system based on the bias words to obtain the text content.
[0116] As an optional embodiment, determining the plurality of second reference videos based on the query text, the meta-information associated with each of the plurality of videos, and the text content includes: Enhance the query text through a pre-trained second large language model to obtain an enhanced query text; Process the text content of the multiple videos through a pre-trained third large language model to generate summary content corresponding to each of the multiple videos; Based on the enhanced query text and the summary content corresponding to each of the multiple videos, perform sparse retrieval to obtain a first retrieval result, where the first retrieval result includes videos with word matches with the enhanced query text; Based on the enhanced query text and the summary content corresponding to each of the multiple videos, perform dense retrieval to obtain a second retrieval result, where the second retrieval result includes videos with semantic matches with the enhanced query text; Based on the first retrieval result and the second retrieval result, determine the multiple second reference videos.
[0117] As an optional embodiment, enhancing the query text through a pre-trained second large language model to obtain an enhanced query text includes: Obtain historical query texts and a list of historical watched videos; Input the historical query text, the historical watched list, and the query text into the second large language model to generate extended query words related to the query text through the second large language model; Fuse the extended query words with the query text to obtain an extended query text; Perform word segmentation on the extended query text to obtain multiple sub-words, and assign corresponding weights to each sub-word; Based on the multiple sub-words and their corresponding weights, determine the enhanced query text.
[0118] As an optional embodiment, performing sparse retrieval based on the enhanced query text and the summary content corresponding to each of the multiple videos to obtain a first retrieval result includes: Extract multiple target elements from the summary content corresponding to each of the multiple videos through a pre-trained fourth large language model, and construct an inverted index based on the multiple target elements and the summary content corresponding to each of the multiple videos, where the inverted index is used to represent the summary content corresponding to each target element; Perform sparse retrieval based on the inverted index and the enhanced query text to obtain the first retrieval result.
[0119] As an optional embodiment, performing dense retrieval based on the enhanced query text and the summary content corresponding to each of the multiple videos to obtain a second retrieval result includes: For each video, obtain associated meta-information; Fuse the meta-information and summary content of each video to construct a video index set; Input the video index set and the enhanced query text into a pre-trained semantic recall model to obtain the second retrieval result through the semantic recall model.
[0120] As an optional embodiment, generating a second search result based on the query text and the multiple second reference videos includes: Obtain the text content associated with each of the multiple second reference videos, and perform chunking processing on each text content to obtain a plurality of text content chunks; wherein, each text content chunk is associated with video information, and the video information includes the meta information of the second reference video corresponding to the text content chunk and the time information of the text content chunk in the corresponding second reference video; Input the plurality of text content chunks and the query text into a pre-trained vector retrieval model to determine a preset number of text content chunks with the highest degree of matching with the query text through the vector retrieval model; Input the query text, the preset number of text content chunks, and their respective associated video information into a fifth large language model that has been pre-supervised and fine-tuned to obtain the second search result through the fifth large language model.
[0121] Embodiment III Figure 15 Schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the video search method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 15 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the video search method. In addition, the memory 10010 may also be used to temporarily store various types of data that have been output or will be output.
[0122] In some embodiments, the processor 10020 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other chip. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0123] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0124] It should be noted that Figure 15 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0125] In this embodiment, the video search method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0126] Embodiment 4 The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video search method in the embodiments are implemented.
[0127] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (e.g., SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the video search method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.
[0128] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.
[0129] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0130] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A video search method, characterized in that, The method includes: Obtain a query text, and determine a query type based on the query text, where the query type includes: information query or question query; When the query type is the information query, determine multiple first reference videos based on the query text and the meta-information associated with each of multiple videos in a preset video library, and generate a first search result, where the first search result includes the multiple first reference videos; When the query type is the question query, determine multiple second reference videos based on the query text, the meta-information associated with each of the multiple videos, and the text content associated with each of the multiple videos; generate a second search result based on the query text and the multiple second reference videos, where the second search result includes a reply text for the query text; Wherein, the text content associated with the video is obtained by performing speech recognition on the video.
2. The method according to claim 1, wherein Determining multiple first reference videos based on the query text and the meta-information associated with each of multiple videos in a preset video library includes: Determine the similarity between the query text and the meta-information associated with each of the multiple videos; Determine the videos with the similarity greater than a first preset threshold as the first reference videos.
3. The method according to claim 1, characterized in that, Before determining multiple second reference videos based on the query text, the meta-information associated with each of the multiple videos, and the text content associated with each of the multiple videos, the method further includes: For each video, obtain the associated meta-information; Input the meta-information into a pre-supervised and fine-tuned first large language model to extract keywords from the meta-information through the first large language model and correct the keywords; Provide the corrected keywords to a pre-configured speech recognition system as bias words; Perform speech recognition on the video through the speech recognition system based on the bias words to obtain the text content.
4. The method according to claim 1, wherein Determining multiple second reference videos based on the query text, the meta-information associated with each of the multiple videos, and the text content associated with each of the multiple videos includes: Enhance the query text through a pre-trained second large language model to obtain an enhanced query text; Process the text content of the multiple videos through a pre-trained third large language model to generate corresponding summary content for each of the multiple videos; Perform sparse retrieval based on the enhanced query text and the corresponding summary content of each of the multiple videos to obtain a first retrieval result, where the first retrieval result includes videos with a word match with the enhanced query text; Perform dense retrieval based on the enhanced query text and the corresponding summary content of each of the multiple videos to obtain a second retrieval result, where the second retrieval result includes videos with a semantic match with the enhanced query text; Determine the multiple second reference videos based on the first retrieval result and the second retrieval result.
5. The method according to claim 4, wherein Enhancing the query text through a pre-trained second large language model to obtain an enhanced query text includes: Obtain historical query texts and a list of historical watched videos; Input the historical query text, the historical viewing list, and the query text into the second large language model to generate extended query terms related to the query text through the second large language model; Fuse the extended query terms with the query text to obtain an extended query text; Perform word segmentation on the extended query text to obtain multiple sub-words, and assign corresponding weights to each sub-word; Based on the multiple sub-words and their respective corresponding weights, determine the enhanced query text.
6. The method according to claim 4, wherein Based on the enhanced query text and the respective summary contents of the multiple videos, perform sparse retrieval to obtain a first retrieval result, including: Extract multiple target elements from the respective summary contents of the multiple videos through a pre-trained fourth large language model, and construct an inverted index based on the multiple target elements and the respective summary contents of the multiple videos, where the inverted index is used to represent the summary content corresponding to each target element; Based on the inverted index and the enhanced query text, perform sparse retrieval to obtain the first retrieval result.
7. The method according to claim 4, wherein Based on the enhanced query text and the respective summary contents of the multiple videos, perform dense retrieval to obtain a second retrieval result, including: Obtain the meta-information associated with each of the multiple videos; Fuse the meta-information and the summary content of each video to construct a video index set; Input the video index set and the enhanced query text into a pre-trained semantic recall model to obtain the second retrieval result through the semantic recall model.
8. The method according to claim 1, characterized in that Generate a second search result based on the query text and the multiple second reference videos, including: Obtain the text content associated with each of the multiple second reference videos, and perform chunking on each text content to obtain multiple text content chunks; where each text content chunk is associated with video information, and the video information includes the meta-information of the second reference video corresponding to the text content chunk and the time information of the text content chunk in the corresponding second reference video; Input the multiple text content chunks and the query text into a pre-trained vector retrieval model to determine a preset number of text content chunks with the highest matching degree to the query text through the vector retrieval model; Input the query text, the preset number of text content chunks, and their respective associated video information into a fifth large language model that has been pre-supervised and fine-tuned to obtain the second search result through the fifth large language model.
9. A video search device, characterized in that, The device includes: An acquisition module, configured to acquire a query text, and determine a query type based on the query text, where the query type includes: information query, or question query; A first search module, configured to, when the query type is the information query, determine multiple first reference videos based on the query text and the meta-information associated with multiple videos in a preset video library, and generate a first search result, where the first search result includes the multiple first reference videos; A second search module, configured to, when the query type is the question query, determine a plurality of second reference videos based on the query text, meta information associated with each video among the plurality of videos, and text content associated with each video among the plurality of videos; and generate a second search result based on the query text and the plurality of second reference videos, where the second search result includes a reply text for the query text. Among them, the text content associated with the video is obtained by performing speech recognition on the video.
10. A computer device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; where: The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to claims 1 to 8 are implemented.