Video query method and computing device

By obtaining the query text, extracting keywords and generating a collection of prompt words, combined with the video search model, the problem of inaccurate search in the existing video surveillance system is solved, and more accurate video content retrieval is achieved.

CN120296201APending Publication Date: 2025-07-11HENAN KUNLUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228115.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the existing video surveillance system, video retrieval relies on manually inputting search conditions, resulting in inaccurate or omission of search results, especially in complex and changeable surveillance video content, which is difficult to accurately retrieve user needs.

Method used

By obtaining the query text, extracting the query keywords, generating a collection of prompt words, and using the video search model to understand the user's intentions, the video content is accurately retrieved, and combined with speech recognition and natural language processing technology, video content that meets user needs is selected.

Benefits of technology

It improves the accuracy and efficiency of video retrieval, can better understand user search intentions, reduce false detection and missed detection, and provide search results that are more in line with user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296201A_ABST
    Figure CN120296201A_ABST
Patent Text Reader

Abstract

The invention provides a video query method and computing equipment. The video query method comprises the following steps: obtaining a query text; wherein the query text is used for indicating a query demand of the user for the queried video; carrying out keyword extraction on the query text, and determining M (a positive integer greater than or equal to 1) query keywords; obtaining N (a positive integer greater than or equal to 1) candidate video contents based on the M query keywords; each candidate video content is a video clip or a frame image in the queried video, and each candidate video content comprises features matched with M query keywords; generating a prompt word set according to the M query keywords and the N candidate video contents; and determining a target video content from the N candidate video contents based on the cue word set and a video search model. Through combination of keyword query and prompt word video search, the search intention of the user can be more accurately understood, so that the video content better meeting the user demand can be searched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video image processing technology, and in particular, to a video query method and a computing device. Background Art

[0002] With the acceleration of the urbanization process and the improvement of public safety awareness, video surveillance systems have been widely used in various public places such as communities, shopping malls, buildings, and stations. Video surveillance not only provides convenience for daily management but also is an important means to maintain public safety. However, in actual applications, the retrieval of surveillance videos faces a series of challenges, especially for front-line operators, these problems are particularly prominent.

[0003] Currently, video retrieval mainly relies on basic information such as query keywords, timestamps, and camera positions for retrieval.

[0004] However, this method requires operators to manually input retrieval conditions, and then the system filters out video segments that meet the conditions according to the conditions. However, due to the complex and variable nature and ambiguity of surveillance video content, the retrieval results may be inaccurate or missed. Summary of the Invention

[0005] Embodiments of this application provide a video query method and a computing device. By querying keywords, video content is pre-screened, and subsequently, based on the query keywords and the video content screened by the query keywords, a set of prompt words for query is obtained. Based on the set of prompt words and a video search model, the user's search intention can be more accurately understood, thereby searching for video content that better meets the user's needs.

[0006] In a first aspect, embodiments of this application provide a video query method, and the method includes:

[0007] Obtain a query text; wherein, the query text is used to indicate the user's query requirement for the video to be queried; extract keywords from the query text to determine M query keywords; wherein, M is a positive integer greater than or equal to 1; based on the M query keywords, obtain N candidate video contents; wherein, N is a positive integer greater than or equal to 1; each candidate video content is a video segment or a frame image in the video to be queried, and each candidate video content includes features matching the M query keywords; generate a set of prompt words according to the M query keywords and the N candidate video contents; determine the target video content from the N candidate video contents based on the set of prompt words and a video search model.

[0008] In this solution, video content is pre-screened by querying keywords. Subsequently, based on the query keywords and the video content screened by the query keywords, a set of prompt words for querying is obtained. Based on the set of prompt words and the video search model, the user's search intent can be understood more accurately, thereby searching for video content that better meets the user's needs.

[0009] In a possible implementation manner, obtaining the query text includes: obtaining the query voice input by the user; converting the query voice into the query text.

[0010] In this solution, considering the arbitrary nature of voice, combining keyword query and video search with prompt words can more accurately understand the user's search intent, thereby searching for video content that better meets the user's needs.

[0011] In a possible implementation manner, extracting keywords from the query text to determine M query keywords includes: performing a word segmentation operation on the query text to obtain the first word segmentation; filtering out the stop words in the first word segmentation to obtain the second word segmentation; determining the word segments with semantic roles in the second word segmentation; and determining M query keywords based on the word segments with semantic roles in the second word segmentation.

[0012] In this solution, by using word segmentation operations, removing stop words, semantic roles, etc., keywords with relatively high reference value can be extracted from the text.

[0013] In a possible implementation manner, the query text includes words with vague semantics related to the target object; the M query keywords include the first query keywords corresponding to the words with vague semantics.

[0014] Extracting keywords from the query text to determine M query keywords includes: obtaining the scene layout data of the video; determining the range data related to the target object based on the scene layout data; and converting the words with vague semantics based on the range data to obtain the first query keywords.

[0015] In this solution, the range related to the object can be analyzed through the scene layout data, thereby converting the words with vague semantics to obtain query keywords with clear semantics.

[0016] In a possible implementation manner, based on the M query keywords, obtaining N candidate video contents includes:

[0017] Match M query keywords with P features of the video to be queried to determine the matching features; P features, where P features include multiple features of each of the W video contents in the video to be queried, and multiple features of each video content include features of multiple modalities, and multiple modalities include vision, text, sound, and / or time, and W and P are positive integers greater than or equal to 2; based on the matching features, determine a query vector; match the query vector with the feature vectors corresponding to each video content to determine N matching feature vectors; the feature vectors are used to indicate the semantics of the corresponding video content, and the feature vectors are constructed based on multiple features of the corresponding video content; use the N video contents corresponding to the N feature vectors as candidate video contents.

[0018] In this solution, a query vector is constructed by using the method of keyword and feature matching, and candidate video contents are screened out by comparing the query vector with the feature vectors of the video contents.

[0019] In an example of this implementation manner, the P features include object features and scene relationship features, where the object features are used to indicate the self-attributes of the objects in the video to be queried, and the scene relationship features are used to indicate the association relationships between different objects in the video to be queried;

[0020] Before matching the M query keywords with the P video features of the video to be queried, the method further includes: extracting the P features based on a hybrid model; the hybrid model includes a first model and a second model, where the first model is used to extract object features; the second model is used to extract scene relationship features.

[0021] In an example of this implementation manner, matching the query vector with the feature vectors corresponding to each video content to determine N matching feature vectors includes:

[0022] Calculate the correlation scores between the M query keywords and each of the W feature vectors; based on the correlation scores of each feature vector, determine the top N most relevant feature vectors as the N matching feature vectors.

[0023] In this solution, screening is achieved by calculating the correlation between vectors.

[0024] In an example of this implementation manner, the method further includes:

[0025] In the case where the correlation scores of at least some of the N feature vectors are the same, sort the N candidate video contents based on the chronological order to obtain the serial numbers of each candidate video content.

[0026] In this solution, in the case where the correlation scores are relevant, the candidate video contents can be sorted based on the chronological order to facilitate presenting the importance of the video contents to the user.

[0027] In a possible implementation, based on M query keywords and N candidate video contents, a set of prompt words is generated, including: matching the M query keywords with the semantic prompt words in the semantic prompt word set, and combining the matched query keywords and semantic prompt words to obtain the prompt words for each of the M query keywords; combining the N candidate video segments with the prompt words for each of the M query keywords to obtain the set of prompt words.

[0028] In this solution, by adding semantics to each query keyword to obtain prompt words, it is convenient to understand the user's query intention. Subsequently, taking the candidate video segments as the information to be verified for the user's query intention, and combining with the prompt words for each query keyword, a set of prompt words is obtained.

[0029] In a second aspect, an embodiment of the present application provides a video query device. The video query device includes several modules, and each module is used to execute each step in the video query method provided in the first aspect of the embodiment of the present application. The division of the modules is not limited herein. For the specific functions executed by each module of the video query device and the beneficial effects achieved, please refer to the functions of each step in the video query method provided in the first aspect of the embodiment of the present application, which will not be elaborated herein.

[0030] Exemplarily, the video query device includes:

[0031] A text acquisition module, configured to acquire a query text; wherein, the query text is used to indicate the user's query requirement for the video to be queried;

[0032] An extraction module, configured to extract keywords from the query text to determine M query keywords; wherein, M is a positive integer greater than or equal to 1;

[0033] A first query module, configured to obtain N candidate video contents based on the M query keywords; wherein, N is a positive integer greater than or equal to 1; each candidate video content is a video segment or a frame image in the video to be queried, and each candidate video content includes features matching the M query keywords;

[0034] A prompt word determination module, configured to generate a set of prompt words according to the M query keywords and the N candidate video contents;

[0035] A second query module, configured to determine the target video content from the N candidate video contents based on the set of prompt words and the video search model.

[0036] In a possible implementation, the text acquisition module is configured to acquire the query voice input by the user; and convert the query voice into a query text.

[0037] In a possible implementation, an extraction module is configured to perform word segmentation on a query text to obtain first word segments; filter out stop words in the first word segments to obtain second word segments; determine word segments with semantic roles in the second word segments; and determine M query keywords based on the word segments with semantic roles in the second word segments.

[0038] In a possible implementation, the query text includes words with semantic ambiguity related to a target object; the M query keywords include first query keywords corresponding to the words with semantic ambiguity.

[0039] The extraction module is configured to obtain scene layout data of a video; determine range data related to the target object based on the scene layout data; and convert the words with semantic ambiguity based on the range data to obtain first query keywords.

[0040] In a possible implementation, a first query module is configured to match the M query keywords with P features of a video to be queried to determine matching features; the P features include multiple features of each of the W video contents in the video to be queried, and the multiple features of each video content include features of multiple modalities, and the multiple modalities include vision, text, sound, and / or time, where W and P are positive integers greater than or equal to 2; determine a query vector based on the matching features; match the query vector with feature vectors corresponding to each video content to determine N matching feature vectors; the feature vectors are used to indicate the semantics of the corresponding video contents, and the feature vectors are constructed based on the multiple features of the corresponding video contents; and use the N video contents corresponding to the N feature vectors as candidate video contents.

[0041] In a possible implementation, the P features include object features and scene relationship features, where the object features are used to indicate the self-attributes of the objects in the video to be queried, and the scene relationship features are used to indicate the association relationships between different objects in the video to be queried.

[0042] The first query module is configured to extract the P features based on a hybrid model; the hybrid model includes a first model and a second model, where the first model is used to extract object features; and the second model is used to extract scene relationship features.

[0043] In an example of this implementation, the first query module is configured to calculate the correlation scores between the M query keywords and each of the W feature vectors; and determine the top N most relevant feature vectors as the N matching feature vectors based on the correlation scores of each feature vector.

[0044] In an example of this implementation manner, the first query module is configured to sort the N candidate video contents based on the chronological order to obtain the serial number of each candidate video content when the correlation scores of at least some of the N feature vectors are the same.

[0045] In a possible implementation manner, the prompt word determination module is configured to match the M query keywords with the semantic prompt words in the semantic prompt word set, combine the matched query keywords and semantic prompt words to obtain the prompt words corresponding to each of the M query keywords; and combine the N candidate video segments with the prompt words corresponding to each of the M query keywords to obtain a set of prompt words.

[0046] In a third aspect, an embodiment of the present application provides a computing device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the method provided in the first aspect.

[0047] In a fourth aspect, an embodiment of the present application provides a computer storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is caused to execute the method provided in the first aspect.

[0048] In a fifth aspect, an embodiment of the present application provides a computer program product including instructions, and when the instructions are run on a computer, the computer is caused to execute the method provided in the first aspect. Description of the Drawings

[0049] Figure 1 is a system architecture diagram of a video query system provided by an embodiment of the present application;

[0050] Figure 2 is a flowchart of a video query method provided by an embodiment of the present application;

[0051] Figure 3 is a schematic diagram of a video query scenario provided by an embodiment of the present application;

[0052] Figure 4 is a schematic diagram of a text processing scenario provided by an embodiment of the present application;

[0053] Figure 5 is a schematic diagram of a feature extraction scenario provided by an embodiment of the present application;

[0054] Figure 6 is a schematic diagram of a video content and feature storage scenario provided by an embodiment of the present application;

[0055] Figure 7 is a schematic diagram of a feature vectorization scenario provided by an embodiment of the present application;

[0056] Figure 8 It is a schematic diagram of another video content and feature storage scenario provided by an embodiment of the present application;

[0057] Figure 9 It is a schematic diagram of yet another video content and feature storage scenario provided by an embodiment of the present application;

[0058] Figure 10 It is an example diagram of a video search scenario based on a prompt word provided by an embodiment of the present application;

[0059] Figure 11 It is a schematic structural diagram of a video query device provided by an embodiment of the present application;

[0060] Figure 12 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0061] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0062] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for instance" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.

[0063] In the description of the embodiments of the present application, the term "and / or" only describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and both A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, multiple systems refer to two or more systems, and multiple terminals refer to two or more terminals.

[0064] In addition, the terms "first" and "second" are only used for indicative purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0065] The following explains some terms in this embodiment. It should be noted that these explanations are for the convenience of those skilled in the art and do not limit the scope of protection required by this application.

[0066] Video preprocessing: This refers to the process of analyzing and processing the original video before video retrieval. It includes steps such as video analysis and feature extraction, scene relationship feature extraction, and vector library construction.

[0067] Object feature extraction: This refers to extracting various features of objects from the video, such as size, color, shape, position, etc.

[0068] Scene relationship feature extraction: This refers to extracting the relationship features between objects from the video, such as the distance and relative position between objects.

[0069] Vector library: This is a database that stores vector representations. In this application, each vector represents the features of a video segment or key frame.

[0070] Speech recognition: This is a technology that converts human speech into text. In this application, the user can input query conditions through speech, and the system uses speech recognition technology to convert the speech into text.

[0071] Natural language processing: This is a technology for processing human languages (such as English, Chinese, etc.), including tasks such as word segmentation, part-of-speech tagging, and named entity recognition. In this application, the system uses natural language processing technology to preprocess the text input by the user.

[0072] Retrieval-augmented generation (RAG): Combines language models and information retrieval technologies. Specifically, when the model needs to generate text or answer questions, it first retrieves relevant information from a large collection of documents, and then uses this retrieved information to guide the generation of text, thereby improving the quality and accuracy of the prediction. In this application, the system uses the RAG model for in-depth retrieval and returns the most relevant video segments.

[0073] Multimodal vector representation: This is a method of integrating various types of information into the same vector representation. In this application, in addition to visual features, sound information, time information, etc. are also incorporated into the vector representation.

[0074] Hierarchical vector library structure: This is a vector library structure that stores features of different types or importance at different levels. In this application, for example, it is first divided by mall area, then by object category, and finally by more detailed features.

[0075] spaCy: A Python library for efficiently processing natural language tasks such as tokenization, part-of-speech tagging, dependency parsing, and named entity recognition. It supports multiple languages, provides a simple and easy-to-use API, and allows users to customize models.

[0076] Stop words: Refers to some common words that are automatically filtered out in information retrieval to save storage space and improve search efficiency.

[0077] Predicate: The core word in a sentence, usually a verb, representing an action, event, or state. SRL revolves around the predicate and analyzes the associated arguments.

[0078] Argument: The participants in an action or event, which can be noun phrases, pronouns, clauses, etc. The goal of SRL is to assign a semantic label to each argument.

[0079] Semantic labels: Markers used to indicate the roles played by arguments in the predicate's behavior, such as "agent", "patient", "time", "location", etc.

[0080] Semantic role labeling (SRL): A natural language processing technique mainly used to analyze the relationship between predicates (usually verbs) and their arguments in a sentence, revealing the structure of the event in the sentence and the relationship between grammatical components, thereby enhancing the understanding of the sentence's semantics. Its main task is, given a sentence and one or more predicates, to identify the arguments associated with these predicates and assign appropriate semantic role labels to them. These labels include core semantic roles (such as agent, patient, etc.) and adjunct semantic roles (such as time, location, manner, reason, etc.).

[0081] Large model: Refers to "large-parameter" models trained using large-scale data and powerful computing capabilities. These models usually have a high degree of generality and generalization ability.

[0082] Next, the video query system applied by the video query method provided in the embodiments of the present application will be introduced. Figure 1 Shows an architecture example diagram of a video query system provided in the embodiments of the present application. The video query method provided in the embodiments of the present application can be applied to a system architecture diagram as shown in Figure 1 the following figure. As shown in Figure 1As shown in the figure, the video query system includes a terminal 110, a camera 120, a data storage device 130, and a video retrieval device 140. Among them, the terminal 110 communicates with the video retrieval service device 140 through a network, and the video retrieval device 140 communicates with the camera 120 and the data storage device 130 through the network respectively. The network can be a wired network and / or a wireless network. It can be understood that the network can use any known network communication protocol to achieve different communications, and the above network communication protocols can be various wired and / or wireless communication protocols. The network can use any known network communication protocol to achieve different communications, and the above network communication protocols can be various wired communication protocols.

[0083] It should be noted that the number of cameras 120 can be one or more. The embodiments of the present application do not make specific limitations on this, and the number of cameras 120 can be designed according to actual needs. The number of data storage devices 130 and video retrieval devices 140 can also be one or more.

[0084] Among them, the terminal 110 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. Exemplary embodiments of the terminal devices involved in this solution include but are not limited to electronic devices equipped with iOS (iPhone operating system), android, Windows, Harmony OS, or other operating systems. The embodiments of the present application do not make specific limitations on the types of electronic devices.

[0085] Among them, the data storage device 130 is used to store the video data collected by the camera 120. The data storage device 130 can be implemented by an independent computing device or a device cluster composed of multiple computing devices. In an optional embodiment, the computing device can be a server. Among them, the server involved in this solution can be a hardware server or can be implanted into a virtualized environment. For example, the server involved in this solution can be a virtual machine running on a hardware server including one or more other virtual machines. In addition, the server involved in this solution can be used to provide cloud services, and it can be a server that can establish a communication connection with other devices and can provide computing functions and / or storage functions for other devices.

[0086] Among them, the video retrieval device 140 can query the videos stored in the data storage device 130 based on the video query request of the terminal 110. The video detection device 140 can be implemented by an independent computing device or a device cluster composed of multiple computing devices. In an alternative embodiment, the computing device can be a server. Among them, the server involved in this solution can be a hardware server or can be implanted into a virtualized environment. For example, the server involved in this solution can be a virtual machine running on a hardware server including one or more other virtual machines. In addition, the server involved in this solution can be used to provide cloud services, which can be a server that can establish a communication connection with other devices and can provide computing functions and / or storage functions for other devices.

[0087] It should be noted that the data storage device 130 and the video retrieval device 140 can be the same computing device.

[0088] Next, in combination with the video query system provided above, a video query method provided in an embodiment of the present application will be introduced in detail.

[0089] Figure 2 It is a schematic flowchart of the video query method provided in an embodiment of the present application. This embodiment can be applied to a computing device, specifically, it can be applied to the video retrieval device 140. As Figure 2 shown, the video query method provided in an embodiment of the present application at least includes the following steps:

[0090] Step 201, the video retrieval device 140 obtains a query text, and the query text is used to indicate the user's query requirements for the video to be queried.

[0091] In some possible implementation manners of this embodiment, the terminal 110 sends a video query request to the video retrieval device 140, and the video query request includes query content.

[0092] Optionally, the terminal 110 may include query software, and the user logs in to the query software to implement the query of the video to be queried. The video to be queried can be one or more; among them, the query software can be database software, and the database software is used to access multiple videos stored in the database, and each video is a video to be queried.

[0093] Optionally, as Figure 3 shown, the terminal 110 has a microphone, and the terminal 110 can collect the user's voice information through the microphone to obtain voice data, use the voice data as query content to obtain a video query request, and send the video query request to the video retrieval device 140.

[0094] Optionally, the user can input text through the terminal 110, use the text as the query content to obtain a video query request, and send the video query request to the video retrieval device 140.

[0095] Next, the video retrieval device 140 obtains the query text based on the query content.

[0096] Optionally, if the query content is voice data, the video retrieval device 140 converts the voice data into a query text based on speech recognition technology.

[0097] Optionally, if the query content is text, the video retrieval device 140 uses the query content as the query text.

[0098] Step 202: The video retrieval device 140 extracts keywords from the query text to obtain M query keywords, where M is a positive integer greater than or equal to 1.

[0099] In some optional implementation manners of this embodiment, as Figure 4 shown, the video retrieval device 140 extracts keywords from the query text through natural language processing technology. The keyword extraction may include word segmentation operation, stop word removal processing, semantic normalization processing, and semantic role-based information filtering, so as to remove meaningless or ambiguous content in the query text and retain the descriptions related to video retrieval, obtaining M query keywords. For example, one query keyword or multiple query keywords. It should be noted that the semantic normalization processing and / or the semantic role-based information filtering are optional steps, which can be added or deleted according to actual needs.

[0100] Among them, the word segmentation operation may include the following content:

[0101] When using word segmentation such as SpaCy, more detailed rules can be adopted. For example, for location-related words, in addition to simply splitting according to the conventional word boundaries, the special layout of the mall can be further subdivided. For example, "mall first floor" can be further subdivided into "mall" and "first floor", so that it can be more accurately matched with the video content information in the video in subsequent processing.

[0102] For compound words, such as "handbag", it can be marked as an integral noun phrase, so that in information filtering, it can be processed as an important semantic unit and will not be mis-split or misjudged.

[0103] Perform special word segmentation on scene-specific vocabulary. Exemplarily, in a mall scene, perform special word segmentation on vocabulary related to promotions and product categories. For example, "buy one get one free" can be split into two meaningful words: "buy one" and "get one free". In a residential community scene, perform precise word segmentation on vocabulary related to property management and community facilities (such as "property fee", "fitness equipment", etc.). In a station scene, perform special word segmentation and tagging on vocabulary related to tickets and train numbers (such as "bullet train ticket", "Train G123", etc.).

[0104] Among them, the processing of removing stop words can include the following content:

[0105] In addition to the regular removal of stop words such as "I", "a", "about" that are of little significance to retrieval, a more targeted stop word list can also be established according to specific scenarios, such as the retrieval scenario of mall surveillance videos. For example, some words that are commonly present in the mall scene but are useless for specific retrieval, like "in the mall" (because other words already contain the general information of the mall) can also be added to the stop word list.

[0106] Among them, the processing of semantic normalization can include the following content:

[0107] For words with similar semantics, in addition to simple simplification, such as simplifying "red-colored" to "red", more in-depth semantic analysis can also be carried out. For example, for the description of colors, similar color words can be unified under a standard color classification. Words like "orange-red" and "scarlet" can both be normalized to "red", and a color code can be assigned to these colors for more efficient comparison during retrieval.

[0108] Among them, the information filtering based on semantic roles can include the following content:

[0109] The semantic role labeling (SRL) technology can be adopted. By performing SRL analysis, semantic roles such as the agent, patient, location, and time in the sentence can be identified. For example, in the text "A red handbag near the elevator on the first floor of the mall", "red handbag" is the patient, and "on the first floor of the mall" and "near the elevator" are locations. Then, according to these semantic roles, some words or phrases that do not meet the semantic role requirements of mall surveillance video retrieval can be filtered out.

[0110] For some words indicating positional relationships, such as "near", it can be converted into a more specific range description according to the actual layout data of the mall. For example, if the area within a radius of 10 meters around the elevator on the first floor of the mall is the key area that is often monitored, "near the elevator" can be converted to "within a radius of 10 meters around the elevator".

[0111] Optionally, in some implementations, the video retrieval device 140 extracts keywords from the query text to determine M query keywords, including: performing a word segmentation operation on the query text to obtain a first word segmentation result; then, referring to the stop word removal process described above, for example, using a stop word list, to filter out the stop words in the first word segmentation result to obtain a second word segmentation result; next, determining each word in the second word segmentation result that has a semantic role, for example, the semantic role can be an agent, a patient, a location, a time, etc.; subsequently, based on each word in the second word segmentation result that has a semantic role, extracting the query keywords in the second word segmentation result to obtain M query keywords. For example, each word with a semantic role is used as a query keyword, and the related words of the word with a semantic role are used as query keywords. For example, the word with the semantic role of patient is "handbag", and the related words can be the size or color of the handbag.

[0112] Optionally, in some implementations, the query text includes words of a semantic model related to the target object. For example, the word is used to indicate the positional relationship with the target object. For example, the word can be a word describing a general position such as "nearby" or "around"; the M query keywords include the first query keywords corresponding to the words with fuzzy semantics; the video retrieval device 140 preprocesses the query text to determine the M query keywords, which may include: obtaining scene layout information, where the scene layout information is used to indicate the positions of the target object and other objects. For example, it can be a scene three-dimensional model, or it can be a scene picture; then, based on the scene layout information, determining the range data related to the target object, where the range data can be used to describe the surrounding range of the target object. For example, the surrounding range can be a circle with a radius of r, and r can be determined based on the minimum distance between the target object and other objects. For example, it can be n times the minimum distance, where n is greater than or equal to 1. Additionally, the range data can also be used to describe the surrounding range of the target object that the user can accept. Exemplarily, the surrounding range can be determined by combining the number of surrounding objects of the target object in the scene. If there are many surrounding objects, the surrounding range is small; if there are few surrounding objects, the surrounding range is large; then, based on the range data of the target object, converting the words with fuzzy semantics to obtain the first query keywords, where the first query keywords can be the combination of the target object and the range data. For example, the location word segmentation is "near the elevator", and the target object is the elevator, then the range data of the elevator can be within a radius of 10 meters around the elevator, and "near the elevator" is converted to "within a radius of 10 meters around the elevator".

[0113] Optionally, the word with fuzzy semantics can be a word segmentation with the semantic role of location (for the convenience of description and distinction, it can be called a location word segmentation), and the location word segmentation is used to indicate the positional relationship with the target object. For example, the location word segmentation is "near the elevator".

[0114] Optionally, in a possible scenario, the user inputs the query content by voice: "I lost a red handbag on the first floor of the mall, probably near the elevator." The video retrieval device 140 uses speech recognition technology to convert the query content input by voice into a text sequence, obtaining the original text "I lost a red handbag on the first floor of the mall, probably near the elevator."

[0115] Next, the video retrieval device 140 performs a word segmentation operation on the original text. For example, the sentence is divided into words such as "I", "on", "the first floor of the mall", "lost", "a", "red", "handbag", "probably", "near", "the elevator".

[0116] Then, stop word removal is performed to remove stop words such as "I", "a", "probably", etc. that have little significance for retrieval.

[0117] Finally, normalization processing is performed on some words with similar semantics. For example, "red" is simplified to "red", obtaining query keywords such as "red handbag, the first floor of the mall, near the elevator".

[0118] Optionally, in another possible scenario, the user inputs the query content by voice: "I lost a red handbag with a certain brand logo near a certain brand store on the first floor of the mall. I last saw it next to the elevator entrance. I want to see the videos after I left, preferably the videos during the promotional activities." The video retrieval device 140 uses speech recognition technology to convert the query content input by voice into a text sequence. The video retrieval device 140 uses natural language processing technology to deeply understand the semantics of the text sequence and perform preprocessing.

[0119] First, a word segmentation operation is performed, marking words such as "a certain brand store" and "a certain brand logo" as a whole vocabulary, and at the same time treating "during the promotional activities" as a special vocabulary unit.

[0120] Special vocabulary marking:

[0121] A certain brand store → [a certain brand store]

[0122] A certain brand logo → [a certain brand logo]

[0123] During the promotional activities → [during the promotional activities]

[0124]

[0125] When performing stop word removal, according to the stop word list, remove stop words such as "I" and "want to see" that have little significance for retrieval, and retain key information such as "a certain brand", "red", "handbag", "next to the elevator entrance", "the first floor", "during the promotional activities", etc. Exemplarily, the stop word list is as follows:

[0126] I, of, the, in, is, it, a, want, to see, behind, me

[0127]

[0128] Next, perform semantic normalization. Transform "next to the elevator entrance" into an area centered on the elevator with a radius of 5 meters according to the layout map of the shopping mall, and unify "a certain brand logo" into the standard logo name of that brand (e.g., "Nike_Logo").

[0129]

[0130] For information filtering based on semantic roles, identify "handbag" as the patient, "on the first floor", "near a certain brand store", "next to the elevator entrance" as locations, and "during the promotion period" as the time. Filter out the words that do not meet the requirements according to these semantic roles, and retain the words related to semantic roles.

[0131]

[0132]

[0133] Filtered result: Shopping mall / Floor 1 / [Nike store] / near / red / has / [Nike_Logo] / handbag / , / [area centered on the elevator with a radius of 5 meters] / , / [during the promotion period] / video.

[0134] Step 203: The video retrieval device 140 obtains N candidate video contents based on M query keywords. Each candidate video content is a video segment or frame image in the queried video, and each candidate video content includes features matching the M query keywords. N is a positive integer greater than or equal to 1.

[0135] In some optional implementation manners of this embodiment, the data storage device 130 is used to store P (a positive integer greater than or equal to 2) features in the queried video. The queried video is divided into W (greater than or equal to 2) video contents. The P features include the features of each of the W video contents, so that the video retrieval device 140 queries the data storage device 130 using the P features and obtains N candidate video contents with features matching the M query keywords. Among them, the features matching the M query keywords can be understood as features semantically the same or similar to the M query keywords. Each feature matching a query keyword can have one or more. Exemplarily, for each video content of the video to be queried, the video retrieval device 140 determines whether the multiple features of the video content match the M query keywords. If so, it is considered that the multiple features of the video content include the features matching the M query keywords.

[0136] Among them, the P features may include multiple features of each feature type. A feature type is a general term for a certain type of feature, such as color, size, dressing style, etc.; a feature is a specific value of a feature type. For example, if the feature type is color, the features are red and green; a feature type has an identifier, which can be the number of the feature type, such as incrementing starting from 0. The multiple feature types can be divided into time type, sound type, region type, object type, and object description type.

[0137] Among them, the features of the time type are used to describe the shooting period of the video content. For example, it can be morning, afternoon, evening, etc. Specifically, the features of the time type can be designed in combination with the actual situation. Optionally, each frame of the image in the video to be queried has a shooting time. Based on the shooting time of each frame of the image in the video to be queried, the shooting time of the video content in the video to be queried can be obtained. Based on the shooting time, the shooting period can be analyzed. For example, if the shooting time is from 9:00 to 12:00, the shooting period can be morning.

[0138] Among them, the features of the sound type are used to describe the sound situation of the video content. For example, specific sound prompts (such as the promotional activities of a certain store in the mall), whether there is noisy ambient sound, etc. Specifically, the features of the sound type can be designed in combination with the actual situation.

[0139] Among them, the features of the region type are used to describe the regions in the video content. For example, the entrance area, elevator area, shopping area, exit area, food area, etc. The specific divided regions can be designed in combination with the actual situation.

[0140] Among them, the object type may include various feature types for distinguishing objects, such as the category of the objects appearing in the video content, the region to which the objects appearing in the video content belong, and / or the relationship between multiple objects appearing in the video content. For example, in a mall surveillance video, the distance relationship between the product shelf and the elevator, the interaction relationship between people and the shelf (such as whether there are behaviors such as approaching and taking products), etc. can be extracted. For example, for some special regions in the mall, such as the entrance and the exit, the connectivity relationship features between the objects and the entrance or the exit can be extracted. For example, record which shelf regions, elevators, etc. are directly connected to the entrance. These connectivity relationships can be used as important reference information in subsequent video retrieval.

[0141] Among them, the object description class can include various types of features for describing an object, such as the attributes of the object appearing in the video content: the size of the object, the position of the object, the color of the object, the area to which the object belongs (such as the entrance area, elevator area, shopping area, exit area, food area, etc.). For example, for a commodity shelf, the features of the object description class not only record its size (approximate dimensions of length, width, and height), but can also analyze the overall distribution of the colors of the commodities on the shelf and the placement density of the commodities on the shelf; for an elevator, the features of the object description class can record the position of the elevator (coordinates relative to the shopping mall) and the operating state of the elevator (ascending, descending, or stopped); for a person, the features of the object description class can also include the dressing style (such as casual wear, formal wear, etc.) and body posture (standing, walking, etc.).

[0142] In addition, when the object is in motion, such as when the object is a person, the object description class is also used to describe the motion trajectory of the object. For example, the features of the object description class can include object features such as the motion direction and motion speed of the object. Optionally, by tracking the position of the object in the video, features such as the motion direction and speed of the object can be obtained.

[0143] It should be noted that a feature is used to indicate the description information for a certain feature type from video content. A feature can be understood as the encoding of the description information of a certain feature type from video content. The description information of a certain feature type may be different in different video contents. Therefore, multiple features can belong to the same feature type. For example, if the feature type is the location of a customer, and the description information is the (x, y) coordinates where the customer is located and the floor where the customer is located, etc., then based on the coordinate system of the shopping mall, the (x, y) coordinates where the customer is located and the floor where the customer is located, etc. can be encoded to obtain a feature. Another example, if the feature type is color, and the description information is "red", a color code can be assigned to "red" to obtain a feature. For example, when the feature type is numerical information such as the size of an object or the distance between an object and an elevator, a feature can be obtained through any real number; when the feature type is a yes / no structured information such as whether it is connected to the entrance, whether there is an act of picking up a commodity, whether there is an act of approaching a commodity, etc., taking whether it is connected to the entrance as an example for description, a feature can be obtained by using 0 (indicating no) or 1 (indicating yes). It should be noted that for multiple features of each video content among W video contents, these features can be features of multiple modalities, and multiple modalities include vision, text, sound, and / or time. Among them, the visual features can be understood as the features obtained by extracting the images in the video content. For example, the above-mentioned region class, object description class, and object class features; the text features can be understood as the feature extraction of the text used to describe the video content. The text used to describe the video content can be pre-input. For example, the text can be used to explain the shooting time of the image and the objects in the image; the sound features can be understood as the features obtained by extracting the sound in the video content such as human voices, broadcast voices. For example, voiceprints, broadcast content, female voices, male voices, etc.; the time features are used to describe the shooting time of the video content.

[0144] In a possible implementation manner, for features of other classes except for the time class, an artificial intelligence model (which can also be called a hybrid model) can be used to extract features from the queried video to obtain P features of multiple feature types. For example, the artificial intelligence model can include a first model and a second model. The first model can extract object features, and the object features are used to indicate the self-attributes of the objects in the queried video. For example, they can be the above-mentioned object description class features and some features in the region class. The second model can extract scene relationship features, and the scene relationship features are used to indicate the association relationships between different objects in the queried video. For example, they can be the features between objects in the above-mentioned region class.

[0145] Optionally, in some examples, such as Figure 5As shown, the camera 120 captures a video and sends it to the video retrieval device 140. The video retrieval device 140 analyzes each frame of the video frame by frame, extracts descriptive information of multiple feature types, and obtains video extraction information for each of the multiple video contents. The video extraction information includes the extraction information for each of the multiple feature types. Subsequently, based on the video extraction information of the video content, multiple features of the video content are obtained and stored in the database. The processing of the video by the video retrieval device 140 is merely an example. In some possible scenarios, the video can also be processed by the data storage device 130.

[0146] Optionally, in one embodiment, frame-by-frame analysis of the video to be processed can be performed using a lightweight model (such as a feature extraction network based on deep learning). Exemplarily, the lightweight model can be MobileNet, which is a deep learning model specifically designed for mobile devices and resource-constrained environments and has efficient feature extraction capabilities.

[0147] Optionally, in one embodiment, when performing frame-by-frame analysis of the video to be processed, more dimensions of video types can be designed according to actual needs.

[0148] Optionally, in some examples, P features can be stored in Q (a positive integer greater than or equal to 2) layers in the database (deployed on the data storage device 130). As Figure 6 shown, the database can include the first layer, the second layer,..., the Qth layer. The first layer includes feature 11, feature 12,..., the second layer includes feature 21, feature 22,..., and the Qth layer includes feature Q1, feature Q2. Each feature in each layer is associated with the video content, so that multiple features of the video content can be extracted from the database.

[0149] Exemplarily, the Q layers include at least one of the following: a time layer, a sound layer, a region layer, an object feature, and an object description layer; where the time layer is used to store time-related features, the sound layer is used to store sound-related features, the region layer is used to store region-related features, the object feature is used to store object-related features, and the object description layer is used to store object description-related features.

[0150] Optionally, in some implementation manners of this embodiment, the data storage device 130 is further configured to store the feature vectors corresponding to W (greater than or equal to 2) video contents in the queried video. The feature vectors are used to indicate the semantics of the corresponding video contents (i.e., the information contained in the video contents, such as people, objects, sounds, etc.), and the feature vectors are obtained based on the multiple features of the corresponding video contents. Correspondingly, the video retrieval device 140 matches the M query keywords with the video to be queried based on the feature vectors of the video content, and there are N candidate video contents with the features matching the M query keywords.

[0151] Exemplarily, as Figure 7 shown, the video retrieval device 140 analyzes each frame in the video frame by frame, extracts description information of multiple feature types, and obtains video extraction information of each of the multiple video contents. The video extraction information includes extraction information of each of the multiple feature types. Subsequently, based on the video extraction information of the video content, multiple features of the video content are obtained, and based on the multiple features of the video content, a feature vector of the video content is obtained.

[0152] Optionally, the feature vector is obtained by encoding multiple features of the corresponding video content using a feature encoding method.

[0153] In an optional example, the feature encoding method may be a one-hot encoding method. Exemplarily, the one-hot encoding method may construct a feature template, which may include a vector representation for indicating the arrangement of multiple feature types. The feature template includes multiple elements arranged in order, and each element represents a feature type. Subsequently, for a video content, during the process of generating the feature vector corresponding to the video content, it is determined whether each feature type exists in the feature template. If not, it can be filled with 0. If it exists, the feature of the feature type is filled in.

[0154] In an optional example, the feature encoding method may include a mapping function, and the mapping function may map multiple features corresponding to the video content to a feature vector of a specific dimension, such as 128.

[0155] It should be noted that, as Figure 8 shown, the database (deployed on the data storage device 130) may store video content, multiple features of the video content, and the feature vector of the video content. The P features are stored in the database (deployed on the data storage device 130) in Q (a positive integer greater than or equal to 2) layers. Exemplarily, as Figure 9 shown, the database may include a first layer, a second layer,..., a Qth layer. The first layer includes feature 11, feature 12,..., the second layer includes feature 21, feature 22,..., and the Qth layer includes feature Q1, feature Q2. Each feature in the first layer is connected to each other feature in the first layer and each feature in the second layer. Each feature in the second layer is connected to each other feature in the second layer and connected to each feature in the third layer,..., each feature in the Q - 1th layer is connected to each other feature in the Q - 1th layer and each feature in the Qth layer. Each feature in the Qth layer is connected to each other feature in the Qth layer and the feature vector, and the feature vector is connected to the video content.

[0156] Optionally, in some implementations, the video retrieval device 140 matches the M query keywords and the video to be queried based on the feature vectors of the video content, which may include: matching the M query keywords with P features of the video to be queried to determine the matching features; then, determining a query vector based on the matching features; then, matching the query vector with W feature vectors to determine N matching feature vectors; W is a positive integer greater than or equal to 2, each feature vector corresponds to video content, and each feature vector is used to indicate the semantics of the video content; the N video contents corresponding to the N feature vectors are used as candidate video contents.

[0157] Optionally, the feature vector is obtained by encoding multiple features of the corresponding video content using a feature encoding method. Correspondingly, the query vector is obtained by encoding the features matched by the M query keywords using the same feature encoding method. In this embodiment, the query vector is used to indicate the features matched with the M query keywords. For example, the features matched by the M query keywords can be understood as features with the same or similar semantics as the M query keywords, and there can be one or more features matched by each query keyword.

[0158] It should be noted that considering the randomness of the user's speech, there may be query keywords in the M query keywords that do not match features. In other words, each candidate video content in the N candidate video contents may include features matched by some of the M keywords.

[0159] Exemplarily, the feature encoding method can be a one-hot encoding method. The query vector can be a vector after combining the features matched by the M query keywords.

[0160] In some optional implementations of this embodiment, the video retrieval device 1400 matches the query vector with W feature vectors to determine the N matching feature vectors, which may include: determining the M features matched by the M query keywords, and using each feature vector containing the M features as a matching feature vector.

[0161] Exemplarily, the feature encoding method may include a mapping function, and the query vector is a function obtained by mapping the features matched by the M query keywords through the mapping function.

[0162] In some optional implementations of this embodiment, the video retrieval device 140 matches the query vector with W feature vectors to determine the N matching feature vectors, which may include: calculating the correlation between the query vector and each of the W feature vectors; determining the N matching feature vectors based on the correlation between the query keywords and each feature vector. For example, using the N feature vectors with the highest correlation as the N matching feature vectors.

[0163] Exemplarily, assume that the M query keywords are: "red handbag", "first floor of the mall", "near the elevator". Then, if a feature vector contains an object that is red, and the shape of the object is similar to that of a handbag, and there is a relatively high possibility that the position is near the elevator on the first floor of the mall, then the relevance of this feature vector to the M query keywords will be relatively high. In this way, multiple feature vectors related to "red handbag", "first floor of the mall", and "near the elevator" may be found. Then, according to the relevance, the top K most relevant feature vectors are selected. For example, assume K = 5, then the 5 feature vectors with the highest relevance will be selected.

[0164] Exemplarily, calculating the relevance between the query vector and the feature vector may include: calculating the distance between the feature vector and the query vector. For example, using distance algorithms such as cosine to calculate the distance between vectors; based on the distance between the feature vector and the query vector, determining the relevance. For example, taking the reciprocal of the distance as the relevance.

[0165] Exemplarily, in the case where the relevance scores of at least some of the N feature vectors in the match are the same, it indicates that there are similar contents in the N candidate video contents. For the convenience of the user to view, optionally, the N candidate video contents can be sorted based on the chronological order to obtain the serial number of each candidate video content; or, the candidate video contents corresponding to the feature vectors with the same relevance score can be sorted based on the chronological order to obtain the serial numbers of the candidate video contents corresponding to the feature vectors with the same relevance score.

[0166] In some optional implementation manners of this embodiment, the video retrieval device 140 may send a query request to the data storage device 140. The query request may include M query keywords; subsequently, receive the N matching feature vectors returned by the data storage device 140. The data storage device 140 may query and obtain the N matching feature vectors in the manner of the P features or feature vectors described in step 203.

[0167] Step 204, the video retrieval device 140 generates a set of prompt words based on the M query keywords and the N candidate video contents.

[0168] The set of prompt words is used to indicate the M query keywords, the semantic prompt words of the M query keywords, and the N candidate video contents. Exemplarily, the set of prompt words may include the respective prompt words (combinations of semantic prompt words and query keywords) of the M query keywords and the respective video vectors of the N candidate video contents.

[0169] In some optional implementation manners of this embodiment, such as Figure 10As shown, the video retrieval device 140 obtains a prompt template, which includes a first blank field for candidate video content and a second blank field for semantic prompt words. The feature vectors of each candidate video content among N candidate video contents are filled in the first blank field, and the query keywords are filled in the second blank field to obtain a set of prompt words corresponding to the prompt template. Subsequently, the set of prompt words is input into the video search model. The semantic prompt words can be one or more. For example, they can be search_object, location, area, floor, building, gate. It should be noted that filling the feature vectors of candidate video content in the first blank field of candidate video content is only an example and does not constitute a specific limitation. In some possible scenarios, candidate video content can also be directly input. Subsequently, the video search model will perform feature extraction on the candidate video content to understand the candidate video content and determine whether the candidate video content meets the user's intention.

[0170] It should be noted that the prompt template is used to help the query model better understand the user's intention, so as to generate an output that meets expectations and has higher quality. As a bridge between the model and the user, the prompt template can significantly improve the performance and accuracy of the model by providing clear guidance or context information.

[0171] Exemplarily, in a shopping mall scenario, if looking for a lost item, the prompt word combination may focus more on information such as the characteristics of the item and the area where it was last seen. The prompt template can be {"search_object":,"location":{"area":,"floor":}}[feature vector 1 of candidate video content 1][feature vector 2 of candidate video content 2]……. The set of prompt words can be {"search_object":"a certain brand red handbag","location":{"area":"near a certain store in a certain shopping area","floor":"the first floor"}}[feature vector 1 of candidate video content 1][feature vector 2 of candidate video content 2]…….

[0172] Exemplarily, in a community scenario, if it is to search for a lost pet, the combination of prompt words will more include information such as the characteristics of the pet and the areas where it often activities. The prompt template can be {"search_object":,"location":{"area":,"building":}}[Feature vector 1 of candidate video content 1][Feature vector 2 of candidate video content 2]……. The set of prompt words can be {"search_object":"White Pomeranian","location":{"area":"Near the community garden","building":"Building 3"}}[Feature vector 1 of candidate video content 1][Feature vector 2 of candidate video content 2]…….

[0173] Exemplarily, in a station scenario, if it is to search for a lost luggage, the combination of prompt words will emphasize information such as the characteristics of the luggage, the characteristics of the passengers, and the possible locations where it may be lost. The prompt template can be {"search_object":,"location":{"area":,"gate":}}[Feature vector 1 of candidate video content 1][Feature vector 2 of candidate video content 2]……. The set of prompt words can be {"search_object":"Black suitcase","location":{"area":"Next to a waiting area in the waiting hall","gate":"A1"}}[Feature vector 1 of candidate video content 1][Feature vector 2 of candidate video content 2]…….

[0174] Step 205, the video retrieval device 140 determines the target video content from the N candidate video contents based on the set of prompt words and the video search model.

[0175] The video retrieval device 140 deeply understands the query intent through the M query keywords in the set of prompt words and the semantic prompt words of the M query keywords, verifies the N candidate video contents based on the query intent, and outputs the target video content with higher relevance to the query text, thereby improving the accuracy of the retrieval. Among them, the target video content can be any one of the N candidate video contents, or any multiple of the N candidate video contents. In some optional implementation manners of this embodiment, the video retrieval device 140 can use the RAG technology to determine the target video content. Optionally, such as Figure 10As shown, the video retrieval device 140 inputs the set of prompt words into the retrieval model, which is used to analyze the query intent based on the semantic prompt words of M keywords and M query keywords, analyze the relevance between each candidate video content and the query intent, and output the target video content. The retrieval model can accurately understand complex query intents, and the returned feature vectors are highly relevant to user needs, improving the accuracy and practicality of retrieval. It should be noted that this embodiment does not intend to limit the structure of the retrieval model. Exemplarily, the retrieval model can be a large model, such as the open-source Qianwen large model.

[0176] It is worth noting that in the scenario where the query text is obtained by speech recognition based on voice data, considering the randomness of the user's speech, each candidate video content among the N candidate video contents may contain features that match some of the keywords in the M keywords. That is to say, there may be query keywords in the M query keywords that do not match the features of the video to be queried, resulting in a relatively low reference value of the N candidate video contents. Therefore, it is considered to construct a set of prompt words by combining the M keywords and the N candidate video contents as the input of the video search model, and use the video search model to fully analyze the user query intent represented by the M query keywords, and perform a relevance analysis of the user intent on the N candidate video contents, thereby improving the accuracy of the query. Optionally, in one implementation, the video retrieval device 140 can determine the target keywords among the M query keywords, where the target keywords are the query keywords that match the features of the video to be queried among the M query keywords, and each candidate video content among the N candidate video contents includes features that match the target keywords. Subsequently, based on the M keywords and the N candidate video contents, a set of prompt words is obtained, and the set of prompt words is input into the video search model to obtain the target video content output by the video search model.

[0177] Exemplarily, the query text can be "I lost a red handbag on the first floor of the mall, probably near the elevator, and heard the discount news of a certain brand at that time". At this time, M query keywords can be extracted as "red handbag, first floor of the mall, near the elevator, heard, a certain brand, discount news". Then, through the red handbag, the first floor of the mall, and near the elevator, N candidate video contents can be queried. Then, based on the feature vectors of the N candidate video contents (used to indicate the picture content and voice content), the red handbag, the first floor of the mall, near the elevator, heard, a certain brand, and discount news, a set of prompt words is constructed. The set of prompt words can include voice content (semantic prompt words): a certain brand, discount news. The set of prompt words is input into the video search model. It should be noted that near the elevator can be converted into an area with a radius of 5 meters centered on the elevator. Then, through the red handbag and the first floor of the mall, N candidate video contents can be queried. Based on the feature vectors of the N candidate video contents (used to indicate the picture content and voice content), the red handbag, the first floor of the mall, the area with a radius of 5 meters centered on the elevator, heard, a certain brand, and discount news, a set of prompt words is constructed, and the set of prompt words is input into the video search model.

[0178] In this solution, through the query keywords, the video content is pre-screened in advance. Subsequently, based on the query keywords and the video content screened by the query keywords, a set of prompt words for query is obtained. Based on the set of prompt words and the video search model, the user's search intention can be understood more accurately, so as to search for video content that better meets the user's needs.

[0179] Based on the video query method provided above, a specific application of the video query method is described. Figure 6 It is a schematic flowchart of a specific application of a video query method provided for the implementation of this application.

[0180] In some possible scenarios, in the monitoring system of a large shopping mall, there is a large amount of monitoring video data. These videos record the activities in various areas of the mall (such as entrances, exits, near elevators, etc.). The owner lost a red handbag on the first floor of the mall and hopes to use the monitoring videos to find the relevant scenarios where the handbag was lost.

[0181] 1. Video preprocessing

[0182] 1.1. Model selection:

[0183] In the monitoring system of this large shopping mall, MobileNet is selected as a lightweight model for video preprocessing. MobileNet is a deep learning model specifically designed for mobile devices and resource-constrained environments, and it has efficient feature extraction capabilities.

[0184] 1.2. Video analysis:

[0185] When analyzing the mall surveillance video frame by frame, for example, for a single frame image, MobileNet can accurately identify various features of customers, product shelves, and elevators. Such as the wearing style of customers (e.g., casual wear, formal wear, etc.), body postures (standing, walking, etc.), the size of the product shelves (approximate dimensions of length, width, and height), the placement density of products on the shelves, the position of the elevator (coordinates relative to the mall), and the operating state of the elevator (ascending, descending, or stopped), etc. object information, thus obtaining rich object information. Based on the object information, the mall surveillance video can be segmented to obtain the video extraction content of multiple frame images and / or multiple video clips respectively.

[0186] Encode the video extraction content of multiple frame images and / or multiple video clips respectively, so as to obtain multiple features of multiple frame images and / or multiple video clips respectively. For example, for the location information of customers, based on the coordinate system of the mall, encode the (x, y) coordinates where the customer is located and information such as the floor where the customer is located, so as to be able to quickly locate and query later.

[0187] 1.3. Vector Library Construction:

[0188] Then convert the multiple features of multiple frame images and / or multiple video clips into vector representations to obtain feature vectors. For example, convert various features of customers (wearing style, posture, etc.) into vectors through a predefined mapping function. The dimension of each vector may be determined according to the complexity of the feature and the design of the model, such as a 128-dimensional vector.

[0189] Store each feature vector into the Milvus vector library. Each feature vector in the vector library corresponds to a frame image and / or a video clip, and is associated with multiple features at the same time.

[0190] 2. User Query

[0191] 2.1. Voice Input:

[0192] The loser inputs through voice: "I lost a red handbag on the first floor of the mall, probably near the elevator."

[0193] 2.2. Speech Recognition:

[0194] Use speech recognition technology to convert speech into text. Speech recognition technology has a high accuracy rate and can recognize various accents and language habits. In this process, it converts the sound signals in the speech into the corresponding text sequence, obtaining the original text "I lost a red handbag on the first floor of the mall, probably near the elevator."

[0195] 2.3. Natural Language Processing:

[0196] First, perform word segmentation. For example, use SpaCy to divide the sentence into words such as "I", "in", "on the first floor of the mall", "lost", "a", "red", "handbag", "around", "in", "near the elevator". Then, perform stop word removal, removing stop words such as "I", "a", "around" that have little significance for retrieval.

[0197] After information filtering and transformation, the query keyword "red handbag near the elevator on the first floor of the mall" is obtained. In this process, normalization processing is performed on some words with similar semantics. For example, "red-colored" is simplified to "red".

[0198] 3. Retrieval and matching

[0199] In the manner of 1.2 and 1.3, convert this query keyword into a vector (for ease of description and distinction, it can be called a query vector), match this query vector with the feature vectors in the vector library, and calculate the correlation between the query vector and each feature vector representing a frame image or video segment in the vector library.

[0200] When calculating the correlation, algorithms such as cosine correlation may be used. If a frame image or video segment represented by a feature vector contains a red object, and the shape of the object is similar to a handbag, and the possibility of its location being near the elevator on the first floor of the mall is relatively high, then the correlation between this feature vector and the query vector will be relatively high. In this way, multiple feature vectors related to "red handbag", "on the first floor of the mall", and "near the elevator" may be found. Then, according to the correlation scores, select the top K most relevant ones. Assuming K = 5, then the 5 feature vectors with the highest correlation will be selected as candidate feature vectors.

[0201] 4. Deep retrieval

[0202] 4.1. Prompt combination:

[0203] According to the preset prompt template, combine the above candidate feature vectors with the query keyword to obtain a prompt. For example, combine 5 candidate feature vectors with "red handbag near the elevator on the first floor of the mall" to form a prompt. The prompt template can be "[Feature vector 1 of candidate video content 1][Feature vector 2 of candidate video content 2]... red handbag near the elevator on the first floor of the mall".

[0204] 4.2. Retrieval by large model:

[0205] Pass the combined prompt to the retrieval model. It can deeply understand and analyze these prompts.

[0206] During the in-depth retrieval process, the retrieval model will further analyze the detailed content in the feature vector based on various information in the prompt words, such as the behaviors of people in the video, the relationships between objects, etc., and verify the matching degree between this detailed content and the query requirements.

[0207] 4.3. Result Return:

[0208] Finally, the video clip that is most likely to show the scene where the owner lost the handbag will be returned.

[0209] In some other possible scenarios, in the monitoring system of a large shopping mall, there is a large amount of monitoring video data. These videos record the activities in various areas of the mall (such as entrances, exits, near elevators, etc.). The owner lost a red handbag on the first floor of the mall and hopes to find the relevant scene of the lost handbag using the monitoring videos.

[0210] 1. Video Preprocessing

[0211] 1.1. Model Selection:

[0212] In the monitoring system of this shopping mall, a hybrid model of MobileNet and ShuffleNet is considered for video preprocessing. MobileNet performs well in extracting the basic visual features of objects, while ShuffleNet has an advantage in processing the object relationship features in complex scenarios. By combining the two, the content in the monitoring videos can be analyzed more comprehensively.

[0213] 1.2. Video Analysis:

[0214] The hybrid model is used to analyze the monitoring videos frame by frame. For a single frame, various features of customers, product shelves, and elevators can be accurately identified, such as the wearing styles of customers (such as casual wear, formal wear, etc.), body postures (standing, walking, bending down to select products, etc.), the number and color of shopping bags carried, facial expressions (such as happy, anxious, etc., which may be related to the shopping experience), the size of the product shelves (the precise dimensions of length, width, and height), the placement density of products on the shelves, the brand logos of products, whether there are promotional signs, the location of the elevators (coordinates relative to the mall and relative positions with surrounding stores), the operating status of the elevators (ascending, descending, or stopped), and the personnel density inside the elevators and other object information.

[0215] At the same time, scene relationships can also be analyzed, such as the distance between customers and shelves, the movement paths of customers between different brand shelves, the interactions between customers and elevators (such as whether waiting at the elevator entrance, whether directly walking towards a certain brand store after coming out of the elevator, etc.), so as to obtain rich scene information.

[0216] Subsequently, according to the object information and scene information, the mall surveillance video can be divided to obtain the video extraction content of each of multiple frame images and / or multiple video segments.

[0217] Encode the video extraction content of each of the multiple frame images and / or multiple video segments respectively, so as to obtain multiple features of each of the multiple frame images and / or multiple video segments. For example, for the location information of a customer, based on the coordinate system of the mall, encode the (x, y) coordinates where the customer is located and information such as the floor where the customer is located, so as to be able to quickly locate and query later.

[0218] 1.3. Vector library construction:

[0219] Then convert the multiple features of each of the multiple frame images and / or multiple video segments into vector representations to obtain feature vectors. For example, convert various features of a customer (such as dressing style, posture, etc.) into vectors through a predefined mapping function. The dimension of each vector may be determined according to the complexity of the feature and the design of the model, such as a 128-dimensional vector.

[0220] Store each feature vector into the Milvus vector library. Each feature vector in the vector library corresponds to a frame image and / or a video segment, and is associated with multiple features at the same time.

[0221] 2. User query stage

[0222] 2.1. User voice input

[0223] The owner of the lost item inputs by voice: "I lost a red handbag with a certain brand logo near a certain brand store on the first floor of the mall. I last saw it next to the elevator. I want to see the video after I left, preferably the video during the promotional event."

[0224] 2.2. Voice recognition

[0225] Use voice recognition technology to convert the voice into text.

[0226] 2.3. Text preprocessing stage

[0227] Use natural language processing technology to deeply understand the semantics of the recognized text and perform preprocessing. First, perform word segmentation operations, mark words such as "a certain brand store" and "a certain brand logo" as a whole vocabulary, and at the same time regard "during the promotional event" as a special vocabulary unit.

[0228] Special vocabulary marking:

[0229] A certain brand store → [a certain brand store]

[0230] A certain brand logo → [a certain brand logo]

[0231] During the promotion period → [During the promotion period]

[0232]

[0233] When performing stop-word removal, remove words that are not very meaningful for retrieval, such as "I" and "want to see", but retain key information such as "a certain brand", "red", "handbag", "next to the elevator entrance", "the first floor", and "during the promotion period".

[0234] For example, the stop-word list is as follows:

[0235] I, of, the, it, a, want, see, behind, I

[0236]

[0237] Perform semantic normalization. Convert "next to the elevator entrance" into an area centered on the elevator with a radius of 5 meters according to the layout diagram of the shopping mall, and unify "a certain brand logo" into the standard logo name of the brand (for example, "Nike_Logo").

[0238]

[0239] For semantic role labeling, identify "handbag" as the patient, "the first floor", "near the exclusive store of a certain brand", and "next to the elevator entrance" as locations, and "during the promotion period" as the time. Filter out words that do not meet the requirements based on these semantic roles. For example, some words describing the subjective thoughts of the owner of the lost item have been removed during the stop-word removal stage.

[0240]

[0241]

[0242] Result after filtering: Shopping mall / the first floor / [Nike exclusive store] / near / red / with / [Nike_Logo] / handbag / , / [area centered on the elevator with a radius of 5 meters] / , / [during the promotion period] / video

[0243] 3. Retrieval and matching stage

[0244] In the way described in 1.2 and 1.3, convert the query keyword into a vector (for the sake of easy description and distinction, it can be called the query vector), match the query vector with the feature vectors in the vector library, and calculate the correlation between the query vector and each feature vector in the vector library representing the frame image or video segment.

[0245] Select the top K most relevant feature vectors as candidate feature vectors according to the relevance. If there are cases of the same relevance, perform a secondary sorting according to the chronological order of the feature vectors, and preferentially select the feature vectors after the owner left, so as to better meet the need of the owner to view the video after they left.

[0246] 4. Deep retrieval

[0247] 4.1. Prompt combination:

[0248] According to the preset prompt template, combine the above candidate feature vectors with the query keywords to obtain the prompt.

[0249] Exemplarily, in a shopping mall scenario, if looking for a lost item, the prompt combination may focus more on information such as the features of the item and the area where it was last seen. The prompt template can be {"search_object": "a red handbag of a certain brand", "location": {"area": "near a certain store in a certain shopping area", "floor": "the first floor"}} [feature vector 1 of candidate video content 1] [feature vector 2 of candidate video content 2]...

[0250] Exemplarily, in a residential community scenario, if looking for a lost pet, the prompt combination will include more information such as the features of the pet and the areas where it often moves. The prompt template can be {"search_object": "a white Pomeranian", "location": {"area": "near the community garden", "building": "Building 3"}} [feature vector 1 of candidate video content 1] [feature vector 2 of candidate video content 2]...

[0251] Exemplarily, in a station scenario, if looking for a lost luggage, the prompt combination will emphasize information such as the features of the luggage, the features of the passenger, and the possible locations where it was lost. The prompt template can be {"search_object": "a black suitcase", "location": {"area": "next to a certain waiting area in the waiting hall", "gate": "A1"}} [feature vector 1 of candidate video content 1] [feature vector 2 of candidate video content 2]...

[0252] 4.2. Retrieval by large model:

[0253] Pass the combined prompt to the retrieval model. It can deeply understand and analyze these prompts.

[0254] During the in-depth retrieval process, the retrieval model further analyzes the detailed content in the feature vectors based on various information in the prompt words, such as the behaviors of people in the video, the relationships between objects, etc., and verifies the matching degree between the detailed content and the query requirements.

[0255] 4.3. Result Return:

[0256] Finally, the video clip most likely to show the scene where the owner lost the handbag is returned.

[0257] Based on the same concept as the method embodiment of the present application, the embodiment of the present application also provides a video query device. The video query device includes several modules, and each module is used to execute each step in the video query method provided by the embodiment of the present application. The division of the modules is not limited here. Those skilled in the art can clearly understand that in practical applications, each step in the video query method provided by the embodiment of the present application can be allocated to different modules according to needs, that is, the internal structure of the device is divided into different modules to complete all or part of the functions described above. Each module in the embodiment can be integrated in a processing unit, or each unit exists physically alone, or two or more modules are integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the modules in the above device can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0258] Exemplarily, the video query device is used to execute the video query method provided by the embodiment of the present application. Figure 8 It is a schematic structural diagram of the video query device provided by the embodiment of the present application. As Figure 11 shown, the video query device provided by the embodiment of the present application includes:

[0259] A text acquisition module 1101, configured to acquire a query text; wherein, the query text is used to indicate the query requirements of the user for the video to be queried;

[0260] An extraction module 1102, configured to extract keywords from the query text to determine M query keywords; wherein, M is a positive integer greater than or equal to 1;

[0261] A first query module 1103, configured to obtain N candidate video contents based on the M query keywords; wherein, N is a positive integer greater than or equal to 1; each candidate video content is a video clip or a frame image in the video to be queried, and each candidate video content includes features matching the M query keywords;

[0262] A prompt word determination module 1104, configured to generate a set of prompt words according to the M query keywords and the N candidate video contents;

[0263] A second query module 1105, configured to determine target video contents from the N candidate video contents based on the set of prompt words and a video search model.

[0264] Based on the same concept as the method embodiment of the present application, an embodiment of the present application further provides a computing device. The computing device can be a server or a terminal device. Among them, the terminal device can be a mobile phone, a tablet computer, a wearable device, etc.

[0265] Figure 12 It is a schematic structural diagram of a computing device provided by an embodiment of the present application.

[0266] As Figure 12 shown, the computing device 1200 includes a processor 1201, a memory 1202, and a network interface 1203.

[0267] The processor 1201 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0268] The memory 1202 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0269] Exemplarily, a computer program can be stored on the memory 1202. When the processor 1201 executes the computer program, the steps in the foregoing embodiments of the video query method are implemented, such as Figure 2 the steps 201 to 205 shown. Alternatively, when the processor 1201 executes the computer program, the functions of the various modules in the foregoing device embodiments are implemented. Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions. The one or more modules / units are stored in the memory 1202 and executed by the processor 1201 to complete this application. For example, the computer program can be divided into a text acquisition module 1101, an extraction module 1102, a first query module 1103, a prompt word determination module 1104, and a second query module 1105. For the specific functions of each module, refer to the foregoing description.

[0270] The network interface 1203 is used for sending and receiving data. For example, it sends the data processed by the processor 1201 to other computing devices, or receives the data sent by other computing devices, etc.

[0271] Of course, for simplicity Figure 9Only some of the components related to this application in the computing device 1200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the computing device 1200 may further include any other appropriate components. Additionally, the computing device may be a desktop computer, a notebook, a palm computer, a cloud server, or other computing devices. Those skilled in the art can understand that Figure 9 merely examples of the computing device 1200, which do not constitute a limitation on the computing device. It may include more or fewer components than those shown, or combine certain components, or different components. For example, the computing device may further include an input device, an output device, a network access device, a bus, etc. Exemplarily, the input device may be a microphone array and may also include, for example, a keyboard, a mouse, etc. Exemplarily, the output device may output various information to the outside and may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0272] In addition to the above methods, devices, and computing devices, embodiments of the present application may further provide a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the video query methods of various embodiments of the present application described in the "Methods" section of this specification. Among them, the computer program product may be written in any combination of one or more programming languages to write computer program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. Among them, the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0273] In addition, embodiments of the present application may further provide a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the video query method according to various embodiments of the present disclosure described in the "method" section of the present specification above. The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0274] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0275] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0276] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, advantages, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purposes of illustration and facilitating understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0277] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," etc. are open-ended terms that mean "including but not limited to" and can be used interchangeably with each other. The word "or" and "and" used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0278] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0279] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

[0280] It can be understood that the various numerical numbers involved in the embodiments of this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application.

Claims

1. A video query method, characterized in that, The method includes: Obtaining a query text; wherein, the query text is used to indicate the user's query requirement for the video to be queried; Performing keyword extraction on the query text to determine M query keywords; wherein, M is a positive integer greater than or equal to 1; Based on the M query keywords, obtaining N candidate video contents; wherein, N is a positive integer greater than or equal to 1; each candidate video content is a video segment or a frame image in the video to be queried, and each candidate video content includes features matching the M query keywords; Generating a set of prompt words based on the M query keywords and the N candidate video contents; Determining the target video content from the N candidate video contents based on the set of prompt words and a video search model.

2. The method according to claim 1, wherein The obtaining of the query text includes: Obtaining the query voice input by the user; Converting the query voice into the query text.

3. The method according to claim 1 or 2, characterized in that, The performing of keyword extraction on the query text to determine M query keywords includes: Performing a word segmentation operation on the query text to obtain a first word segmentation; Filtering out the stop words in the first word segmentation to obtain a second word segmentation; Determining the word segments with semantic roles in the second word segmentation; Based on the word segments with semantic roles in the second word segmentation, determining the M query keywords.

4. The method according to any one of claims 1 to 3, characterized in that The query text includes words with semantic ambiguity related to the target object; The M query keywords include first query keywords corresponding to the words with semantic ambiguity; The performing of keyword extraction on the query text to determine M query keywords includes: Obtaining the scene layout data of the video; Based on the scene layout data, determining the range data related to the target object; Based on the range data, converting the words with semantic ambiguity to obtain the first query keywords.

5. The method according to any one of claims 1 to 4, characterized in that, The obtaining of N candidate video contents based on the M query keywords includes: Matching the M query keywords with P features of the video to be queried to determine the matching features; the P features include multiple features of each of the W video contents in the video to be queried, and the multiple features of each video content include features of multiple modalities, and the multiple modalities include vision, text, and sound and / or time, and W and P are positive integers greater than or equal to 2; Based on the matching features, determining a query vector; Matching the query vector with the feature vectors corresponding to each video content to determine N matching feature vectors; the feature vectors are constructed based on the multiple features of the corresponding video content; the feature vectors are used to indicate the semantics of the corresponding video content; and taking the N video contents corresponding to the N feature vectors as candidate video contents.

6. The method according to claim 5, wherein The P features include object features and scene relationship features, the object features are used to indicate the self-attributes of the objects in the video to be queried, and the scene relationship features are used to indicate the association relationships between different objects in the video to be queried; Before matching the M query keywords with the P video features of the video to be queried, the method further includes: extracting P features based on a hybrid model; wherein the hybrid model includes a first model and a second model, the first model is used to extract the object features; the second model is used to extract the scene relationship features.

7. The method according to claim 5 or 6, characterized in that, Matching the query vector with the feature vectors corresponding to each video content to determine N matching feature vectors, including: Calculating the correlation scores between the M query keywords and each of the W feature vectors; Based on the correlation scores of each feature vector, determining the top N most relevant feature vectors as the N matching feature vectors.

8. The method according to claim 7, characterized in that, The method further includes: When the correlation scores of at least some of the N feature vectors are the same, sorting the N candidate video contents based on the chronological order to obtain the serial numbers of each of the candidate video contents.

9. The method according to any one of claims 1-8, characterized in that, Generating a set of prompt words based on the M query keywords and the N candidate video contents, including: Matching the M query keywords with the semantic prompt words in the semantic prompt word set, and combining the matching query keywords and semantic prompt words to obtain the prompt words for each of the M query keywords; Combining the N candidate video segments with the prompt words for each of the M query keywords to obtain a set of prompt words.

10. A computing device, characterized in that, The computing device includes a processor and a memory; the processor is coupled to the memory; wherein the memory is used to store programs; The processor is used to execute the programs to implement the method according to any one of claims 1 to 9.