Video-based retrieval processing method and device, electronic equipment and storage medium

By breaking down videos into video segments and combining them with structured storage based on visual features and multidimensional keywords, the problem of inaccurate search results in traditional video retrieval methods is solved. This enables accurate retrieval and efficient management of long videos, and is applicable to multiple fields such as smart homes, smart cities, public safety, video content creation, intelligent teaching aids, autonomous driving, and medical surgery.

CN122432381APending Publication Date: 2026-07-21BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-07
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional video retrieval methods rely on text descriptions, resulting in inaccurate search results and failing to meet the need for precise retrieval of long videos.

Method used

A video-based retrieval processing method is adopted, which decomposes videos into independent video segments, combines visual feature pairs and search keywords for structured storage, and uses a multimodal large model to generate retrieval results, including keyframes, text descriptions, temporal indexes and multidimensional keywords, to achieve three-dimensional storage and accurate retrieval of video segments.

Benefits of technology

It improves the accuracy and reliability of search results, reduces storage resource consumption, adapts to the segmented management needs of long videos, supports real-time interaction and deployment on resource-constrained devices, and enhances the robustness of the system in complex video scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432381A_ABST
    Figure CN122432381A_ABST
Patent Text Reader

Abstract

The disclosure provides a video-based retrieval processing method and device, electronic equipment and storage medium, and relates to the fields of artificial intelligence such as video understanding, computer vision, natural language processing, deep learning and large models. The method can include: obtaining a retrieval request; determining a target memory unit matched with the retrieval request from each candidate memory unit according to unit content of each candidate memory unit stored in a bottom database, any candidate memory unit corresponding to at least one video segment in the same video, different videos corresponding to at least one candidate memory unit, and the unit content including a visual feature pair and a retrieval keyword, the visual feature pair including a key frame and a text description of the corresponding video segment; and generating a retrieval result corresponding to the retrieval request according to the target memory unit. The scheme can improve the accuracy of the retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to video understanding, computer vision, natural language processing, deep learning and large models, and especially to video-based retrieval and processing methods, devices, electronic devices and storage media. Background Technology

[0002] Video understanding refers to the technology of automatically analyzing and interpreting video content. Traditionally, videos, especially long videos exceeding a predetermined duration, are sampled at equal intervals or keyframes are extracted. The extracted image frames are then converted into corresponding text descriptions. When a search request is received, the video information matching the request can be determined based on the text description, and search results can be generated and returned. However, this method relies solely on text descriptions, which can easily lead to inaccurate video information, thus reducing the accuracy of the search results. Summary of the Invention

[0003] This disclosure provides video-based retrieval processing methods, apparatus, electronic devices, and storage media.

[0004] A video-based retrieval and processing method, comprising:

[0005] Obtain the search request;

[0006] Based on the content of each candidate memory unit stored in the underlying database, the target memory unit that matches the retrieval request is determined from each candidate memory unit. Each candidate memory unit corresponds to at least one video segment in the same video, and different videos correspond to at least one candidate memory unit. The content of the unit includes: visual feature pairs and retrieval keywords. The visual feature pairs include: keyframes of the corresponding video segment and text descriptions.

[0007] The search results corresponding to the search request are generated based on the target memory unit.

[0008] A video-based retrieval and processing device includes: a retrieval module and a storage module;

[0009] The retrieval module is used to obtain a retrieval request, determine the target memory unit that matches the retrieval request from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database, each candidate memory unit corresponds to at least one video segment in the same video, and different videos correspond to at least one candidate memory unit, the unit content includes: visual feature pairs and retrieval keywords, the visual feature pairs include: keyframes of the corresponding video segment and text description, and generate retrieval results corresponding to the retrieval request based on the target memory unit;

[0010] The storage module is used to store the underlying database.

[0011] An electronic device, comprising:

[0012] At least one processor; and

[0013] A memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described above.

[0015] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above.

[0016] A computer program product includes a computer program / instructions that, when executed by a processor, implement the method described above.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0018] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0019] Figure 1 Flowcharts of embodiments of the video-based retrieval and processing method described in this disclosure;

[0020] Figure 2 This is a flowchart illustrating an embodiment of the processing method for any video segment described in this disclosure;

[0021] Figure 3 This is a flowchart of an embodiment of the method for determining the target memory unit as described in this disclosure;

[0022] Figure 4 This is a flowchart illustrating an embodiment of the method for obtaining a target multimodal large model as described in this disclosure;

[0023] Figure 5 This is a schematic diagram of the composition structure of the first embodiment 500 of the video-based retrieval and processing device described in this disclosure;

[0024] Figure 6 This is a schematic diagram of the composition structure of the second embodiment 600 of the video-based retrieval processing device described in this disclosure;

[0025] Figure 7 This is a schematic diagram of the composition structure of the third embodiment 700 of the video-based retrieval processing device described in this disclosure;

[0026] Figure 8 This is a schematic diagram of the composition structure of the fourth embodiment 800 of the video-based retrieval processing device described in this disclosure;

[0027] Figure 9 A schematic block diagram of an electronic device 900 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0028] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0029] Furthermore, it should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0030] Figure 1 Flowcharts of embodiments of the video-based retrieval and processing method described in this disclosure. (See also...) Figure 1 As shown, the specific implementation methods are as follows.

[0031] In step 101, a search request is obtained.

[0032] In step 102, based on the content of each candidate memory unit stored in the underlying database, the target memory unit that matches the retrieval request is determined from each candidate memory unit. Each candidate memory unit corresponds to at least one video segment in the same video, and different videos correspond to at least one candidate memory unit. The content of the unit includes: visual feature pairs and retrieval keywords. The visual feature pairs include: keyframes of the corresponding video segment and text descriptions.

[0033] In step 103, the search results corresponding to the search request are generated based on the target memory unit.

[0034] Using the above-described method embodiments, candidate memory units can be used to structurally store the video segment content of each video. By dividing the video segments, it can adapt to the processing needs of various types of videos, especially long videos. That is, long videos can be broken down into independent small units for targeted processing. Moreover, the unit content of the candidate memory unit can include "visual feature pairs + search keywords". The visual feature pairs include both keyframes of the video segment and text descriptions of the video segment, taking into account both low-level visual detail information and high-level semantic information. At the same time, search keywords can also provide more accurate matching dimensions for retrieval, thereby making up for the lack of information in a single text description. Accordingly, by combining these unit contents, the target memory unit that matches the retrieval request can be more accurately determined, thereby improving the accuracy and reliability of the retrieval results.

[0035] The videos, text descriptions, and search results described in the embodiments of this disclosure are not targeted at any specific user and are not intended to reflect the personal information of any specific user. The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, involved in the technical solutions of this disclosure comply with relevant laws and regulations and do not violate public order and good morals.

[0036] In some embodiments of this disclosure, the content of each candidate memory unit may further include: a time-domain index, which includes the start and end timestamps of the corresponding video segment. In addition, the search keywords may include keywords of four different dimensions: time, person, object, and event, extracted from the corresponding video segment.

[0037] In other words, the underlying database may include candidate memory units corresponding to each video. Each candidate memory unit may correspond to at least one video segment within the same video. The content of each candidate memory unit may include: visual feature pairs, temporal indexes, and search keywords. The visual feature pairs may include keyframes from the corresponding video segment and a text description of the video segment. The temporal index records the start and end timestamps of the corresponding video segment (e.g., from frame x to frame x+30). The search keywords include structured tags extracted from four different dimensions: time, people, objects, and events. The text description may be dynamic / action-based.

[0038] For example, taking the video clip of "an elderly man taking white pills in the living room at 2 pm" as an example, the text description may include: The elderly man is sitting on the sofa in the living room, holding a white medicine bottle in his right hand, pouring out a white pill and putting it in his mouth, and then picking up a water glass to drink water. Search keywords may include: 1) Time: 2 pm; 2) Person: elderly male; 3) Object: medicine bottle, white pill, water glass; 4) Event: taking medicine, drinking water.

[0039] As can be seen, by using the candidate memory units described in this disclosure, low-level visual detail information can be preserved through keyframes, action evolution information can be recorded through dynamic text descriptions, precise temporal positioning can be achieved through temporal indexing, and precise indexing of time, people, objects, and events can be achieved through multi-dimensional keywords. This enables three-dimensional and structured storage of video clip content, laying a good foundation for subsequent processing.

[0040] Each candidate memory unit in the underlying database can be pre-generated and stored. Accordingly, in some embodiments of this disclosure, for any video, it can be divided into a series of video segments according to a predetermined duration, and the following first process can be performed on each video segment in chronological order from first to last: generating a decision result corresponding to the video segment, and in response to determining that a new candidate memory unit needs to be added based on the decision result, generating the candidate memory unit corresponding to the video segment and storing it in the underlying database.

[0041] The video can be a real-time video stream (such as a video stream captured in real time by a camera in a smart home monitoring scenario) or an offline video file (such as classroom playback video, historical surveillance recordings, etc.).

[0042] The specific value of the predetermined duration can be determined according to actual needs, such as 10 seconds.

[0043] In the above processing method, video segments can be divided according to a predetermined duration and processed sequentially, thereby achieving refined and orderly video decomposition, adapting to the segmented management needs of long videos, and adding candidate memory units as needed based on decision results, thereby actively filtering out valuable video segments and reducing the occupation of storage resources.

[0044] In some embodiments of this disclosure, for any video segment, the decision result may include: the temporal index, text description, and search keywords of the video segment, as well as decision action and thought process information for generating the decision action. Accordingly, in response to determining that a new candidate memory unit needs to be added based on the decision result, the method of generating a candidate memory unit corresponding to the video segment and storing it in the underlying database may include: in response to the decision action being an add action, generating a candidate memory unit corresponding to the video segment and storing it in the underlying database based on the decision result and the video segment.

[0045] In other words, for any video segment, the Chain of Thought (CoT) output format can be used to output the corresponding decision result for that video segment. This makes the thinking process explicit, improves the interpretability of the decision result, makes it easier for developers to trace the root of the decision logic, quickly locate and solve misjudgment problems, and greatly improves the robustness of the system in complex video scenarios.

[0046] For any video segment, if the decision action in the decision result is determined to be a new action, a candidate memory unit corresponding to the video segment can be generated based on the temporal index, text description, and search keywords in the decision result, as well as the keyframes extracted from the video segment, and the generated candidate memory unit can be stored in the underlying database.

[0047] In some embodiments of this disclosure, for any video segment, the decision action may further include: an update action, a delete action, and a skip action. Accordingly, in response to the decision action being an update action or a delete action, the decision result may further include: operation instruction information, used to indicate the pending memory units that need to be updated or deleted in the candidate memory units corresponding to the video to which the video segment belongs. Further, the first processing may further include: in response to the decision action being an update action, updating the pending memory units according to the decision result and the video segment; in response to the decision action being a delete action, deleting the pending memory units; in response to the decision action being a skip action, directly processing the next video segment.

[0048] Based on the above introduction, Figure 2 This is a flowchart illustrating an embodiment of the processing method for any video segment (let's say video segment y) described in this disclosure. Figure 2 As shown, the specific implementation methods are as follows.

[0049] In step 201, the decision result corresponding to video segment y is generated.

[0050] For example, the decision result corresponding to video segment y can be generated by combining video segment y and the candidate memory units corresponding to the video to which video segment y belongs.

[0051] The decision results may include <time>、、 <memory> 、 <think>and <action>Information such as the following format can be used:

[0052] <time> Temporal index of video clip y< / time> Text description and search keywords for video clip y <memory> Operation instructions (i.e., the memory content related to the current video segment)< / memory> <think> Information on the thinking process< / think> <action> Decision-making actions< / action> .

[0053] Decision-making actions can include the following four types:

[0054] Adding new actions: If it is determined that video clip y includes a brand new target (such as a new character) or a significant event, the corresponding candidate memory unit needs to be added to the underlying database;

[0055] Update action: If new information about known people or known objects is identified from video clip y, then the relevant candidate memory units are updated accordingly.

[0056] Deletion actions: such as identifying and removing redundant or outdated candidate memory units in the underlying database;

[0057] Skip Action: If it is determined that video segment y is meaningless noise or redundant background content, skip it directly.

[0058] In step 202, it is determined whether the decision action in the decision result is a new action. If so, step 203 is executed; otherwise, step 204 is executed.

[0059] In step 203, candidate memory units corresponding to video segment y are generated based on the decision result and video segment y, and stored in the underlying database.

[0060] In step 204, it is determined whether the decision action in the decision result is an update action, a deletion action, or a skip action. If it is an update action, then step 205 is executed; if it is a deletion action, then step 206 is executed; if it is a skip action, then step 207 is executed.

[0061] In step 205, the memory units to be processed in the underlying database are updated based on the decision results and the video segment y.

[0062] The memory units to be processed here refer to the candidate memory units that need to be updated among the candidate memory units corresponding to the video to which the stored video segment y belongs.

[0063] Updating the memory unit to be processed can refer to updating one or more of the following: updating the temporal index, updating the visual feature pairs, and updating the search keywords.

[0064] For example, if the video segments x and y corresponding to the memory unit to be processed are adjacent, the temporal index of the memory unit can be updated to the start timestamp of video segment x and the end timestamp of video segment y. If video segments x and y are not adjacent, the start and end timestamps of video segment y can be added to the temporal index of the memory unit to be processed. Furthermore, updating visual feature pairs can refer to adding keyframes and corresponding text descriptions from video segment y to the visual feature pairs of the memory unit to be processed, and updating search keywords can refer to adding the search keywords corresponding to video segment y to the search keywords of the memory unit to be processed.

[0065] In step 206, the unprocessed memory units in the underlying database are deleted.

[0066] The memory units to be processed here refer to the candidate memory units that need to be deleted from the candidate memory units corresponding to the video to which the stored video segment y belongs.

[0067] In step 207, the next video segment is processed.

[0068] Traditional methods often suffer from "missing some information" or "significant omissions" when dealing with hours-long videos due to memory limitations or limited context window length. The solution described in this disclosure introduces the ability to proactively identify the value of different video segments. This allows for dynamic planning of storage based on the value of each segment, avoiding the storage of large amounts of worthless information in the underlying database and reducing memory usage. This enables accurate and comprehensive storage of all information from long videos with low memory consumption, effectively solving the problems of "missing some information" or "significant omissions" in long video understanding. Furthermore, it not only reduces server-side storage costs but also makes this solution suitable for edge deployment on resource-constrained smart devices (such as cameras, robots, and mobile phones), laying a solid foundation for large-scale product adoption.

[0069] It should be noted that in practical applications, the underlying database can be updated at any time according to the actual situation, such as updating the underlying database based on newly added videos or video clips.

[0070] In addition, it can obtain retrieval requests and determine the target memory unit that matches the retrieval request from among the candidate memory units based on the unit content of each candidate memory unit stored in the underlying database.

[0071] In some embodiments of this disclosure, the target question corresponding to the retrieval request can be determined. In response to the target question not meeting the requirements for fast retrieval, the target memory unit can be determined from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database.

[0072] For example, if the retrieval request is a visual retrieval request, then the corresponding target question can be generated based on the image carried in the visual retrieval request (such as a scene captured by surveillance cameras). As another example, if the retrieval request is a text retrieval request, then the text description information carried in the text retrieval request (i.e., the natural language question posed by the user) can be directly determined as the target question.

[0073] Correspondingly, a complete full-database search process can be initiated only when the target question does not meet the requirements for rapid retrieval. That is, the target memory unit is determined from each candidate memory unit based on the content of each candidate memory unit stored in the underlying database, thus taking into account both the efficiency of retrieval and the comprehensiveness of full-database retrieval.

[0074] Therefore, in some embodiments of this disclosure, at least one predicted question can be generated in advance for each candidate memory unit in the underlying database. Accordingly, the similarity between the target question and each predicted question can be obtained. In response to each similarity being less than or equal to a first threshold, it can be determined that the target question does not meet the requirements for fast retrieval.

[0075] Correspondingly, in some embodiments of this disclosure, in response to the similarity between the target question and at least one predicted question being greater than a first threshold, it can be determined that the target question meets the requirements for fast retrieval. Further, the candidate memory unit corresponding to the predicted question with a similarity greater than the first threshold can be determined as the target memory unit, or the candidate memory unit corresponding to the predicted question with the highest similarity can be determined as the target memory unit. The specific value of the first threshold can be determined according to actual needs.

[0076] In the traditional approach, the search operation is usually triggered passively after a search request is received. However, the search process may take a long time, making it difficult to meet the application requirements of real-time interaction.

[0077] By adopting the processing method described in this disclosure, at least one corresponding prediction question can be generated in advance for each candidate memory unit in the underlying database. In this way, if there is a prediction question that matches the target question (a prediction question with a similarity greater than the first threshold or a prediction question with the highest similarity), the candidate memory unit corresponding to the matching prediction question can be directly determined as the target memory unit, thereby greatly shortening the retrieval time and improving the response speed.

[0078] In addition, in some embodiments of this disclosure, the prediction problem generated for any candidate memory unit may include: the prediction problem generated for the video segment corresponding to the candidate memory unit, and the prediction problem generated for the key focus frames determined from each key frame of the candidate memory unit.

[0079] The key frame of interest can refer to the key frame corresponding to the "moment of surprise", that is, the key frame with a sudden increase in information. There are no restrictions on how to identify key frames of interest.

[0080] By adopting the above processing method, prediction questions can be generated from both the overall video segment and the local focus frames, thereby covering full information and core details. This makes the generated prediction questions more comprehensive and accurate, and thus increases the probability of a successful match between the target question and the prediction question.

[0081] Based on the above introduction, Figure 3 This is a flowchart illustrating an embodiment of the method for determining a target memory cell as described in this disclosure. Figure 3 As shown, the specific implementation methods are as follows.

[0082] In step 301, the target question corresponding to the search request is determined.

[0083] In step 302, the similarity between the target question and the predicted question corresponding to each candidate memory unit in the underlying database is obtained.

[0084] Specifically, for each candidate memory unit in the underlying database, at least one corresponding prediction question can be generated in advance.

[0085] In step 303, it is determined whether there is a similarity greater than the first threshold. If yes, step 304 is executed; otherwise, step 305 is executed.

[0086] In step 304, the candidate memory unit corresponding to the prediction question with the highest similarity is determined as the target memory unit.

[0087] In step 305, the target memory unit is determined from the candidate memory units based on the cell contents of each candidate memory unit stored in the underlying database.

[0088] In some embodiments of this disclosure, the method of determining the target memory unit from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database may include: in response to a visual retrieval request, determining the corresponding image keywords based on the image carried in the visual retrieval request, and determining the target memory unit based on the image, the image keywords, and the unit content of each candidate memory unit; in response to a text retrieval request, determining the corresponding text keywords based on the text description information carried in the text retrieval request, and determining the target memory unit based on the text keywords and the unit content of each candidate memory unit.

[0089] For example, image keywords and text keywords can be extracted from four dimensions: time, people, objects, and events. The extraction results for each dimension may be empty or not.

[0090] Furthermore, when determining the target memory unit based on the image, image keywords, and the content of each candidate memory unit, for any candidate memory unit, visual similarity matching can be performed between the image features of the image and each keyframe in the candidate memory unit. Text similarity matching can also be performed between the image keywords and the text description and search keywords of the candidate memory unit. The comprehensive score of the candidate memory unit can then be determined by combining the visual and text similarity matching results. The candidate memory unit with the highest comprehensive score or a comprehensive score greater than a second threshold can be identified as the target memory unit. Similarly, when determining the target memory unit based on text keywords and the content of each candidate memory unit, text similarity matching can be performed between the text keywords and the text description and search keywords of the candidate memory unit. The comprehensive score of the candidate memory unit can then be determined based on the text similarity matching results. The candidate memory unit with the highest comprehensive score or a comprehensive score greater than a second threshold can be identified as the target memory unit. In addition, if the image keywords or text keywords include time, the target memory unit can be determined by combining the time domain index of the candidate memory unit. For example, the time domain matching evaluation result (whether the time in the image keywords or text keywords is within the time range corresponding to the time domain index) can be introduced. Accordingly, the comprehensive score can be generated by combining the time domain matching evaluation result.

[0091] As can be seen, by adopting the above processing method, the target memory unit can be determined according to the corresponding method for visual retrieval request and text retrieval request, thereby improving the accuracy of the obtained target memory unit and meeting the usage needs of different scenarios.

[0092] In addition, in some embodiments of this disclosure, the method of generating a corresponding decision result for any video segment may include: using a target multimodal large language model (MLLM) to generate text descriptions, search keywords, decision actions, and thinking process information in the decision result corresponding to the video segment. Furthermore, the method of generating at least one corresponding prediction question for each candidate memory unit in the underlying database may include: using the target multimodal large language model to generate at least one corresponding prediction question for each candidate memory unit.

[0093] With the rapid development of artificial intelligence technology, Large Language Models (LLMs) and Multimodal Large Models have been increasingly widely used. Large Language Models achieve a deep understanding of human language through pre-training on large-scale corpora, while Multimodal Large Models go a step further, capable of integrating data from multiple modalities such as text, images, and videos for comprehensive reasoning.

[0094] The solution described in this disclosure can utilize a target multimodal large model to generate text descriptions, search keywords, decision-making actions, and thought process information for video clips in one stop. It can also generate prediction questions for candidate memory units, thereby achieving intelligent unification of memory management decision-making and retrieval prediction, and reducing the difficulty of system development and deployment.

[0095] The target multimodal large model can be pre-trained. Accordingly, Figure 4 This is a flowchart illustrating an embodiment of the method for obtaining a target multimodal large model as described in this disclosure. Figure 4 As shown, the specific implementation methods are as follows.

[0096] In step 401, the basic multimodal large model is determined.

[0097] For example, an existing multimodal large model can be identified as the basic multimodal large model.

[0098] In step 402, the basic multimodal large model is fine-tuned using supervised fine-tuning (SFT) and reinforcement learning (RL) using the corresponding training sample set to obtain the initial multimodal large model.

[0099] For example, a specialized memory planning training dataset of approximately 80k can be constructed through processes such as data cleaning, long video slicing (to obtain video segments), and joint annotation of large models. The training dataset can then be used to perform supervised fine-tuning and reinforcement learning fine-tuning on the basic multimodal large model to obtain the initial multimodal large model.

[0100] The inputs and outputs during training can be shown below:

[0101] Input: Video clip + historical memory information (such as memory units corresponding to video clips that belong to the same video clip and are located before this video clip).

[0102] Output (JSON / tag format):

[0103] {

[0104] "think": "Information about the thought process"

[0105] "action": "Add / Update / Delete / Skip",

[0106] "caption": "Text description + search keywords"

[0107] "predicted_questions": ["Question 1 that might be asked", "Question 2"]

[0108] }

[0109] JSON stands for JavaScript Object Notation. JavaScript is a lightweight, interpreted or just-in-time compiled high-level programming language. It is a dynamic scripting language that supports multiple programming paradigms such as object-oriented, imperative, and declarative programming.

[0110] Taking supervised fine-tuning as an example, full-parameter fine-tuning mode can be used during training, and the training cycle can be 2-3 epochs. In addition, low-rank adaptation (LoRA) and other efficient parameter fine-tuning techniques can be added as alternatives to reduce computational pressure.

[0111] In step 403, the initial multimodal large model is evaluated for performance. In response to the performance evaluation being passed, the initial multimodal large model is determined as the target multimodal large model.

[0112] After obtaining the initial multimodal large model, its performance can be evaluated from aspects such as memory accuracy, retrieval recall, and inference quality. If the performance evaluation is passed, the initial multimodal large model can be identified as the target multimodal large model; otherwise, the initial multimodal large model can be further optimized.

[0113] By adopting the above processing method and using a two-stage training strategy of "supervised fine-tuning + reinforcement learning", the model first masters basic memory planning ability through supervised fine-tuning, and then optimizes long-term decision-making strategy through reinforcement learning, thereby improving the performance of the obtained target multimodal large model and thus improving the accuracy of its output results.

[0114] Furthermore, if necessary, the target multimodal large model can be used to generate image keywords corresponding to visual retrieval requests and text keywords corresponding to text retrieval requests. The content in the training dataset can be adjusted accordingly, such as adding user instructions (image or text description information carried in the retrieval request) to the input and adding image keywords or text keywords to the output.

[0115] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this disclosure. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure. Furthermore, for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0116] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.

[0117] Figure 5 This is a schematic diagram of the structural composition of the first embodiment 500 of the video-based retrieval and processing device described in this disclosure. Figure 5 As shown, it includes: a retrieval module 501 and a storage module 502.

[0118] The retrieval module 501 is used to obtain a retrieval request, determine the target memory unit that matches the retrieval request from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database, each candidate memory unit corresponds to at least one video segment in the same video, and different videos correspond to at least one candidate memory unit, the unit content includes: visual feature pairs and retrieval keywords, the visual feature pairs include: keyframes of the corresponding video segment and text description, and generate retrieval results corresponding to the retrieval request based on the target memory unit.

[0119] Storage module 502 is used to store the underlying database.

[0120] In some embodiments of this disclosure, the content of each candidate memory unit may further include: a time-domain index, which includes the start and end timestamps of the corresponding video segment. In addition, the search keywords may include keywords of four different dimensions: time, people, things, and events extracted from the corresponding video segment.

[0121] Figure 6 This is a schematic diagram of the structural composition of the second embodiment 600 of the video-based retrieval processing device described in this disclosure. Figure 6 As shown, it includes: a retrieval module 501, a storage module 502, and a memory management module 503.

[0122] The memory management module 503 can divide any video into a series of video segments according to a predetermined duration, and can perform the following first processing on each video segment in chronological order from first to last: generate the decision result corresponding to the video segment, and in response to determining the candidate memory unit that needs to be added based on the decision result, generate the candidate memory unit corresponding to the video segment and store it in the underlying database.

[0123] In some embodiments of this disclosure, for any video segment, the decision result may include: the temporal index, text description, and search keywords of the video segment, as well as the decision action and the thought process information for generating the decision action. Accordingly, the memory management module 503 may generate and store the candidate memory unit corresponding to the video segment in the underlying database in response to determining that a new candidate memory unit needs to be added based on the decision result.

[0124] In some embodiments of this disclosure, for any video segment, the decision action may further include: an update action, a deletion action, and a skip action. Accordingly, in response to the decision action being an update action or a deletion action, the decision result may further include: operation instruction information, used to indicate the pending memory units that need to be updated or deleted in the candidate memory units corresponding to the video to which the video segment belongs. Further, in response to the decision action being an update action, the memory management module 503 may update the pending memory units according to the decision result and the video segment; in response to the decision action being a deletion action, it may delete the pending memory units; and in response to the decision action being a skip action, it may directly process the next video segment.

[0125] In some embodiments of this disclosure, the retrieval module 501 determines the target memory unit that matches the retrieval request from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database. This may include: determining the target question corresponding to the retrieval request; and, in response to the target question not meeting the requirements for fast retrieval, determining the target memory unit from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database.

[0126] Accordingly, Figure 7 This is a schematic diagram of the structural composition of the third embodiment 700 of the video-based retrieval processing device described in this disclosure. Figure 7 As shown, it includes: a retrieval module 501, a storage module 502, a memory management module 503, and a prediction module 504.

[0127] The prediction module 504 can generate at least one corresponding prediction question for each candidate memory unit in the underlying database. Accordingly, the prediction module 504 can obtain the similarity between the target question and each prediction question. In response to each similarity being less than or equal to a first threshold, it can be determined that the target question does not meet the requirements for fast retrieval, and the retrieval module 501 can be notified.

[0128] In some embodiments of this disclosure, the prediction problem generated by the prediction module 504 for any candidate memory unit may include: a prediction problem generated for the video segment corresponding to the candidate memory unit, and a prediction problem generated for the key frames of interest determined from each key frame of the candidate memory unit.

[0129] In addition, in some embodiments of this disclosure, the prediction module 504 can determine that the target question meets the fast retrieval requirements in response to the similarity between the target question and at least one predicted question being greater than a first threshold. It can also determine the candidate memory unit corresponding to the predicted question with a similarity greater than the first threshold as the target memory unit, or determine the candidate memory unit corresponding to the predicted question with the highest similarity as the target memory unit.

[0130] In addition, in some embodiments of this disclosure, the retrieval module 501 determines the target memory unit from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database in the following ways: in response to a visual retrieval request, it determines the corresponding image keywords based on the image carried in the visual retrieval request, and determines the target memory unit based on the image, the image keywords, and the unit content of each candidate memory unit; in response to a text retrieval request, it determines the corresponding text keywords based on the text description information carried in the text retrieval request, and determines the target memory unit based on the text keywords and the unit content of each candidate memory unit.

[0131] In some embodiments of this disclosure, the memory management module 503 generates decision results corresponding to video segments in a manner that may include: using a target multimodal large model to generate text descriptions, search keywords, decision actions, and thought process information in the decision results corresponding to the video segments. Additionally, the prediction module 504 generates at least one corresponding prediction question for each candidate memory unit in the underlying database in a manner that may include: generating at least one corresponding prediction question for each candidate memory unit using the target multimodal large model. Further, if necessary, the retrieval module 501 may also use the target multimodal large model to generate image keywords corresponding to visual retrieval requests and text keywords corresponding to text retrieval requests.

[0132] Accordingly, Figure 8 This is a schematic diagram of the structural composition of the fourth embodiment 800 of the video-based retrieval processing device described in this disclosure. Figure 8 As shown, it includes: a retrieval module 501, a storage module 502, a memory management module 503, a prediction module 504, and a model acquisition module 505.

[0133] The model acquisition module 505 can obtain the target multimodal large model in the following ways: determine the basic multimodal large model, use the corresponding training sample set to perform supervised fine-tuning and reinforcement learning fine-tuning on the basic multimodal large model to obtain the initial multimodal large model, perform performance evaluation on the initial multimodal large model, and in response to the performance evaluation passing, determine the initial multimodal large model as the target multimodal large model.

[0134] In practical applications, the retrieval module 501, memory management module 503, and prediction module 504 can be represented as a retrieval agent, a memory management agent, and a predictive agent, respectively. All three agents can be driven by a large multimodal target model, each performing its own function while coordinating with each other. The memory management agent dynamically maintains candidate memory units, the retrieval agent performs accurate parsing and matching of visual / text retrieval, and the predictive agent generates prediction questions to achieve proactive memorization and rapid response. This multi-agent collaboration overcomes the technical bottlenecks of traditional video understanding, enabling efficient management and accurate retrieval of videos, especially long videos.

[0135] The specific workflow of each of the above device embodiments can be found in the relevant descriptions in the foregoing method embodiments, and will not be repeated here.

[0136] The solution described in this disclosure can be applied to various scenarios such as smart home monitoring, smart city and public safety retrieval, video content creator assistant, intelligent teaching aids and online education, autonomous driving behavior analysis, and medical surgical video archiving. It has wide applicability and lays a solid technical foundation for providing more personalized and efficient intelligent video services.

[0137] For example, the smart home care scenario could be a home intelligent monitoring system for elderly people living alone or infants. The camera can collect real-time, 24 / 7 video streams, sending video clips to the system at fixed intervals (e.g., every 10 seconds). If it detects that the elderly person took a white pill in the living room at 2 PM, a new action can be triggered. The generated candidate memory unit can include a keyframe of taking the pill, a text description of "the elderly person sitting on the living room sofa, holding a white pill bottle in their right hand, pouring out a white pill and putting it in their mouth, then picking up a water glass to drink water," and can generate search keywords such as "Time: 2 PM; Person: Elderly male; Objects: Pill bottle, white pill, water glass; Event: Taking medication, drinking water." Additionally, a predictive question corresponding to the candidate memory unit, such as "Did you take your medicine today?", can be pre-generated. This allows children to ask after get off work, "Did my dad take his medicine this afternoon?" Correspondingly, it can achieve second-level location tracking, directly retrieving the corresponding candidate memory unit and generating a voice reply: "Yes, the elderly person took medication in the living room at 2 PM (with a keyframe image as evidence)."

[0138] For example, the smart city and public safety retrieval scenario can be an urban security scenario. Law enforcement officers can input a text retrieval request such as "find a person wearing red clothes and riding a blue shared bicycle". Correspondingly, key frames that match visual features can be quickly matched from thousands of distributed candidate memory units, and the trajectory path of the target can be planned according to the timeline, which greatly improves the efficiency of handling cases.

[0139] The solutions described in this disclosure can be applied to the field of artificial intelligence, particularly in areas such as video understanding, computer vision, natural language processing, deep learning, and large-scale models. Artificial intelligence is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. Artificial intelligence hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0140] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0141] Figure 9 A schematic block diagram of an electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0142] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0143] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0144] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as those described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the methods described in this disclosure can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the methods described herein by any other suitable means (e.g., by means of firmware).

[0145] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0147] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0149] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0150] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0151] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0152] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.< / action> < / think> < / memory> < / time>

Claims

1. A video-based retrieval and processing method, comprising: Obtain the search request; Based on the content of each candidate memory unit stored in the underlying database, the target memory unit that matches the retrieval request is determined from each candidate memory unit. Each candidate memory unit corresponds to at least one video segment in the same video, and different videos correspond to at least one candidate memory unit. The content of the unit includes: visual feature pairs and retrieval keywords. The visual feature pairs include: keyframes of the corresponding video segment and text descriptions. The search results corresponding to the search request are generated based on the target memory unit.

2. The method according to claim 1, wherein, The unit content also includes: a time domain index, which includes the start timestamp and end timestamp of the corresponding video segment; The search keywords include keywords extracted from the corresponding video clips based on four different dimensions: time, people, objects, and events.

3. The method according to claim 2, further comprising: Before obtaining the retrieval request, for any given video, it is divided into a series of video segments according to a predetermined duration, and the following first processing is performed on each video segment in chronological order from first to last: Generate a decision result corresponding to the video segment, and in response to determining that a new candidate memory unit needs to be added based on the decision result, generate the candidate memory unit corresponding to the video segment and store it in the underlying database.

4. The method according to claim 3, wherein, The decision result includes: the temporal index of the video segment, the text description and the search keywords, as well as the decision action and the thought process information that generated the decision action; The step of responding to determining that a new candidate memory unit needs to be added based on the decision result, generating a candidate memory unit corresponding to the video segment and storing it in the underlying database includes: responding to the decision action being a new action, generating a candidate memory unit corresponding to the video segment and storing it in the underlying database based on the decision result and the video segment.

5. The method according to claim 4, wherein, The decision-making actions also include: update actions, delete actions, and skip actions; In response to the decision action being the update action or the deletion action, the decision result further includes: operation instruction information, used to indicate the pending memory units that need to be updated or deleted in the candidate memory units corresponding to the video to which the stored video segment belongs; The first processing further includes: updating the memory unit to be processed according to the decision result and the video segment in response to the decision action being the update action; deleting the memory unit to be processed in response to the decision action being the delete action; and directly processing the next video segment in response to the decision action being the skip action.

6. The method according to claim 3, wherein, The step of determining the target memory unit that matches the retrieval request from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database includes: Identify the target question corresponding to the search request; In response to the fact that the target problem does not meet the requirements for fast retrieval, the target memory unit is determined from each candidate memory unit based on the content of each candidate memory unit stored in the underlying database.

7. The method according to claim 6, further comprising: Before obtaining the retrieval request, at least one prediction question is generated for each candidate memory unit in the underlying database. The response to the target question does not meet the requirements for fast retrieval, including: The similarity between the target question and each predicted question is obtained respectively. In response to each similarity being less than or equal to a first threshold, it is determined that the target question does not meet the fast retrieval requirements.

8. The method according to claim 7, wherein, The prediction problems generated for any candidate memory unit include: the prediction problem for the video segment corresponding to the candidate memory unit, and the prediction problem for the key focus frames determined from each key frame of the candidate memory unit.

9. The method according to claim 7, further comprising: In response to the similarity between the target question and at least one predicted question being greater than the first threshold, the target question is determined to meet the fast retrieval requirements, and the candidate memory unit corresponding to the predicted question with a similarity greater than the first threshold is determined as the target memory unit, or the candidate memory unit corresponding to the predicted question with the highest similarity is determined as the target memory unit.

10. The method according to claim 1, wherein, The step of determining the target memory unit that matches the retrieval request from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database includes: In response to the retrieval request being a visual retrieval request, the corresponding image keywords are determined based on the image carried in the visual retrieval request, and the target memory unit is determined based on the image, the image keywords, and the content of each candidate memory unit; In response to the retrieval request being a text retrieval request, the corresponding text keywords are determined based on the text description information carried in the text retrieval request, and the target memory unit is determined based on the text keywords and the unit content of each candidate memory unit.

11. The method according to claim 7, wherein, The generation of the decision result corresponding to the video segment includes: using a target multimodal large model to generate the text description, the search keywords, the decision action, and the thinking process information in the decision result corresponding to the video segment; The step of generating at least one prediction question for each candidate memory unit in the underlying database includes: generating at least one prediction question for each candidate memory unit using the target multimodal large model.

12. The method according to claim 11, wherein, The methods for obtaining the target multimodal large model include: Determine the basic multimodal large model; The basic multimodal large model is then subjected to supervised fine-tuning and reinforcement learning fine-tuning using the corresponding training sample set to obtain the initial multimodal large model. The initial multimodal large model is subjected to performance evaluation, and in response to the performance evaluation being passed, the initial multimodal large model is determined as the target multimodal large model.

13. A video-based retrieval processing apparatus, comprising: Retrieval module and storage module; The retrieval module is used to obtain a retrieval request, determine the target memory unit that matches the retrieval request from each candidate memory unit based on the unit content of each candidate memory unit stored in the underlying database, each candidate memory unit corresponds to at least one video segment in the same video, and different videos correspond to at least one candidate memory unit, the unit content includes: visual feature pairs and retrieval keywords, the visual feature pairs include: keyframes of the corresponding video segment and text description, and generate retrieval results corresponding to the retrieval request based on the target memory unit; The storage module is used to store the underlying database.

14. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.

16. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method of any one of claims 1-12.