Offline video pre-analysis method, retrieval method, system and equipment based on large and small model collaboration and storage medium

By collaboratively analyzing offline videos with large and small models, combined with semantic understanding and vector databases, the problem of low retrieval efficiency in massive video data is solved, and fast and accurate video content retrieval is achieved.

CN120804365AActive Publication Date: 2025-10-17CHENGDU KOALA URAN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511293586.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-17
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently utilize video content in massive video data, traditional retrieval methods are difficult to meet user needs, and large models have high computing resource requirements and high dependence on training data, while the effectiveness of small model algorithm lists is limited.

Method used

The collaborative analysis method of large and small models is adopted. Through semantic understanding of offline videos, pre-analysis and retrieval are performed by combining the algorithm lists of large and small models. The large model is used to cover user unknown algorithms, and the small model is used for precise analysis. The results are stored in a vector database.

Benefits of technology

The efficiency of video analysis and retrieval has been improved, and users can quickly return relevant results in subsequent searches, reducing the need for full video analysis and improving retrieval speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804365A_ABST
    Figure CN120804365A_ABST
Patent Text Reader

Abstract

The invention discloses an off-line video pre-analysis method, a retrieval method, a retrieval system, equipment and a storage medium based on large and small model collaboration, and the method comprises the steps: obtaining a union set of an input algorithm list set and a recommendation algorithm list set, and obtaining a total calculation method list set; obtaining a small model algorithm list set, and obtaining an intersection of the small model algorithm list set and the total calculation method list set to obtain a first algorithm list set; obtaining a second algorithm list set according to the total calculation method list set and the first algorithm list set; issuing the first algorithm list set to the small model to analyze the offline video, generating an alarm after a corresponding algorithm is analyzed, and storing the alarm in a small model alarm library; and performing second algorithm list set analysis on the offline video through the large model, performing character description on an analysis result, converting the character description into a vector, and storing the vector into a vector database. According to the method, a large model and small model collaborative analysis mode is adopted, the advantages of the small model and the large model are fully utilized, and the video analysis and retrieval performance is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to an offline video pre-analysis method, a retrieval method, a system, a device and a storage medium based on size model cooperation. BACKGROUND

[0002] With the explosive growth of video data, how to effectively utilize the video data has become a major technical difficulty. At present, the related video content is usually queried by searching the video name, or some text description information input in advance by the user, and the text description information is retrieved based on the traditional processing mode. It is difficult to realize the retrieval and search of the video content in the massive video, and the effective utilization rate of the video is low.

[0003] With the development of artificial intelligence technology, VLM large models, CV models and small models are applied to video retrieval. The small model is a model obtained by data annotation, training and optimization processing for related algorithms, and has better effect than the VLM large model in the accuracy and recall rate of the algorithm, but the small model can analyze the effective algorithm list, and often cannot meet the user's retrieval demand. The large model has obvious advantages over the small model in semantic understanding and processing capacity, but has the problems of high demand for computing resources and high dependence on training data. SUMMARY

[0004] The purpose of the present application is to provide an offline video pre-analysis method, a retrieval method, a system, a device and a storage medium based on size model cooperation, to solve the above problems existing in the retrieval and analysis of the video by the large model and the small model.

[0005] The present application is realized by the following technical scheme: The offline video pre-analysis method based on size model cooperation comprises the following steps: Performing semantic understanding on the structured information input for the offline video to obtain an input algorithm list set {A1}; Analyzing the offline video by using a large model to obtain a recommended algorithm list set {A2}; Taking the union of the input algorithm list set {A1} and the recommended algorithm list set {A2} to obtain a total algorithm list set {A}; Obtaining a small model algorithm list set {B}, taking the intersection of the small model algorithm list set {B} and the total algorithm list set {A} to obtain a first algorithm list set, represented as {A}∩{B}; and obtaining a second algorithm list set according to the total algorithm list set {A} and the first algorithm list set {A}∩{B}, represented as {A}-{A}∩{B}; The first algorithm list set is sent to the small model to analyze the offline video, and an alarm is generated after the corresponding algorithm is analyzed and stored in the small model alarm library. The second algorithm list set analysis of the offline video is performed by the large model, a text description is made on the analysis result, the text description is converted into a vector, and the vector is stored in a vector database.

[0006] In some embodiments of the present application, the step of analyzing the offline video by using the large model to obtain the recommended algorithm list set {A2} comprises: n pictures are extracted at equal intervals from the offline video, and one picture is extracted at a position of a first set time before and after each picture; The large model interface is called n times by the prompt method, and the three pictures extracted at intervals are analyzed by the large model to obtain the recommended algorithm list set {A2}.

[0007] In some embodiments of the present application, the step of performing the second algorithm list set analysis of the offline video by using the large model comprises: Three pictures are extracted at equal intervals from the first frame of the offline video, and the large model is called to perform the second algorithm list set analysis on the extracted pictures; From the time of the third picture, the third set time is slid, and the operation of calling the large model to perform the second algorithm list set analysis on the extracted pictures is repeated; the operation is repeated in turn until the analysis of the entire offline video is completed.

[0008] In some embodiments of the present application, the retrieval of the offline video processed by using the offline video pre-analysis method based on the large model cooperation comprises the following steps: The retrieval content input by the user is analyzed by the large model, the retrieval content is split and standardized into a new retrieval format, and is expressed as:

first algorithm list set

other algorithm list set

first algorithm list set

other algorithm list set

other algorithm list set

[0009] On the other hand, the present application also provides an offline video pre-analysis system based on the large model cooperation, which is used to execute the offline video pre-analysis method based on the large model cooperation, and comprises: The first module is used to call the large model to perform semantic understanding on the structured information input for the offline video to obtain an input algorithm list set {A1}; The second module is used to call the large model to analyze the offline video to obtain a recommended algorithm list set {A2}; The third module is configured to obtain a total algorithm list set {A} by taking a union of the input algorithm list set {A1} and the recommended algorithm list set {A2}; The fourth module is configured to obtain a small model algorithm list set {B}, take an intersection of the small model algorithm list set {B} and the total algorithm list set {A} to obtain a first algorithm list set, denoted as {A}∩{B}, and obtain a second algorithm list set according to the total algorithm list set {A} and the first algorithm list set {A}∩{B}, denoted as {A}-{A}∩{B}; The fifth module is configured to distribute the first algorithm list set to the small model to analyze the offline video, generate an alarm after analyzing the corresponding algorithm, and store the alarm in a small model alarm database. The sixth module is configured to call the large model to analyze the offline video according to the second algorithm list set, describe the analysis result in words, convert the words into a vector, and store the vector in a vector database.

[0010] In another aspect, the present application also provides an offline video retrieval system based on the cooperation of large and small models, which is used to execute the offline video retrieval method based on the cooperation of large and small models, and includes: The seventh module is configured to call the large model to analyze the retrieval content input by a user in one sentence, split and standardize the retrieval content into a new retrieval format; The eighth module is configured to retrieve in the small model alarm database by using the first algorithm list set to obtain a first retrieval result. The ninth module is configured to convert the other algorithm list set into a vector, and retrieve in the vector database by taking the other algorithm list set as a retrieval condition to obtain a second retrieval result. The tenth module is configured to aggregate the first retrieval result and the second retrieval result and output.

[0011] In another aspect, the present application also provides an electronic device, which includes: a processor; and a memory configured to store executable instructions of the processor; The processor is configured to execute the offline video pre-analysis method based on the cooperation of large and small models by executing the executable instructions.

[0012] In another aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the offline video pre-analysis method based on the cooperation of large and small models.

[0013] In another aspect, the present application also provides an electronic device, which includes: a processor; and a memory configured to store executable instructions of the processor; The processor is configured to perform the offline video retrieval method based on the size model cooperation via executing the executable instructions.

[0014] In another aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the offline video retrieval method based on the size model cooperation.

[0015] Compared with the prior art, the present application has the following advantages and beneficial effects: The present application adopts the way of large model and small model cooperative analysis, and the small model cannot analyze the algorithm, so the large model is used to analyze the algorithm, which fully utilizes the advantages of the small model and the large model, optimizes the performance of video analysis and retrieval, and through the algorithm preprocessing of the offline video by using the large model and the small model respectively, the user can quickly return the retrieval result of the algorithm concerned by the user by asking the question during the subsequent retrieval, without analyzing the video from the beginning to the end according to the user's question every time, which greatly improves the retrieval speed. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings in the embodiments will be briefly introduced below, and it should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0017] Figure 1 The present application is an offline video pre-analysis method flowchart in the embodiment.

[0018] Figure 2 The present application is an offline video retrieval method flowchart in the embodiment. DETAILED DESCRIPTION

[0019] In order to make the purposes, technical solutions and advantages of the present application clearer, specific embodiments of the present application are further described in detail below in combination with the drawings. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that, for the convenience of description, only parts related to the present application are shown in the drawings, but not all. Before discussing the example embodiments in more detail, it should be mentioned that some example embodiments are described as processes or methods depicted as flowcharts. Although the flowchart describes each operation (or step) as a sequential process, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0020] In the vast amount of video data, quickly and accurately retrieving the video data required by the user is a hot direction in the field of video processing technology. With the development of artificial intelligence technology, artificial intelligence technology is applied to the processing and retrieval of video data. In the retrieval processing of video, the model needs to quickly analyze the algorithm (including target, event, scene) concerned by the user, and then retrieve the related content.

[0021] When the user uploads an offline video, the user enters the algorithm (including target, event, scene) concerned by himself through a sentence, so that the large model can focus on the corresponding algorithm during analysis. Otherwise, since the large model cannot output the algorithm without attention during analysis, the user may not be able to retrieve the related video. For example, when the user inputs "retrieve the video with the situation of riding a bike without a helmet", if the large model does not pay attention to the algorithm of "riding a bike without a helmet" in advance, the large model may only analyze the situation of riding a bike when analyzing the video, and often cannot retrieve the content of "riding a bike without a helmet" in the video.

[0022] The small model is processed by data labeling, training and optimization for related algorithms, and the accuracy and recall rate of the algorithm are better, but the list of algorithms that can be analyzed by the small model is limited.

[0023] Therefore, the present application adopts the cooperative analysis mode of large model and small model, and the large model analyzes the algorithms that cannot be analyzed by the small model, so as to cover more algorithms that the user wants to analyze. The present application analyzes the known algorithm of the user, quickly analyzes the video and quickly recommends the algorithm through the large model, combines the small model algorithm, and then classifies all the obtained algorithms, so that the small model and the large model respectively analyze the offline video based on the classified algorithms, to facilitate subsequent retrieval.

[0024] For example, the list of algorithms that can be analyzed by the small model is: fighting, crowd gathering, and vehicle, which correspond to the algorithm set {B}.

[0025] When a user uploads an offline video, the list of algorithms of interest is: congestion, fighting, and urban flooding, which correspond to the algorithm set {A1}.

[0026] The recommended list of algorithms for the video analyzed by the large model is: congestion and garbage overflow, which correspond to the algorithm set {A2}.

[0027] The algorithm set {A1} is merged with the algorithm set {A2} to obtain the algorithm set {A}: congestion, fighting, urban flooding, and garbage overflow.

[0028] The small model processes the algorithms corresponding to {A}∩{B}, i.e., the small model only processes the "fighting" algorithm; the rest is processed by the large model, i.e., the large model processes the algorithms corresponding to {A}-{A}∩{B}, which are "congestion, urban flooding, and garbage overflow", to pre-analyze the offline video.

[0029] The above-mentioned collaborative analysis method of the large model and the small model can cover the algorithms (targets, events, and scenes) that the user wants to analyze.

[0030] Referring to Figure 1 In some embodiments of the present application, the offline video pre-analysis method based on the collaborative analysis of the large and small models includes the following steps: S11, the structured information entered by the user when uploading the offline video, such as a one-sentence description of the algorithm content of interest, is analyzed for semantic understanding by calling the large model interface through the prompt method, and a list of entered algorithms {A1} is obtained; S12, according to the length of the offline video, n pictures (such as 10 pictures) are extracted at equal intervals from the offline video, and one picture is extracted at a position 1s before and after each picture; The large model interface is called n times through the prompt method, and the three pictures extracted at intervals are analyzed by the large model to obtain a list of recommended algorithms {A2}; S13, the union of the list of entered algorithms {A1} and the list of recommended algorithms {A2} is obtained to obtain a total list of algorithms {A}; S14, the list of small model algorithms {B} is obtained, and the intersection of the list of small model algorithms {B} and the total list of algorithms {A} is obtained to obtain a first list of algorithms, denoted as {A}∩{B}; According to the total list of algorithms {A} and the first list of algorithms {A}∩{B}, a second list of algorithms is obtained, denoted as {A}-{A}∩{B}; S15, the first algorithm list set is issued to the small model to analyze the offline video, and after the algorithm corresponding to the first algorithm list set is analyzed, an alarm is generated and stored in the small model alarm library; S16, starting from the first frame of the offline video, every interval of the first set time (such as 2s) extracts 3 pictures, and the extracted pictures are analyzed by the large model through the prompt method to analyze the second algorithm list set; From the time of the third picture (the last picture of the last extraction), slide the second set time (such as 10s), repeat the step of analyzing the extracted pictures by the large model through the prompt method to analyze the second algorithm list set; and then repeat the operation in turn until the analysis of the entire offline video is completed; The result analyzed by the large model each time is described in words, the obtained word description is converted into a vector by calling embedding, and stored in a vector database.

[0031] At this point, the pre-analysis of the offline video is completed by the cooperative operation of the small model and the large model.

[0032] Based on the pre-analysis of the offline video described above, the small model alarm library and the vector database are obtained, and on this basis, some embodiments of the present application provide an offline video retrieval method, with reference to Figure 2 , comprising the following steps: S21, the user inputs a sentence of retrieval content, identifies the retrieval intention of the user by calling the large model through the prompt method, splits and standardizes the key content of the retrieval into a new retrieval format, and is expressed as:

first algorithm list set

other algorithm list set

other algorithm list set

first algorithm list set

first algorithm list set

other algorithm list set

other algorithm list set

[0033] After the offline video is preprocessed by the large model and the small model, the user can quickly return the result through retrieval by related questions in subsequent retrieval, without analyzing the video from beginning to end according to the user's question each time, which greatly improves the retrieval speed.

[0034] In another aspect, the present application also provides an offline video pre-analysis system based on large-small model cooperation, which is used to execute the offline video pre-analysis method based on large-small model cooperation, comprising: A first module is configured to call a large model to perform semantic understanding on structured information input for an offline video, to obtain an input algorithm list set {A1}; A second module is configured to call a large model to analyze the offline video, to obtain a recommended algorithm list set {A2}; A third module is configured to take a union set of the input algorithm list set {A1} and the recommended algorithm list set {A2}, to obtain a total algorithm list set {A}; A fourth module is configured to obtain a small model algorithm list set {B}, to take an intersection set of the small model algorithm list set {B} and the total algorithm list set {A}, to obtain a first algorithm list set, denoted as {A}∩{B}; and to obtain a second algorithm list set according to the total algorithm list set {A} and the first algorithm list set {A}∩{B}, denoted as {A}-{A}∩{B}; A fifth module is configured to distribute the first algorithm list set to a small model to analyze the offline video, to generate an alarm after analyzing the corresponding algorithm, and to store the alarm in a small model alarm database; A sixth module is configured to call a large model to analyze the offline video according to the second algorithm list set, to perform a text description on the analysis result, to convert the text description into a vector, and to store the vector in a vector database.

[0035] In another aspect, the present application also provides an offline video retrieval system based on large-small model cooperation, which is used to execute the offline video retrieval method based on large-small model cooperation, comprising: A seventh module is configured to call a large model to analyze retrieval content input by a user in one sentence, to split and standardize the retrieval content into a new retrieval format; An eighth module is configured to retrieve in a small model alarm database by using the first algorithm list set, to obtain a first retrieval result; A ninth module is configured to convert other algorithm list sets into vectors, to retrieve in a vector database by using the other algorithm list sets as retrieval conditions, to obtain a second retrieval result; A tenth module is configured to aggregate the first retrieval result and the second retrieval result and to output.

[0036] In another aspect, the present application also provides an electronic device, comprising: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the offline video pre-analysis method based on large-small model cooperation by executing the executable instructions.

[0037] In another aspect, the present application also provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the offline video pre-analysis method based on size model collaboration.

[0038] In another aspect, the present application also provides an electronic device comprising: a processor; and, a memory for storing executable instructions of the processor; wherein the processor is configured to execute the offline video retrieval method based on size model collaboration via executing the executable instructions.

[0039] In another aspect, the present application also provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the offline video retrieval method based on size model collaboration.

[0040] The above description is only the preferred embodiment of the present application, and does not limit the present application in any form. Any simple modification or equivalent change of the above embodiment according to the technical essence of the present application falls within the protection scope of the present application.

Claims

1. Offline video pre-analysis method based on collaboration of large and small models, characterized by: The following steps are involved: Perform semantic understanding on the structured information recorded for offline videos to obtain a list of recorded algorithms {A1}; Use a large model to analyze offline videos and obtain a list of recommended algorithms {A2}; Take the union of the input algorithm list set {A1} and the recommended algorithm list set {A2} to obtain the total algorithm list set {A}; Obtain the small model algorithm list set {B}, take the intersection of the small model algorithm list set {B} and the total algorithm list set {A} to obtain the first algorithm list set, expressed as {A}∩{B}; and obtain the second algorithm list set, expressed as {A} - {A}∩{B}, based on the total algorithm list set {A} and the first algorithm list set {A}∩{B}; The first algorithm list set is sent to the small model to analyze the offline video. After the corresponding algorithm is analyzed, an alarm is generated and stored in the small model alarm library; The offline video is analyzed by the second algorithm list set through the large model, the analysis results are described in text, the text description is converted into a vector, and stored in the vector database.

2. The offline video pre-analysis method based on large and small model collaboration according to claim 1 is characterized in that: The step of using a large model to analyze offline videos to obtain a recommended algorithm list set {A2} includes: Extract n pictures at equal intervals from the offline video, and extract one picture at the first set time interval before and after each picture; The large model interface is called n times through the prompt method. The three images extracted at intervals are analyzed by the large model to obtain a list of recommended algorithms {A2}.

3. The offline video pre-analysis method based on large and small model collaboration according to claim 1 is characterized in that: The step of performing a second algorithm list set analysis on the offline video using the large model includes: Starting from the first frame of the offline video, three pictures are continuously extracted at intervals of the second set time, and the extracted pictures are analyzed by the large model using the second algorithm list set; Starting from the moment of the third picture, slide the third set time and repeat the operation of calling the large model to perform the second algorithm list set analysis on the extracted pictures; repeat in sequence until the analysis of the entire offline video is completed.

4. Offline video retrieval method based on collaboration of large and small models, characterized by: The method for retrieving an offline video processed by the offline video pre-analysis method based on large and small model collaboration according to any one of claims 1 to 3 comprises the following steps: The large model analyzes the search content of a sentence input by the user, splits the search content and standardizes it into a new search format, expressed as: [first algorithm list set] or\and [other algorithm list sets]; Use the [first algorithm list set] to search in the small model alarm library to obtain the first search result; Convert the [other algorithm list set] into a vector, use the [other algorithm list set] as the search condition, search in the vector database, and obtain the second search result; The first search result and the second search result are summarized and output.

5. Offline video pre-analysis system based on collaboration of large and small models, characterized by: The method for performing offline video pre-analysis based on large and small model collaboration according to any one of claims 1 to 3 comprises: The first module is used to call the large model to perform semantic understanding on the structured information recorded for the offline video, and obtain a list set of recorded algorithms {A1}; The second module is used to call the large model to analyze the offline video and obtain a list of recommended algorithms {A2}; The third module is used to obtain the union of the input algorithm list set {A1} and the recommended algorithm list set {A2} to obtain the total algorithm list set {A}; The fourth module is used to obtain the small model algorithm list set {B}, take the intersection of the small model algorithm list set {B} and the total algorithm list set {A} to obtain the first algorithm list set, expressed as {A}∩{B}; and obtain the second algorithm list set based on the total algorithm list set {A} and the first algorithm list set {A}∩{B}, expressed as {A} - {A}∩{B}; The fifth module is used to send the first algorithm list set to the small model to analyze the offline video. After the corresponding algorithm is analyzed, an alarm is generated and stored in the small model alarm library; The sixth module is used to call the large model to perform a second algorithm list set analysis on the offline video, to describe the analysis results in text, to convert the text description into a vector, and to store it in a vector database.

6. Offline video retrieval system based on large and small model collaboration, characterized by: The method for performing the offline video retrieval method based on large and small model collaboration as claimed in claim 4 comprises: The seventh module is used to call the large model to analyze the search content of the user's input sentence, split the search content and standardize it into a new search format; The eighth module is used to use the first algorithm list set to search in the small model alarm library to obtain a first search result; The ninth module is used to convert the [other algorithm list set] into a vector, and use the [other algorithm list set] as a search condition to search in the vector database to obtain a second search result; The tenth module is used to summarize the first search result and the second search result and output them.

7. An electronic device, characterized in that include: processor; as well as, a memory for storing executable instructions of the processor; The processor is configured to execute the offline video pre-analysis method based on large and small model collaboration according to claims 1-3 by executing the executable instructions.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the offline video pre-analysis method based on large and small model collaboration described in claims 1-3 is implemented.

9. An electronic device, characterized in that include: processor; as well as, a memory for storing executable instructions of the processor; The processor is configured to execute the offline video retrieval method based on large-small model collaboration according to claim 4 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the offline video retrieval method based on large and small model collaboration according to claim 4 is implemented.

Citation Information

Patent Citations

  • Movable target retrieval device and retrieval method for large and small model fusion

    CN119739891A

  • Design method of large model and small model combined image processing system

    CN119785187A

  • Video dotting placement analysis system, analysis method and storage medium

    US20190050890A1