Intelligent video retrieval method and system based on AI
By using an AI-based intelligent video retrieval method that combines image recognition and speech semantic analysis, the problem of low efficiency in traditional video retrieval technology is solved, achieving efficient and accurate video content retrieval and adapting to the retrieval needs of various application scenarios.
Patent Information
- Application Number
- CN202511313965.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-12-19
AI Technical Summary
Traditional video retrieval technologies are inefficient, struggle to deeply understand the semantic information of videos, and are unable to efficiently and accurately retrieve the key content that users need from massive amounts of video data, especially in cases of fuzzy video queries or cross-modal retrieval needs.
An AI-based intelligent video retrieval method is adopted, which segments and describes video content through image recognition and speech semantic analysis, generates content tags and evaluates their usability, establishes an auxiliary retrieval table, optimizes the retrieval index structure, and achieves semantic-level query matching.
It improves the efficiency and relevance of video retrieval, enhances the adaptability to natural language queries, and can quickly filter out high-value video content to meet the retrieval needs of different application scenarios.
Smart Images

Figure CN121166971A_ABST
Abstract
Description
Technical Field
[0001] Several embodiments of this specification relate to the field of information technology, specifically to an AI-based intelligent video retrieval method and system. Background Technology
[0002] With the rapid development of video surveillance technology, especially the widespread application of high-definition cameras, intelligent sensing devices, and 5G networks, the volume of video data has shown a significant growth trend. According to relevant statistics, hundreds of millions of hours of video content are generated and stored every day, covering multiple fields such as public safety, traffic management, smart homes, and industrial production. Faced with such massive, high-dimensional, and unstructured video information streams, how to efficiently and accurately retrieve the key content needed by users has become a crucial and highly challenging research topic in the field of information processing. Traditional video retrieval technologies mainly rely on manually annotated metadata (such as video titles, timestamps, location information, user-added tags, etc.) or text matching methods based on simple keywords. However, these methods are inefficient, difficult to cover massive amounts of video content, and easily affected by subjective factors. On the other hand, relying solely on keyword matching cannot deeply understand the semantic information of the video, making it difficult to capture complex scenes, actions, or emotional expressions in the footage, resulting in low retrieval accuracy and difficulty in handling fuzzy video queries or cross-modal retrieval needs. Therefore, it is necessary to research technologies that can achieve video retrieval more quickly and conveniently. Summary of the Invention
[0003] This specification describes an AI-based intelligent video retrieval method and system through several embodiments.
[0004] Firstly, embodiments of this specification provide an AI-based intelligent video retrieval method, including the following steps:
[0005] Read the video file, identify the content of the video file, segment the video file according to the content, and generate a content description for each segment;
[0006] Based on the content description and preset application scenarios, generate content tags;
[0007] Based on the content description of the segment, the usability of the video file in the preset application scenario is generated;
[0008] Video files whose availability is higher than a preset reference value are designated as the preferred set.
[0009] Based on the segmented content description, for each video file that does not belong to the preferred set, match at least one video file that belongs to the preferred set to obtain a matching set of video files belonging to the preferred set;
[0010] An auxiliary table is established, which records the filename, storage address, content description, content tags, and matching set of the video files in the preferred set;
[0011] Receive a search request description, and identify search tags and search scenarios based on the search request description;
[0012] The matching rows in the auxiliary table are obtained based on the search tags, and the search results are obtained based on the matching rows.
[0013] Secondly, embodiments of this specification provide an AI-based intelligent video retrieval system, including:
[0014] The reading module reads the video file, identifies the content of the video file, segments the video file according to the content, and generates a content description for each segment.
[0015] The generation module generates content tags based on the content description and preset application scenarios, and generates the usability of the video file in the preset application scenarios based on the segmented content descriptions.
[0016] The marking module marks video files whose availability is higher than a preset reference value and designates them as the preferred set;
[0017] The matching module, based on the content description of the segments, matches at least one video file that belongs to the preferred set for each video file that does not belong to the preferred set, thereby obtaining a matching set of video files belonging to the preferred set;
[0018] The auxiliary module establishes an auxiliary table, which records the filenames, storage addresses, content descriptions, content tags, and matching sets of the video files in the preferred set.
[0019] The retrieval module receives a retrieval request description and identifies retrieval tags and retrieval scenarios based on the retrieval request description.
[0020] The results module obtains matching rows in the auxiliary table based on the search tags, and obtains search results based on the matching rows.
[0021] Thirdly, embodiments of this specification provide an electronic device, including a processor and a memory;
[0022] The processor is connected to the memory;
[0023] The memory is used to store executable program code;
[0024] The processor runs a program corresponding to the executable program code stored in the memory to perform the method described in any of the above aspects.
[0025] Fourthly, embodiments of this specification provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the above aspects.
[0026] Fifthly, embodiments of this specification provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in any of the above aspects.
[0027] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:
[0028] In several embodiments of this specification, an AI-based intelligent video retrieval method and system are provided. This method combines image recognition and speech semantic analysis for multimodal content understanding, segmenting and describing video content. It overcomes the semantic gaps caused by relying solely on metadata or keyword matching, improving the accuracy and completeness of content representation. By setting content tags according to preset application scenarios and generating usability assessments, it can filter high-value videos to form a preferred set and establish an auxiliary retrieval table containing matching relationships, optimizing the retrieval index structure and improving retrieval efficiency and result relevance. It automatically identifies scenarios and tags based on user search requests, achieving semantic-level query matching and enhancing adaptability to natural language queries.
[0029] Other features and advantages of various embodiments of this specification will be further revealed in the following detailed description and accompanying drawings. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a schematic diagram of intelligent video retrieval provided for an embodiment of this specification.
[0032] Figure 2 This is a schematic diagram of the intelligent video retrieval method provided in the embodiments of this specification.
[0033] Figure 3 This is a schematic diagram of video file segmentation provided for embodiments of this specification.
[0034] Figure 4 This is a schematic diagram of video retrieval provided for an embodiment of this specification.
[0035] Figure 5 This is a schematic diagram of an intelligent video retrieval system provided in the embodiments of this specification.
[0036] Figure 6 A schematic diagram of an electronic device provided in an embodiment of this specification. Detailed Implementation
[0037] The technical solutions of the embodiments of this specification will be explained and described below with reference to the accompanying drawings. However, the following embodiments are only preferred embodiments of this specification and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments in the implementation methods without creative effort are all within the protection scope of this specification.
[0038] The terms "first," "second," "third," etc., in the description, claims, and accompanying drawings are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0039] In the following description, terms such as “inner,” “outer,” “upper,” “lower,” “left,” and “right” are used only to facilitate the description of the embodiments and to simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this specification.
[0040] All data involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0041] Before introducing the technical solutions described in this manual, the application scenarios and related technologies of the technical solutions will be introduced.
[0042] With the development of the internet, digital media, and surveillance technology, the amount of video content has exploded. From social media and online education platforms to the film and entertainment industry and industrial monitoring, a massive amount of video content is generated daily. Faced with such a vast video resource library, how to quickly and accurately retrieve videos that meet specific needs has become a new research topic. This is of great significance for fully utilizing video resources. It can be used to solve scenarios requiring efficient processing and retrieval of large amounts of video data, such as video content management, copyright protection, advertising analysis, and on-site safety monitoring. For example, when conducting road monitoring, it is not only necessary to acquire and store road surveillance video, but also to identify events such as road traffic flow and traffic violations, providing data support for road traffic management.
[0043] This manual provides a video retrieval method and system based on big data. Please refer to the appendix. Figure 1 This method and system divides the video file 10 into multiple segments 11, processes each segment 11 to obtain a content description 12, content tags 22, and effect description. It also obtains a preferred set 32 based on the usability 31 of the video file 10 in a preset application scenario 21, thus classifying the video file 10. Video files 10 that record events such as production line stoppages or violations are prioritized for retrieval, facilitating faster retrieval results 51. Furthermore, by receiving retrieval requests through the retrieval request description 4130, it enables retrieval via natural language. This allows even those lacking relevant technical expertise and knowledge to utilize the video file 10 resources.
[0044] This manual first provides an AI-based intelligent video retrieval method; please refer to the appendix. Figure 2 The steps include:
[0045] Step S1) Read video file 10, identify the content of video file 10, divide video file 10 into segments 11 according to the content, and generate content description 12 for each segment 11.
[0046] Segmenting video file 10 into segments 11 is a key step in achieving efficient, accurate, and intelligent video retrieval. Videos typically contain multiple scenes, events, or topic transitions. Processing them as a whole can easily confuse different content, leading to semantic ambiguity. Segmentation 11 breaks down long videos into several segments with independent semantics.
[0047] The steps of identifying the content of video file 10 and segmenting video file 10 into segments 11 based on the content include:
[0048] Read the frame images of video file 10, identify people and objects in the frame images, and obtain person and object annotations. Perform visual content recognition on the extracted frame images, use computer vision technology to detect and identify people and objects appearing in the frames, and obtain the corresponding person annotation information (such as person identity, location, etc.) and object annotation information (such as item category, location, etc.).
[0049] Please see the appendix Figure 3 The video file 10 is divided into sub-video files 10 of a preset first length. Based on the person and object annotations included in the sub-video file 10, an annotation vector 14 is generated for each sub-video file 10. The video is divided into multiple sub-video segments of fixed length (e.g., 5 seconds or 10 seconds). For each sub-video segment, a structured annotation vector 14 is constructed based on the person and object annotation information it contains to represent the main visual content of the segment.
[0050] Clustering is performed using the labeled vector 14 to obtain the clusters of sub-video files 10. If two adjacent sub-video files 10 on the timeline belong to different clusters, an image segmentation marker 13 is added between the two sub-video files 10. Clustering analysis is performed on the labeled vectors 14 of all sub-video segments using a clustering algorithm (such as K-means or DBSCAN) to group sub-videos with similar visual content into the same category. If two adjacent sub-video segments on the timeline belong to different clusters, an image segmentation marker 13 is inserted between them to indicate that the video content has changed significantly at this point.
[0051] The audio content of the video file 10 is identified using speech recognition technology. The semantics of the audio content are then determined, and topic segments 11 are obtained through semantic analysis. Audio segment markers are added between two topics. The audio tracks in the video are extracted synchronously, and the audio content is converted into text information using Automatic Speech Recognition (ASR) technology. Natural Language Processing (NLP) technology is then applied to perform semantic analysis on the text, identifying topic transition points. Whenever a topic changes, an audio segment marker is inserted at the corresponding time point.
[0052] Based on the image segmentation markers 13 and the audio segmentation markers, the video file 10 is segmented 11. Taking into account the positions of the image segmentation markers 13 and the audio segmentation markers, the system ultimately determines the overall segmentation boundary 11 of the video, dividing the original video into several logically independent video segments. Each segment represents a relatively consistent time interval.
[0053] On the other hand, the steps for generating content description 12 for each segment 11 include:
[0054] Extract the positions of people and objects from the frame images to obtain their relative positions. In each video segment 11, first analyze the frame images to determine the positions of people and objects in the frame and record their relative positional relationships.
[0055] The system uses multiple preset action templates to match the sequence of relative positions of the people and objects along the time axis, identifying the actions included in each segment 11. Using multiple preset action templates (which cover various possible action types, such as walking, carrying, and operating machinery), the system compares and matches the obtained sequence of relative positions of people and objects along the time axis. In this way, the system can identify the specific actions occurring within each video segment 11.
[0056] Based on the semantics of each segment 11, including people, objects, actions, and corresponding audio content, a content description 12 is generated. The content description 12 is automatically generated based on the identified people, objects, their relationships, actions, and the semantic information of the corresponding audio content. The content description 12 includes not only visual elements but also the results of audio analysis (such as background dialogue or ambient sounds) to provide a richer content description for subsequent steps.
[0057] Step S2) Generate content tags 22 based on the content description 12 and the preset application scenario 21.
[0058] The steps for generating content tags 22 based on the content description 12 and the preset application scenario 21 include: extracting keywords from the content description 12, calculating the matching degree between the keywords and the preset application scenario 21, deleting keywords with a matching degree lower than a preset threshold, and obtaining content tags 22 based on the remaining keywords.
[0059] Extract keywords from the content description 12 of segment 11, such as person names, object categories, actions, and scene types (e.g., "meeting," "speech," "handshake," "office," etc.). Calculate the semantic matching degree between these keywords and preset application scenarios (e.g., "security monitoring," "education and training," "advertising analysis," "media content management," etc.). The matching degree is evaluated using a pre-trained language model or a keyword weight library based on scene definitions. Keywords with matching degrees below a preset threshold are filtered out, retaining those highly relevant to the current application scenario, thus forming a preliminary set of content tags 22. Optimize the tags according to the semantic preferences of different application scenarios. For example, in the "security monitoring" scenario, behavioral keywords such as "stranger appearance" and "climbing over fences" are given higher weights and are more likely to become content tags 22; while in the "e-commerce advertising" scenario, keywords such as "product display," "price mention," and "user reviews" are more representative. Content tags 22 can be expressed in a structured form, such as "subject-behavior-object" triples (e.g., "salesperson-introduction-new mobile phone"), facilitating subsequent semantic retrieval and logical reasoning.
[0060] Step S3) Based on the content description 12 of the segment 11, generate the usability 31 of the video file 10 under the preset application scenario 21.
[0061] In this embodiment, the availability 31 under the preset application scenario 21 is obtained based on the length of the content description 12, the number of content tags 22, and the degree of preference rating.
[0062] Specifically, the steps for obtaining the usability 31 under the preset application scenario 21 based on the length of the content description 12, the number of content tags 22, and the degree of preference rating include:
[0063] Read the preset length range, each length range is associated with a length score, and obtain the length score according to the length range corresponding to the length of the content description 12;
[0064] Read the preset tag quantity range, each tag quantity range is associated with a tag score, and obtain the tag score based on the tag quantity range corresponding to the quantity of the content tag 22;
[0065] Identify the semantics of the content description 12, obtain the tendency and degree of tendency based on the semantics, and obtain a degree of tendency rating based on the degree of tendency;
[0066] Calculate the weighted product of the length score, tag score, and tendency rating. Based on the comparison between the weighted product and the preset usability 31 interval, obtain the usability 31 under the preset application scenario 21.
[0067] The three dimensions of content description 12—length, number of tags, and tendency rating—quantify the usability 31 of video segment 11 in specific application scenarios. The text length of content description 12 is used to determine the level of detail. Too short a length may indicate insufficient information, while too long a length may be redundant. Therefore, different length ranges are set (e.g., less than 100 words, 100-300 words, and more than 300 words), each with a corresponding score to reflect content completeness. The number of content tags 22 reflects the richness and searchability of the video content. Too few tags indicate sparse information, while too many may introduce noise. Reasonable tag number ranges are set according to different application scenarios, and corresponding scores are assigned. Semantic analysis is used to identify the tendency (e.g., positive, neutral, negative) and tendency rating of content description 12, such as high positive, medium positive, low positive, and high negative, medium negative, low negative. Length score, tag score, and tendency rating are all numerical values, and are first normalized before being weighted and multiplied.
[0068] For example, the content description 12 of segment 11 is: "A red car did not stop 2 seconds after the red light turned on, crossed the stop line at the intersection, and quickly entered the intersection while pedestrians were crossing the zebra crossing." In terms of length score, content description 12 is complete, includes time, action, object, and environmental information, and has a moderate word count, making it a comprehensive and detailed description, thus earning a high length score.
[0069] Regarding the number of tags, several key tags were extracted, including: motor vehicles running red lights, pedestrians crossing the street, intersections, and failure to yield. The number of tags is relatively large, covering the core elements of violations, and corresponding to high tag scores.
[0070] For the tendency rating, the semantics clearly express a negative tendency of "dangerous driving" and "violation of traffic regulations," and the behavior is highly risky with a strong tendency, therefore it is rated as "high negative tendency," and the tendency rating is expressed numerically. The final usability score of 31 is "high," expressed as a value of 0.93, with a maximum value of 1.
[0071] Another example is shown in the appendix. Figure 4 The content description of section 11 is as follows: "Multiple lanes of the main road are congested with vehicles, the traffic flow is slow, some vehicles are queued for more than 200 meters, and the queue lasts for about 5 minutes. No obvious accidents or obstacles are observed."
[0072] The description is clear, covering the scope, duration, and state characteristics of congestion, and its length is reasonable, earning a high score in length evaluation. Key tags extracted include: traffic congestion, slow traffic flow, long queues, and no accidents. The information dimensions are relatively comprehensive, and the number of tags is moderate, resulting in a high tag score. The semantics reflect a negative state of "poor traffic flow," but without any emergencies, indicating a "moderately negative" tendency, earning a "moderate" tendency rating. After normalization and weighted multiplication, the score is above average. The usability of this segment in the "traffic congestion monitoring" scenario is determined to be "moderately high" (31), represented by a value of 0.72, suitable for traffic situation analysis, traffic light optimization, or issuing congestion warnings.
[0073] Step S4) Mark the video file 10 whose availability 31 is higher than the preset reference value and record it as the preferred set 32.
[0074] Step S5) According to the content description 12 of segment 11, for each video file 10 that does not belong to the preferred set 32, match at least one video file 10 that belongs to the preferred set 32 to obtain a matching set of video files 10 belonging to the preferred set 32.
[0075] Video files 10 with higher availability (31) are included in the preferred set (32) and will be retrieved first. Then, for each video file 10 not belonging to the preferred set (32), a video file 10 belonging to the preferred set is associated with it. When a video file 10 belonging to the preferred set (32) is retrieved, its associated video files 10 not belonging to the preferred set can be retrieved more quickly. Through this association, the number of video files 10 that need to be retrieved can be significantly reduced, improving retrieval efficiency.
[0076] Step S6) Establish auxiliary table 33, which records the filename, storage address, content description 12, content tag 22, and matching set of the video files 10 in the preferred set 32. Auxiliary table 33 helps to improve the retrieval efficiency of video files 10 in the preferred set 32.
[0077] Table 1 Auxiliary Table 33
[0078]
[0079] By constructing auxiliary table 33 as shown in Table 1, retrieval efficiency can be effectively improved. For example, when performing a road congestion check, relevant videos can be quickly filtered based on specific content tags 22 (such as "congestion lasting more than 5 minutes"), and potential risks can be assessed based on their effect descriptions. Furthermore, by utilizing the matching set function, starting from video file 10 in a preferred set 32, more relevant video resources can be discovered, providing more comprehensive support for decision-making.
[0080] Step S7) Receive search request description 41, and identify search tag 43 and search scenario 42 according to the search request description 41.
[0081] The steps of receiving a search request description 41 and identifying search tags 43 and search scenarios 42 based on the search request description 41 include:
[0082] Identify the semantics of the search request description 41, and identify the search scenario 42 based on the semantics of the search request description 41;
[0083] According to the preset tag library, the search tag 43 corresponding to the search scenario 42 is obtained. The tag library stores multiple preset application scenarios 21 and multiple tags corresponding to each preset application scenario 21.
[0084] The search request description 41 is presented in natural language and semantically analyzed. Natural language processing techniques are used to understand its intended meaning and context, identifying the user's search scenario 42. Examples include "traffic violation analysis," "crowd gathering monitoring," or "equipment operation status check."
[0085] Based on the preset tag library, search for the search tag 43 that matches the search scenario 42. For example, the search tag 43 corresponding to the "traffic violation" scenario includes: running a red light, driving against traffic, illegal parking, etc.; the search tag 43 for the "congestion monitoring" scenario includes: slow traffic, long queues, traffic obstruction, etc.
[0086] For example, the user's search request description 41 is: "What was the traffic situation on Road A from Avenue B to Road C yesterday afternoon?"
[0087] The search request description 41 does not specify a particular event type (such as smooth traffic, sudden braking, congestion, accident, etc.), making the expression rather vague and the search scope too broad. First, natural language understanding is used to extract key information: "yesterday afternoon" (time), "the section of Road A from Avenue B to Road C" (spatial area), and "traffic conditions" (state description). Among these, "traffic conditions" is a high-level semantic expression and may encompass various situations, such as smooth traffic, frequent lane changes, cutting in line, jaywalking, traffic stagnation, short-term congestion, minor scrapes, etc.
[0088] Based on semantic analysis, the retrieval request description 41 is mapped to several preset possible application scenarios, such as: traffic violation monitoring, traffic congestion identification, and emergency event warning. Through identification, the most relevant scenario is selected as "traffic anomaly monitoring," which is used as the retrieval scenario 42. According to a preset tag library, retrieval tags 43 associated with the "traffic anomaly monitoring" scenario are searched, including but not limited to: frequent lane changes, forced lane cutting, pedestrians running red lights, non-motorized vehicles going against traffic, traffic congestion, localized congestion, brief stoppages, sudden braking, low-speed driving, and large fluctuations in traffic flow.
[0089] Step S8) Obtain the matching row in the auxiliary table 33 according to the search tag 43, and obtain the search result 51 according to the matching row.
[0090] Finally, based on search tag 43, the time and location of the segments 11 recorded in auxiliary table 33 are retrieved, and it is compared to see if their "content tag 22" field contains one or more search tags 43. If a match is found, the row is marked as a "matching row". All matching rows are obtained, and the search results 51 that meet the search constraints are expanded according to the matching set of video files 10 corresponding to the matching rows, thus completing this search.
[0091] On the other hand, this specification provides an AI-based intelligent video retrieval system; please refer to the appendix. Figure 5 ,include:
[0092] The reading module 100 reads the video file 10, identifies the content of the video file 10, divides the video file 10 into segments 11 according to the content, and generates a content description 12 for each segment 11.
[0093] The generation module 200 generates content tags 22 based on the content description 12 and the preset application scenario 21, and generates the usability 31 of the video file 10 under the preset application scenario 21 based on the content description 12 of the segment 11.
[0094] The marking module 300 marks video files 10 whose availability 31 is higher than a preset reference value, and records them as the preferred set 32;
[0095] The matching module 400 matches at least one video file 10 belonging to the preferred set 32 for each video file 10 that does not belong to the preferred set 32 according to the content description 12 of the segment 11, thereby obtaining a matching set of video files 10 belonging to the preferred set 32;
[0096] The auxiliary module 500 establishes an auxiliary table 33, which records the file name, storage address, content description 12, content tag 22 and matching set of the video file 10 in the preferred set 32;
[0097] The retrieval module 600 receives a retrieval request description 41 and identifies a retrieval tag 43 and a retrieval scenario 42 based on the retrieval request description 41.
[0098] The result module 700 obtains the matching rows in the auxiliary table 33 based on the search tag 43, and obtains the search result 51 based on the matching rows.
[0099] Please see Figure 6 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this specification.
[0100] like Figure 6 As shown, the electronic device 1100 may include: at least one processor 1101, at least one network interface 1104, a user interface 1103, a memory 1105, and at least one communication bus 1102. The communication bus 1102 can be used to connect and communicate with the various components mentioned above. The user interface 1103 may include buttons, and optionally may include standard wired or wireless interfaces. The network interface 1104 may include, but is not limited to, a Bluetooth module, an NFC module, or a Wi-Fi module. The processor 1101 may include one or more processing cores. The processor 1101 connects to various parts within the electronic device 1100 using various interfaces and lines, and performs various functions of the routing device and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1105, and by calling data stored in the memory 1105. Optionally, the processor 1101 may be implemented using at least one hardware form of DSP, FPGA, or PLA. The processor 1101 may integrate one or more combinations of CPU, GPU, and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content that the display screen needs to show; and the modem is used for wireless communication.
[0101] It is understandable that the aforementioned modem may not be integrated into the processor 1101, but may be implemented using a separate chip.
[0102] The memory 1105 may include RAM or ROM. Optionally, the memory 1105 may include a non-transitory computer-readable medium. The memory 1105 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1105 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 1105 may also be at least one storage device located remotely from the aforementioned processor 1101. As a computer storage medium, the memory 1105 may include an operating system, a network communication module, a user interface module, and application programs. The processor 1101 may be used to call the application programs stored in the memory 1105 and execute the methods in the above-described embodiments.
[0103] This specification also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform multiple steps as described in the above embodiments. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in the computer-readable storage medium.
[0104] This specification also provides a computer program product, including a computer program that, when executed by a processor, implements the multiple steps described in the above embodiments.
[0105] Where there is no conflict, the technical features in this embodiment and implementation scheme can be combined arbitrarily.
[0106] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes multiple computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating multiple available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0107] When implemented through hardware or firmware, the aforementioned method flow is programmed into the hardware circuit to obtain the corresponding hardware circuit structure and achieve the corresponding function. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit, whose logic function is determined by the user programming the device. Designers can program a digital system onto a PLD themselves, eliminating the need for chip manufacturers to design and fabricate dedicated integrated circuit chips. Furthermore, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, similar to the software compiler used in program development. The original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There is not just one HDL, but many. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of the aforementioned hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logic method flow can be easily obtained.
[0108] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Any modifications and improvements made by those skilled in the art to the technical solutions of this specification without departing from the spirit of this specification should fall within the protection scope defined by the claims of this specification.
Claims
1. An AI-based intelligent video retrieval method, characterized in that, Including the following steps: Read the video file, identify the content of the video file, segment the video file according to the content, and generate a content description for each segment; Based on the content description and preset application scenarios, generate content tags; Based on the content description of the segment, the usability of the video file in the preset application scenario is generated; Video files whose availability is higher than a preset reference value are designated as the preferred set. Based on the segmented content description, for each video file that does not belong to the preferred set, match at least one video file that belongs to the preferred set to obtain a matching set of video files belonging to the preferred set; An auxiliary table is established, which records the filename, storage address, content description, content tags, and matching set of the video files in the preferred set; Receive a search request description, and identify search tags and search scenarios based on the search request description; The matching rows in the auxiliary table are obtained based on the search tags, and the search results are obtained based on the matching rows.
2. The AI-based intelligent video retrieval method according to claim 1, characterized in that, The steps of identifying the content of a video file and segmenting the video file according to the content include: Read the frame images of the video file, identify people and objects in the frame images, and obtain person and object annotations; The video file is divided into sub-video files of a preset first length, and a label vector is generated for each sub-video file based on the person and object labels included in the sub-video files. Clustering of the labeled vectors is used to obtain the clusters of the sub-video files. If two adjacent sub-video files on the timeline belong to different clusters, image segmentation markers are added between the two sub-video files. The audio content of the video file is identified using speech recognition technology, the semantics of the audio content are identified, topic segments are obtained through semantic analysis, and audio segmentation markers are added between two topics. The video file is segmented based on the image segmentation markers and the audio segmentation markers.
3. The AI-based intelligent video retrieval method according to claim 2, characterized in that, The steps for generating content descriptions for each segment include: Extract the positions of people and objects in the frame image to obtain their relative positions; Multiple preset action templates are matched with the sequence of relative positions of the person and object along the time axis to identify the actions included in each segment; Based on the semantics of each segment, including people, objects, actions, and corresponding audio content, a content description is generated.
4. The AI-based intelligent video retrieval method according to any one of claims 1 to 3, characterized in that, The steps of generating content tags based on the content description and preset application scenarios, and generating usability under the preset application scenarios based on the content description, include: Extract keywords from the content description, calculate the matching degree between the keywords and the preset application scenario, delete keywords with a matching degree lower than a preset threshold, and obtain content tags based on the remaining keywords; Based on the length of the content description, the number of content tags, and the degree of preference rating, the usability in the preset application scenario is obtained.
5. The AI-based intelligent video retrieval method according to claim 4, characterized in that, The steps for determining usability in a preset application scenario based on the length of the content description, the number of content tags, and the degree of preference rating include: Read the preset length range, each length range is associated with a length score, and obtain the length score based on the length range corresponding to the length of the content description; Read the preset tag quantity range, each tag quantity range is associated with a tag score, and obtain the tag score based on the tag quantity range corresponding to the number of content tags; Identify the semantics of the content description, obtain the tendency and degree of tendency based on the semantics, and obtain a degree of tendency rating based on the degree of tendency. Calculate the weighted product of the length score, tag score, and tendency rating. Based on the comparison between the weighted product and the preset usability division interval, obtain the usability under the preset application scenario.
6. The AI-based intelligent video retrieval method according to any one of claims 1 to 3, characterized in that, The steps of receiving a search request description and identifying search tags and search scenarios based on the search request description include: Identify the semantics of the search request description, and identify the search scenario based on the semantics of the search request description; Based on a preset tag library, the search tags corresponding to the search scenario are obtained. The tag library stores multiple preset application scenarios and multiple tags corresponding to each preset application scenario.
7. An AI-based intelligent video retrieval system, characterized in that, include: The reading module reads the video file, identifies the content of the video file, segments the video file according to the content, and generates a content description for each segment. The generation module generates content tags based on the content description and preset application scenarios, and generates the usability of the video file in the preset application scenarios based on the segmented content descriptions. The marking module marks video files whose availability is higher than a preset reference value and designates them as the preferred set; The matching module, based on the content description of the segments, matches at least one video file that belongs to the preferred set for each video file that does not belong to the preferred set, thereby obtaining a matching set of video files belonging to the preferred set; The auxiliary module establishes an auxiliary table, which records the filenames, storage addresses, content descriptions, content tags, and matching sets of the video files in the preferred set. The retrieval module receives a retrieval request description and identifies retrieval tags and retrieval scenarios based on the retrieval request description. The results module obtains matching rows in the auxiliary table based on the search tags, and obtains search results based on the matching rows.
8. An electronic device, characterized in that, Including the processor and memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code stored in the memory to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Video retrieval method, device and equipment and storage medium
CN111506771A
Video tag processing method and device, computer equipment and storage medium
CN118093936A
Method and apparatus for searching for information inside video
WO2021221209A1