Method, computer device, and computer program for driving scene retrieval
Patent Information
- Application Number
- KR1020250056279
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2045-04-29
Smart Images

Figure 112025048653681-PAT00024_ABST
Abstract
Description
Technology Field
[0001] The following description concerns driving video search technology. Background Technology
[0002] In autonomous driving, where extensive driving video data is continuously recorded, finding appropriate scenarios is essential for system verification. However, identifying suitable scenarios from vast video databases is time-consuming. To accelerate this, developers often rely on filtering based on metadata such as objects, time, or driving strategies. Nevertheless, annotating metadata on a per-video basis is labor-intensive, and search conditions are limited by predefined object classes, which limits their ability to adapt to dynamic and complex real-world driving environments.
[0003] With the emergence of deep representation models, text-to-video search is being used as a popular solution. Text-to-video search serves as a powerful tool for navigating vast video databases. This is particularly useful in autonomous driving for searching video using text queries to simulate and evaluate driving systems according to desired scenarios.
[0004] Traditional ranking-based text-video search methods rank videos based on how similar their features are to text query features. However, errors occur due to mismatches between the language domain and the video domain. While most approaches attempt to mitigate these mismatches, additional matching modules or fine-tuning are essential in the downstream domain to reduce data distribution inconsistencies.
[0005] Even if the two modality spaces are finely aligned, the information in the video can increase indefinitely, potentially reducing the similarity between the video and the query features. For example, if a query includes 'pedestrian crossing a crosswalk, bus, narrow road, rainy day,' videos that correspond to 'pedestrian crossing, bus, narrow road' but not 'rainy day' should not be searched, even if they have high similarity. Conversely, videos corresponding to 'pedestrian crossing, bus, narrow road, rainy day, bicycle, pedestrian crossing a tree' should be searched because they satisfy all specified conditions, even though their similarity scores may be low due to additional content such as 'bicycle' or 'tree.'
[0006] As an example of technology for autonomous driving testing, Korean registered patent No. 10-2579590 (registered on September 13, 2023) discloses a technology for generating an evaluation scenario for autonomous vehicle driving ability based on the Road Traffic Act. The problem to be solved
[0007] A driving video search framework using inclusive text matching can be provided.
[0008] Compressed captions for driving video search can be generated using a vision-language model (VLM) and a large language model (LLM).
[0009] It can provide positive and negative data curation strategies and an attention-based scoring mechanism optimized for driving video search. means of solving the problem
[0010] A driving video retrieval method for a computer device comprising at least one processor, wherein the at least one processor searches for a driving video corresponding to a query text based on object-level similarity between a query and a video caption, as an inclusive text-to-video retrieval that searches for a video satisfying all conditions of a query for a query text given by a user.
[0011] According to one aspect, the searching step can search for positive videos containing all query conditions by evaluating videos in a database through object-level text matching for the query text.
[0012] According to another aspect, the searching step may include a step of calculating a similarity score with the most similar object in the image caption for each object included in the query text, and calculating the average of the similarity scores for the entire set of objects as the final matching score.
[0013] According to another aspect, the method further comprises the step of generating a training dataset for a model for comprehensive text-video search by the at least one processor, wherein the generating step may include: generating an object caption from an input video; and generating a positive object and a negative object using a partial object included in the object caption.
[0014] According to another aspect, the step of generating the object caption may generate a video caption as metadata related to the input image using a vision-language model (VLM).
[0015] According to another aspect, the step of generating the object caption may generate a video caption from the input video using a language prompt in which a category describing the driving video is defined.
[0016] According to another aspect, the step of generating the positive object and the negative object may generate positive text and negative text in the object caption using a large language model (LLM).
[0017] According to another aspect, the step of generating the positive object and the negative object may include: generating the positive object from a partial element according to a key characteristic extracted from the object caption; and generating a negative object set composed of objects that cannot be inferred from a positive object set composed of the positive object.
[0018] According to another aspect, the method may further include the step of training a model for comprehensive text-video search through contrastive learning using the positive object and the negative object by the at least one processor.
[0019] According to another aspect, the learning step may train a model using a first loss function that induces a similarity score of 1 for the positive object and a second loss function that induces a similarity score of 0 for the negative object.
[0020] According to another aspect, the learning step described above may further train the model using a third loss function that induces a similarity score between identical object sets to 1.
[0021] A computer program stored on a computer-readable recording medium is provided to execute the above driving video search method on the computer device.
[0022] The present invention provides a computer device comprising at least one processor implemented to execute a command readable by a computer device, wherein the at least one processor processes a process of searching for a driving video corresponding to said query text based on object-level similarity between a query and a video caption as an inclusive text-to-video retrieval that searches for a video satisfying all conditions of a query for a query text given by a user. Effects of the invention
[0023] According to embodiments of the present invention, driving video retrieval performance can be improved by providing inclusive text-to-video retrieval that searches for videos without query conditions, regardless of the amount of unqueried information among the information included in the video.
[0024] According to embodiments of the present invention, by generating compressed captions for driving videos through VLM and LLM, text-to-video search is transformed into a more efficient text-to-text search problem, thereby resolving modality mismatch issues and high annotation costs. Brief explanation of the drawing
[0025] FIG. 1 is a drawing illustrating an example of a network environment according to an embodiment of the present invention. FIG. 2 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. FIG. 3 illustrates an example of searching for driving images through comprehensive text matching in an embodiment of the present invention. FIG. 4 illustrates an example of a video caption generation process in an embodiment of the present invention. FIG. 5 illustrates a driving video search framework using comprehensive text matching in one embodiment of the present invention. Specific details for implementing the invention
[0026] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0028] Embodiments of the present invention relate to driving video search technology.
[0029] Embodiments including those specifically disclosed in this specification can improve driving video search performance by providing a comprehensive text-video search that searches for videos without query conditions, regardless of the amount of information included in the video that does not correspond to the query.
[0030] A driving image search device according to embodiments of the present invention may be implemented by at least one computer device, and a driving image search method according to embodiments of the present invention may be performed through at least one computer device included in the driving image search device. At this time, a computer program according to an embodiment of the present invention may be installed and run on the computer device, and the computer device may perform a driving image search method according to embodiments of the present invention under the control of the run computer program. The above-described computer program may be stored on a computer-readable recording medium to be combined with the computer device to execute the driving image search method on the computer.
[0031] FIG. 1 is a diagram illustrating an example of a network environment according to an embodiment of the present invention. The network environment of FIG. 1 illustrates an example including a plurality of electronic devices (110, 120, 130, 140), a plurality of servers (150, 160), and a network (170). FIG. 1 is an example for explaining the invention, and the number of electronic devices or servers is not limited to that shown in FIG. 1. Furthermore, the network environment of FIG. 1 is merely an example of one of the environments applicable to the present embodiments, and the environments applicable to the present embodiments are not limited to the network environment of FIG. 1.
[0032] Multiple electronic devices (110, 120, 130, 140) may be fixed terminals or mobile terminals implemented as computer devices. Examples of multiple electronic devices (110, 120, 130, 140) include smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, etc. For example, FIG. 1 shows the shape of a smartphone as an example of an electronic device (110), but in embodiments of the present invention, the electronic device (110) may substantially refer to one of various physical computer devices capable of communicating with other electronic devices (120, 130, 140) and / or servers (150, 160) via a network (170) using a wireless or wired communication method.
[0033] The communication method is not limited and may include not only communication methods utilizing communication networks (e.g., mobile communication networks, wired internet, wireless internet, broadcasting networks) that the network (170) may include, but also short-range wireless communication between devices. For example, the network (170) may include any one or more networks such as a PAN (personal area network), LAN (local area network), CAN (campus area network), MAN (metropolitan area network), WAN (wide area network), BBN (broadband network), and the Internet. Additionally, the network (170) may include any one or more network topologies such as a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree or hierarchical network, but is not limited thereto.
[0034] Each of the servers (150, 160) may be implemented as a computer device or multiple computer devices that communicate with multiple electronic devices (110, 120, 130, 140) through a network (170) to provide commands, code, files, content, services, etc. For example, the server (150) may be a system that provides services (e.g., autonomous driving safety verification services, etc.) to multiple electronic devices (110, 120, 130, 140) connected through the network (170).
[0035] FIG. 2 is a block diagram illustrating an example of a computer device according to an embodiment of the present invention. Each of the plurality of electronic devices (110, 120, 130, 140) or servers (150, 160) described above can be implemented by the computer device (200) illustrated in FIG. 2.
[0036] As illustrated in FIG. 2, such a computer device (200) may include memory (210), a processor (220), a communication interface (230), and an input / output interface (240). The memory (210) is a computer-readable recording medium and may include a non-perishable mass storage device such as RAM (random access memory), ROM (read only memory), and a disk drive. Here, a non-perishable mass storage device such as a ROM and a disk drive may be included in the computer device (200) as a separate permanent storage device distinct from the memory (210). Additionally, an operating system and at least one program code may be stored in the memory (210). These software components may be loaded into the memory (210) from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. In another embodiment, software components may be loaded into memory (210) via a communication interface (230) rather than a computer-readable recording medium. For example, software components may be loaded into memory (210) of a computer device (200) based on a computer program installed by files received through a network (170).
[0037] The processor (220) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (220) via memory (210) or a communication interface (230). For example, the processor (220) may be configured to execute instructions received according to program code stored in a recording device such as memory (210).
[0038] The communication interface (230) may provide a function for the computer device (200) to communicate with other devices (e.g., storage devices described above) through the network (170). For example, requests, commands, data, files, etc. generated by the processor (220) of the computer device (200) according to program code stored in a recording device such as memory (210) may be transmitted to other devices through the network (170) under the control of the communication interface (230). Conversely, signals, commands, data, files, etc. from other devices may be received by the computer device (200) through the communication interface (230) of the computer device (200) via the network (170). Signals, commands, data, etc. received through the communication interface (230) may be transmitted to the processor (220) or memory (210), and files, etc. may be stored in a storage medium (the permanent storage device described above) that the computer device (200) may further include.
[0039] The input / output interface (240) may be a means for interfacing with an input / output device (250). For example, the input device may include a device such as a microphone, keyboard, or mouse, and the output device may include a device such as a display or speaker. As another example, the input / output interface (240) may be a means for interfacing with a device in which the functions for input and output are integrated into one, such as a touchscreen. The input / output device (250) may be composed of a computer device (200) and a single device.
[0040] Additionally, in other embodiments, the computer device (200) may include fewer or more components than the components of FIG. 2. However, it is not necessary to clearly illustrate most of the prior art components. For example, the computer device (200) may be implemented to include at least some of the input / output devices (250) described above, or may include other components such as a transceiver, a database, etc.
[0041] Below, specific embodiments of the technology for searching driving images will be described.
[0042] Traditional driving video search methods do not consider situations where videos that do not meet specific criteria must be excluded, whereas videos containing additional relevant information are allowed.
[0043] When evaluating safety-critical systems such as autonomous driving, test conditions often act as difficult, non-negotiable constraints. Instead of searching for the "most relevant" scenarios in a test database, it is important to identify scenarios that comprehensively satisfy all specified conditions while allowing for additional information to ensure diversity in the test environment.
[0044] The present invention can provide a driving video search framework (hereinafter referred to as 'CARIM') that uses comprehensive text matching as a comprehensive text-video search environment for searching videos without question conditions, regardless of the amount of unqueried information contained in the video.
[0045] In this embodiment, a comprehensive text-video search operation can be provided with a new dataset curation strategy, a contrastive learning objective, and an evaluation metric that ensures all searched videos include all elements specified in the query without any omission.
[0046] In addition, it can provide an approach to effectively handle hard negative queries by preventing the search for videos that do not satisfy any specified conditions, despite queries showing high similarity.
[0047] In addition, it can support an environment for controlling search results through a flexible and user-friendly search system where users can input natural language queries and adjust thresholds to process the search range.
[0048] FIG. 3 illustrates an example of searching for driving images through comprehensive text matching in an embodiment of the present invention.
[0049] Most text-to-video search models focus on bridging the gap between text and video modalities. Some studies train on vast amounts of image-to-text pairs to match the two domains into a single embedding space, but this is not suitable for specific domains such as autonomous driving.
[0050] There are technologies such as a technology that searches for multi-view bird's-eye view features based on natural language text derived from video text understanding in driving footage, and a technology that searches for similar driving footage based on images and a predefined set of action vectors.
[0051] In addition, general text matching searches for the most relevant documents to a query by ranking documents in a database. However, this approach can search for and rank documents even if they do not fully contain the information requested in the query.
[0052] The present invention aims to generate captions from videos in the autonomous driving category and use them to perform text-video search in the form of text-to-text matching, and in particular, can perform text-video search through a comprehensive text matching method.
[0053] Referring to FIG. 3, a ranking-based text-video search retrieves negative video information (301) that has high similarity to the query but is not satisfied when a text query is given. In contrast, CARIM according to the present invention can retrieve positive video information (302) that includes all query conditions by evaluating the video through object-level dense text matching.
[0054] Below, we define the comprehensive text-video search task and describe the dataset generation, model architecture, and training strategy, including training objectives.
[0055] A comprehensive text-video search operation can be defined as follows.
[0056] Video database V={v1, v2, ..., v n Given}, each video v i is object O i ={o1, o2, ..., o m Includes a set of}. Query text condition C={c1, c2, ..., c k In the set}, the goal is to search for all videos V'⊆V as follows.
[0057]
[0058] All k conditions in the query are O i m objects must satisfy this, where k≤m. Also, O i The existence of additional mk objects in should not affect the search of V'.
[0059] FIG. 4 illustrates an example of a video caption generation process in an embodiment of the present invention.
[0060] Referring to FIG. 4, in step (S401), the processor (220) can generate video captions as metadata associated with the given input video using a VLM. To generate captions from the input video, major categories describing the driving video, including the main subject, action, weather conditions, time, and road structure, are defined, and a language prompt is designed to obtain video captions corresponding to these categories. Then, the driving video v in the VLM i It instructs to list m objects representing. In this embodiment, it is necessary to compress each image into a limited number of elements to extract key information from the driving video. For example, the goal of the present invention is to determine whether a pedestrian is crossing a crosswalk rather than capturing details that are not important to the design of an autonomous driving system, such as the color of a shirt, whether a hat is worn, or gender.
[0061] The present invention utilizes VLM to generate captions from video and uses them to perform text-video search in the form of text-to-text matching. This approach eliminates the need to match text and video modalities, thereby avoiding the need for processes and manual annotation work for matching various data types or input sources.
[0062] In step (S402), the processor (220) uses the LLM to create the image caption O generated in step (S401). i Extracting key features from partial elements Partial(O iIt can generate ). For example, if the object is "pedestrian crossing the street," the corresponding sub-objects could be "pedestrian" and "street." Sub-objects act as additional descriptors for images, allowing related videos to be retrieved using single-word queries instead of natural language representations. Sub-objects can compensate if the VLM fails to generate sufficient captions. Object O i and the corresponding partial object is a set of positive objects. It forms, which can be defined as follows.
[0063]
[0064] In step (S403), the processor (220) positive object set O P A set of negative objects O consisting of objects that cannot be inferred from N It can create. Negative objects are It has. Since negative objects can semantically overlap with positive objects, it filters out cases where similar words appear simultaneously in two sets, such as "pedestrian" and "person," so that the two sets are semantically distinguished.
[0065] In this embodiment, positive and negative text can be generated from image captions by utilizing LLM for model training with contrast learning applied.
[0066] In this embodiment, six objects are randomly selected from each set. and It is saved as the final caption indicated by . Among language models of similar size, LLaVA-NeXT-34B is used due to its low rate of incorrect descriptions and the ability to consistently extract the same image description structure from various images, and LLaMA3 can be used to generate partial objects and negative objects. To facilitate post-processing, prompt engineering can be applied to extract objects in a bullet-point format.
[0067] The process of generating the training dataset is as follows.
[0068] The processor (220) is the final caption and You can generate positive query pairs and negative query pairs using .
[0069] First, all possible m positive objects Combinations are generated. Here, 1≤k≤m acts as positive queries for the video. Based on these positive queries, negative query pairs can be generated using two strategies called balanced sampling and adaptive negative injection.
[0070] At this point, in the balanced sampling strategy In this process, the same number of negative objects as the corresponding positive query is sampled. Through this, the model is trained to distinguish whether a query is positive or negative when the number of objects is the same. Additionally, the adaptive negative injection strategy constructs more complex negative queries by adding negative objects to positive queries. If a positive query contains k objects, l negative objects are added to the positive objects for every l > 0 greater than k, such that k + l ≤ m. Strong negative queries are generated when l is close to 1 (i.e., the query consists mostly of positive objects). Conversely, easy negative queries are formed when l is close to m, increasing the proportion of negative objects. Through this approach, the model can classify a query as negative regardless of how many non-positive objects exist in the query.
[0071] In driving videos, similar objects such as cars, pedestrians, crosswalks, and traffic lights appear in almost every video. To train effectively in such scenarios, relying solely on global video and text features is insufficient; a more granular understanding and differentiation between videos are required. In this embodiment, object-level attention operations can be performed to calculate object-level similarity between a query and a video caption and to learn the video caption object that needs to be matched with each query object.
[0072] FIG. 5 illustrates a driving image search framework (CARIM) using comprehensive text matching in one embodiment of the present invention.
[0073] Referring to Fig. 5, the CARIM model receives a video caption embedding and a query text as input. The transformed query is optimized to be similar to the key feature, and the object-level query-key similarity is averaged to obtain a final matching score.
[0074] In detail, query Q i Create object-level text embeddings for and all video caption embeddings key K i and value V i It is transformed into features. Then, the query focuses on key features and follows an attention mechanism (e.g., scaled dot-product attention mechanism) to obtain an attention score A (Q i , K i Creates ).
[0075] [Mathematical Formula 1]
[0076]
[0077] Here, d kis a feature dimension. The attention score is the major index I that the query focuses on most strongly. MAX,i and the corresponding maximum principal feature f mk,i It is used to identify.
[0078] [Mathematical Formula 2]
[0079]
[0080] [Mathematical Formula 3]
[0081]
[0082] The arg max function guides each query object to optimally match with the most relevant caption in the video. At the same time, it multiplies the attention score and value feature to obtain the attention value A(Q i , K i , V i Extracts ). Applyes residual connections by adding query features, and passes the result through the feed-forward network Φ to obtain the transformed query feature f. tq,i You can obtain.
[0083] [Mathematical Formula 4]
[0084]
[0085] Transformed query features f to ensure that query features are transformed into key features that the query should focus on during the attention process tq,i and maximum key features f mk,i The cosine similarity σ between i Calculate the positive pair (v i , Q i In the case of ), since all objects in the query match one of the key features, the similarity is made to approach 1. Final score s i It is calculated as the average of the cosine similarity for all k objects in the query.
[0086] [Mathematical Formula 5]
[0087]
[0088] [Mathematical Formula 6]
[0089]
[0090] Through the process described above, the CARIM model learns the key features that each query object should pay attention to, and the output score represents the attention rate. Queries that are completely included in the caption receive a score close to 1, and the score decreases in proportion to the number of negative objects, approaching 0 for queries that do not match completely.
[0091] Finally, similarity scores for all videos are calculated for the text query, and videos with scores exceeding a threshold h are retrieved. If no videos exceed h, the input query is considered a negative query for all videos, and no videos are returned. For each object in the query, an attention-based score is calculated for all video captions, and then the average is calculated. This allows for a granular evaluation of whether each query condition is satisfied within the video. The final score represents the proportion of query conditions satisfied in the video and ensures that the model can evaluate query-caption similarity and retrieve videos regardless of the number of objects in the captions.
[0092] The CARIM model training process is as follows.
[0093] The training objective of the CARIM model is based on the concept of contrastive learning, with two main loss functions (i.e., positive loss L pos and negative loss L neg It is to optimize ).
[0094] [Mathematical Formula 7]
[0095]
[0096]
[0097] Here, B is a mini-batch set of videos, and s brepresents the similarity score between the b-th query object and the corresponding video learning pair.
[0098] Positive Loss L pos optimizes the model when given positive pairs to generate a score of 1. The model is guided to assign a score of 1 as long as every query object matches at least one caption, regardless of the number of video captions. Conversely, the negative loss L neg It sets the similarity score to 0 if there is even one negative object in the query. It penalizes higher similarity scores for non-matching objects.
[0099] Positive Loss L pos and negative loss L neg When this balance is achieved, the final score is distributed between 1 and 0 based on the ratio of positive and negative objects in the query. For example, if a query contains 4 objects and 3 of them match the video caption, the similarity score for the 3 matching objects is optimized to 1, and the score for the non-matching objects is trained to 0, resulting in a final score of 0.75. Based on these results, a threshold is set to 0.99 to search for videos with a score close to 1 above the threshold as the final output.
[0100] In other words, positive loss L pos It is trained to induce a similarity score of 1 for images included in all object conditions of the query, and learns to converge the similarity score to 1 for positive pairs where the query and the image are exact matches. Negative loss L neg It is trained to induce a similarity score of 0 for images where even one of the object conditions in the query is mismatched, so that if the query contains even one negative object, the entire query is considered mismatched and the similarity score is made close to 0.
[0101] In addition, to address the imbalance in the data distribution, Self-Alignment Loss Lself This can be applied. While the caption generation strategy effectively generates hard positive and negative pairs, full-length positive pairs containing all objects in the set of positive objects can only be generated once per video. On the other hand, negative pairs are generated by adding negative captions to positive queries until the number of objects reaches a fixed number k, resulting in a much larger number of negative pairs of length k. This imbalance causes the model to overfit to predict lower similarity scores for longer queries. To mitigate the overfitting problem, additional positive queries of length k are generated during training by utilizing the actual captions of the videos. For these queries, the similarity score is L self It is explicitly calculated as 1 by, allowing the model to learn how to correctly handle long queries without confusion.
[0102] The calculation result is as follows.
[0103] [Mathematical Formula 8]
[0104]
[0105] Here, o b is video v i It is a set of objects.
[0106] In other words, while positive queries can be generated from various combinations, negative queries are generated with a fixed length k, leading to an imbalance in the training data and potentially causing the model to become confused by long queries. To address this issue, the entire positive set constructed based on ground truth captions is used as a query again, and the model is trained so that the similarity score for that query becomes 1. That is, the model can be trained so that the similarity between sets of the same objects becomes 1.
[0107] The total loss function L is minimized as the sum of the three objective functions mentioned above.
[0108] [Mathematical Formula 9]
[0109]
[0110] These embodiments relate to CARIM, an approach for driving scenario search, which can capture object-level similarity and ensure conditional completeness in text-to-video search tasks. High-quality captions can be generated for videos using VLM and LLM for effective text-to-video matching. Consequently, the performance of VLM and LLM directly impacts search results. By utilizing VLM and LLM, flexible annotation for videos is allowed, enabling the addition of annotations at desired depths and amounts. By using high-performance open-source VLM and LLM, bottlenecks can be resolved and search performance can be improved by leveraging a wider range of video and text sources.
[0111] As such, according to embodiments of the present invention, driving video search performance can be improved by providing a comprehensive text-video search that searches for videos without query conditions, regardless of the amount of information included in the video that does not correspond to the query.
[0112] Furthermore, according to embodiments of the present invention, by generating compressed captions for driving videos through VLM and LLM, text-to-video search can be transformed into a more efficient text-to-text search problem, thereby resolving modality mismatch issues and high annotation costs.
[0113] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.
[0114] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0115] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may continuously store a program executable by a computer, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or several hardware combined, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Additionally, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0116] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0117] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
Claim 1 A driving video retrieval method for a computer device comprising at least one processor, wherein the at least one processor performs an inclusive text-to-video retrieval to search for a driving video corresponding to a query text based on object-level similarity between a query and a video caption, wherein the at least one processor performs an inclusive text-to-video retrieval to search for a video satisfying all conditions of a query for a query text given by a user, and further includes the step of generating a training dataset for a model for the inclusive text-to-video retrieval by the at least one processor, wherein the generating step includes: generating an object caption from an input video; and generating a positive object and a negative object using a partial object included in the object caption, wherein the step of generating the positive object and the negative object generates positive text and negative text from the object caption using a large language model (LLM). Claim 2 A driving video search method according to claim 1, wherein the searching step is characterized by evaluating a video in a video database through object-level text matching with respect to the query text to search for a positive video that includes all query conditions. Claim 3 A driving video search method according to claim 1, wherein the searching step comprises the step of calculating a similarity score with the most similar object in the video caption for each object included in the query text, and calculating the average of the similarity scores for the entire set of objects as the final matching score. Claim 4 delete Claim 5 A driving video search method according to claim 1, wherein the step of generating the object caption is characterized by generating a video caption as metadata related to the input video using a vision-language model (VLM). Claim 6 A driving video retrieval method for a computer device comprising at least one processor, wherein the at least one processor includes the step of retrieving a driving video corresponding to a query text based on object-level similarity between a query and a video caption as an inclusive text-to-video retrieval that searches for a video satisfying all conditions of a query for a query text given by a user, and further includes the step of generating a training dataset for a model for the inclusive text-to-video retrieval by the at least one processor, wherein the generating step comprises: generating an object caption in an input video; and generating a positive object and a negative object using a partial object included in the object caption, wherein the step of generating the object caption generates a video caption in the input video using a language prompt in which a category describing the driving video is defined. Claim 7 delete Claim 8 A driving video retrieval method for a computer device comprising at least one processor, wherein the at least one processor performs an inclusive text-to-video retrieval to search for a driving video corresponding to a query text based on object-level similarity between a query and a video caption, wherein the at least one processor performs an inclusive text-to-video retrieval to search for a video satisfying all conditions of a query for a query text given by a user, and further comprises the step of generating a training dataset for a model for the inclusive text-to-video retrieval by the at least one processor, wherein the generating step includes: generating an object caption from an input video; and generating a positive object and a negative object using a partial object included in the object caption, wherein the step of generating the positive object and the negative object includes: generating a partial element according to a key characteristic extracted from the object caption as the positive object; and generating a negative object set composed of objects that cannot be inferred from a positive object set composed of the positive object. Claim 9 A driving video retrieval method for a computer device comprising at least one processor, wherein the at least one processor performs an inclusive text-to-video retrieval to search for a driving video corresponding to a query text based on object-level similarity between a query and a video caption, wherein the at least one processor performs an inclusive text-to-video retrieval to search for a video satisfying all conditions of a query for a query text given by a user; further comprising the step of generating a training dataset for a model for the inclusive text-to-video retrieval by the at least one processor, wherein the generating step includes: generating an object caption from an input video; and generating a positive object and a negative object using a partial object included in the object caption; and further comprising the step of training a model for the inclusive text-to-video retrieval by the at least one processor through contrastive learning using the positive object and the negative object. Claim 10 A driving image search method according to claim 9, wherein the learning step is characterized by training a model using a first loss function that induces a similarity score of 1 for the positive object and a second loss function that induces a similarity score of 0 for the negative object. Claim 11 A driving image search method according to claim 10, wherein the learning step further involves training a model using a third loss function that induces a similarity score between identical object sets to 1. Claim 12 A computer program stored on a computer-readable recording medium to execute a driving video search method of any one of paragraphs 1 to 3, 5, 6, and 8 to 11 on the computer device. Claim 13 A computer device comprising at least one processor implemented to execute a readable command on a computer device, wherein the at least one processor processes a process of searching for a driving video corresponding to said query text based on object-level similarity between a query and a video caption as an inclusive text-to-video retrieval that searches for a video satisfying all conditions of a query for a query text given by a user, wherein the at least one processor generates a training dataset for a model for said inclusive text-to-video retrieval, generates an object caption from an input video, and generates a positive object and a negative object using a partial object included in said object caption, and wherein the at least one processor generates a positive text and a negative text from said object caption using LLM. Claim 14 A computer device comprising at least one processor implemented to execute a readable command on a computer device, wherein the at least one processor processes a process of searching for a driving video corresponding to said query text based on object-level similarity between the query and the video caption as inclusive text-to-video retrieval for a query text given by a user, which searches for a video satisfying all conditions of the query, and wherein the at least one processor evaluates a video in a video database through object-level text matching for said query text to search for a positive video including all query conditions. Claim 15 delete Claim 16 A computer device comprising at least one processor implemented to execute a readable command on a computer device, wherein the at least one processor processes a process of searching for a driving video corresponding to said query text based on object-level similarity between a query and a video caption as an inclusive text-to-video retrieval that searches for a video satisfying all conditions of a query for a query text given by a user, wherein the at least one processor generates a training dataset for a model for said inclusive text-to-video retrieval, generates an object caption from an input video, and generates positive objects and negative objects using partial objects included in said object caption, and wherein the at least one processor generates a video caption as metadata related to said input video using a vision-language model (VLM). Claim 17 delete Claim 18 A computer device comprising at least one processor implemented to execute a readable command on a computer device, wherein the at least one processor processes a process of searching for a driving video corresponding to said query text based on object-level similarity between a query and a video caption as inclusive text-to-video retrieval for a query text given by a user, which searches for a video satisfying all conditions of the query, and wherein the at least one processor generates a training dataset for a model for said inclusive text-to-video retrieval by generating an object caption from an input video and generating a positive object and a negative object using a partial object included in said object caption, and wherein the at least one processor learns the model for said inclusive text-to-video retrieval through contrastive learning using said positive object and said negative object. Claim 19 A computer device according to claim 18, wherein at least one processor trains a model using a first loss function that induces a similarity score of 1 for the positive object and a second loss function that induces a similarity score of 0 for the negative object. Claim 20 A computer device according to claim 19, wherein the at least one processor further trains a model using a third loss function that induces a similarity score between identical object sets to 1.
Citation Information
Patent Citations
Method and procedure for driving autonomous test scenario using traffic accident image based on operational environment information in road traffic
KR1020210050150A
Method and system for retrieval of semantic in video
KR1020230032317A