Multi-modality-based vehicle re-identification apparatus and method

The multimodality-based vehicle re-identification system addresses misidentification issues by integrating license plate and environmental context with external characteristics, ensuring accurate vehicle recognition despite lighting and weather variations.

WO2026095724A1PCT designated stage Publication Date: 2026-05-07DELTAX CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DELTAX CO LTD
Filing Date
2025-10-31
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing vehicle re-identification technologies face challenges in accurately identifying vehicles due to exterior similarities and variations caused by lighting and weather conditions, leading to potential misidentification.

Method used

A multimodality-based vehicle re-identification system that integrates license plate information and environmental context with external vehicle characteristics, using a combination of cameras, processors, and machine learning models like YOLO and ResNet to generate a multimodality integrated vector for reliable identification.

Benefits of technology

Ensures high reliability in vehicle re-identification performance by leveraging comprehensive vehicle profiles, even under changing lighting and weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017776_07052026_PF_FP_ABST
    Figure KR2025017776_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a multi-modality-based vehicle re-identification apparatus and method including all of appearance information of a vehicle, license plate information of the vehicle, and environmental context information. The multi-modality-based vehicle re-identification apparatus according to the present disclosure comprises: a camera for capturing a visual image of the outside of a vehicle in real time; a communication unit for supporting wired or wireless communication between the vehicle re-identification apparatus and an external electronic device; a memory storing one or more instructions; and a processor for executing the one or more instructions stored in the memory, wherein the instructions, when executed by the processor, may instruct the processor to perform the operations of: detecting the vehicle and a surrounding environment in the visual image captured by the camera; cropping an image of the detected vehicle; extracting license plate information of the vehicle from the cropped image; extracting appearance information of the vehicle from the cropped image; generating multi-modality integrated information by combining the extracted appearance information and license plate information, and information on the detected surrounding environment; and re-identifying the vehicle on the basis of the multi-modality integrated information.
Need to check novelty before this filing date? Find Prior Art

Description

Multimodality-based vehicle re-identification device and method

[0001] The present disclosure relates to a multimodality-based vehicle re-identification device and method, and more specifically, to a multimodality-based vehicle re-identification device and method that includes all of the vehicle's external information, license plate information, and environmental context information.

[0002] Existing Re-ID technologies primarily identify vehicles by utilizing external characteristics, but there is a possibility of generating incorrect identification results due to factors such as exterior similarity, various lighting conditions, and changes in viewpoint.

[0003] For example, if the vehicle manufacturer or model is identical, external characteristics (color, shape, model, etc.) may be almost indistinguishable from the exterior, and consequently, there is a possibility that autonomous driving systems or vehicle tracking systems equipped with existing Re-ID technology may fail to identify the vehicle.

[0004] In addition, the vehicle's external characteristics may vary depending on changes in lighting or weather conditions; for example, the vehicle's color or outline may appear different when the lighting is dim, or the exterior may appear blurry due to rain, snow, or fog.

[0005] Therefore, autonomous driving systems or vehicle tracking systems applying existing vehicle re-identification (Re-ID) technology contained the possibility of misidentification caused by various environmental factors.

[0006] The present disclosure was devised in consideration of these conventional problems and aims to provide a multi-modality-based vehicle re-identification device and method that integrally utilizes various data, such as license plate information and environmental context information, as well as external characteristics of the vehicle.

[0007] In one embodiment of the present disclosure, a multimodality-based vehicle re-identification device mounted on a vehicle is provided, wherein the vehicle re-identification device comprises: a camera that captures a visual image of the exterior of the vehicle in real time; a communication unit that supports wired or wireless communication between the vehicle re-identification device and an external electronic device; a memory that stores one or more instructions; and a processor that executes the one or more instructions stored in the memory, wherein when the instructions are executed by the processor, the processor may perform the following operations: detecting a vehicle and surrounding environment within a visual image captured by the camera; cropping the detected vehicle image; extracting license plate information of the vehicle from the cropped image; extracting exterior information of the vehicle from the cropped image; generating multimodality integrated information by combining the extracted exterior information and license plate information and the detected surrounding environment information; and re-identifying the vehicle based on the multimodality integrated information.

[0008] In one embodiment, the operation of generating the multimodality integrated information by combining the extracted appearance information, the license plate information, and the detected surrounding environment information may include: the operation of converting each feature included in the extracted appearance information, the license plate information, and the detected surrounding environment information into each high-dimensional vector; the operation of generating a multimodality integrated vector by combining each high-dimensional vector; and the operation of storing the multimodality integrated vector as the multimodality integrated information.

[0009] In one embodiment, the operation of generating a multimodality integrated vector by combining each of the high-dimensional vectors may be performed by any one of the following: a concatentation-based combination method that generates the multimodality integrated vector by connecting each of the high-dimensional vectors; a weighted sum combination method that generates the multimodality integrated vector by assigning weights to each of the high-dimensional vectors and then summing them; a multi-layer neural network-based combination method that generates the multimodality integrated vector by receiving each of the high-dimensional vectors as input values ​​and combining them non-linearly through a multi-layer neural network; or an attention-based combination method that generates the multimodality integrated vector by learning the importance of each of the high-dimensional vectors and then assigning weights to the more important vector.

[0010] In one embodiment, the operation of detecting a vehicle and surrounding environment within a visual image captured by the camera can be performed using a YOLO (You Only Look Once) model that is pre-trained to detect a vehicle and surrounding environment within an image through an object recognition learning dataset.

[0011] In one embodiment, the operation of extracting license plate information of the vehicle from the cropped image can be performed through OCR (Optical Character Recognition) processing.

[0012] In one embodiment, the operation of extracting external information of the vehicle from the cropped image can be performed using a ResNet (Residual Network) model that is pre-trained to extract external information of the vehicle within the image through a vehicle recognition learning dataset.

[0013] In another embodiment of the present disclosure, a multimodality-based vehicle re-identification method used in a vehicle is provided, the method may include: detecting a vehicle and a surrounding environment within a visual image captured by a camera mounted on the vehicle; cropping the detected vehicle image; extracting license plate information of the vehicle from the cropped image; extracting exterior information of the vehicle from the cropped image; generating multimodality integrated information by combining the extracted exterior information, the license plate information, and the detected surrounding environment information; and re-identifying the vehicle based on the multimodality integrated information.

[0014] In one embodiment, the step of generating the multimodality integrated information by combining the extracted appearance information, the license plate information, and the detected surrounding environment information may include: a step of converting each feature included in the extracted appearance information, the license plate information, and the detected surrounding environment information into each high-dimensional vector; a step of generating a multimodality integrated vector by combining each high-dimensional vector; and a step of storing the multimodality integrated vector as the multimodality integrated information.

[0015] In one embodiment, the step of generating a multimodality integrated vector by combining each of the high-dimensional vectors may be performed by any one of a concatenation-based combination method that generates the multimodality integrated vector by connecting each of the high-dimensional vectors, a weighted sum combination method that generates the multimodality integrated vector by assigning weights to each of the high-dimensional vectors and then summing them, a multilayer neural network-based combination method that generates the multimodality integrated vector by receiving each of the high-dimensional vectors as input values ​​and non-linearly combining them through a multilayer neural network, or an attention-based combination method that generates the multimodality integrated vector by learning the importance of each of the high-dimensional vectors and then assigning weights to the more important vector.

[0016] In one embodiment, the step of detecting a vehicle and surrounding environment within a visual image captured by the camera may be performed using a YOLO (You Only Look Once) model that has been pre-trained to detect a vehicle and surrounding environment within an image through an object recognition learning dataset.

[0017] In one embodiment, the step of extracting license plate information of the vehicle from the cropped image can be performed through Optical Character Recognition (OCR) processing.

[0018] In one embodiment, the step of extracting exterior information of the vehicle from the cropped image may be performed using a ResNet (Residual Network) model that has been pre-trained to extract exterior information of the vehicle within the image through a vehicle recognition learning dataset.

[0019] According to the apparatus and method for performing multi-modality-based vehicle re-identification according to the present disclosure, even if the external characteristics of the vehicle change due to changes in lighting, angle, or weather conditions, etc., high reliability of re-identification performance can be guaranteed through a comprehensive unique profile of the vehicle.

[0020] FIG. 1 is a block diagram of a vehicle re-identification device according to one embodiment of the present disclosure.

[0021] FIG. 2 is a drawing for explaining the specific usage state of a vehicle re-identification device according to one embodiment of the present disclosure.

[0022] FIG. 3 is a block diagram of a vehicle re-identification method according to one embodiment of the present disclosure.

[0023] FIG. 4 is a drawing for explaining a specific usage state of a vehicle re-identification method according to one embodiment of the present disclosure.

[0024] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. First, it should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the present disclosure, if it is determined that a detailed description of related known components or functions could obscure the essence of the present disclosure, such detailed description will be omitted.

[0025] Throughout the specification, when a part is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "…part," "…unit," and "module" as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.

[0026] In addition, to facilitate understanding of the present disclosure, the terms used in this specification are explained as follows.

[0027] As used in this specification, the term "vehicle appearance information" refers to information regarding various vehicle appearance elements used to visually identify a vehicle, such as vehicle color, vehicle shape, vehicle angle, additional accessories and appearance variations, vehicle condition, etc.

[0028] As used in this specification, the term "vehicle environmental context information (or vehicle surrounding environment information)" refers to information regarding elements of the physical environment and situations surrounding the vehicle, such as roads, traffic signals, signs, surrounding vehicles, lighting, weather, etc., while the vehicle is driving.

[0029] As used in this specification, the term "vehicle license plate information" refers to a unique number or character assigned to a vehicle in the country or region where the vehicle is registered.

[0030] As used in this specification, the term "multimodality integrated information" refers to information combining vehicle appearance information, license plate information, and environmental context information used to uniquely identify a vehicle.

[0031] Hereinafter, an apparatus and method for detecting traffic accidents in a vehicle emergency rescue system according to the present disclosure will be described with reference to the drawings.

[0032] [Multi-modality based vehicle re-identification device (10)]

[0033] FIG. 1 is a block diagram of a multi-modality-based vehicle re-identification device (10) (hereinafter abbreviated as vehicle re-identification device (10)) according to one embodiment of the present disclosure, and FIG. 2 is a diagram showing a specific usage state in which the vehicle re-identification device (10) is used.

[0034] Referring to FIG. 1 and FIG. 2 together, the vehicle re-identification device (10) includes a camera (110), a memory (120), a communication unit (130), an input / output unit (140), and a control unit (150).

[0035] The camera (110) can capture visual images of the outside of the vehicle in real time, such as other vehicles (20) or surrounding environments (30, 40).

[0036] The memory (120) is configured to store instructions executed by the control unit (150). Additionally, the memory (120) can store data or output stored data according to the control of the control unit (150), for example, it can store a visual image of the vehicle interior captured by the camera (110) or output stored data according to the control of the control unit (150). In particular, the memory (120) according to the present disclosure stores a process and / or program regarding a vehicle re-identification method that can be executed by the control unit (150).

[0037] The communication unit (130) can support wired or wireless communication between the vehicle re-identification device (10) and an external electronic device (e.g., CCTV (closed circuit television), smartphone or security system, etc.).

[0038] The communication unit (130) can communicate with an external electronic device using wired communication technologies such as LAN (Local Area Network), WAN (Wide Area Network), Ethernet and / or ISDN (Integrated Services Digital Network) and / or wireless communication technologies such as WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Bluetooth, NFC (Near Field Communication), RFID (Radio Frequency Identification), infrared communication (IrDA, infrared Data Association), LTE (Long Term Evolution), and LTE-Advanced.

[0039] The communication unit (130) may include a communication processor, a communication circuit, an antenna, and / or a transceiver, etc.

[0040] The input / output unit (140) includes means for inputting data to control the control unit (150) and output means for outputting information processed by the control unit (150). For example, the input means may include a key pad, a dome switch, a touch pad (contact capacitive type, pressure resistive type, infrared detection type, surface ultrasonic conduction type, integral tension measurement type, piezo effect type, etc.), a jog wheel, a jog switch, etc., and the output means may include a display made of a liquid crystal screen.

[0041] The control unit (150) can control the overall operation of the vehicle re-identification device (10). For example, the control unit (150) may be implemented as at least one of a processing device such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), a PLD (Programmable Logic Device), an FPGA (Field Programmable Gate Array), a CPU (Central Processing Unit), a microcontroller, and / or a microprocessor.

[0042] In particular, the control unit (150) according to the present disclosure may execute a process including the operation of detecting a vehicle (20) and surrounding environment (30, 40) within a visual image captured by a camera (110), the operation of cropping the detected vehicle image, the operation of extracting license plate (210) information of the vehicle from the cropped image, the operation of extracting external information of the vehicle from the cropped image, the operation of generating multimodality integrated information by combining the external information of the vehicle, license plate information, and surrounding environment information, and the operation of re-identifying the vehicle based on the multimodality integrated information.

[0043] To this end, the control unit (150) may include an object detection module (151), a cropping module (152), a license plate information extraction module (153), an appearance information extraction module (154), a multi-modality integration module (155), and a re-identification module (156).

[0044] Hereinafter, the operations executed by each module (151-156) of the control unit (150) will be described.

[0045] The object detection module (151) can receive at least one image through the camera (110) or the communication unit (130). For example, this image may be an image including a vehicle and the surrounding environment, or it may be a video frame.

[0046] The object detection module (151) can perform the operation of detecting a vehicle (20) and surrounding environment (30, 40) within an input image.

[0047] In particular, the object detection module (151) according to the present disclosure can quickly and efficiently detect a vehicle (20) and surrounding environment (30, 40) within an input image (or video frame) by using a YOLO (You Only Look Once) model that has been pre-trained through an object recognition learning dataset.

[0048] In particular, the object detection module (151) according to the present disclosure employs a YOLO model that has been pre-trained with object recognition sample images, and it has been confirmed that when using a YOLO model that has been pre-trained with approximately 1,000 object recognition images, the vehicle and surrounding environment within the image (or video frame) can be detected very accurately.

[0049] In another embodiment, the object detection module (151) may use a Faster-RCNN model instead of a YOLO model to detect the vehicle and surrounding environment within an image (or video frame).

[0050] The YOLO model applied to the object detection module (151) according to the present disclosure may be trained to detect a vehicle (20) and surrounding environment (30, 40) within an image (or video frame) through the operation of dividing an input image (or video frame) into a grid of size SxS, the operation of predicting B bounding boxes and confidence and class probability for each box in each grid cell, the operation of filtering the predicted bounding boxes according to confidence scores, the operation of removing duplicate boxes using Non-Maximum Suppression (NMS) and finally leaving only boxes with high confidence, and the operation of outputting the remaining bounding boxes and class labels.

[0051] The cropping module (152) can perform the operation of cropping the vehicle area within an image (or video frame) based on the bounding box of the vehicle detected through the object detection module (151).

[0052] The vehicle image cropped by the cropping module (152) is subsequently transmitted to the license plate information extraction module (153) and the appearance information extraction module (154).

[0053] That is, the license plate information extraction module (153) and the appearance information extraction module (154) do not analyze the entire image, but process only the cropped specific area transmitted from the cropping module (152), thereby allowing for efficient use of computing resources and increased data processing speed, and thus supporting real-time processing performance of the autonomous driving system or vehicle tracking system.

[0054] The license plate information extraction module (153) can perform the operation of extracting license plate (210) information of a vehicle from a cropped vehicle image (or video frame) transmitted from the cropping module (152).

[0055] In particular, the license plate information extraction module (153) according to the present disclosure can perform the operation of extracting license plate (210) information of a vehicle from a cropped vehicle image (or video frame) through OCR (Optical Character Recognition) processing.

[0056] The external information extraction module (154) can perform the operation of extracting external features (e.g., color, shape, angle, etc.) of a vehicle from a cropped vehicle image (or video frame) transmitted from the cropping module (152).

[0057] In particular, the external information extraction module (154) according to the present disclosure can extract external information of a vehicle within a cropped vehicle image (or video frame) by using a ResNet (Residual Network) model that has been pre-trained through a vehicle recognition learning dataset.

[0058] For example, the external information extraction module (154) can generate a visual profile of the vehicle through a high-dimensional feature vector as an external feature of the vehicle.

[0059] The multimodality integration module (155) can perform the operation of generating multimodality integration information unique to the vehicle by combining environmental context information extracted by the object detection module (151), vehicle license plate information extracted by the license plate information extraction module (153), and vehicle appearance information extracted by the appearance information extraction module (154).

[0060] Specifically, the multimodality integration module (155) can perform the operation of converting each feature included in the environmental context information extracted by the object detection module (151), the vehicle license plate information extracted by the license plate information extraction module (153), and the vehicle appearance information extracted by the appearance information extraction module (154) into each high-dimensional vector, the operation of combining each of these high-dimensional vectors to generate a multimodality integration vector, and the operation of storing the multimodality integration vector as multimodality integration information.

[0061] In one embodiment, the multimodality integration module (155) may use any one of the following: a concatentation-based combination method that generates a multimodality integration vector by connecting each high-dimensional vector; a weighted sum combination method that generates a multimodality integration vector by assigning weights to each high-dimensional vector and then summing them; a multi-layer neural network-based combination method that generates a multimodality integration vector by receiving each high-dimensional vector as input and non-linearly combining them through a multi-layer neural network; or an attention-based combination method that generates a multimodality integration vector by learning the importance of each high-dimensional vector and then assigning weights to the more important vector.

[0062] For example, the multimodality integration module (155) can generate a multimodality integration vector using a concatentation-based combination method.

[0063] In this case, for example, the multimodality integration module (155) can simply connect the respective feature vectors related to environmental context information, vehicle license plate information, and vehicle appearance information to generate a single multimodality integration vector. This can be expressed as a formula as follows:

[0064]

[0065] Here, V fused is the vehicle's multimodality integration vector, and V appearnce is a feature vector related to the vehicle's external information, and V license plate is a feature vector related to vehicle license plate information, and V environment is a feature vector related to the vehicle's environmental context information, and ";" signifies a connection between vectors.

[0066] In another example, the multimodality integration module (155) may generate a multimodality integration vector using a weighted sum combination method.

[0067] In this case, for example, the multimodality integration module (155) can generate a multimodality integration vector by assigning weights to each feature vector related to environmental context information, vehicle license plate information, and vehicle appearance information, and then summing them. This is used when certain information is more important than other information, and can be expressed as a formula as follows:

[0068]

[0069] Here, V fused is the vehicle's multimodality integration vector, and V appearnce is a feature vector related to the vehicle's external information, and V license plate is a feature vector related to vehicle license plate information, and V environment is a feature vector related to the environmental context information of the vehicle, and w1, w2, and w3 are weights representing the importance of each piece of information.

[0070] In another example, the multimodality integration module (155) may generate a multimodality integration vector using a multi-layer neural network-based combination method.

[0071] In this case, for example, the multimodality integration module (155) receives each feature vector related to environmental context information, vehicle license plate information, and vehicle appearance information as input and combines them non-linearly through a multilayer neural network, and the multilayer neural network learns the correlation between each piece of information.

[0072] In another example, the multimodality integration module (155) may also generate a multimodality integration vector using an attention-based combination method.

[0073] In this case, for example, the multimodality integration module (155) can learn the importance of each feature vector related to environmental context information, vehicle license plate information, and vehicle appearance information, and based on this, give weights to the information that is more important in a specific situation to generate a flexible multimodality integration vector. This can be expressed as a formula as follows:

[0074]

[0075] Here, V fused is the vehicle's multimodality integration vector, and is an attention value representing the importance of each feature vector.

[0076] The re-identification module (156) performs the operation of re-identifying the vehicle based on multimodality integration information unique to the vehicle.

[0077] For example, the re-identification module (156) performs a re-identification operation tailored to the situation by increasing the importance of the vehicle's external information during the day and increasing the importance of the vehicle's license plate information or environmental context information at night.

[0078] [Multi-modality-based Vehicle Re-identification Method]

[0079] FIG. 3 is a block diagram of a multi-modality-based vehicle re-identification method (hereinafter abbreviated as vehicle re-identification method) according to one embodiment of the present disclosure, and FIG. 4 is a diagram showing a specific usage state in which the vehicle re-identification method is used.

[0080] Referring together to FIG. 3 and FIG. 4, a vehicle re-identification method according to one embodiment of the present disclosure comprises: a first step process (S310) of receiving an image (or video frame) of the outside of the vehicle through a camera; a second step process (S320) of detecting the vehicle and the surrounding environment within the visual image; a third step process (S330) of cropping the detected vehicle image; a fourth step process (S340) of extracting license plate information of the vehicle from the cropped image; a fifth step process (S350) of extracting external information of the vehicle from the cropped image; a sixth step process (S360) of generating multi-modality integrated information by combining the external information of the vehicle, license plate information, and surrounding environment information; and a seventh step process (S370) of re-identifying the vehicle based on the multi-modality integrated information.

[0081] In one embodiment, the vehicle re-identification method according to the present disclosure may further include, prior to the first step (S310) of receiving an image (or video frame) outside the vehicle through a camera, the step of preparing a YOLO (You Only Look Once) model that is pre-trained to detect the vehicle and surrounding environment within the image (or video frame) through an object recognition learning dataset.

[0082] In one embodiment, the vehicle re-identification method according to the present disclosure may further include, prior to the first step (S310) of receiving an image (or video frame) outside the vehicle through a camera, the step of preparing a ResNet (Residual Network) model that is pre-trained to extract external information of the vehicle within the image (or video frame) through a vehicle recognition learning dataset.

[0083] In one embodiment, the second step (S320) of detecting a vehicle and surrounding environment within a visual image can be performed by a pre-trained YOLO (You Only Look Once) model.

[0084] In one embodiment, the process of the fourth step (S340) of extracting vehicle license plate information from a cropped image can be performed through Optical Character Recognition (OCR) processing.

[0085] In one embodiment, the fifth step (S350) of extracting external information of a vehicle from a cropped image can be performed by a pre-trained ResNet model.

[0086] In one embodiment, the process of step 6 (S360) for generating multimodality integrated information by combining the vehicle's external information, license plate information, and surrounding environment information may include the process of converting features included in the extracted external information and license plate information and the detected surrounding environment information into respective high-dimensional vectors, the process of generating a multimodality integrated vector by combining each high-dimensional vector, and the process of storing the multimodality integrated vector as multimodality integrated information.

[0087] In one embodiment, the process of the step of generating a multimodality integrated vector by combining each high-dimensional vector may be performed by any one of the following: a concatenation-based combination method that generates a multimodality integrated vector by connecting each high-dimensional vector; a weighted sum combination method that generates a multimodality integrated vector by assigning weights to each high-dimensional vector and then summing them; a multilayer neural network-based combination method that generates a multimodality integrated vector by receiving each high-dimensional vector as input and non-linearly combining them through a multilayer neural network; or an attention-based combination method that generates a multimodality integrated vector by learning the importance of each high-dimensional vector and then assigning weights to the more important vector.

[0088] [Industrial Use]

[0089] The vehicle re-identification device and method according to the present disclosure can be used primarily in autonomous driving systems or vehicle tracking systems.

[0090] In autonomous driving systems or vehicle tracking systems, it is very important to distinguish similar vehicles in complex traffic situations.

[0091] According to the vehicle re-identification device and method of the present disclosure, by combining vehicle license plate information, exterior features, and environmental context, each vehicle can be accurately re-identified even in a complex traffic environment.

[0092] For example, according to the vehicle re-identification device and method of the present disclosure, even if the external characteristics of the vehicle change due to changes in lighting or weather conditions, the vehicle can be accurately re-identified in any situation through auxiliary information such as license plate information and environmental context information.

[0093] The devices and methods described above may be implemented as hardware components, software components, and / or combinations of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on the operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0094] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or in combination. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or transmitted signal wave, so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0095] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based on the above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0096] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

1. As a multi-modality-based vehicle re-identification device mounted on a vehicle, A camera that captures a visual image of the exterior of the above vehicle in real time; A communication unit that supports wired or wireless communication between the above-mentioned vehicle re-identification device and an external electronic device; Memory for storing one or more instructions; and It includes a processor that executes one or more instructions stored in the memory, When the above instructions are executed by the processor, the processor, Operation of detecting a vehicle and surrounding environment within a visual image captured by the above camera; The operation of cropping the above-detected vehicle image; The operation of extracting license plate information of the vehicle from the cropped image above; The operation of extracting external information of the vehicle from the cropped image above; An operation to generate multimodality integrated information by combining the extracted appearance information, the license plate information, and the detected surrounding environment information; and operation of re-identifying the vehicle based on the above multimodality integration information A multi-modality-based vehicle re-identification device that enables the performance of 2. In Paragraph 1, The operation of generating the multimodality integrated information by combining the extracted appearance information, the license plate information, and the detected surrounding environment information is, An operation of converting each feature included in the extracted appearance information, the license plate information, and the detected surrounding environment information into respective high-dimensional vectors; The operation of generating a multimodality integration vector by combining each of the above high-dimensional vectors; and The operation of storing the above multimodality integration vector as the above multimodality integration information A multi-modality based vehicle re-identification device including 3. In Paragraph 2, The operation of generating a multimodality integration vector by combining each of the above high-dimensional vectors is, A concatentation-based combination method that generates the multimodality integration vector by connecting each of the above high-dimensional vectors, or A weighted sum combination method that generates the multimodality integration vector by assigning weights to each of the above high-dimensional vectors and summing them, or A multi-layer neural network-based combination method that generates the multi-modality integration vector by receiving each of the above high-dimensional vectors as input values ​​and combining them non-linearly through a multi-layer neural network, or An attention-based combination method that generates the multimodality integration vector by learning the importance of each of the above high-dimensional vectors and then assigning weights to the more important vectors. A multi-modality-based vehicle re-identification device performed by any one of the following.

4. In Paragraph 1, A multi-modality-based vehicle re-identification device in which the operation of detecting a vehicle and surrounding environment within a visual image captured by the above camera is performed using a YOLO (You Only Look Once) model pre-trained to detect a vehicle and surrounding environment within an image through an object recognition learning dataset.

5. In Paragraph 4, A multi-modality-based vehicle re-identification device in which the operation of extracting license plate information of the vehicle from the cropped image is performed through OCR (Optical Character Recognition) processing.

6. In Paragraph 5, A multi-modality-based vehicle re-identification device, wherein the operation of extracting external information of the vehicle from the cropped image is performed using a ResNet (Residual Network) model pre-trained to extract external information of the vehicle within the image through a vehicle recognition learning dataset.

7. As a multimodality-based vehicle re-identification method used in a vehicle, A step of detecting the vehicle and the surrounding environment within a visual image captured by a camera mounted on the vehicle; A step of cropping the detected vehicle image above; A step of extracting license plate information of the vehicle from the cropped image above; A step of extracting external information of the vehicle from the cropped image above; A step of generating multimodality integrated information by combining the extracted appearance information, the license plate information, and the detected surrounding environment information; and A step of re-identifying the vehicle based on the above multimodality integration information A method including 8. In Paragraph 7, The step of generating the multimodality integration information by combining the extracted appearance information, the license plate information, and the detected surrounding environment information is: A step of converting each feature included in the extracted appearance information, the license plate information, and the detected surrounding environment information into respective high-dimensional vectors; A step of generating a multimodality integration vector by combining each of the above high-dimensional vectors; and Step of storing the above multimodality integration vector as the above multimodality integration information A multi-modality based vehicle re-identification method including 9. In Paragraph 8, The step of generating a multimodality integration vector by combining each of the above high-dimensional vectors is: A concatenation-based combination method that generates the multimodality integration vector by connecting each of the above high-dimensional vectors, or A weighted sum combination method that generates the multimodality integration vector by assigning weights to each of the above high-dimensional vectors and summing them, or A multilayer neural network-based combination method that generates the multimodality integration vector by receiving each of the above high-dimensional vectors as input values ​​and combining them non-linearly through a multilayer neural network, or An attention-based combination method that generates the multimodality integration vector by learning the importance of each of the above high-dimensional vectors and then assigning weights to the more important vectors. A multi-modality-based vehicle re-identification method performed by any one of the following.

10. In Paragraph 9, A multi-modality-based vehicle re-identification method, wherein the step of detecting a vehicle and surrounding environment within a visual image captured by the camera is performed using a YOLO (You Only Look Once) model pre-trained to detect a vehicle and surrounding environment within an image through an object recognition learning dataset.

11. In Paragraph 10, A multi-modality-based vehicle re-identification method in which the step of extracting license plate information of the vehicle from the cropped image is performed through Optical Character Recognition (OCR) processing.

12. In Paragraph 11, A multi-modality-based vehicle re-identification method, wherein the step of extracting external information of the vehicle from the cropped image is performed using a ResNet (Residual Network) model that is pre-trained to extract external information of the vehicle within the image through a vehicle recognition learning dataset.