Device and Method for Detecting and Tracking Object Based on Multi-Domain
Patent Information
- Application Number
- KR1020250024051
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2026-09-01
Smart Images

Figure PAT00024_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an object tracking and detection device and method, and more specifically, to a method and device for tracking and detecting objects using multiple domains. Background Technology
[0003] Recent research on multi-object tracking primarily utilizes the tracking-by-detection method, which identifies objects of interest using an object detector and generates trajectories through data association. Data association is addressed by techniques such as minimum cost flow, partial filtering, Hungarian allocation, and graph cut, but these methods mostly rely on heuristic approaches and have the problem of being sensitive to local optimization.
[0004] To address this problem, an end-to-end framework combining an attention mechanism with appearance and motion modeling was proposed; however, relying solely on RGB-based single-domain information still presented a challenge in achieving satisfactory multi-object estimation performance in complex environments with significant lighting variations.
[0005] Meanwhile, although techniques for tracking objects using multiple domains have been proposed, most of the proposed techniques involve selecting only one domain from the multiple domains, and there was a problem in that they could not fully utilize the information of the multiple domains. The problem to be solved
[0007] The present invention proposes an apparatus and method for tracking and detecting objects robustly to changes in the surrounding environment, such as time zones and weather, by utilizing information from multiple domains. means of solving the problem
[0009] According to one aspect of the present invention, a multi-domain-based object tracking and detection method performed in a computing device including a processor and memory is provided, comprising: a step of acquiring individual domain features of each sensor by inputting each image acquired by a plurality of sensors into a neural network model pre-set for each sensor (a); a step of acquiring common domain features of each sensor by inputting each image acquired by the plurality of sensors into a single neural network pre-set (b); a step of calculating sensor-specific domain features by combining the individual domain features and the common domain features (c); a step of acquiring an environment variable vector, which is a multi-dimensional vector, for each sensor by inputting the domain features into an environment variable inference model (d); a step of calculating a reliability score for each sensor by applying a pre-set importance to each vector element value of the sensor-specific environment variable vector, and calculating a fusion weight for each sensor based on the reliability score (f); a step of acquiring fusion features by inputting the fusion weight of each sensor and the domain features of each sensor into a transformer model (g); and a step of performing object tracking and detection by inputting the fusion features into an object tracking / detection model (h).
[0010] The above Transformer model performs the cross-attention operation by setting the highest fusion weight as the sensor's domain feature Q (query) and the combination of the remaining sensors as a K (Key) - V (Value) combination.
[0011] The above transformer model outputs fused features through the weighted summation of the self-attention vector obtained from the self-attention operation of the domain features for each sensor and the cross-attention vector obtained from the cross-attention operation of the domain features for each sensor.
[0012] The single neural network of step (b) above includes a vision transformer.
[0013] The importance set for each vector element value constituting the above sensor-specific environment variable vector is pre-set based on the characteristics of each sensor.
[0014] The reliability score of each of the above sensors is calculated as follows:
[0015]
[0016] In the above mathematical formula is the confidence score of each sensor, i is the index representing the sensor, env is the index representing each dimension element of the environment variable vector, and K represents the vector element value constituting the environment variable vector.
[0017] The above fusion weights are calculated as follows:
[0018]
[0019] In the above mathematical formula, W is the fusion weight.
[0020] According to another aspect of the present invention, a multi-domain-based object tracking and detection device comprises: a processor; and a memory connected to the processor, wherein the processor comprises the steps of: (a) inputting each image acquired by a plurality of sensors into a neural network model pre-set for each sensor to acquire individual domain features of each sensor; (b) inputting each image acquired by the plurality of sensors into a single neural network pre-set to acquire common domain features of each sensor; (c) combining the individual domain features and the common domain features to compute sensor-specific domain features; (d) inputting the domain features into an environment variable inference model to acquire an environment variable vector, which is a multi-dimensional vector, for each sensor; (f) applying a pre-set importance to each vector element value of the sensor-specific environment variable vector to compute a reliability score for each sensor and compute a fusion weight for each sensor based on the reliability score; and (g) inputting the fusion weight of each sensor and the domain features of each sensor into a transformer model to acquire fusion features. A multi-domain-based object tracking and detection device is provided, which performs the step (h) of inputting the above-mentioned fusion features into an object tracking / detection model to perform object tracking and detection. Effects of the invention
[0022] According to the object tracking and detection of the present invention, there is an advantage in that objects can be tracked and detected robustly against changes in the surrounding environment, such as time of day and weather, by utilizing information from multiple domains. Brief explanation of the drawing
[0024] FIG. 1 is a block diagram illustrating the schematic structure of a multi-domain-based object tracking and detection device according to one embodiment of the present invention. FIG. 2 is a diagram illustrating an example of the detailed structure of a domain feature extraction model according to an embodiment of the present invention. FIG. 3 is a diagram showing an example of the operational structure of an environment variable inference model according to an embodiment of the present invention. FIG. 4 is a diagram showing the operation structure of a transformer model according to an embodiment of the present invention. FIG. 5 is a diagram showing the learning structure of an object tracking and detection device according to an embodiment of the present invention. FIG. 6 is a flowchart showing the overall flow of a multi-domain-based object detection method according to an embodiment of the present invention. Specific details for implementing the invention
[0025] In order to fully understand the present invention, the operational advantages of the present invention, and the objectives achieved by the implementation of the present invention, reference must be made to the accompanying drawings illustrating preferred embodiments of the present invention and the contents described in the accompanying drawings.
[0026] The present invention will be described in detail below by explaining preferred embodiments with reference to the attached drawings. However, the present invention may be implemented in various different forms and is not limited to the embodiments described. Furthermore, to clearly explain the present invention, parts unrelated to the description are omitted, and the same reference numerals in the drawings indicate the same components.
[0027] Throughout the specification, when a part is described as “comprising” a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as “...part,” “...unit,” “module,” and “block” as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.
[0028] FIG. 1 is a block diagram illustrating the schematic structure of a multi-domain-based object tracking and detection device according to one embodiment of the present invention.
[0029] Referring to FIG. 1, a multi-domain-based object tracking and detection device according to one embodiment of the present invention includes an individual domain feature extraction model (100), a common domain feature extraction model (200), an environment variable inference model (300), a sensor-specific fusion weight calculation module (400), a transformer model (500), and an object tracking / detection model (600).
[0030] The individual domain feature extraction model (100) receives a multi-domain image obtained from a multi-domain and functions to calculate and output individual domain features of each sensor. The present invention is based on the premise of using multiple domains, where multiple domains mean using multiple sensors. In this embodiment, domains and sensors may be used interchangeably, and sensors and domains have the same meaning.
[0031] The individual domain feature extraction model (100) extracts features of each sensor through neural network operations on images obtained from each of the multiple sensors. Here, the sensors may include RGB cameras, thermal cameras, depth cameras, LiDAR sensors, etc.
[0032] FIG. 2 is a diagram illustrating an example of the detailed structure of a domain feature extraction model according to an embodiment of the present invention.
[0033] Referring to FIG. 2, the individual domain feature extraction model (100) may include a first neural network model (110), a second neural network model (120), a third neural network model (130), and a fourth neural network model (140). In the individual domain feature model (100), the number of neural network models is equal to the number of domains (sensors) used, and different neural networks may be used for each domain to extract individual features of each domain.
[0034] The first neural network model (110) receives an RGB image from an RGB camera and extracts features of the RGB image through neural network operations.
[0035] The second neural network model (120) receives a thermal image from a thermal camera and extracts features of the thermal image through neural network operations.
[0036] The third neural network model (130) receives a depth image from a depth camera and extracts features of the depth image through neural network operations.
[0037] The fourth neural network model (140) receives point cloud data from the LiDAR sensor and extracts features of the point cloud data through neural network operations.
[0038] Here, for the first to fourth neural networks, a neural network suitable for each domain image may be selected. For example, a CNN (Convolutional Neural Network) model may be used as the first neural network model (110) for extracting features for RGB images. In this embodiment, individual domain features F domain , i Defined as such, where i is an index representing the domain (sensor).
[0039] Referring again to FIG. 1, the common feature extraction model (200) calculates common features for each domain image through neural network operations. Unlike individual domain feature extraction, a single neural network is used for the common feature extraction model (200), and the single neural network extracts features for each domain image. According to one embodiment of the present invention, a Vision Transformer may be used as the single neural network of the common feature extraction model (200).
[0040] In this embodiment, the common features of each domain output from the common feature extraction model (200) are F common,i Defined as.
[0041] The individual domain features output through the individual domain feature extraction model (100) and the common features output through the common feature extraction model (200) are combined and input into the environment variable inference model (300).
[0042] Individual domain features (F domain , i ) and common domain features (F common,i The combination of ) is individual domain features (F) as shown in the following mathematical formula 1. domain , i ) and common domain features (F common,i It can be expressed as the sum of ).
[0043]
[0044] In this embodiment, F i is defined as a domain characteristic.
[0045] The environment variable inference model (300) receives domain features as input and outputs an environment variable vector for each domain feature through neural network operation. According to one embodiment of the present invention, a general classification model may be used as the environment variable inference model (300).
[0046] The dimensions of the environment variable vector are pre-set, and the vector element values of each dimension are set to represent a specific environment. For example, the environment variable vector may be a 3-dimensional vector and may be a 3-dimensional vector consisting of values representing weather, lighting, and time. Specifically, the environment variable inference model (300) may be trained to represent the degree of clearness for the weather as a value from 0 to 1, the illuminance as a value from 0 to 1, and the time as a value from 0 to 1.
[0047] The operation of outputting an environment variable vector in the environment variable inference model (300) can be defined as follows in Equation 2.
[0048]
[0049] In mathematical equation 2 above, W envrepresents the pre-trained weights of the environment variable inference model, and b env is the pre-trained bias of the environment variable inference model, and Softmax represents the softmax operation, represents an environment variable vector, and i is an index representing the domain.
[0050] As can be seen from the above mathematical formula 2, the environment variable vector is output for each domain, and as in the previous example, when four domains (sensors) are used, four environment variable vectors are output for each domain through the environment variable inference model (300).
[0051] FIG. 3 is a diagram showing an example of the operational structure of an environment variable inference model according to one embodiment of the present invention.
[0052] Referring to FIG. 3, domain features of each domain are input to the environment variable inference model (300), and the environment variable inference module performs neural network operations based on pre-learned weights and biases.
[0053] Figure 3 illustrates a structure in which an environment variable inference module (300) outputs an environment variable vector containing information representing the probability of the weather, the probability of the illuminance, and the probability of the time through neural network operation.
[0054] For example, the environment variable inference module can be trained so that it is close to 1 when the weather is clear and close to 0 when the weather is cloudy, close to 1 when the illuminance is high and close to 0 when the illuminance is low, and close to 1 when the time is day and close to 0 when the time is night. As previously explained, the number of classes (vector dimensions) representing the environment can be changed as needed.
[0055] The sensor-specific fusion weight calculation module (400) performs the function of calculating the fusion weight of each sensor. Here, the fusion weight refers to the weight applied when fusing each domain feature.
[0056] The domain-specific fusion weight calculation module (400) calculates the reliability score of each sensor before calculating the fusion weight. For the calculation of the reliability score, a preset importance is used for each vector element of the environment variable vector. As previously explained, the environment variable vector consists of multiple dimensions, and the importance for each vector element value (scalar) is preset, and the reliability score for each sensor is calculated using the preset importance. The importance for each vector element value (scalar) is preset based on the characteristics of the sensor.
[0057] For example, RGB cameras are sensors that are sensitive to illumination and time. Accordingly, for RGB cameras, a relatively high weight is set for illumination and time, while a relatively low weight is set for weather.
[0058] In addition, since thermal imaging cameras are not sensitive to illuminance and time but can be sensitive to weather, a relatively high weight is set for weather, and a relatively low weight is set for illuminance and time.
[0059] The reliability score for each sensor is calculated through weighted summation by applying a pre-set vector element importance as a weight to each vector element value constituting the environment variable vector.
[0060] The confidence score for each domain can be calculated as shown in the following mathematical formula 3.
[0061]
[0062] In the above mathematical formula is the confidence score of each sensor, i is the index representing the sensor, env is the index representing each dimension element of the environment variable vector, and K represents the vector element value constituting the environment variable vector.
[0063] The sensor-specific fusion weight calculation module (400) calculates sensor-specific fusion weights based on the calculated sensor-specific reliability scores. The sensor-specific fusion weights are calculated through normalization of the domain-specific reliability scores of each domain.
[0064] The fusion weight for each sensor can be calculated as shown in the following mathematical formula 4.
[0065]
[0066] In the above mathematical formula 4, W is the fusion weight, and i is an index representing the sensor (domain).
[0067] For the transformer model (500), the domain features (F) for each sensor are i Fusion weights for each sensor and ) are input. The transformer model (500) generates fusion features through self-attention and cross-attention operations of domain features based on fusion weights.
[0068] FIG. 4 is a diagram showing the operation structure of a transformer model according to one embodiment of the present invention.
[0069] Referring to Fig. 4, the Transformer model performs a self-attention operation for each domain feature. The self-attention operation is performed using only domain features, and fusion weights are not used in the self-attention operation.
[0070] According to one embodiment of the present invention, the self-attention operation can be performed as shown in the following mathematical formula 5.
[0071]
[0072] In the above mathematical formula 5 is the query, key, and value matrix of sensor i, and It is defined as. Also, d k represents the dimension of the matrix key.
[0073] The transformer model (500) performs cross-attention operations based on fusion weights. According to a preferred embodiment of the present invention, the fusion feature of the sensor (domain) with the highest fusion weight is fixed as Q for the cross-attention operation, and the fusion features of the remaining sensors (domains) are used as a combination of K and V to perform the cross-attention operation. For example, if an RGB camera, a thermal imaging camera, a depth camera, and a LiDAR sensor are used, and the fusion weight of the RFG camera is the highest, the fusion feature of the RGB camera is fixed as Q for the cross-attention operation, and the fusion features of the remaining sensors are used as a KV combination. Therefore, when four sensors are used, a total of three cross-attention operation results are output as vectors. If five sensors are used, six cross-attention operation results of 4C2 may be output as vectors.
[0074] The self-attention vector, which is the result of the self-attention operation of the Transformer model (500), and the cross-attention vector, which is the result of the cross-attention operation, are weighted and summed, and the weights for summing the self-attention vector and the cross-attention vector are determined by learning.
[0075] The cross-attention operation of the transformer model (500) can be performed as shown in the following mathematical formula 6.
[0076]
[0077] In the above mathematical formulas 5 and 6, is a self-attention vector, and is a cross-attention vector.
[0078] The weighted sum of the cross attention vector and the cell attention vector is performed as shown in the following mathematical formula 7.
[0079]
[0080] In the above mathematical formula, W i and W ijis a pre-learned weight. According to Equation 7, the fusion feature (F final ) is calculated.
[0081] The present invention makes it possible to derive fusion features that appropriately reflect the characteristics of multiple domains by calculating fusion weights for each sensor (domain) and reflecting the fusion weights in the cross-attention operation of a transformer. Most existing feature extraction structures using multiple domains were based on a method of selecting the most appropriate domain, and had the problem of failing to properly fuse the meaningful features of each domain. The present invention makes it possible to obtain improved fusion features compared to existing ones by deriving sensor-specific fusion weights based on the reliability of each sensor and reflecting them in the cross-attention operation.
[0082] The object tracking / detection model (600) receives the final fusion feature as input and tracks and detects the object through neural network operations.
[0083] FIG. 5 is a diagram showing the learning structure of an object tracking and detection device according to one embodiment of the present invention.
[0084] Referring to FIG. 5, the difference between the output of the object tracking / detection model (600) and the label is set as a loss, and the weights of the object tracking / detection model (600), the transformer model (500), and the environment variable inference model (300) are learned by backpropagating the loss.
[0085] The domain feature extraction model (100) and common feature extraction model (200) utilize existing commercial neural network models that have been trained and are not trained separately.
[0086] According to one embodiment of the present invention, the loss for training the object tracking / detection model (600), the transformer model (500), and the environment variable inference model (300) is calculated as shown in the following mathematical formula 8.
[0087]
[0088] In the above mathematical formula, λdet and λ track is a pre-set hyperparameter, and L det is the loss function for object detection, and L track is a loss function for object tracking.
[0089] FIG. 6 is a flowchart showing the overall flow of a multi-domain-based object detection method according to an embodiment of the present invention.
[0090] Referring to FIG. 6, first, each image obtained from a plurality of sensors is input into a separate neural network to obtain individual domain features for each sensor (step 600).
[0091] Images obtained from each sensor are input into the same neural network (e.g., Vision Transformer) to obtain common domain features for each sensor (step 610).
[0092] Domain features for each sensor are obtained by combining individual domain features and common domain features (step 620). For example, domain features for each sensor can be obtained by summing the individual domain features and common domain features of each sensor.
[0093] Domain features for each sensor are input into an environment variable inference model to obtain an environment variable vector for each sensor (step 630). The environment variable vector for each sensor has a preset dimension, and the environment variable inference model is trained so that each vector element represents the probability of a specific environment.
[0094] For each vector element of the environment variable vector for each sensor, a reliability score for each sensor is calculated using a preset importance (step 640). The reliability score for each sensor can be calculated as shown in Equation 3 above.
[0095] The reliability score of each sensor is normalized to calculate the fusion weight (step 650). The fusion weight is calculated as shown in Equation 4.
[0096] The domain features and fusion weights of each sensor are input into the transformer model to obtain fusion features (step 660). The transformer model performs self-attention and cross-attention operations. During the cross-attention operation, the transformer model is configured to fix the domain feature of the sensor with the highest fusion weight as Q (Query) and set the combination of the remaining sensors as the KV combination to perform cross-attention.
[0097] Fusion features are input to an object tracking / detection model, and the object tracking / detection model performs object tracking and detection in multi-domain images through neural network operations (step 670).
[0098] According to one embodiment of the present invention, the device of the present invention may be implemented as a computing device including a processor and memory, and the device of the present invention may be performed by the computing device.
[0099] The present invention has been described with reference to the embodiments illustrated in the drawings, but this is merely illustrative, and those skilled in the art will understand that various modifications and equivalent alternative embodiments are possible therefrom.
[0100] Therefore, the true scope of technical protection of the present invention should be determined by the technical concept of the appended claims.
Claims
Claim 1 A multi-domain-based object tracking and detection method performed in a computing device including a processor and memory, comprising: a step of acquiring individual domain features of each sensor by inputting each image acquired by a plurality of sensors into a neural network model pre-set for each sensor; a step of acquiring common domain features of each sensor by inputting each image acquired by the plurality of sensors into a single neural network pre-set; a step of calculating sensor-specific domain features by combining the individual domain features and the common domain features; a step of acquiring a multi-dimensional vector of environment variables for each sensor by inputting the domain features into an environment variable inference model (d); a step of calculating a reliability score for each sensor by applying a pre-set importance to each vector element value of the sensor-specific environment variable vector, and calculating a fusion weight for each sensor based on the reliability score (f); a step of acquiring fusion features by inputting the fusion weight of each sensor and the domain features of each sensor into a transformer model (g); and a step of performing object tracking and detection by inputting the fusion features into an object tracking / detection model (h). Claim 2 A multi-domain-based object tracking and detection method according to claim 1, wherein the transformer model performs a cross-attention operation by setting the highest fusion weight during the cross-attention operation as the domain feature of the sensor as Q (query) and the combination of the remaining sensors as a K (Key) - V (Value) combination. Claim 3 In paragraph 2, the transformer model outputs a fused feature through the weighted summation of a self-attention vector formed by a self-attention operation of the domain features for each sensor and a cross-attention vector formed by a cross-attention operation of the domain features for each sensor. This is a multi-domain-based object tracking and detection method. Claim 4 In claim 1, the single neural network of step (b) is a multi-domain-based object tracking and detection method including a vision transformer. Claim 5 A multi-domain-based object tracking and detection method according to claim 1, wherein the importance set for each vector element value constituting the environment variable vector for each sensor is pre-set based on the characteristics of each sensor. Claim 6 A multi-domain-based object tracking and detection method according to claim 5, wherein the reliability score of each sensor is calculated as follows: In the above mathematical formula is the confidence score of each sensor, i is the index representing the sensor, env is the index representing each dimension element of the environment variable vector, and K represents the vector element value constituting the environment variable vector. Claim 7 In claim 6, the fusion weight is calculated as follows in the multi-domain-based object tracking and detection method. In the above mathematical formula, W is the fusion weight. Claim 8 A multi-domain-based object tracking and detection device comprises: a processor; and a memory connected to the processor, wherein the processor comprises the steps of: (a) inputting each image acquired by a plurality of sensors into a neural network model pre-set for each sensor to acquire individual domain features of each sensor; (b) inputting each image acquired by the plurality of sensors into a single neural network pre-set to acquire common domain features of each sensor; (c) combining the individual domain features and the common domain features to compute sensor-specific domain features; (d) inputting the domain features into an environment variable inference model to acquire an environment variable vector, which is a multidimensional vector, for each sensor; (f) applying a pre-set importance to each vector element value of the sensor-specific environment variable vector to compute a reliability score for each sensor and compute a fusion weight for each sensor based on the reliability score; and (g) inputting the fusion weight of each sensor and the domain features of each sensor into a transformer model to acquire fusion features. A multi-domain-based object tracking and detection device that performs the step (h) of inputting the above-mentioned fusion features into an object tracking / detection model to perform object tracking and detection. Claim 9 In claim 8, the above-mentioned transformer model is a multi-domain-based object tracking and detection device that performs cross-attention operations by setting the highest fusion weight during cross-attention operations as the domain feature of the sensor as Q (query) and the combination of the remaining sensors as a K (Key) - V (Value) combination. Claim 10 In claim 9, the transformer model outputs fusion features through the weighted summation of a self-attention vector formed by a self-attention operation of the domain features for each sensor and a cross-attention vector formed by a cross-attention operation of the domain features for each sensor. This is a multi-domain-based object tracking and detection device. Claim 11 In claim 8, the single neural network of step (b) is a multi-domain-based object tracking and detection device including a vision transformer. Claim 12 A towing cable according to claim 11, wherein a plurality of first-1 dummy members made of a plurality of plastic material are disposed in a plurality of inter-conductor spacing spaces between the second sheath and the third sheath. Claim 13 In Clause 12, a multi-domain-based object tracking and detection device in which the reliability score of each sensor is calculated as follows: In the above mathematical formula is the confidence score of each sensor, i is the index representing the sensor, env is the index representing each dimension element of the environment variable vector, and K represents the vector element value constituting the environment variable vector. Claim 14 In Clause 13, the fusion weight is a multi-domain-based object tracking and detection device calculated as follows: In the above mathematical formula, W is the fusion weight.