Calculation device
Through a multimodal model based on contrast learning, combined with 1-to-many pairing training and the merging processing of face feature acquisition models, the problem of low accuracy of multimodal data processing in on-board systems is solved, and high-accuracy data processing in various application scenarios is achieved.
Patent Information
- Application Number
- CN202510174748.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
In applications such as driving behavior monitoring and in-vehicle data query, it is difficult to achieve high-accuracy multimodal data processing.
The multimodal model based on contrast learning is adopted, and the parameters of the image encoder and text encoder of the CLIP model are trained through a 1-to-many pairing method, and the output of the face feature acquisition model is merged into the model to improve the accuracy of the model.
It realizes multimodal data processing with high accuracy in various application scenarios, especially driving behavior monitoring and in-vehicle data query.
Smart Images

Figure CN120106157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to applications of artificial intelligence (AI), and more particularly to a computing device for training a multimodal model based on contrastive learning, and a computing device for performing inference using the trained multimodal model based on contrastive learning. Background Art
[0002] The in-vehicle system can be regarded as a computer system on the vehicle. With the development of vehicle networks, many new applications are constantly added to the in-vehicle system, such as entertainment applications, navigation applications, communication applications, etc. In addition, as drivers have higher and higher requirements for driving safety, the current in-vehicle system will also include an Advanced Driver Assistance System (ADAD), which can continuously detect driving information inside and outside the vehicle for the driver and issue warnings or prompts in real time. For example, the in-vehicle monitoring camera can provide real-time images of driving. Therefore, the in-vehicle system can detect dangerous driving behaviors (such as talking on the phone, fatigue driving, etc.) through the image output of the in-vehicle monitoring camera. In recent years, driven by deep learning, artificial intelligence has achieved excellent performance in many fields. How to apply artificial intelligence technology to in-vehicle systems (especially advanced driver assistance systems) to significantly improve the functions and performance of existing in-vehicle systems has become an important issue. Summary of the invention
[0003] Therefore, one of the objectives of the present invention is to provide a computing device for training a contrastive learning-based multimodal model and a computing device for performing inference using the trained contrastive learning-based multimodal model.
[0004] In one embodiment of the present invention, a computing device is disclosed. The computing device includes a storage device and a processor. The storage device is used to store a program code, wherein the program code includes a multimodal model based on contrastive learning. The processor is used to load and execute the program code, wherein the multimodal model based on contrastive learning performs the following inference operations: obtaining a plurality of first vectors corresponding to a plurality of first type data respectively; receiving an input data, wherein the input data includes a second type data; generating a second vector corresponding to the input data; generating a plurality of output values based on the plurality of first vectors and the second vector; and determining, based on the plurality of output values, that a plurality of first type data in the plurality of first type data are simultaneously matched to the second type data.
[0005] In one embodiment of the present invention, a computing device is disclosed. The computing device includes a storage device and a processor. The storage device is used to store a program code, wherein the program code includes a multimodal model based on contrastive learning. The processor is used to load and execute the program code, wherein the multimodal model based on contrastive learning performs the following inference operations: obtaining a plurality of first vectors corresponding to a plurality of first type data respectively; receiving an input data, wherein the input data includes a plurality of second type data; generating a second vector corresponding to the input data; generating a plurality of output values based on the plurality of first vectors and the second vector; and determining, based on the plurality of output values, whether a first type data among the plurality of first type data is matched to the input data.
[0006] In one embodiment of the present invention, a computing device is disclosed. The computing device includes a storage device and a processor. The storage device is used to store a program code, wherein the program code includes a multimodal model based on contrastive learning, and the multimodal model based on contrastive learning includes a first type data encoder and a second type data encoder. The processor is used to load and execute the program code, wherein the multimodal model based on contrastive learning performs the following training operations: receiving a plurality of first type training data and a plurality of second type training data; inputting the plurality of first type training data to the first type data encoder; inputting the plurality of second type training data to the second type data encoder; training the parameters of the first and second type data encoders at least according to the outputs of the first and second type data encoders, so that the output of the multimodal model based on contrastive learning indicates that a first type training data among the plurality of first type training data is matched to a plurality of second type training data among the plurality of second type training data.
[0007] The contrastive learning-based multimodal model disclosed in the present invention uses a one-to-many pairing to train the parameters of the image encoder and the text encoder of the CLIP model, so the trained model can be applied to a variety of application scenarios (e.g., driving behavior monitoring, in-vehicle data query, etc.). In addition, when the contrastive learning-based multimodal model includes a pre-trained facial feature extraction model, the training operation can merge the output of the facial feature extraction model and the output of the image encoder, and train the parameters of the image encoder and the text encoder of the CLIP model based on the merged result, so that the trained model can have a higher accuracy when performing inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 Schematic diagram of an artificial intelligence system according to an embodiment of the present invention.
[0009] Figure 2It is a schematic diagram of the training operation of the multimodal model based on contrastive learning disclosed in the present invention.
[0010] Figure 3 It is a schematic diagram of merging the output of the image encoder and the output of the facial feature extraction model.
[0011] Figure 4 It is a schematic diagram of the first application of performing inference operations using the multimodal model based on contrastive learning disclosed in the present invention.
[0012] Figure 5 is a schematic diagram of a second application of performing inference operations using the multimodal model based on contrastive learning disclosed in the present invention.
[0013] The reference numerals are as follows: 100: artificial intelligence system; 102, 104: computing device; 112: training data source; 114, 124: processor; 116, 126: storage device; 122: input data source; 202: text encoder; 204: image encoder; 206: facial feature extraction model; D_L: training data; D_IN: input data; PROG: program code; MD_CL: multimodal model based on contrastive learning; MD_CLIP: contrastive language-image pre-training model; MD_FFE: pre-trained facial feature extraction model; TXT_1, TXT_2, TXT_N, TXT IN _1. TXT IN _2: Text; IMG_1, IMG_2, IMG_N, IMG_i, IMG_IN: Image; I 1 -I N , I_2, I i : image vector; I_1: facial feature vector; I_3: merge vector; T 1 -T N : text vector. DETAILED DESCRIPTION
[0014] Figure 1 1 is a schematic diagram of an artificial intelligence system 100 according to an embodiment of the present invention. The artificial intelligence system 100 includes a plurality of computing devices 102 and 104, wherein the computing device 102 is responsible for training the artificial intelligence model according to a large amount of labeled data in the training phase, and after training to a certain extent, deploying the trained artificial intelligence model to a hardware platform (e.g., computing device 104), and therefore, the computing device 104 is responsible for processing a large amount of unlabeled data according to the trained artificial intelligence model in the inference phase.
[0015] like Figure 1As shown, the computing device 102 includes a processor 114 and a storage device 116. The storage device 116 is used to store a program code PROG. The program code PROG includes a contrastive learning-based multimodal model MD_CL. In one embodiment, the contrastive learning-based multimodal model MD_CL adopts a contrastive language-image pre-training (Contrastive Language-Image Pre-training, hereinafter referred to as "CLIP") model MD_CLIP. However, the present invention is not limited thereto. In another embodiment, in addition to the contrastive language-image pre-training model MD_CLIP, the contrastive learning-based multimodal model MD_CL may also include a pre-trained facial feature extractor model MD_FFE. The contrastive learning-based multimodal model MD_CL may include encoders for different types of data (for example, an image encoder and a text encoder of the CLIP model). The processor 114 is used to load and execute the program code PROG to perform a training operation of the multimodal model MD_CL based on contrastive learning. In addition, the training data source 112 is used to provide training data (eg, labeled training data) D_L required for the training operation.
[0016] In this embodiment, the training operation of the contrastive learning-based multimodal model MD_CL may include: receiving a plurality of first-type training data (e.g., image data) and a plurality of second-type training data (e.g., text data) from a training data source 112, inputting the plurality of first-type training data (e.g., image data) to a first-type data encoder (e.g., an image encoder of the CLIP model MD_CLIP), and inputting the plurality of second-type training data (e.g., text data) to a second-type data encoder (e.g., a text encoder of the CLIP model MD_CLIP); and training the parameters of the first and second-type data encoders (e.g., the image encoder and the text encoder of the CLIP model MD_CLIP) at least based on the outputs of the first and second-type data encoders (e.g., the image encoder and the text encoder of the CLIP model MD_CLIP), so that the output (e.g., a plurality of probability values) of the contrastive learning-based multimodal model MD_CL indicates that a first-type training data among the plurality of first-type training data can be simultaneously paired with a plurality of second-type training data among the plurality of second-type training data. In other words, the contrastive learning-based multimodal model MD_CL disclosed in the present invention uses a 1-to-N pairing to train the parameters of the first and second type data encoders (e.g., the image encoder and text encoder of the CLIP model MD_CLIP). In addition, when the contrastive learning-based multimodal model MD_CL includes a pre-trained facial feature extraction model MD_FFE, the training operation can merge (concatenate) the output of the facial feature extraction model MD_FFE and the output of the first type data encoder (e.g., the image encoder of the CLIP model MD_CLIP), and train the parameters of the first and second type data encoders (e.g., the image encoder and text encoder of the CLIP model MD_CLIP) according to the concatenation result.
[0017] See also Figure 2 , Figure 2 A schematic diagram of the training operation of a multimodal model based on contrastive learning disclosed in the present invention. Figure 2 The multimodal model based on contrastive learning shown can be constructed by Figure 1 The program code PROG shown is implemented, so Figure 2The multimodal model based on contrastive learning shown includes a text encoder 202, an image encoder 204, and a facial feature extraction model 206, wherein the facial feature extraction model 206 is an optional component. In other words, for some applications, the additional use of the facial feature extraction model 206 can improve the accuracy of inference, while for other applications, the facial feature extraction model 206 can be omitted. In short, as long as the parameters of the image encoder 204 and the text encoder 202 of the CLIP model MD_CLIP are trained using a one-to-many pairing, they all fall within the scope of the present invention.
[0018] The training data D_L (which has labeled training data) provided by the training data source 112 may include a plurality of different texts TXT_1, TXT_2, ..., TXT_N and a plurality of different images IMG_1, IMG_2, ..., IMG_N. The text encoder 202 converts the plurality of different texts TXT_1-TXT_N into a plurality of text embedding vectors T 1 ~T N The image encoder 204 converts the multiple different images IMG_1 to IMG_N into multiple image embedding vectors. 1 ~I N will be obtained from the output of the image encoder 204. In one example, the multimodal model MD_CL based on contrastive learning omits the pre-trained facial feature extraction model MD_FFE (e.g., facial feature extraction model 206), and the image encoder 204 converts the multiple different images IMG_1-IMG_N into multiple image vectors I 1 ~I N In another example, the contrastive learning-based multimodal model MD_CL additionally includes a pre-trained facial feature extraction model MD_FFE (eg, facial feature extraction model 206), and each image vector output by the image encoder 204 is further merged to generate a plurality of image vectors I 1 ~I N The corresponding image vector I in i (i={1,2,…,N}).
[0019] Figure 3It is a schematic diagram of merging the output of the image encoder and the output of the facial feature extraction model. The image encoder 204 converts an image IMG_i (i={1,2,…,N}) from a plurality of different images IMG_1~IMG_N into an image vector I_2. In addition, the pre-trained facial feature extraction model 206 performs facial feature extraction on the same image IMG_i to generate a facial feature vector I_1. In this embodiment, the facial feature vector I_1 is a high-dimensional facial feature, and its content may include features such as facial contour, facial features, color, expression, etc. learned by the model. Basically, the image encoder 204 is for the global features of the image IMG_i, while the pre-trained facial feature extraction model 206 is for the local features of the image IMG_i. Therefore, the present invention can combine global features and local features to train the contrastive learning based multimodal model MD_CL (especially, the image encoder 204 and the text encoder 202 of the CLIP model MD_CLIP) to improve the accuracy of the contrastive learning based multimodal model MD_CL. Figure 3 As shown, the image vector I_2 and the facial feature vector I_1 are combined to generate a combined vector I_3, and the image vector I_2 is obtained according to the combined vector I_3. i In this embodiment, the merged vector I_3 can be filtered by a multilayer perception (MLP) 302 to generate the final image vector I_3 by filtering the important features of the merged vector I_3. i However, this is only an example and not a limitation of the present invention. In other embodiments, the multilayer perceptron 302 may be omitted and the merged vector I_3 may be directly used as the final image vector I i .
[0020] The training operation of the multimodal model MF_CL based on contrastive learning is based on multiple text vectors T 1 ~T N And multiple image vectors I 1 ~I N To train the parameters of the image encoder 204 and the text encoder 202. This embodiment uses a one-to-many pairing to train the parameters of the image encoder 204 and the text encoder 202 of the CLIP model MD_CLIP, so that each set of paired single image vectors and multiple text vectors have a high similarity (e.g., cosine similarity) and any pair of image vectors and text vectors that are not paired with each other have a low similarity (e.g., cosine similarity). For example, the loss function provides feedback based on the result of the cosine similarity operation, so that the training operation continuously optimizes the parameters of the image encoder 204 and the text encoder 202. Figure 2 As shown, after the parameters of the image encoder 204 and the text encoder 202 are optimized, the image vector I 1 With multiple text vectors T 1 、T 3 Each inner product of will have a high similarity (that is, the image vector I 1 Paired to multiple text vectors T 1 、T 3 ), and the image vector I 1 The inner products of the other text vectors will have very low similarity. Similarly, the image vector I 2 With multiple text vectors T 2 、T 3 Each inner product of will have a high similarity (that is, the image vector I 2 Paired to multiple text vectors T 2 、T 3 ), and the image vector I 2 The inner products of the other text vectors will have very low similarity; the image vector I 3 With multiple text vectors T 1 、T 3 Each inner product of will have a high similarity (that is, the image vector I 3 Paired to multiple text vectors T 1 、T 3 ), and the image vector I 3 The inner products with the other text vectors will have very low similarity; and the image vector I N With multiple text vectors T 1 、T 3 、T N The inner products of the image vectors have a high similarity (that is, the image vectors I N Paired to multiple text vectors T 1 、T 3 、T N ), and the image vector I N Each inner product with the remaining text vectors will have a very low similarity.
[0021] After the contrastive learning-based multimodal model MD_CL (especially, the CLIP model MD_CLIP adopted by the contrastive learning-based multimodal model MD_CL) is trained to a certain extent, the trained artificial intelligence model can be deployed to a hardware platform for inference operations. Figure 1As shown, the trained artificial intelligence model obtained by the computing device 102 in the training phase will be deployed to another computing device 104, and the computing device 104 will subsequently process unknown data according to the trained artificial intelligence model in the inference phase to generate meaningful predictions or decisions. For example, the computing device 104 can be a computing device that executes an on-board system, which is used to perform driving behavior monitoring, in-vehicle data query and other functions through the trained artificial intelligence model. However, this is only an example and not a limitation of the present invention. In fact, any device that uses the trained artificial intelligence model obtained by the computing device 102 to perform inference operations falls within the scope of the present invention.
[0022] like Figure 1 As shown, the computing device 104 includes a processor 124 and a storage device 126. The storage device 126 is used to store program code PROG, wherein the program code PROG includes the trained artificial intelligence model obtained by the computing device 102 (that is, the trained multimodal model MD_CL based on contrastive learning, which includes the CLIP model MD_CLIP obtained at the end of the training phase (or, including the CLIP model MD_CLIP obtained at the end of the training phase and the facial feature extraction model MD_FFE that has been trained before the start of the training phase)). The processor 124 is used to load and execute the program code PROG to perform the inference operation of the multimodal model MD_CL based on contrastive learning. In addition, the input data source 122 is used to provide the input data D_IN to be processed by the inference operation.
[0023] In one embodiment, the inference operation of the contrastive learning based multimodal model MD_CL is applied to driving behavior monitoring, and therefore, the input data source 122 may be an in-vehicle device, and an input data of the contrastive learning based multimodal model MD_CL includes a captured image of the in-vehicle monitoring device (e.g., an in-vehicle monitoring camera). The inference operation of the contrastive learning-based multimodal model MD_CL may include: obtaining a plurality of first vectors (e.g., text vectors) respectively corresponding to a plurality of first type data (e.g., predetermined texts); receiving an input data (the input data includes a second type data (e.g., an image from an in-vehicle monitoring device (e.g., an in-vehicle monitoring camera)); generating a second vector (e.g., an image vector) corresponding to the input data (e.g., image data); generating a plurality of output values (e.g., probability values) based on the plurality of first vectors and the second vector; and determining that a plurality of first type data (e.g., a plurality of texts) among the plurality of first type data are simultaneously paired to the second type data (e.g., a single image). In addition, when the contrastive learning-based multimodal model MD_CL includes a pre-trained facial feature extraction model MD_FFE, the inference operation may merge the output of the facial feature extraction model MD_FFE and the output of an encoder (e.g., an image encoder of the CLIP model MD_CLIP), and generate the second vector according to the merged result.
[0024] See also Figure 4 , Figure 4 It is a schematic diagram of the first application of performing inference operations using the multimodal model based on contrastive learning disclosed in the present invention. Figure 4 The contrastive learning-based multimodal model shown can be implemented by the trained contrastive learning-based multimodal model MD_CL, so, Figure 4 The multimodal model based on contrastive learning shown may include a text encoder 202, an image encoder 204, and a facial feature extraction model 206, wherein the facial feature extraction model 206 is an optional element. In other words, for some applications, the additional use of the facial feature extraction model 206 can improve the accuracy of inference, while for other applications, the facial feature extraction model 206 can be omitted.
[0025] In this embodiment, the input data D_L provided by the input data source 122 includes an input image IMG_IN (eg, a captured image from an in-vehicle monitoring device (eg, an in-vehicle monitoring camera)). The inference operation obtains the corresponding text vectors T of a plurality of different texts TXT_1 -TXT_N. 1 ~T N (For example, the corresponding text vectors T of different texts TXT_1 to TXT_N used in the training phase 1 ~T NHowever, the present invention is not limited thereto, and a plurality of different texts TXT_1-TXT_N can also be input by the user and converted into corresponding text vectors T by the trained text encoder 204. 1 ~T N In addition, the trained image encoder 204 converts the input image IMG_IN into an image vector, wherein the image vector (eg, I 1 ) is obtained from the output of the image encoder 204. In one example, the multimodal model MD_CL based on contrastive learning omits the facial feature extraction model MD_FFE (e.g., the facial feature extraction model 206), and the image encoder 204 converts the input image IMG_IN into an image vector (e.g., I 1 In another example, the contrastive learning-based multimodal model MD_CL additionally includes a facial feature extraction model MD_FFE (e.g., facial feature extraction model 206), and the image vector generated by the image encoder 204 according to the conversion of the input image IMG_IN is further merged (e.g., Figure 3 The merging process shown in the figure) is used to generate the image vector (eg, I 1 For example, the image encoder 204 converts the input image IMG_IN into an image vector I_2, and the facial feature extraction model 204 extracts facial features from the input image IMG_IN to generate a facial feature vector I_1. Finally, the image vector I_2 and the facial feature vector I_1 are combined to generate a combined vector I_3, and the image vector (e.g., I_2) to be used for similarity determination is obtained based on the combined vector I_3. 1 In one embodiment, the merged vector I_3 can be filtered by the multi-layer perceptron 302 to obtain important features of the merged vector I_3 to generate an image vector (eg, I_3) to be subsequently subjected to similarity determination. 1 However, this is only an example and not a limitation of the present invention. In other embodiments, the multilayer perceptron 302 may be omitted and the merged vector I_3 may be directly used as the image vector (e.g., I_4) to be used for subsequent similarity determination. 1 ).
[0026] Since the image vector I 1 Can be paired to multiple text vectors T 1 、T 3 , so the image vector I 1 With text vector T 1 The inner product of and image vector I 1 With text vector T 3 The inner product of will get a high similarity (for example, cosine similarity), and the image vector I1 With other text vectors T 2 , T 5 ~T N The inner product of will result in a very low similarity (e.g., cosine similarity). Thus, among the multiple output values (e.g., probability values) generated by the multimodal model MD_CL based on contrastive learning, the image vector I 1 With text vector T 1 Pairing and image vector I 1 With text vector T 3 The pairing will have a larger output value. By sorting the multiple output values generated by the contrastive learning-based multimodal model MD_CL to obtain the largest K (for example, K=2) output values, the computing device 104 executing the vehicle-mounted system can determine whether the input image IMG_IN provided by the in-vehicle monitoring camera contains captured image content of dangerous driving behavior (for example, yawning, drinking, talking on the phone, etc.) according to the K texts corresponding to the K (for example, K=2) output values, and issue a warning when a dangerous driving behavior is detected.
[0027] In another embodiment, the inference operation of the contrastive learning-based multimodal model MD_CL is applied to in-vehicle record query, so the input data source 122 can be an in-vehicle device, for example, the input data source 122 includes multiple images provided by the in-vehicle surveillance camera and text input provided by the user interface, and one of the input data for in-vehicle record query based on the contrastive learning-based multimodal model MD_CL is user input (for example, multiple keywords entered by the user). The inference operation of the contrastive learning-based multimodal model MD_CL may include: obtaining a plurality of first vectors (e.g., image vectors) respectively corresponding to a plurality of first type data (e.g., a plurality of images provided by an in-vehicle surveillance camera); receiving an input data (the input data includes a plurality of second type data (e.g., a plurality of texts in user input); generating a second vector (e.g., text vector) corresponding to the input data (e.g., text data); generating a plurality of output values (e.g., probability values) based on the plurality of first vectors and the second vector; and determining that a first type data among the plurality of first type data is matched to the input data (which includes a plurality of second type data). In addition, when the contrastive learning-based multimodal model MD_CL includes a facial feature extraction model MD_FFE, the inference operation may merge the output of the facial feature extraction model MD_FFE and the output of an encoder (e.g., an image encoder of a CLIP model MD_CLIP), and generate each first vector among the plurality of first vectors according to the merged result.
[0028] See also Figure 5 , Figure 5is a schematic diagram of a second application of performing inference operations using the multimodal model based on contrastive learning disclosed in the present invention. Figure 5 The contrastive learning-based multimodal model shown can be implemented by the trained contrastive learning-based multimodal model MD_CL, so, Figure 5 The multimodal model based on contrastive learning shown may include a text encoder 202, an image encoder 204, and a facial feature extraction model 206, wherein the facial feature extraction model 206 is an optional element. In other words, for some applications, the additional use of the facial feature extraction model 206 can improve the accuracy of inference, while for other applications, the facial feature extraction model 206 can be omitted.
[0029] In this embodiment, the input data D_L provided by the input data source 122 includes a plurality of texts "TXT IN _1","TXT IN _2" (for example, multiple keywords from the in-vehicle user interface). The inference operation obtains the corresponding image vectors I of multiple different images IMG_1~IMG_N (for example, multiple images actually recorded by the in-vehicle monitoring camera) 1 ~I N In one example, the multimodal model MD_CL based on contrastive learning omits the facial feature extraction model MD_FFE (eg, the facial feature extraction model 206), and the trained image encoder 204 directly converts the multiple different images IMG_1-IMG_N into the image vector I 1 ~I N In another example, the contrastive learning-based multimodal model MD_CL additionally includes a facial feature extraction model MD_FFE (e.g., facial feature extraction model 206). The image vectors generated by the trained image encoder 204 according to the conversion of the multiple different images IMG_1 to IMG_N are further merged (e.g., Figure 3 The merging process shown in FIG. 1 is used to generate the image vector I to be used for subsequent similarity judgment. 1 ~I N For example, the trained image encoder 204 converts each input image IMG_IN of the plurality of different images IMG_1 to IMG_N into an image vector I_2, and the facial feature extraction model 204 extracts facial features from the input image IMG_IN to generate a facial feature vector I_1. Finally, the image vector I_2 and the facial feature vector I_1 are merged to generate a merged vector I_3, and the image vector I_2 to be used for subsequent similarity determination is obtained based on the merged vector I_3. 1 ~I NIn one embodiment, the merged vector I_3 is filtered by the multi-layer perceptron 302 to generate an image vector for subsequent similarity determination. However, this is only an example and not a limitation of the present invention. In other embodiments, the multi-layer perceptron 302 may be omitted and the merged vector I_3 may be directly used as the image vector for subsequent similarity determination.
[0030] The text encoder 202 converts the input text "TXT IN _1+TXT IN _2" into a text vector (e.g., T 1 ). Since the text "TXT IN _1","TXT IN _2" may be paired to multiple images, so if image IMG_1 can be paired to multiple texts "TXT IN _1","TXT IN _2", then the text vector (for example, T 1 ) and the image vector I 1 The inner product of will result in a high similarity (e.g., cosine similarity), while the text vector (e.g., T 1 ) and other image vectors I 2 ~I N The inner product of will result in a very low similarity (e.g., cosine similarity). Thus, among the multiple output values (e.g., probability values) generated by the multimodal model MD_CL based on contrastive learning, the image vector I 1 With the text vector (e.g., T 1 ) will have the largest output value, so the computing device 104 executing the vehicle-mounted system can use the image corresponding to the maximum output value (for example, IMG_1) as the query result.
[0031] The above description is only a preferred embodiment of the present invention, but is not intended to limit the scope of the present invention. Any technical personnel in this field may make further improvements and changes on this basis without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope defined by the claims of this application.
Claims
1. A computing device, characterized in that: include: A storage device for storing program code, wherein the program code includes a multimodal model based on contrastive learning; as well as A processor is used to load and execute the program code, wherein the multimodal model based on contrastive learning performs the following inference operations: Acquire a plurality of first vectors respectively corresponding to a plurality of first type data; receiving input data, wherein the input data includes second type data; generating a second vector corresponding to the input data; generating a plurality of output values according to the plurality of first vectors and the second vector; as well as According to the plurality of output values, it is determined that a plurality of first type data among the plurality of first type data are simultaneously matched to the second type data.
2. The computing device according to claim 1, wherein: The multimodal model based on contrastive learning adopts a contrastive language-image pre-training model, the multiple first-type data are respectively multiple different texts, and the second-type data are images.
3. The computing device according to claim 2, wherein: The multimodal model based on contrastive learning includes: an image encoder for converting the image into an image vector; and A facial feature extraction model is used to extract facial features from the image to generate a facial feature vector; The operation of generating the second vector corresponding to the input data includes: Merging the image vector and the facial feature vector to generate a merged vector; and The second vector is generated according to the combined vector.
4. The computing device according to claim 2, wherein: The input data is a captured image of an in-vehicle monitoring device.
5. A computing device, characterized in that: include: A storage device for storing program code, wherein the program code includes a multimodal model based on contrastive learning; as well as A processor is used to load and execute the program code, wherein the multimodal model based on contrastive learning performs the following inference operations: Acquire a plurality of first vectors respectively corresponding to a plurality of first type data; receiving input data, wherein the input data comprises a plurality of second type data; generating a second vector corresponding to the input data; generating a plurality of output values according to the plurality of first vectors and the second vector; as well as According to the plurality of output values, it is determined whether the first type of data among the plurality of first type of data is matched to the input data.
6. The computing device according to claim 5, wherein: The contrastive learning-based multimodal model adopts a contrastive language-image pre-training model, the plurality of first-type data are respectively a plurality of different images, and the plurality of second-type data are respectively a plurality of different texts.
7. The computing device according to claim 6, wherein: The multimodal model based on contrastive learning includes: an image encoder for converting each of the plurality of different images into an image vector; and A facial feature extraction model, used for extracting facial features from each of the plurality of different images to generate a facial feature vector; The step of generating a plurality of first vectors respectively corresponding to a plurality of first type data comprises: Merging the image vector and the facial feature vector to generate a merged vector; and A first vector among the plurality of first vectors is generated according to the combined vector.
8. The computing device according to claim 6, wherein: The plurality of different images are captured images of an in-vehicle monitoring device, and the input data is input by a user.
9. A computing device, characterized in that: include: A storage device for storing program code, wherein the program code includes a multimodal model based on contrastive learning, and the multimodal model based on contrastive learning includes a first type of data encoder and a second type of data encoder; and A processor is used to load and execute the program code, wherein the multimodal model based on contrastive learning performs the following training operations: receiving a plurality of first type training data and a plurality of second type training data; Inputting the plurality of first type training data into the first type data encoder; Inputting the plurality of second type training data into the second type data encoder; The parameters of the first type data encoder and the second type data encoder are trained at least based on the outputs of the first type data encoder and the second type data encoder, so that the output of the multimodal model based on contrastive learning indicates that the first type training data among the multiple first type training data are paired with the multiple second type training data among the multiple second type training data.
10. The computing device according to claim 9, wherein: The multimodal model based on contrastive learning adopts a contrastive language-image pre-training model, the first type data encoder is an image encoder, the second type data encoder is a text encoder, the multiple first type training data are respectively multiple different images, and the multiple second type training data are respectively multiple different texts.
11. The computing device according to claim 10, wherein: The multimodal model based on contrastive learning also includes: A pre-trained facial feature extraction model, used to extract facial features from the plurality of different images; The contrastive learning-based multimodal model trains the parameters of the first type data encoder and the second type data encoder according to the combined result of the output of the first type data encoder, the output of the pre-trained facial feature extraction model, and the output of the second type data encoder.