Model training method, electronic equipment and computer readable storage medium
By conducting multi-modal search model training, using image and text annotation data to generate learning prompts, the adaptability and performance problems of the multi-modal search model in traffic scenarios are solved, and efficient graphic and text retrieval is achieved.
Patent Information
- Application Number
- CN202510813458.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multimodal retrieval model cannot adapt to variability and large-scale data needs, and has poor performance, especially in traffic scenarios with poor graphics and text retrieval.
By obtaining the training data set, including image annotation data and description text annotation data of multiple types of objects of interest, the training data set is used to train the initial multimodal search model, generate learning prompts, and perform image encoding training and text-image alignment training to generate the target multimodal search model.
It improves the performance of multimodal models in specific scenarios, enhances the model's adaptability to complex environments, improves the accuracy and retrieval efficiency of graphics and text matching, especially in traffic scenarios, which can quickly process large amounts of data.
Smart Images

Figure CN120336859A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of data processing and large model, and in particular, to a model training method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Multimodal retrieval technology is a retrieval method that integrates multiple information modalities (such as images, texts, audios, videos, etc.), allowing users to use one modality (such as text) to retrieve information of another or multiple modalities (such as images or videos). When dealing with complex and multi-information-source data, manual screening is inefficient and error-prone, while multimodal retrieval technology can be used for efficient target positioning.
[0003] Currently, traditional multimodal retrieval models are limited by manually annotated training data and fixed description patterns, and it is difficult to adapt to variability and large-scale data requirements. In addition, although traditional multimodal retrieval models can handle Chinese image-text matching, they are not optimized for specific scenarios (such as traffic scenarios). Although some multimodal retrieval models can be applied to specific scenarios, there are still deficiencies in image-to-text and text-to-image search performance.
[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a model training method, an electronic device, and a computer-readable storage medium to at least solve the technical problems in the related art that multimodal retrieval models cannot adapt to variability and large-scale data requirements, and the model performance is poor.
[0006] According to one aspect of the embodiments of this application, a model training method is provided, including: obtaining a training data set, where the training data set includes: image annotation data of multiple types of attention objects and description text annotation data associated with the image annotation data; training an initial multimodal retrieval model using the training data set to generate learnable prompts; performing image encoding training on the initial multimodal retrieval model using the training data set and the learnable prompts to generate an intermediate multimodal retrieval model; performing text-image alignment training on the intermediate multimodal retrieval model using the training data set to generate a target multimodal retrieval model, where the target multimodal retrieval model is used to perform multimodal retrieval on a target retrieval request to obtain a target retrieval result.
[0007] According to another aspect of the embodiments of this application, a data processing method is further provided, including: obtaining a target retrieval request; performing multimodal retrieval on the target retrieval request using the target multimodal retrieval model to obtain a target retrieval result; where the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0008] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining a target retrieval request, where the target retrieval request is used to query a specified participating object in an urban traffic scenario; performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model to obtain a specified participating object retrieval result; where the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0009] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining a data processing request through a first application programming interface, where the request data carried in the data processing request includes: a target retrieval request; returning a data processing response through a second application programming interface, where the response data carried in the data processing response includes: a target retrieval result, the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0010] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining a current input data processing dialogue request, where the request data carried in the data processing dialogue request includes: a target retrieval request; in response to the data processing dialogue request, returning a data processing dialogue reply, where the information carried in the data processing dialogue reply includes: a target retrieval result, the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above; displaying the target retrieval result in a graphical user interface.
[0011] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: in response to an input instruction acting on an operation interface, displaying a target retrieval request on the operation interface; in response to a processing instruction acting on the operation interface, displaying a target retrieval result on the operation interface; where the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0012] According to another aspect of the embodiments of the present application, there is also provided a data processing system, including: a client for sending a target retrieval request; a server connected to the client for performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model to obtain a target retrieval result; the client is further used for outputting the target retrieval result; where the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0013] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory storing an executable program; a processor connected to the memory through a bus, and configured to run the program, wherein when the program runs, it executes any one of the above-mentioned model training methods or data processing methods.
[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute any one of the above-mentioned model training methods or data processing methods.
[0015] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, which when executed by a processor, implements any one of the above-mentioned model training methods or data processing methods.
[0016] In the embodiments of the present application, by obtaining a training data set, wherein the training data set includes: image annotation data of multiple types of attention objects and description text annotation data associated with the image annotation data, then using the training data set to train an initial multi-modal retrieval model to generate learnable prompts, and then using the training data set and the learnable prompts to perform image encoding training on the initial multi-modal retrieval model to generate an intermediate multi-modal retrieval model, and finally using the training data set to perform text-image alignment training on the intermediate multi-modal retrieval model to generate a target multi-modal retrieval model, wherein the target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result, thereby achieving the purpose of improving the performance of the multi-modal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multi-modal model to complex environments and improving the performance of the multi-modal model in an open scenario, and further solving the technical problem that the multi-modal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance.
[0017] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0019] Figure 1 is a schematic diagram of an application scenario of a model training method according to an embodiment of the present application;
[0020] Figure 2It is a flowchart of a model training method according to an embodiment of the present application;
[0021] Figure 3 It is an overall flowchart according to an embodiment of the present application;
[0022] Figure 4 It is an example diagram of data augmentation according to an embodiment of the present application;
[0023] Figure 5 It is a flowchart of a data processing method according to an embodiment of the present application;
[0024] Figure 6 It is a flowchart of a data processing method according to an embodiment of the present application;
[0025] Figure 7 It is a flowchart of a data processing method according to an embodiment of the present application;
[0026] Figure 8 It is a flowchart of a data processing method according to an embodiment of the present application;
[0027] Figure 9 It is a flowchart of a data processing method according to an embodiment of the present application;
[0028] Figure 10 It is a schematic structural diagram of a data processing system according to an embodiment of the present application;
[0029] Figure 11 It is a schematic structural diagram of a model training device according to an embodiment of the present application;
[0030] Figure 12 It is a schematic structural diagram of a data processing device according to an embodiment of the present application;
[0031] Figure 13 It is a schematic structural diagram of another data processing device according to an embodiment of the present application;
[0032] Figure 14 It is a schematic structural diagram of another data processing device according to an embodiment of the present application;
[0033] Figure 15 It is a schematic structural diagram of another data processing device according to an embodiment of the present application;
[0034] Figure 16 It is a schematic structural diagram of another data processing device according to an embodiment of the present application;
[0035] Figure 17 It is a structural block diagram of a computing device according to an embodiment of the present application;
[0036] Figure 18It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0037] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0038] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0039] The technical solution provided by the present application is mainly implemented by using large model technology. Here, the large model refers to a deep learning model with a large number of model parameters, which usually can contain hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. The large model can also be called a foundation model. Through large-scale pre-training of the large model with unlabeled corpus, a pre-trained model with more than hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0040] It should be noted that when the large model is actually applied, the pre-trained model can be fine-tuned through a small number of samples so that the large model can be applied to different tasks. For example, the large model can be widely used in natural language processing (NLP), computer vision, speech processing and other fields, and can be specifically applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, etc. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiment of the present application, the data processing of the target multimodal retrieval model obtained by training in the present application in a multimodal retrieval scenario is taken as an example for explanation.
[0041] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0042] Transformer: It is a deep learning model architecture, which mainly consists of two parts: encoder and decoder, both of which are built based on a multi-layer self-attention mechanism.
[0043] Vision Transformer (ViT): is a model that applies the Transformer architecture to computer vision tasks. It can be understood as an image encoder based on the Transformer architecture. In visual tasks, ViT effectively captures and processes image features by dividing the image into a series of patches and inputting them into the Transformer model as a sequence.
[0044] Triplet loss: A triplet loss function, the core idea is to learn a feature space that can distinguish different categories by comparing the distance between a positive sample (anchor), a positive sample (positive), and a negative sample (negative). During the training process, the model will try to make the distance between the anchor and the positive as small as possible, and the distance between the anchor and the negative as large as possible, so as to widen the gap between different categories and reduce the difference between samples of the same type. That is, the learning process will shorten the distance within the class and increase the distance between classes.
[0045] Soft ID Loss: In traditional ID Loss, the model usually uses hard labels (i.e., explicit 0 or 1 labels) to learn to distinguish samples of different identities, which works well when the differences between categories in the dataset are obvious. However, in some application scenarios, such as person re-identification, the appearances between different individuals may be very similar, and traditional hard-label methods may ignore these subtle differences, resulting in poor generalization ability of the model. Soft ID Loss alleviates this problem by introducing soft labels (i.e., label smoothing or labels in the form of probability distributions). Soft ID Loss represents the label of each identity as a probability distribution rather than a binary label. This probability distribution not only contains the estimation of the identity of the current sample but also allows the model to consider the possible ambiguity and similarity between identities. In Soft ID Loss, the goal of the model is to maximize the probability of the current identity while maintaining a relatively high probability for other similar identities, so as to learn the features of each individual more carefully during the training process, improve the discrimination between classes and within classes, and thus be more suitable for scenarios considering similarity.
[0046] Vision-and-Language Model (VLM): A multi-modal language large model, which can be understood as an enhanced multi-modal (image and natural language text) matching and generation hybrid model. It is a deep learning model that combines computer vision and natural language processing capabilities and is designed to handle multi-modal tasks, that is, tasks that involve both image and text information at the same time. Text-image retrieval: Retrieving pictures that match the description through text description and sorting them according to the degree of similarity.
[0047] Text-to-Image Retrieval: It refers to retrieving pictures that match a text description through a text description and sorting them according to the degree of similarity. When querying, the text-to-image retrieval system needs to understand the semantic information in the text and map it to the image feature space to find pictures that match the description. Exemplarily, in an intelligent transportation system, if there is a text description saying: "The target object is a man wearing a red hat and a blue jacket", the text-to-image retrieval system should be able to find all pedestrian pictures that match this description from a large number of surveillance pictures and sort them according to the degree of similarity.
[0048] Image-to-Image Retrieval: Given an image, it searches for other similar images in the database and sorts them according to the degree of similarity. When querying, the user submits an image as the query condition. The image retrieval system analyzes features such as the color, texture, shape, objects, and scenes of the image and compares them with the features of the images stored in the database to find similar images and sort them according to the degree of similarity. Exemplarily, in the scenario of pedestrian re-identification (ReID), when the user submits an image of a specific pedestrian, the image retrieval system can return all the images of the same pedestrian in the database, even if these images are different in terms of angle, lighting, occlusion, etc.
[0049] Recall@k: In the field of retrieval, recall@k is a metric that measures the proportion of relevant items correctly retrieved by the system among the top k retrieval results. Specifically for non-motor vehicle text search recall@14, it refers to the proportion of non-motor vehicle images that are truly relevant to the query text among the top 14 retrieval results returned by the system in the task of text searching for non-motor vehicle images.
[0050] In the current traffic scenario, there are ubiquitous monitoring devices. When some cases require finding a certain person or the course of the case, these monitoring data can be used as the data source for case analysis. However, in many cases, it is difficult to utilize the monitoring data. Firstly, there are many camera devices. For example, in a certain area, there are 5 million to 30 million people flowing through in a day, that is, 3,000 to 20,000 people per minute. Assuming that the target pedestrian passes a single camera in 10 seconds, even if we know the approximate time and detailed location of the case occurrence, at least 250 frames need to be checked for one camera, and each frame may have dozens of pedestrians. In this situation, it is very difficult to find a person or the course of the case manually.
[0051] In the industry, it is relatively common to calculate the attributes of a target with the help of traditional small models and rely on these attributes for positioning. However, the categories of traditional small models are relatively limited, so they are not very flexible when the demand increases and in the retrieval of such open scenarios. In addition, the link of traditional small models may be relatively long. They not only need to rely on detection but may also rely on some fine-grained classification models. In the monitoring scenario, due to the large number of camera device models, the generalization ability of small models has always been relatively weak. Suppose there is such a scenario where one needs to find a person among a large amount of monitoring data. The known information includes time and appearance description. The information obtained from the verbal description of the informant is only relatively rough. If one searches manually, one needs to manually search and compare in the monitoring according to the information described by the informant, and the workload is extremely large. Suppose the target picture is found, but often the face of the target is not clear and cannot be compared well with the identity information database, then this manual process has to be repeated, resulting in a significant reduction in the detection efficiency.
[0052] Therefore, the demand for image search and text search of multi-modal large models in the traffic system becomes very meaningful. Multi-modal large models have the characteristics of open semantics and strong end-to-end generalization performance, that is, they can support infinite categories for image search and text search, and are end-to-end with excellent generalization performance, capable of handling the diversity of monitoring.
[0053] Currently, a model based on the dual-encoder architecture of vision Transformer and language understanding Transformer has been proposed, and the model performance is enhanced through two-stage pre-training. However, this model is not for encoding pictures and texts in specific scenarios. For example, in the traffic scenario, its effect in retrieving pedestrians is poor. In addition, its ability to search for people by text is also relatively weak.
[0054] A model that can be applied to the work of pedestrian / vehicle re-identification has also been proposed. This model can achieve the image search ability of searching for people by people and searching for vehicles by vehicles. However, this model only considers the performance of image search, does not consider the performance of text search, and the text search effect is poor, and it also has no Chinese processing ability.
[0055] In addition, in the traffic industry, most of the image search and text search still stay in the traditional solution of using small models. Generally, the target is obtained through detection, and then the intra-class similarity is improved and the inter-class similarity is reduced through re-identification. In addition, some text search capabilities are introduced by enhancing the retrieval effect through multiple attributes. However, the traditional solution has the problem of difficult acquisition of training data. It is necessary to label instance identification (ID) and attribute tags, and it cannot implement text search in an open semantic scenario, and cannot handle the current large-scale traffic monitoring data. Most of them can only be applied to fixed closed scenarios.
[0056] It can be seen that the current multi-modal retrieval models cannot meet the needs of Chinese image-text retrieval in the current traffic scenarios, and there are few industry-wide image-text retrieval methods for specific targets in traffic data.
[0057] The multi-modal retrieval models in the related technologies have the following defects.
[0058] Defect 1: High data annotation cost. Existing multi-modal retrieval models usually rely on manually annotated training data, which is time-consuming, laborious, and costly, and cannot be applied to large-scale open scenarios. Moreover, due to the limitations of manual annotation, only a limited number of training data can be obtained, which is not sufficient for the model to learn rich enough feature representations, especially when dealing with fine-grained categories such as pedestrians and non-motor vehicles.
[0059] Defect 2: Weak model generalization ability. Traditional small-model-based solutions have insufficient generalization ability in open scenarios because small models may be optimized to handle predefined categories and cannot well adapt to unseen instances or background changes. In addition, even if traditional solutions perform well in multi-modal matching, they are not optimized for specific scenarios (such as traffic scenarios), so their performance in specific target retrieval (such as pedestrians and non-motor vehicles) may not be satisfactory.
[0060] Defect 3: Insufficient performance in image retrieval from text. Traditional models perform poorly in the task of image retrieval from text, especially when retrieving specific targets such as non-motor vehicles in open scenarios, and the recall rate may be very low.
[0061] No effective solutions have been proposed to address the above defects before this application.
[0062] According to the embodiments of the present application, a model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0063] Considering that the number of model parameters of large models is huge and the computing resources of mobile terminals are limited, the above method provided by the embodiments of the present application can be applied to the application scenarios as Figure 1 shown, but not limited thereto. In scenarios such as Figure 1In the application scenario shown, the large model is deployed in server 10. Server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, client devices 20 can include, but are not limited to: smartphones, tablets, laptops, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. Client devices 20 can interact with users through a graphical user interface to implement the invocation of the large model, thereby implementing the method provided in the embodiments of the present application.
[0064] In the embodiments of the present application, the system composed of the client device and the server can perform the following steps: The client device performs steps such as obtaining a training data set, where the training data set includes: image annotation data of various types of objects of interest and description text annotation data associated with the image annotation data, and sending the training data set to the server. The server performs steps such as training an initial multi-modal retrieval model using the training data set to generate learnable prompts, then performing image encoding training on the initial multi-modal retrieval model using the training data set and the learnable prompts to generate an intermediate multi-modal retrieval model, and then performing text-image alignment training on the intermediate multi-modal retrieval model using the training data set to generate a target multi-modal retrieval model. Here, the target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result, and returning the target multi-modal retrieval model to the client device. It should be noted that in the case where the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be performed in the client device.
[0065] It should be noted that with the rapid development of high-performance computing units, in other application scenarios, the above method provided in the embodiments of the present application can also be applied to a model all-in-one machine. In an optional embodiment, multiple models are built into the model all-in-one machine, and users can select one model for adjustment according to their needs to obtain their own model. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the adjusted model to perform the above method provided in the embodiments of the present application. In another optional embodiment, a trained model is built into the large model all-in-one machine. Thus, the high-performance computing unit built into the model all-in-one machine can directly call this model to perform the above method provided in the embodiments of the present application.
[0066] Furthermore, when a user needs to train their own model, they can also upload their own dataset through the client. This dataset is sent from the client to the server, enabling the server to adjust the pre-trained model with this dataset to obtain the user's own model, which is then deployed to the production environment. To facilitate the user's model adjustment requirements, the server can provide complete adjustment tools, development frameworks, and processes, and support multiple adjustment strategies, enabling the adjusted model to better adapt to different domain applications and achieve high customization.
[0067] Under the above operating environment, the present application provides a model training method as Figure 2 shown. Figure 2 It is a flowchart of a model training method according to an embodiment of the present application. As Figure 2 shown, the method may include the following steps:
[0068] Step S21, obtain a training dataset, where the training dataset includes: image annotation data of multiple types of objects of interest and descriptive text annotation data associated with the image annotation data;
[0069] Step S22, use the training dataset to train an initial multi-modal retrieval model to generate learnable prompts;
[0070] Step S23, use the training dataset and the learnable prompts to perform image encoding training on the initial multi-modal retrieval model to generate an intermediate multi-modal retrieval model;
[0071] Step S24, use the training dataset to perform text-image alignment training on the intermediate multi-modal retrieval model to generate a target multi-modal retrieval model, where the target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result.
[0072] In an embodiment of the present application, first, a training dataset is obtained. The training dataset is a data set used to train the target multi-modal retrieval model. The training dataset includes both image data and corresponding text description data, that is, the training dataset includes image annotation data of multiple types of objects of interest and descriptive text annotation data associated with the image annotation data.
[0073] Multiple types of objects of interest can be understood as any elements that are important in image recognition, object detection, or multi-modal information retrieval. For example, multiple types of objects of interest may include people of different genders, ages, dressing styles, and holding items (such as umbrellas, mobile phones, backpacks), and may also include various species such as cats, dogs, birds, and insects. In addition, it may also include various items such as clothes, electronic products, artworks, and food packaging, which are not limited here.
[0074] Exemplarily, in a traffic scenario, multiple types of objects of interest may include pedestrians (e.g., pedestrians with different clothing, ages, genders, and carrying different items), non-motor vehicles (e.g., transportation vehicles such as bicycles, electric vehicles, scooters, etc.), motor vehicles (such as cars, motorcycles), and other possible objects (traffic signs, road facilities, etc.), which are specifically determined by the application scenario and requirements and are not limited herein.
[0075] Image annotation data can be understood as an image that adds information on an image or a set of images to indicate objects, attributes, or relationships in the image. In a traffic scenario, the image annotation data can be image annotation data for objects of interest such as pedestrians and non-motor vehicles. After annotation processing, each object is assigned an ID or other form of label to represent that each object is actually the same entity in different images.
[0076] Exemplarily, when performing annotation processing on an image, object bounding box annotation can be used, that is, by drawing a rectangular box in the image to define the position of the target object, such as pedestrians, vehicles, animals, etc., and attaching corresponding category labels, such as "pedestrian", "car", etc., which are not limited herein. Pixel-level annotation, also known as segmentation annotation, can also be used to finely annotate the shape and contour of the object, and is usually used for semantic segmentation and instance segmentation tasks, such as road segmentation, building contour annotation, etc., which are not limited herein. In addition, key point annotation can be used, that is, for certain specific objects, such as the joint positions in human pose estimation, or the annotation of key parts such as eyes, nose, mouth, etc. in facial expression recognition, which are not limited herein.
[0077] Descriptive text annotation data can be understood as the text description corresponding to the image annotation data, which is used to provide detailed information about the image content. For example, the appearance characteristics and clothing descriptions of pedestrians in the image. The text description helps the model understand the specific content of the image, thereby improving the accuracy of multimodal matching.
[0078] Exemplarily, the descriptive text annotation data can provide detailed text information about the image content, such as "A man wearing a blue jacket and a red hat is walking on the street". It can also list the attributes of the objects in the image, such as color, material, action, etc. For example, "Pedestrian - blue jacket, red hat, walking; Bicycle - red body, white wheels, stationary". In addition, it can also provide the relationships between objects, that is, describe the relationships between the objects in the image, such as "The pedestrian is riding a bicycle". Providing this text description information helps the model understand the complexity of the scenario.
[0079] In this application, by obtaining a dataset containing a large number of images of objects of interest and descriptive texts, it provides the basic input and output data for subsequent model training, ensuring that the model is exposed to objects in various forms and scenarios, thereby improving the generalization ability of the model and enabling the model to effectively learn in tasks such as image-text matching and person re-identification (ReID).
[0080] After that, the initial multi-modal retrieval model is trained using the training dataset to generate learnable prompts. Among them, the initial multi-modal retrieval model can be understood as a pre-trained model that can initially handle multi-modal matching between images and texts, but may not be optimized for specific tasks (such as person re-identification in traffic scenarios).
[0081] A learnable prompt can be understood as a parameterized sequence attached to the input text, used to guide the model to better understand and process the input, especially in multi-modal tasks to guide the model to focus on key features in the image. That is, in the training of the multi-modal retrieval model, the learnable prompt will be used as part of the model input to supplement or guide how the model processes the input data, which can help the model better understand and capture cross-modal information connections, thereby improving the performance of the model in image-text retrieval tasks. It can be understood that the learnable prompt is a set of dynamic parameters that can be optimized and adjusted through the training process.
[0082] In this application, the obtained training dataset is used to train the initial multi-modal retrieval model to generate learnable prompts, that is, to learn for the learnable prompt. It can be understood that the focus of this stage of training is to introduce the learnable prompt, a kind of prompt information that can be learned and adjusted during the training process. The learnable prompt guides the model to focus on the key visual features of the object through iterative learning based on image annotation data and descriptive text annotation data. The learnable prompt, as an additional adjustable parameter through iterative learning, co-evolves with the model, gradually strengthening the response to annotation details, enabling the image encoder to better capture and retain the visual information unique to the object during the encoding stage, rather than relying solely on static text descriptions, effectively improving detail retention and recognition accuracy. That is, the learnable prompt can help the model retain more visual detail information related to the object during the image encoding stage, providing a better image feature representation for subsequent ReID tasks.
[0083] Thus, by introducing the learnable prompt, the model can retain more details about specific objects in the image while learning the text description, which helps improve the accuracy of subsequent retrieval tasks. At the same time, the existence of the learnable prompt makes the behavior of the model more interpretable, and it is possible to analyze which features the model focuses on during the matching process.
[0084] Next, the initial multi-modal retrieval model is trained for image encoding using the training dataset and the learnable prompt to generate an intermediate multi-modal retrieval model. Among them, in the multi-modal retrieval task, image encoding training refers to optimizing the image processing branch of the model so that it can efficiently and accurately extract the key features in the image.
[0085] The intermediate multi-modal retrieval model can be understood as an intermediate version of the model after preliminary training of the initial multi-modal retrieval model. Compared with the initial multi-modal retrieval model, it has further improved in image encoding and feature retention, but the training of the entire multi-modal pair has not been completed.
[0086] In this application, the initial multi-modal retrieval model is trained for image encoding using the training dataset and the learnable prompt to generate an intermediate multi-modal retrieval model. It can be seen that the training at this stage focuses on training the image encoder (such as ViT) of the model. It should be noted that at this time, the text encoder of the model remains frozen and the parameters are no longer updated. During the training process, the model uses both the learnable prompt and the detailed description text as inputs, which helps the image encoder to retain more visual details while also being more closely aligned with the text description.
[0087] Thus, through the training at this stage, the model can more effectively encode the visual information in the image, especially the detailed features of multiple types of objects of interest (such as pedestrians and non-motor vehicles) in a specific scenario (such as a traffic scenario), which helps improve the performance of image-to-text and image-to-image search. In addition, by freezing the text encoder and focusing on the training of the image encoder, the training efficiency can be effectively improved, and at the same time, the performance degradation caused by over-training of the text encoder can be prevented.
[0088] Finally, the intermediate multi-modal retrieval model is trained for text-image alignment using the training dataset to generate the target multi-modal retrieval model. Among them, in the multi-modal retrieval model, the image features and text features should be able to correspond to each other to form an effective association. Text-image alignment training can be understood as aligning the image features with the corresponding text features. Exemplarily, a contrastive loss function is usually used to ensure that the embedding vectors of the image and the corresponding text are similar, while the embedding vectors of the non-corresponding text are different.
[0089] The target multi-modal retrieval model can be understood as the final model version after all training stages. The target multi-modal retrieval model has shown excellent performance in image-to-text search, text-to-image search, and ReID tasks, and can adapt to the requirements of specific scenarios (such as traffic scenarios) to achieve efficient and stable multi-modal retrieval.
[0090] The target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result. Among them, the target retrieval request can be understood as a request sent by a user to a retrieval system (i.e., the target multi-modal retrieval model) in order to find a specific entity, object, or scene. Exemplarily, the target retrieval request may be in the form of an image query. For example, a user uploads one or a set of images, and the retrieval system needs to find other images that match the objects in these images. The target retrieval request may also be in the form of a text description query. For example, a user provides a piece of text describing the characteristics of a target object, and the retrieval system finds the corresponding object in the image database according to the text description. In addition, the target retrieval request may also be in the form of a composite query. The user provides both an image and a text description at the same time, and the retrieval system needs to comprehensively retrieve both pieces of information to find the images that meet the description.
[0091] The target retrieval result can be understood as the result obtained from text-to-image or image-to-text search, that is, a set of images that the target multi-modal retrieval model finds in the database and best matches the target retrieval request.
[0092] It can be seen that the target multi-modal retrieval model trained in this application can quickly process a large amount of image data, significantly shortening the time from request to result return, which helps to quickly respond in emergency situations and improves the retrieval efficiency. And through multi-stage training and model optimization, the accuracy of the target multi-modal retrieval model in cross-validation of images and texts has been greatly improved. Even in an open scenario, it can accurately identify the target, enhancing the retrieval accuracy. In addition, whether it is a pure image query, a pure text query, or a composite query, the target multi-modal retrieval model can handle them properly, meeting the diverse query needs of users, especially in the field of traffic scenarios, so it can adapt to complex query requirements.
[0093] In this application, a training dataset is used to train the intermediate multi-modal retrieval model for text and image feature alignment, thereby generating the target multi-modal retrieval model. The purpose of the training at this stage is to further enhance the multi-modal matching ability of the model, ensuring that the model can not only extract effective features from images, but also accurately correspond these features to text descriptions. It should be noted that at this time, the image encoder of the model is frozen, and only the text encoder is trained to ensure that the integrity of the image features is not affected, while adjusting the text encoder so that it can generate a text representation that is more aligned with the image features.
[0094] Thus, by performing text-image alignment training, the model can more accurately understand and match images with text descriptions, significantly improving the retrieval ability of text-to-image search. At the same time, significant performance improvements have also been achieved in specific tasks in specific scenarios. Additionally, by using the augmented training data (i.e., a dataset containing a large number of images of objects of interest and descriptive text), the robustness of the model to different text formats and image variations can be further enhanced, making the model more stable and reliable in actual deployment. Furthermore, after the above multi-stage training, the generated target multi-modal retrieval model reaches a better state in terms of comprehensive performance and can efficiently handle multi-modal retrieval tasks in specific scenarios (such as traffic scenarios), meeting the requirements of practical applications.
[0095] It can be seen that in this application, a training dataset is obtained. The training dataset includes: image annotation data of various types of objects of interest and descriptive text annotation data associated with the image annotation data. Then, the initial multi-modal retrieval model is trained using the training dataset to generate learnable prompts. After that, the initial multi-modal retrieval model is trained for image encoding using the training dataset and the learnable prompts to generate an intermediate multi-modal retrieval model. Finally, the intermediate multi-modal retrieval model is trained for text-image alignment using the training dataset to generate a target multi-modal retrieval model. The target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result. Thus, by introducing learnable prompts, the model can retain more visual detail information during the training process, avoiding the problem of image feature simplification that may be caused by direct contrast training. At the same time, through the learning of learnable prompts, the adaptability of the model to complex environments is enhanced, and the performance of the model in open scenarios is improved.
[0096] The above model training method provided by the embodiments of this application can be but is not limited to being applied to application scenarios involving multi-modal retrieval in fields such as urban security, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services.
[0097] By adopting the embodiment of the present application, a training data set is obtained. The training data set includes: image annotation data of multiple types of objects of interest and descriptive text annotation data associated with the image annotation data. Then, the initial multi-modal retrieval model is trained using the training data set to generate learnable prompts. After that, the initial multi-modal retrieval model is trained for image encoding using the training data set and the learnable prompts to generate an intermediate multi-modal retrieval model. Finally, the intermediate multi-modal retrieval model is trained for text-image alignment using the training data set to generate a target multi-modal retrieval model. The target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result. Thus, the purpose of improving the performance of the multi-modal model in a specific scenario is achieved, thereby realizing the technical effect of enhancing the adaptability of the multi-modal model to complex environments and improving the performance of the multi-modal model in an open scenario. Furthermore, the technical problem that the multi-modal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance is solved.
[0098] In an alternative embodiment, the multiple types of objects of interest include: objects of vital signs. In step S21, obtaining the training data set includes the following method steps:
[0099] Step S211, obtaining first image data from a preset image storage area, where the display content of the first image data includes: objects of vital signs;
[0100] Step S212, generating first annotation data and second annotation data based on the first image data, where the first annotation data is re-identification annotation data of the objects of vital signs, and the second annotation data is descriptive text annotation data of the objects of vital signs.
[0101] In the embodiment of the present application, the multiple types of objects of interest include: objects of vital signs. The objects of vital signs can be understood as biological or target objects closely related to biological activities, such as pedestrians in a traffic scenario.
[0102] When obtaining the training data set, first image data can be obtained from a preset image storage area, where the display content of the first image data includes: objects of vital signs. The preset image storage area can be understood as a specified storage space or database, that is, a picture library, for storing the collected original image data. In a traffic scenario, the preset image storage area is used to store the image set of the objects of vital signs recorded in the traffic scenario.
[0103] The first image data can be understood as the image extracted from the preset image storage area. The first image data contains at least one object of vital signs, such as a pedestrian.
[0104] In this application, by obtaining the first image data from the preset image storage area, it can be understood as screening and collecting image frames containing vital sign objects from the preset image storage area. In a traffic scenario, it can be understood as using image processing and computer vision technologies to screen and collect image frames containing vital sign objects.
[0105] Thus, by selecting images containing specific vital sign objects from the preset image storage area, the pertinence and effectiveness of the training data are ensured, the interference of irrelevant images is avoided, and the efficiency of model learning is improved. In addition, the preliminary data screening lays a solid foundation for subsequent annotation and model training.
[0106] After that, the first annotation data and the second annotation data are generated based on the first image data. Among them, the first annotation data is the re-identification (ReID) annotation data of vital sign objects, such as the image annotation data of each person in the first image data. It can be understood that the re-identification annotation data assigns a unique ID to each person in the image, and all images belonging to the same person can be recognized even when taken at different times and locations. That is to say, the first annotation data is essentially to mark and distinguish diverse instance displays of the same vital sign object in the image, by identifying different occurrence instances of the same vital sign object in the image, so that the model can recognize and match the images of the same vital sign object in different time and location scenes.
[0107] Exemplarily, assume there is a series of images captured by surveillance cameras. In these images, a specific pedestrian appears in multiple different frames, but the position, pose, and lighting conditions are different each time. The first annotation data will mark all the images of this pedestrian as the same ID (such as ID_001), so that the model can learn that despite the changing external conditions, the images belonging to the same person should be assigned to the same ID.
[0108] The second annotation data is the descriptive text annotation data of vital sign objects, which can include text descriptions of the object's appearance, attributes, behaviors, etc., to assist the model in learning how to understand and locate the corresponding image content from the text description.
[0109] Exemplarily, continuing with the above-mentioned pedestrian as an example, the second annotation data is the descriptive text annotation data of the pedestrian, which may contain the text description: "ID_001 is a male, wearing a blue jacket, a black hat, carrying a red backpack, and holding a coffee cup in his left hand". This description provides specific appearance details to help the model find the pedestrian image that matches the description in the text-to-image retrieval task.
[0110] In the present application, by generating the first annotation data and the second annotation data based on the first image data, it involves the identification and detailed characteristic description of the vital signs objects in the image. Thus, by automatically identifying and marking different instances of the same vital signs object, high-precision image retrieval can be maintained even in a changing environment, that is, the retrieval capability of the model is enhanced by the generation of ReID annotation data. At the same time, the text description provides additional semantic information, which helps the model understand the deep meaning of the image, thereby showing better performance in the image and text retrieval task, that is, the generation of descriptive text annotation data enriches the semantic understanding ability of the model. In addition, the combined training of the first annotation data and the second annotation data promotes the model's comprehensive understanding ability of visual information and text information, making the model more intelligent and accurate when processing image and text inputs, overcoming the limitations of single modality training.
[0111] In an optional embodiment, in step S212, generating first annotation data based on the first image data includes the following method steps:
[0112] Step S2121, based on the time information and geographic information of the first image data, performing spatiotemporal clustering processing on the vital sign objects to obtain a clustering result;
[0113] Step S2122: Perform confidence reflow processing on the clustering result to obtain first labeled data.
[0114] In the embodiment of the present application, when generating the first annotation data based on the first image data, the spatiotemporal clustering processing can be performed on the vital signs objects based on the time information and geographic information of the first image data to obtain a clustering result. In the field of computer vision, spatiotemporal clustering processing is a technology for grouping objects in an image based on time and space information, and by utilizing the similarity of objects at different time and space positions, objects that may belong to the same identity are grouped together.
[0115] The clustering result can be understood as the result obtained after spatiotemporal clustering processing, that is, the vital signs objects in the image are grouped into different categories, and each group may correspond to a specific identity.
[0116] In this application, based on the spatiotemporal information (i.e., time information and geographic information) of the first image data, the vital signs objects in the image are clustered in the spatiotemporal domain. The time information can be the timestamp when the image was taken, and the geographic information can be the coordinates or identification information of the location where the image was taken. By clustering the vital signs objects in the image in the spatiotemporal domain (for example, clustering a large number of pedestrian targets acquired in the spatiotemporal domain), the retrieval system can identify a set of images of the same object that appear continuously at similar times and locations, and provide a set of filtered candidate objects for subsequent ReID annotation.
[0117] Thus, through clustering, the model only needs to label the images within each group once, rather than labeling each object in all images one by one. This greatly reduces the labeling workload, improves efficiency, and reduces the labeling difficulty. At the same time, clustering images within the same spatio-temporal domain can increase the probability that objects of the same identity are correctly labeled with the same ID, thereby improving the overall labeling quality and accuracy.
[0118] After that, confidence backflow processing is performed on the clustering results to obtain the first labeled data. Among them, confidence backflow processing is a step to evaluate and screen the clustering results. By calculating the credibility of each clustering group, the groups with high confidence (i.e., the set of objects very likely to belong to the same identity) are used as the basis for subsequent labeling.
[0119] The first labeled data can be understood as assigning a unique identifier to each high-confidence clustering group after spatio-temporal domain clustering and confidence backflow processing, for subsequent model training to identify and match different instances of the same identity.
[0120] In this application, by performing confidence backflow on the clustering results, the clustering groups with high confidence are screened out, that is, the samples with high confidence are backflowed as the source of the first labeled data. The confidence calculation is based on the similarity of the images within the clustering group. High confidence indicates that the images within the group are very likely to belong to the same entity. Through this screening step, the reliability of the labeled data is ensured. Thus, through confidence backflow processing, it helps to exclude those images with unclear object recognition or low similarity in the clustering group, thereby reducing mislabeling, improving the purity of the labeled data, and avoiding mislabeling. At the same time, using the high-confidence clustering groups as the first labeled data provides high-quality training samples for the model, which can help the model more accurately learn the re-identification features of the object and improve the model performance.
[0121] In an optional embodiment, in step S212, generating the second labeled data based on the first image data includes the following method steps:
[0122] Step S2123, determining multiple local regions corresponding to the vital sign objects;
[0123] Step S2124, using a multi-modal language model to split and describe the multiple local regions to obtain a splitting result;
[0124] Step S2125, combining and describing the splitting result to obtain the second labeled data.
[0125] In the embodiments of the present application, when generating the second annotation data based on the first image data, multiple local regions corresponding to the vital sign objects can be determined first. Here, the local region can be understood as different parts or attribute regions of the vital sign objects in the image. For example, the head, upper body, backpack, lower body, shoes of a pedestrian, and riding tools, etc.
[0126] In the present application, multiple local regions corresponding to the vital sign objects in the first image data are determined. Thus, by decomposing the vital sign objects into multiple local regions, the model can analyze and extract the features of each region more meticulously, such as color, shape, and specific identifiers, to achieve refined feature extraction, which helps with subsequent description generation and feature matching. At the same time, the determination of the local regions helps the model accurately identify the key features of the target object even when dealing with images with occlusion, angle changes, and inconsistent lighting conditions, enhancing the robustness of the model.
[0127] Then, a multimodal language model is used to split and describe the multiple local regions to obtain a splitting result. Among them, the multimodal language model is a deep learning model that can process and understand multiple information forms (such as images and texts) and can generate corresponding text descriptions based on the image content.
[0128] In the multimodal scenario, the split description can be understood as a process of separately describing different local regions of the vital sign objects to obtain more delicate feature information.
[0129] The splitting result can be understood as the text information obtained after the multimodal language model describes the local regions, which contains the feature descriptions of each local region.
[0130] In the present application, a multimodal language model is used to split and describe each local region to generate text describing the features of the region, that is, the splitting result. For example, for the backpack of a pedestrian, the model will generate a description such as "a black backpack with white stripes". Thus, the split description of the local regions can ensure the accuracy of the text description, avoid the feature blur or information loss that may be brought by the overall description, and improve the description accuracy. At the same time, generating text descriptions for each local region increases the diversity and depth of the training data, enriches the training data, and helps the model learn more comprehensive target features.
[0131] Finally, the splitting result is combined and described to obtain the second annotation data. Here, the combined description can be understood as integrating the description information of multiple local regions in the splitting result to form a text describing a complete object, so as to provide a comprehensive and detailed description for subsequent model training and retrieval.
[0132] In this application, the splitting results of all local regions are merged to form a complete description of the vital-signs object, obtaining the second annotation data. For example, the descriptions of the head, upper body, backpack, lower body, and shoes are merged into a complete description of a pedestrian. Thus, the merged description ensures that all important features of the vital-signs object are included in the text, providing a comprehensive perspective for the model and helping the model accurately identify the target in image-text retrieval. At the same time, through the merged description, the generated text description maintains consistency in format and semantics, which is beneficial for the model to learn and understand, and improves the training efficiency.
[0133] In an alternative embodiment, the multiple types of objects of interest include: tool objects. In step S21 of obtaining the training data set, the following method steps are further included:
[0134] Step S213: Obtain second image data from a preset image storage area, where the displayed content of the second image data includes: tool objects;
[0135] Step S214: Generate third annotation data based on the second image data, where the third annotation data is the descriptive text annotation data of the tool objects.
[0136] In the embodiments of this application, the multiple types of objects of interest include: tool objects. In a traffic scenario, tool objects can be understood as non-biological entities closely related to the activities of pedestrians, such as non-motor vehicles, motor vehicles, and other items, such as bicycles, electric vehicles, etc. These objects play an important role in understanding the behaviors and identity characteristics of pedestrians.
[0137] When obtaining the training data set, second image data can be first obtained from a preset image storage area, where the displayed content of the second image data includes: tool objects. The second image data can be understood as the image data containing tool objects selected from the preset image storage area during the process of obtaining the training data set. These images can provide the visual representation of tool objects in the actual environment for training the model to recognize and describe these objects.
[0138] In this application, by obtaining the second image data containing tool objects from the preset image storage area, that is, the retrieval system will automatically browse the preset image storage area and screen out the specific frames with tool objects in those images, preparing data for subsequent annotation and model training. Thus, it can ensure that the training data set contains images of tool objects, laying a foundation for the model to understand the visual features of such objects. At the same time, the screened image data helps the model focus on learning the key features of tool objects, avoiding the waste of computing resources caused by a large number of irrelevant images, and improving the training efficiency.
[0139] Then, third annotation data is generated based on the second image data, where the third annotation data is descriptive text annotation data of tool objects, that is, descriptive text annotation data of non-motor vehicles or motor vehicles. The third annotation data may include detailed information such as the attributes, appearance, and location of the tool objects, and is used to train the model to understand and generate accurate descriptions of the tool objects.
[0140] In this application, for the obtained second image data, a multi-modal language model can be used to generate descriptive text of tool objects, that is, the third annotation data. Thus, the third annotation data provides a detailed description of the tool objects, including not only the visual features of the objects, but also their relationships with pedestrians, such as "riding a red motorcycle" and "carrying a black backpack", which helps the model understand and match objects more accurately. At the same time, by combining the text description with the image data, the model can learn how to convert visual information into language information, thereby improving the model's understanding and matching capabilities in image-text retrieval tasks.
[0141] In an alternative embodiment, in step S214, generating the third annotation data based on the second image data includes the following method steps:
[0142] Step S2141, obtaining the basic description of the tool object;
[0143] Step S2142, using a multi-modal language model to perform perspective-based supplementary description on the basic description to obtain a supplementary result;
[0144] Step S2143, combining the basic description and the supplementary result for description to obtain the third annotation data.
[0145] In the embodiment of this application, when generating the third annotation data based on the second image data, the basic description of the tool object can be obtained first. The basic description of the tool object can be understood as a preliminary text description of the tool object, usually including basic attribute information such as the category, color, and shape of the tool object. For example, "This is a bicycle".
[0146] In this application, the retrieval system obtains the basic description information of the tool object from the second image data, that is, identifies and marks the category and some obvious appearance features of the tool object in the image. Thus, the obtained basic description provides preliminary identification information of the tool object, which is the basis for subsequent detailed description and feature supplementation, ensuring the accuracy and consistency of the description.
[0147] After that, a multi-modal language model is used to supplement the basic description from different perspectives to obtain a supplementary result. Among them, the supplementary description from different perspectives can be understood as follows: considering that tool objects have different characteristic manifestations from different perspectives, the multi-modal language model will give additional description information for different perspectives. For example, the description of a bicycle from the back will highlight the basket and the rear light, while the description from the front will mention the handlebars and the wheels.
[0148] The supplementary result can be understood as the text information obtained by the multi-modal language model to supplement the description of the tool object from different perspectives. These information enrich the basic description and provide more details.
[0149] In this application, a multi-modal language model is used to supplement the basic description from different perspectives. For example, if it is a bicycle, the model will provide a description such as "there is a basket in the front and a red reflector in the back" according to the perspective of the image, so as to obtain a supplementary result. Thus, through the supplementary description from different perspectives, the model can capture and describe the unique features of tool objects from different perspectives, which is crucial for fine-grained image recognition and matching, and helps to improve the accuracy of image-text retrieval.
[0150] Finally, the basic description and the supplementary result are combined to obtain the third annotation data. Among them, the combined description can be understood as integrating the perspective description information in the basic description and the supplementary result to form a complete description containing various perspective features as the third annotation data.
[0151] In this application, after obtaining the basic description and the supplementary perspective description, the basic description and the supplementary result are combined to form a comprehensive and detailed text description, that is, the description is combined to generate a high-quality image-text description, and the third annotation data is obtained. For example, the combined description may be "This is a blue bicycle, there is a basket in the front and a red reflector in the back". Thus, through the combined description, it is ensured that the text contains all the key features and details of the tool object, including the basic attributes and the supplementary information from different perspectives, which helps the model to understand the target object more comprehensively and improve its performance and generalization ability in image-text retrieval.
[0152] In an optional embodiment, the model training method further includes the following method steps:
[0153] Step S215, perform data augmentation processing on the second annotation data and the third annotation data to obtain a data augmentation result, where the data augmentation processing includes at least one of the following: randomly adjusting the word order, randomly erasing the text, and back-translating the text.
[0154] In the embodiments of the present application, data augmentation processing can also be performed on the second annotation data and the third annotation data to obtain a data augmentation result. Among them, data augmentation processing is a common technical means in the field of machine learning, aiming to improve the generalization ability and robustness of the model and prevent overfitting by artificially modifying the training data. In the present application, the data augmentation processing includes at least one of the following: randomly adjusting the word order, randomly erasing text, and text back-translation.
[0155] Randomly adjusting the word order can be understood as randomly changing the order of words. For example, changing "the blue backpack" to "backpack, the blue one".
[0156] Randomly erasing text can be understood as randomly deleting some words from the text to simulate the information loss situation that may occur in actual applications. For example, changing "carrying a blue backpack" to "carrying blue".
[0157] Text back-translation can be understood as converting the original text into another language and then re-translating it back to the original language to generate synonymous text with slightly different styles or expressions. For example, first translating the Chinese description "This is a blue bike" into English "The bike is blue", and then translating it back to Chinese "This bike is blue".
[0158] The data augmentation result can be understood as the modified data obtained by augmenting the original text annotation data through processing methods such as randomly adjusting the word order, randomly erasing text, and text back-translation, aiming to improve the robustness and generalization ability of the model.
[0159] In the present application, after obtaining the second annotation data (pedestrian description data) and the third annotation data (tool object description data), in order to improve the robustness and generalization ability of the model, data augmentation processing will be performed on the above text annotation data to obtain a data augmentation result. Thus, through randomly adjusting the word order, that is, changing the order of words in the description, the model can learn multiple expressions of the same entity. Even if the description order is different, it can correctly identify and match the target, which helps to improve the adaptability and retrieval accuracy of the model when facing diverse description texts. At the same time, through randomly erasing text, the model learns to still be able to perform effective identification and retrieval in the case of partial feature loss, which is very useful for dealing with fuzzy, occluded or incomplete information, and improves the robustness and flexibility of the model. In addition, through translation and back-translation, the diversity of text expressions is introduced, allowing the model to come into contact with different expression styles and language structures, thereby enhancing its ability to understand and generate text, especially the adaptability to cross-language application scenarios.
[0160] In an alternative embodiment, the initial multi-modal retrieval model includes an initial image encoder and an initial text encoder. In step S22, the initial multi-modal retrieval model is trained using a training dataset to generate learnable prompts, including the following method steps:
[0161] Step S221: Train the initial image encoder using first annotated data and train the initial text encoder using second annotated data to obtain a first loss, where the first loss is used to determine the contrast loss between the first annotated data and the second annotated data;
[0162] Step S222: In response to both the initial image encoder and the initial text encoder being in a frozen state, generate learnable prompts based on the first loss.
[0163] In the embodiment of the present application, the initial multi-modal retrieval model includes an initial image encoder and an initial text encoder. Among them, the initial image encoder is a part of the initial multi-modal retrieval model, such as ViT, which is used to extract features from image data. The initial text encoder is another part of the initial multi-modal retrieval model, which is used to extract features from text data, such as RoBERTa.
[0164] When training the initial multi-modal retrieval model using a training dataset to generate learnable prompts, the initial image encoder can be trained using first annotated data and the initial text encoder can be trained using second annotated data to obtain a first loss. Among them, the first loss is used to determine the contrast loss between the first annotated data and the second annotated data, that is, the first loss can be understood as the loss value calculated during training, which is used to measure the matching degree between the features generated by the image encoder and the text encoder, ensuring that the model can understand the association between images and texts. Exemplarily, the first loss can be an img2text loss function or a text2img loss function, which is not limited here.
[0165] In the present application, the initial image encoder and the initial text encoder are respectively trained by using first annotated data (i.e., pedestrian image data with ID tags) and second annotated data (i.e., pedestrian description text data) to minimize the contrast loss between the two, that is, the first loss. Exemplarily, this process may involve fine-tuning on a pre-trained model so that the model can learn how to map images and corresponding text descriptions to a common feature space.
[0166] Thus, by minimizing the first loss, the model can learn how to align image features and text descriptions in the feature space, and can initially establish the connection between images and texts even in the early training stage of the model. At the same time, the above initial training process helps the model start to understand the multimodal relationship between images and texts, laying a foundation for subsequent more advanced model training and performance improvement.
[0167] After that, in response to both the initial image encoder and the initial text encoder being in a frozen state, a learnable prompt is generated based on the first loss. It can be understood that when both the initial image encoder and the initial text encoder are frozen (i.e., the model weights are no longer updated), the present application generates a learnable prompt based on the first loss (i.e., the contrast between the image and the text, that is, during the training stage, when the parameters of the initial image encoder and the initial text encoder are fixed, the difference metric between the model's performance on the training dataset and the actual labels). That is, without changing its core representation ability, the model learns how to generate effective input prompts according to the first loss to improve the feature matching effect in the next step of training.
[0168] Exemplarily, in the state where the initial image encoder and the initial text encoder are frozen, the difference between the image and the text encoding is calculated through the first loss. The first loss, as a signal, only drives the parameter update of the learnable prompt, and the goal is to minimize the matching error between the image and the text. The learnable prompt gradually adapts to and supplements the visual features in the image that are not fully expressed by the text description through iterative learning, especially those details that are crucial for the ReID task. As the training progresses, the learnable prompt can better capture the nuances of the image, improve the visual consistency and retrieval performance of the model, and enhance the accuracy of image-text retrieval.
[0169] Thus, the learnable prompt can guide the model to more accurately understand the association between images and texts, start to optimize feature matching even in the early training stage, without affecting the existing knowledge of the model, and achieve precise guidance for training. At the same time, generating a learnable prompt when the encoder is frozen helps to retain more detailed information, prevent the loss of important visual and text features during the training process, and thus enhance the recognition and retrieval capabilities of the model in subsequent training.
[0170] In an optional embodiment, the initial multimodal retrieval model includes: an initial image encoder and an initial text encoder. In step S23, the initial multimodal retrieval model is trained for image encoding using the training dataset and the learnable prompt to generate an intermediate multimodal retrieval model, including the following method steps:
[0171] Step S231, using the first annotated data to train the initial image encoder, and using the second annotated data and the learnable prompt to train the initial text encoder to obtain a second loss, wherein the second loss is used to determine the relative distance loss between objects of different vital signs classes and the identity classification loss of objects of different vital signs classes;
[0172] Step S232, in response to the initial text encoder being in a frozen state, adjusting parameters of the initial image encoder based on the second loss to generate a target image encoder;
[0173] Step S233, using the initial text encoder and the target image encoder to determine an intermediate multimodal retrieval model.
[0174] In an embodiment of the present application, when the initial multimodal retrieval model is trained for image encoding using a training data set and a learnable prompt to generate an intermediate multimodal retrieval model, the initial image encoder can be trained using the first annotated data, and the initial text encoder can be trained using the second annotated data and a learnable prompt to obtain a second loss.
[0175] Among them, the second loss is used to determine the relative distance loss between objects of different vital signs and the identity classification loss of objects of different vital signs. That is, the second loss is used to measure the relative distance between objects of different vital signs (such as different pedestrians) and the accuracy of identity classification. The second loss usually consists of two parts: relative distance loss and identity classification loss. In the multimodal retrieval model, the relative distance loss focuses on how to shorten the distance of the same instance (such as different images of the same person) in the feature space, while widening the distance of different instances (such as different people), while the identity classification loss focuses more on the model's classification accuracy of individual identities. Exemplarily, the second loss can be a loss function that meets the ReID task, such as the triplet loss loss function and the soft id loss loss function, which is not limited here.
[0176] It can be understood that the relative distance loss is used to learn an embedding space in which samples from the same vital signs class objects (such as different images of the same person or the same person at different time points) should be mapped to close positions, while samples from different vital signs class objects should be mapped to farther positions. That is, the relative distance loss ensures the above goals by minimizing the distance between positive sample pairs (samples of the same object) and maximizing the distance between negative sample pairs (samples of different objects). Exemplarily, based on image annotation data and descriptive text annotation data, the relative distance loss is calculated by tripletloss, and images and texts of the same identity are selected as positive sample pairs, and those of different identities are negative sample pairs. The goal is to minimize the distance between positive sample pairs and maximize the distance to negative samples.
[0177] The identity classification loss allows the model to be more tolerant near the decision boundary by introducing label smoothing technology, thereby better handling the fuzzy recognition problem between different objects. When calculating the identity classification loss, by slightly dispersing the true distribution of the labels (i.e., reducing the confidence of a certain category and smoothing it to other categories), the model is prompted to learn a more robust classification boundary. Exemplarily, the identity classification loss is based on the description text annotation data, supervises the model classification through soft id loss, and uses label smoothing technology to enable the model to not only focus on exact matches but also recognize similar but different identities, improving the classification robustness.
[0178] In this application, the initial image encoder is trained using the first annotation data, and at the same time, the initial text encoder is trained using the second annotation data combined with learnable prompts to minimize the second loss, thereby ensuring that the model can accurately understand the relative distances between different vital sign category objects and the identity classification of different vital sign category objects. Thus, by minimizing the relative distance loss, the model learns how to distinguish different vital sign category objects in the feature space, which is crucial for improving the accuracy of image retrieval and enhancing the instance discrimination. At the same time, by minimizing the identity classification loss, it helps the model to more accurately identify and classify different individuals, such as the identity of pedestrians, thereby enhancing the ability of text to retrieve images and improving the identity recognition rate.
[0179] After that, in response to the initial text encoder being in a frozen state, the parameters of the initial image encoder are adjusted based on the second loss to generate a target image encoder. The target image encoder can be understood as a more accurate image feature extraction component obtained by optimizing the parameters of the initial image encoder based on the second loss.
[0180] It can be understood that when adjusting the parameters of the initial image encoder, the relative distance loss is mainly used. Because in ReID or related multi-modal matching tasks, once the text encoder is frozen, the feature representation it outputs no longer changes. At this time, the relative distance loss is used to train and adjust the parameters of the image encoder to optimize the image feature representation, ensuring that images of the same identity are close to each other in the feature space, while images of different identities are far from each other. The identity classification loss is usually used in the stage when the text encoder is also activated and participates in the training together to improve the classification performance.
[0181] In this application, the initial text encoder is frozen, which means its parameters are no longer updated. Then, the parameters of the initial image encoder are adjusted according to the second loss. That is, RoBERTa is frozen but the ViT part is unfrozen for learning of the image encoder. Thus, the model can focus on optimizing the performance of the image encoder to ensure that the image features can accurately reflect the relative distances and identity information of vital sign objects in the feature space. Thereby, by adjusting the parameters of the image encoder, the model can better capture and represent the key features in the image, especially those details that contribute to identity recognition and classification, so as to improve the retrieval accuracy and efficiency.
[0182] Finally, the initial text encoder and the target image encoder are used to determine an intermediate multi-modal retrieval model. Thus, by combining the frozen initial text encoder and the target image encoder, an intermediate multi-modal retrieval model is generated. This intermediate multi-modal retrieval model combines the optimized image feature representation and the original text understanding ability, and is an important transitional version in the multi-modal retrieval task.
[0183] Therefore, the generated intermediate multi-modal retrieval model is a more mature and effective model version, which can initially achieve cross-modal retrieval between images and texts, laying a solid foundation for subsequent training and optimization.
[0184] In an alternative embodiment, the intermediate multi-modal retrieval model includes: a target image encoder and an initial text encoder. In step S24, the intermediate multi-modal retrieval model is trained for text-image alignment using a training data set to generate a target multi-modal retrieval model, including the following method steps:
[0185] Step S241, at least the first labeled data is used to train the target image encoder, and the data augmentation result is used to train the initial text encoder to obtain a third loss, where the third loss is used to determine the contrast loss between the first labeled data and the data augmentation result;
[0186] Step S242, in response to the target image encoder being in a frozen state, the parameters of the initial text encoder are adjusted based on the third loss to generate a target text encoder;
[0187] Step S243, the target text encoder and the target image encoder are used to determine the target multi-modal retrieval model.
[0188] In the embodiment of this application, the intermediate multi-modal retrieval model includes: a target image encoder and an initial text encoder, that is, the intermediate multi-modal retrieval model includes a target image encoder optimized by training and an initial text encoder that has not been optimized.
[0189] When training the intermediate multi-modal retrieval model for text-image alignment using a training dataset to generate a target multi-modal retrieval model, the target image encoder can be trained at least using first annotation data, and the initial text encoder can be trained using the data augmentation result to obtain a third loss.
[0190] Among them, the third loss is used to determine the contrast loss between the first annotation data and the data augmentation result. It can be understood that the third loss is a comprehensive metric used to evaluate and guide the model's ability to perform feature alignment and matching between image and text modalities. The third loss combines the contrast loss and the ID classification loss, ensuring that the model can not only distinguish the differences between different instances but also accurately identify and match the associations between images and text descriptions. That is, the third loss is a loss function used to evaluate and guide the feature alignment degree between the target image encoder and the initial text encoder. The matching degree between the image and the text description is measured through the contrast loss, ensuring that the model can perform accurate cross-modal retrieval from text to image and from image to text.
[0191] In this application, the first annotation data is used to further train the target image encoder to ensure that it can more accurately extract features from images and distinguish different vital sign objects. At the same time, the data augmentation result is used to train the initial text encoder to improve its ability to process diverse text descriptions, ensuring that the model can perform accurate text feature extraction even if the text description has variations. That is, the matching degree between the image and the text description is measured through the third loss function, thereby optimizing the model's text-image alignment ability.
[0192] Thus, the feature extraction ability of the target image encoder can be further strengthened, making it more accurate when processing images in specific scenarios (such as traffic scenarios). At the same time, the robustness of the initial text encoder is enhanced, enabling it to better handle changes in text descriptions and improving the generalization performance of the model.
[0193] After that, in response to the target image encoder being in a frozen state, the parameters of the initial text encoder are adjusted based on the third loss to generate a target text encoder. The target text encoder can be understood as an optimized version evolved from the initial text encoder through parameter adjustment based on the third loss. The target text encoder can better transform text descriptions into the feature space to match the output features of the target image encoder.
[0194] In this application, the target image encoder remains frozen, which means its parameters are no longer updated. That is, after removing the learnable prompt, the ViT is frozen and the RoBERTa is unfrozen. At this time, the parameters of the initial text encoder are adjusted based on the third loss to generate the target text encoder. This adjustment process ensures that the features generated by the text encoder can be aligned with the output of the image encoder in the feature space, thereby improving the image-text retrieval performance of the model.
[0195] Thereby, the feature representation ability of the text encoder can be improved, so that the features it outputs can match the features output by the image encoder more closely, optimizing the image-text pairs of the model. In addition, by freezing the target image encoder, the stability of image feature extraction is ensured, and at the same time, the text encoder is allowed to make necessary adjustments, improving the overall coordination and consistency of the model.
[0196] Finally, the target multimodal retrieval model is determined using the target text encoder and the target image encoder. By combining the target text encoder and the target image encoder, the target multimodal retrieval model is generated. The target multimodal retrieval model integrates the retrieval functions from image to text and from text to image, and is the final model version after multi-stage training and optimization.
[0197] Thereby, the generated target multimodal retrieval model realizes comprehensive multimodal retrieval capabilities, can handle the image-text retrieval requirements in complex specific scenarios (such as traffic scenarios), and improves the practicality of the model. At the same time, through multi-stage training, the generated target multimodal retrieval model can process image and text data more precisely, improving the retrieval efficiency and accuracy in practical applications.
[0198] Figure 3 is the overall flowchart according to the embodiments of the present application. As Figure 3 shown, it is mainly divided into a data processing stage, a model training stage (including three stages), and a model inference stage.
[0199] The data processing stage (automatic data processing and generating labeled data) is mainly divided into four parts: pedestrian ReID labeled data generation, pedestrian description text data generation, non-motor vehicle description text data generation, and description text data enhancement.
[0200] Regarding the generation of pedestrian ReID annotation data, for ReID training data, IDs need to be annotated for each instance. Most of the current research in the ReID field obtains annotation data through manual annotation methods. Therefore, it is often difficult to automatically and cost-effectively construct a large amount of high-quality training data for the ReID task. However, in the application scenario of this application, the time and geographical information of video frames can be obtained. Therefore, it becomes possible to combine detection algorithms to automate the construction of ReID. This application clusters a large number of pedestrian targets in the spatio-temporal domain and then returns high-confidence samples as ReID annotation data, and finally obtains more than 500,000 ReID annotation data at low cost.
[0201] Specifically, when generating pedestrian ReID annotation data, first locate pedestrians in the picture through a detection algorithm, then cluster them in combination with time and geographical information, and return high-confidence samples as ReID annotation data, so as to generate a large number of training samples at low cost.
[0202] Regarding the generation of pedestrian description text data, if directly described by the vlm model, due to the high information density of the prompt words, there is often text inertia, and the generated text description will have hallucinations and the effect is not ideal. Therefore, the pedestrian is split and described into the head, upper body, backpack, lower body, riding tool, etc. Separate text prompts are designed for the above local areas, and the vlm is used to split and describe them. At the same time, a binary judgment is made on the belongings and their existence. Finally, the description is merged through the prompt words containing the deduplication and merging logic to generate high-quality picture text descriptions.
[0203] Specifically, when generating pedestrian description text data, first split and describe the pedestrian, design specific text prompts, then generate descriptions of each part through a multimodal language model, and then merge them to generate the final high-quality description text.
[0204] Regarding the generation of non-motor vehicle description text data, like the generation of pedestrian description text, there is text inertia, and sometimes the information in the picture will be ignored, resulting in hallucinations. Therefore, when generating non-motor vehicle description text data, this application first designs global basic description prompts to generate the basic description of the non-motor vehicle, and then uses the vlm to judge the perspective of the non-motor vehicle, including the front, side, and back. Exclusive prompts for each perspective are designed. For example, when looking at the back, the trunk of the non-motor vehicle can be seen, while the front includes the non-motor vehicle's headlights, basket, and front end. Supplementary descriptions of each perspective are generated through the vlm, and finally the description is merged through the prompt words containing the deduplication and merging logic to generate high-quality picture text descriptions.
[0205] Specifically, when generating non-motor vehicle description text data, first design the global basic description and prompt words for different perspectives, and then generate the descriptions of non-motor vehicles through a multi-modal language model and merge them to obtain high-quality description text.
[0206] When performing data augmentation on the description text, the present application enhances the description text through strategies such as random adjustment of word order, random erasure of text, and text back-translation. On the one hand, data can be expanded through text data augmentation, and on the other hand, the robustness of the model text encoder can be improved.
[0207] Specifically, when performing data augmentation on the description text, through strategies such as random adjustment of word order, random erasure of text, and text back-translation, the generated description text is enhanced, thereby improving the robustness and adaptability of the model.
[0208] Exemplarily, Figure 4 is a data augmentation example diagram according to an embodiment of the present application. As Figure 4 shown, taking "The man has short black hair and is wearing a white short-sleeved shirt, black trousers, and black shoes" as an example, through random adjustment of word order, this example is adjusted to "The man has short black hair, black shoes, and is wearing a white short-sleeved shirt, black trousers". Then perform random erasure of text, erase "has" in the example, and get "The man short black hair, is wearing a white short-sleeved shirt, black trousers, and black shoes". Finally, perform text back-translation, translate "The man has short black hair and is wearing a white short-sleeved shirt, black trousers, and black shoes" into English "The man had short black hair and was wearing a white short-sleeved shirt, black trousers and black shoes", and then translate the English back into Chinese to get "The man has short black hair and is wearing a white short-sleeved shirt, black trousers and black shoes".
[0209] Thus, it intuitively demonstrates how to generate training data at low cost and high quality, not only reducing the burden of manual annotation, but also improving the efficiency and quality of model training, providing a solid foundation for subsequent multi-stage training.
[0210] Model training stage (the first stage): In the first stage of model training, this application uses ViT as the image encoder and RoBERTa as the text encoder to construct a graph-text alignment model. In the first stage, this application freezes ViT and RoBERTa, introduces learnable prompt and the detailed description corresponding to the picture, and while retaining the information corresponding to the detailed description text for the image, it can retain more information about pedestrian instances in the image through learnable prompt, providing a basis for image information for subsequent learning. In this stage, this application uses the contrastive loss function of image and text (such as img2text, text2img loss function) as the objective function to learn for learnable prompt.
[0211] Specifically, the key point of the first stage of model training is to learn learnable prompt through the contrastive loss function, while freezing the ViT and RoBERTa encoders and only optimizing learnable prompt. It can be seen that through the training of the first stage, the model can better retain the visual details of pedestrian instances, providing a richer and more accurate feature basis for the ReID training in the subsequent stage. At the same time, through the optimization of the contrastive loss function, the effectiveness of learnable prompt in image retrieval and text retrieval is enhanced.
[0212] Model training stage (the second stage): In the second stage of model training, this application introduces the ReID technology, constructs a data loading module with instance ID as the unit for training samples, and freezes RoBERTa, but unfreezes the ViT part for learning of the image encoder. In addition, this application also loads the learnable prompt and detailed description learned in the first stage of training for splicing as the input of the text encoder, and designs a loss function that conforms to the ReID task (such as triplet loss, soft id loss) for training. Through the training of this stage, the visual consistency of the model is greatly enhanced, and the indicators in the image retrieval task are greatly improved.
[0213] Specifically, the main goal of the second stage of model training is to introduce the ReID technology, improve the visual consistency of the model, and enhance the image retrieval ability. On the basis of the model in the first stage, the second stage model adds ReID training samples, uses the triplet loss and soft id loss functions for training, and unfreezes the ViT encoder at the same time to enhance the discrimination ability of the image encoder for pedestrian instances. It can be seen that the training in the second stage further optimizes the performance of the image encoder, greatly improves the indicators in the image retrieval task, especially in terms of visual consistency, enabling the model to more accurately identify and retrieve pedestrian instances.
[0214] Model training stage (the third stage): Since the above two training stages introduce learnable prompt training, and the learnable prompt is only for the pedestrian instance ID in the training stage, the learnable prompt of the current instance cannot be obtained during the inference process, and it is also impossible to use the learnable prompt as input when performing text search. Therefore, in the third stage of model training, in order to retain the image retrieval ability of the image encoder, after removing the learnable prompt, the ViT is frozen. At the same time, in order to align the features generated by the text encoding after removing the learnable prompt to the image, the RoBERTa is unfrozen in this application, so as to perform the third stage of training on the model. In addition, this application also uses the data after data augmentation for training in this stage to improve the robustness of the entire model. Since the first two training stages have well retained the encoding ability of the image encoder for image details, the text can be well aligned to the image in this stage, greatly improving the overall performance of text search. At the same time, the training data also includes high-quality description annotations of non-motor vehicles, and the recall and other indicators are also greatly improved when performing text search for non-motor vehicles.
[0215] Specifically, the third stage of model training focuses on aligning the output of the text encoder with the image encoder to improve the overall performance of text search. Based on the first two stages, the third stage model removes the learnable prompt, freezes the optimized ViT encoder, and only trains the RoBERTa encoder. It uses the description text after data augmentation for training to ensure that the text encoder can generate feature representations aligned with the image features. It can be seen that by adjusting the text encoder, the model can more accurately match the description text with the image during text search, greatly improving the performance of text search. Especially in the recall index of text search for non-motor vehicles, a significant improvement is achieved.
[0216] Model inference stage: This application finds that if the input description text is kept in the same format as the training during text search, the retrieval effect can be improved. Therefore, in the inference stage, this application introduces vlm to rewrite the originally custom description text input to make it more in line with the text description format during training.
[0217] Specifically, in the model inference stage, the format of the input description text is adjusted by using a multi-modal language model to ensure that it is consistent with the text description format during training, thereby improving the retrieval effect. It can be seen that through text rewriting, the model can better understand and match the input description text during inference, improving the accuracy and stability of retrieval. Especially for the specific target image-text retrieval in an open scenario, a more stable and reliable solution is provided.
[0218] In summary, Figure 3 The process integrating all the above stages is a full - process technical roadmap from automatic data processing and annotation to final model inference.
[0219] When generating pedestrian ReID annotation, first use the detection algorithm to locate and intercept pedestrian images from video frames, then cluster the pedestrian images in the time domain based on time and geographical location information, and finally select high - confidence samples from the clustering results to label a unique ID for each pedestrian instance. Thus, high - confidence pedestrian instance ID annotation data is automatically generated for subsequent ReID task training. It realizes the construction of a large - scale ReID training dataset at low cost and high efficiency, and enhances the model's ability to identify pedestrian instances.
[0220] When generating description text annotation data for pedestrians, first divide the pedestrian image into local features such as head, upper body, lower body, backpack, hat, riding tool, etc., then design text prompt words for each local feature, then use a multi - modal language model to generate local feature descriptions according to the prompt words, and finally perform deduplication and merging processing on the generated descriptions to form a complete description text. Thus, high - quality description text data about pedestrian appearance features is automatically generated for training the text encoder. It improves the accuracy and detail of the text description, and promotes the model's understanding and retrieval ability of pedestrian features.
[0221] When generating description text annotation data for non - motor vehicles, first design global basic description prompt words for non - motor vehicles, then judge and identify the shooting angle of the non - motor vehicle, then design exclusive prompt words according to the angle, use a multi - modal language model to generate supplementary descriptions, and finally merge the global basic description and the angle supplementary description to generate the final non - motor vehicle description text. Thus, the description text of non - motor vehicles is generated for improving the text - based search performance of non - motor vehicle retrieval. Through multi - angle supplementary descriptions, the comprehensiveness and accuracy of non - motor vehicle retrieval are improved, providing strong support for specific target retrieval in traffic scenarios.
[0222] When enhancing the data of picture description text, first adopt the word order random adjustment strategy to change the appearance order of words in the description text, then use the text random erasure strategy to randomly delete a part of the words in the description, and finally implement the text back - translation strategy to translate the text into another language and then back - translate it into the original text to introduce semantic variants. Thus, the description text dataset is expanded, and the robustness and generalization ability of the model are enhanced. Data enhancement enriches the diversity of training samples, reduces the model's dependence on specific text formats, improves the model's ability to process different text descriptions, and enhances the stability and generalization performance of the model in practical applications.
[0223] In summary, this application leverages the capabilities of the current vlm to automatically generate a large number of high-quality training samples by designing reasonable prompt words, step-by-step descriptions, and spatiotemporal constrained clustering that conforms to the data distribution. At the same time, it expands the data through a variety of text data enhancement methods to improve the robustness of the model. At the same time, this application introduces learnable prompts and detailed description texts into the model to allow the image-text alignment model to retain more visual information of specific target categories. In addition, this application introduces ReID technology to improve the visual consistency of the model and enhances the performance of pedestrian retrieval. In the training stage, this application breaks down the tasks in stages to improve the interpretability of the model and the certainty of the effect. In the reasoning stage, this application further improves the stability of image-text retrieval by rewriting the description text.
[0224] It can be seen that this application combines the actual situation of specific traffic task scenarios and can support specific target and Chinese image and text retrieval. In addition, this application can construct text pairs and ReID tags at low cost without manual participation. Through performance improvement, industry adaptation and low-cost annotation, image and text retrieval can be truly implemented, providing traffic scenarios with image and text retrieval capabilities for specific targets in open scenarios, while also retaining most of the open image and text matching capabilities.
[0225] This application takes into account the images acquired based on the detection, and automatically generates the required large amount of annotation data for pedestrians and non-motor vehicles respectively. It can provide a large amount of high-quality annotation data for the transportation industry, and also provides a low-cost and high-quality automatic annotation solution for image and text retrieval in traffic scenes.
[0226] This application found that due to the large semantic gap between images and texts, direct image-text alignment training will cause serious loss of image information, and can only retain coarse-grained information aligned with the text, resulting in greatly reduced results when retrieving pedestrian images. Therefore, this application introduces learnable prompts to retain as much visual detail information as possible in the image-text dual-tower structure. However, this application found that although the fixed text input combined with learnable prompts can retain more details in the visual part, these details may not necessarily provide better pedestrian visual information, and the model is more easily disturbed by the background. Therefore, this application introduces dynamic text descriptions to explicitly retain the visual detail information related to the model and pedestrians, improve the model's attention to the target, and provide a foundation for subsequent stages.
[0227] After the above transformation, the model can better retain the visual information of pedestrians and can also recall better during image and text retrieval. However, it was found that when searching for pedestrian instances by image search, the visual consistency of the model was not good enough, resulting in less than ideal overall indicators. Therefore, this application introduces ReID technology to enhance the discrimination of inter-class and intra-class samples in pedestrian retrieval and improve the visual consistency of the model.
[0228] It can be seen that this application automatically constructs more than 500,000 ReID annotation data through a spatio-temporal clustering strategy that conforms to the task scenario. Through the splitting, description, and summarization capabilities of the multi-modal model, reasonable prompts are set to generate controllable fine-grained image descriptions, and more than 5 million image-text pairs of data are automatically constructed. This application splits the task through a multi-stage training method, improves the retrieval metrics and effects, and also enhances the interpretability of the model, making the task more controllable. This application introduces learnable prompts and fine-grained image descriptions in the training stage to retain more detailed information of the visual ViT. The performance of pedestrian image retrieval is significantly enhanced through ReID technology. This application combines fine-grained image descriptions and text enhancement methods in the training stage to improve the image-text retrieval effect. Additionally, in the retrieval stage, the inference effect is improved through text standardization, thereby significantly enhancing the overall performance of text search.
[0229] It is easy to understand that the beneficial effects of the model training method provided by this application include the following points.
[0230] Beneficial effect (1), automatic generation of low-cost and high-quality data, that is, this application can automatically construct a large amount of ReID annotation data and image-text pairs of pedestrians and non-motor vehicles, significantly reducing the data annotation cost and improving the quality of training data at the same time, which is the basis for the improvement of model performance.
[0231] Beneficial effect (2), significant performance improvement, that is, in the traffic open scenario, the mean Average Precision (mAP) metrics of image search and text search in this application are respectively improved by 7.7% and 4%. Specifically, for the text search recall@14 of non-motor vehicles, it jumps from 17.1% to 75.7%, indicating that the model's retrieval ability and accuracy for targets have been significantly improved.
[0232] Beneficial effect (3), enhanced visual consistency. This application optimizes the model by introducing ReID technology, strengthens the intra-class similarity and inter-class difference in image retrieval, improves the performance of visual consistency, and makes the retrieval results more stable and accurate.
[0233] Beneficial effect (4), improved robustness and generalization ability. This application increases the model's adaptability to the diversity of text descriptions through data augmentation technology, and improves the robustness and generalization performance of the model in different scenarios and different text formats.
[0234] Beneficial effect (5), multi-modal ability enhancement. This application combines technologies such as ReID, learnable prompts, text rewriting, and data augmentation to strengthen the multi-modal matching and generation ability of the model. It not only performs well in image-text retrieval but also retains the general ability of the model in the open scenario.
[0235] Beneficial effect (6), specific target retrieval optimization, especially for pedestrians and non-motor vehicles in traffic scenarios. The present application provides a more refined retrieval strategy, which can handle the specific target retrieval requirements in open scenarios and improve the efficiency and safety of intelligent transportation.
[0236] Beneficial effect (7), applicable to large-scale data retrieval. Through automated data processing and model optimization, the present application can effectively meet the retrieval requirements of large-scale surveillance video data, greatly shortening the case handling time and providing strong support for public safety emergency response.
[0237] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0238] In addition, it should also be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0239] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0240] According to the embodiments of the present application, there is also provided a Figure 5 data processing method as shown in Figure 5 is a flowchart of a data processing method according to the embodiments of the present application, as shown in Figure 5 shown, the method includes:
[0241] Step S51, obtaining a target retrieval request;
[0242] Step S52: Perform multimodal retrieval on the target retrieval request using the target multimodal retrieval model to obtain a target retrieval result. The target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0243] In the embodiment of the present application, the target retrieval request can be understood as a retrieval requirement initiated by a user, including an appearance description or an image of a specific target, with the purpose of finding a pedestrian or non-motor vehicle image that matches the request description or image in a database.
[0244] By obtaining the target retrieval request, and then performing multimodal retrieval on the target retrieval request using the target multimodal retrieval model to obtain a target retrieval result. For specific details, refer to the description of the foregoing embodiments, and no further elaboration will be provided here.
[0245] The data processing method provided in the embodiment of the present application can be but is not limited to being applied to application scenarios involving multimodal retrieval in fields such as urban security, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services.
[0246] Exemplarily, the data processing method provided in the embodiment of the present application can be applied to an urban security scenario. For example, when an emergency occurs, such as looking for a missing person, a text description of the target person can be quickly provided, and then the data processing method of the present application can be used to automatically retrieve matching images in surveillance videos based on the text description, greatly shortening the time to find clues. Another example is that for a large number of surveillance videos, traditional video analysis methods are inefficient. Through the data processing method of the present application, multimodal description and retrieval of pedestrians or non-motor vehicles in the video can be directly performed without manual frame-by-frame inspection, significantly improving the retrieval speed and efficiency of surveillance videos. Another example is that the data processing method of the present application can not only perform retrieval based on historical data, but also perform real-time analysis and early warning, identify and mark targets that match the description, such as pedestrians with abnormal behaviors or non-motor vehicles with specific styles, providing immediate intelligent support for urban security.
[0247] By adopting the embodiment of the present application, by obtaining the target retrieval request, and then performing multimodal retrieval on the target retrieval request using the target multimodal retrieval model to obtain a target retrieval result. The target multimodal retrieval model is generated according to the model training method described in any one of the above, thereby achieving the purpose of improving the performance of the multimodal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario, and further solving the technical problem that the multimodal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance.
[0248] In an alternative embodiment, the information carried in the target retrieval request includes: a custom description text, and the data processing method further includes the following method steps:
[0249] Use a multi-modal language model to rewrite the format of the custom description text to obtain a target description text;
[0250] Use the target multi-modal retrieval model to perform multi-modal retrieval on the target retrieval request to obtain a target retrieval result, including: using the target multi-modal retrieval model to perform graphic and text retrieval on the target description text to obtain a target retrieval image.
[0251] In the embodiment of the present application, the information carried in the target retrieval request includes: a custom description text, which can be understood as a personalized description provided by the user about the target to be retrieved, and may include information such as color, shape, size, position, etc., but the format and content may be inconsistent with the description text in the training data, and are not limited here.
[0252] At the beginning of the retrieval, the retrieval system receives the custom description text submitted by the user, and then standardizes and rewrites the format of the custom description text through the deployed target multi-modal language model to ensure that the custom description text matches the format of the description text encountered during model training. Exemplarily, the rewriting process includes grammar correction, word replacement, and adjustment of sentence structure, etc. The purpose is to improve the consistency between the text description and image features, thereby improving the accuracy of multi-modal retrieval.
[0253] Thereby, it is ensured that the format of the custom description text is consistent with the description text during training, and the retrieval error caused by the description format difference is eliminated. At the same time, through the optimization of the text description, the model's understanding of the description content is enhanced, and even in the face of complex descriptions, it can be more accurately transformed into image features. In addition, the standardized text description accelerates the model's retrieval process, improving the response speed and user experience.
[0254] After completing the text format rewriting, the retrieval system inputs the target description text into the target multi-modal retrieval model for retrieval. The model extracts the semantic features of the text and maps these features to the image feature space to find the most matching image and generate a target retrieval result. This target retrieval result is usually a series of images sorted according to the matching degree between the target description text and the image, which is convenient for users to quickly locate the most relevant results.
[0255] Thus, based on the optimized feature extraction and matching mechanism, the model can more accurately identify the targets in the image and improve the accuracy of retrieval. At the same time, since the open-scene data is used during model training, it can well adapt to various complex environmental factors in actual retrieval and improve the reliability in real-world applications. In addition, through the previous training, especially the introduction of ReID technology and learnable prompt, the model has significant performance advantages in image retrieval, especially in the retrieval of pedestrians and non-motor vehicles, greatly improving the mAP indicators of image search and text search.
[0256] In an optional embodiment, the information carried in the target retrieval request includes: a custom image. In step S51, the target multimodal retrieval model is used to perform multimodal retrieval on the target retrieval request to obtain a target retrieval result, including the following method steps:
[0257] Use the target multimodal retrieval model to perform image retrieval on the custom image to obtain a target retrieval image.
[0258] In the embodiment of the present application, the information carried in the target retrieval request includes: a custom image. The custom image can be understood as a specific image uploaded by the user and used as the query basis for retrieval. The custom image can be an image of a pedestrian, a non-motor vehicle, or an object in any other traffic scene.
[0259] When using the target multimodal retrieval model to perform multimodal retrieval on the target retrieval request to obtain a target retrieval result, the target multimodal retrieval model can be used to perform image retrieval on the custom image to obtain a target retrieval image. The retrieval system receives the custom image submitted by the user as part of the retrieval request. Subsequently, the custom image is input into the target multimodal retrieval model for image retrieval. The task of the model is to search the images in the database to find other images that belong to the same instance object (such as the same pedestrian) as the custom image, that is, to perform an image search for image operation.
[0260] Thus, the target multimodal retrieval model optimizes the ReID technology in the training stage, significantly enhancing the recognition of specific instances in image retrieval, so that it can quickly and accurately find the object images that match the query image. At the same time, since the model retains more visual details during training and enhances visual consistency through ReID technology, it can more carefully compare image features and reduce the possibility of fuzzy matching. In addition, the model is not only effective for fixed training data sets, but also has good adaptability to new images in open scenes, which means that even in the face of unknown backgrounds or lighting conditions, it can maintain high retrieval performance.
[0261] For specific descriptions, refer to the descriptions of the foregoing embodiments, and details are not elaborated here.
[0262] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and will not be elaborated here.
[0263] According to an embodiment of the present application, there is also provided a Figure 6 data processing method as shown below. Figure 6 FIG. is a flowchart of a data processing method according to an embodiment of the present application. As Figure 6 shown below, the method includes:
[0264] Step S61, obtaining a target retrieval request, where the target retrieval request is used to query a specified participating object in an urban traffic scenario;
[0265] Step S62, performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model to obtain a specified participating object retrieval result; where the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0266] The solution of the embodiment of the present application can be applied to querying in an urban traffic scenario. The target retrieval request is used to query a specified participating object in an urban traffic scenario, that is, the target retrieval request can be a query request initiated by a user or a system, aiming to search for specific participating objects in an urban traffic scenario, such as pedestrians, non-motor vehicles, etc.
[0267] In a traffic scenario, the specified participating object retrieval result can be understood as the retrieval result of a participating object (such as a pedestrian, a non-motor vehicle, etc.) for a specific description or image query, and usually includes a set of object information that is relatively matched with the target description or image.
[0268] For specific descriptions, refer to the descriptions of the foregoing embodiments, and will not be elaborated here.
[0269] The above data processing method provided by the embodiment of the present application can be but is not limited to being applied to application scenarios involving multimodal retrieval in fields such as urban security, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services.
[0270] By adopting the embodiments of the present application, a target retrieval request is obtained, where the target retrieval request is used to query specified participating objects in the urban traffic scenario, and then a target multimodal retrieval model is used to perform multimodal retrieval on the target retrieval request to obtain a specified participating object retrieval result; the target multimodal retrieval model is generated according to the model training method described in any one of the above, thereby achieving the purpose of improving the performance of the multimodal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario, and further solving the technical problems that the multimodal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance.
[0271] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, which will not be elaborated here.
[0272] According to the embodiments of the present application, there is also provided a Figure 7 data processing method as shown. Figure 7 is a flowchart of a data processing method according to the embodiments of the present application. As Figure 7 shown, the method includes:
[0273] Step S71, obtaining a data processing request through a first application programming interface, where the request data carried in the data processing request includes: a target retrieval request;
[0274] Step S72, returning a data processing response through a second application programming interface, where the response data carried in the data processing response includes: a target retrieval result, and the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0275] In the embodiments of the present application, the above first application programming interface (Application Programming Interface, API) and the second application programming interface can be either the same application programming interface or different application programming interfaces. In an alternative embodiment, the interface parameters in the above first application programming interface and the second application programming interface may include but are not limited to: interface global identifier, interface signature key, interface timestamp, interface request identifier, system call credential identifier, etc. The above first application programming interface can use GET or POST as the interface request method to obtain a file processing request. The above second application programming interface can use the JSON format to feedback the file processing response.
[0276] A data processing request can be understood as a request from an external system, which contains specific instructions or requirements for data processing. The goal of the request is to perform multimodal retrieval.
[0277] A data processing response can be understood as the processing result of the system for an external request, which contains specific outputs or information feedback after analyzing the request data.
[0278] For specific descriptions, refer to the descriptions of the foregoing embodiments, and details are not elaborated here.
[0279] The above data processing method provided by the embodiments of this application can be but is not limited to being applied to application scenarios involving multimodal retrieval in fields such as urban security, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services.
[0280] By adopting the embodiments of this application, a data processing request is obtained through a first application programming interface. Among them, the request data carried in the data processing request includes: a target retrieval request. Then, a data processing response is returned through a second application programming interface. Among them, the response data carried in the data processing response includes: a target retrieval result, which is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model. The target multimodal retrieval model is generated according to the model training method described in any one of the above. Thus, the purpose of improving the performance of the multimodal model in a specific scenario is achieved, thereby realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in open scenarios. Furthermore, the technical problem in the related art that the multimodal retrieval model cannot adapt to variability and large-scale data requirements and has poor model performance is solved.
[0281] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and details are not elaborated here.
[0282] According to the embodiments of this application, there is also provided a Figure 8 data processing method as shown. Figure 8 is a flowchart of a data processing method according to the embodiments of this application, as Figure 8 shown, and the method includes:
[0283] Step S81, obtain the current input data processing dialogue request. Among them, the request data carried in the data processing dialogue request includes: a target retrieval request;
[0284] Step S82: In response to a data processing dialogue request, return a data processing dialogue reply, where the information carried in the data processing dialogue reply includes: a target retrieval result, which is obtained by performing multimodal retrieval on a target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0285] Step S83: Display the target retrieval result within the graphical user interface.
[0286] In the embodiments of the present application, the data processing dialogue request is a common request form in an interactive system, usually initiated through a dialogue, requesting the system to process specific data. In the present application, the data processing dialogue request is used to request the system to perform a target retrieval.
[0287] The data processing dialogue reply is the system's response to the data processing dialogue request, containing the processing result of the requested data, that is, the target retrieval result.
[0288] The graphical user interface is a visual interface for users to interact with a computer system, usually including elements such as buttons, menus, windows, and icons, which simplifies the difficulty of users operating complex systems.
[0289] For specific descriptions, refer to the descriptions of the foregoing embodiments, and details are not elaborated here.
[0290] The above data processing method provided by the embodiments of the present application can be, but is not limited to, applied to application scenarios involving multimodal retrieval in fields such as urban security, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services.
[0291] By adopting the embodiments of the present application, by obtaining the current input data processing dialogue request, where the request data carried in the data processing dialogue request includes: a target retrieval request, and then in response to the data processing dialogue request, returning a data processing dialogue reply, where the information carried in the data processing dialogue reply includes: a target retrieval result, which is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above, and finally displaying the target retrieval result within the graphical user interface, the purpose of improving the performance of the multimodal model in a specific scenario is achieved, thereby realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario, and further solving the technical problems in the related art that the multimodal retrieval model cannot adapt to variability and large-scale data requirements and the model performance is poor.
[0292] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and will not be elaborated here.
[0293] According to an embodiment of the present application, there is also provided a Figure 9 data processing method as shown. Figure 9 is a flowchart of a data processing method according to an embodiment of the present application. As Figure 9 shown, the method includes:
[0294] Step S91, in response to an input instruction acting on the operation interface, display a target retrieval request on the operation interface;
[0295] Step S92, in response to a processing instruction acting on the operation interface, display a target retrieval result on the operation interface; wherein, the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0296] In the embodiment of the present application, the operation interface can be understood as a visual platform for interaction between the system and the user, usually a graphical user interface, where the user can input instructions, query, or perform other operations.
[0297] The input instruction can be understood as an instruction issued by the user through the operation interface for submitting a target retrieval request, such as clicking a button, filling out a form, or using a voice command to submit a retrieval requirement.
[0298] The processing instruction can be understood as an instruction for the user to cause the system to perform a specific task on the operation interface, such as a processing instruction for the target retrieval request, that is, requiring the system to start performing a retrieval task.
[0299] For specific descriptions, refer to the descriptions of the foregoing embodiments, and will not be elaborated here.
[0300] The above data processing method provided by the embodiment of the present application can be but is not limited to being applied to application scenarios involving multimodal retrieval in fields such as urban security, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services.
[0301] By adopting the embodiments of the present application, by responding to an input instruction acting on an operation interface, a target retrieval request is displayed on the operation interface, and then by responding to a processing instruction acting on the operation interface, a target retrieval result is displayed on the operation interface; wherein, the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above, thereby achieving the purpose of improving the performance of the multimodal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario, and further solving the technical problems in the related art that the multimodal retrieval model cannot adapt to variability and large-scale data requirements and the model performance is poor.
[0302] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and will not be elaborated here.
[0303] According to the embodiments of the present application, there is also provided a Figure 10 data processing system as shown. Figure 10 FIG. is a schematic structural diagram of a data processing system according to the embodiments of the present application, as Figure 10 shown, the method includes:
[0304] A client for sending a target retrieval request;
[0305] A server, connected to the client, for performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model to obtain a target retrieval result;
[0306] The client is further configured to output the target retrieval result; wherein, the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0307] In the embodiments of the present application, the data processing system includes a client and a server. For specific descriptions, refer to the descriptions of the foregoing embodiments, and will not be elaborated here.
[0308] The above data processing system provided by the embodiments of the present application can be but is not limited to being applied to application scenarios involving multimodal retrieval in fields such as urban security, e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services.
[0309] By adopting the embodiments of the present application, through the data processing system, the purpose of improving the performance of the multimodal model in a specific scenario is achieved, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario, and further solving the technical problems in the related art that the multimodal retrieval model cannot adapt to variability and large-scale data requirements and the model performance is poor.
[0310] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and will not be elaborated here.
[0311] According to an embodiment of the present application, there is also provided an apparatus embodiment for implementing the above model training method. Figure 11 It is a schematic structural diagram of a model training apparatus according to an embodiment of the present application, as Figure 11 shown, the apparatus includes:
[0312] A first acquisition module 1101, configured to acquire a training data set, where the training data set includes: image annotation data of multiple types of objects of interest and description text annotation data associated with the image annotation data;
[0313] A first training module 1102, configured to train an initial multi-modal retrieval model by using the training data set to generate learnable prompts;
[0314] A second training module 1103, configured to perform image encoding training on the initial multi-modal retrieval model by using the training data set and the learnable prompts to generate an intermediate multi-modal retrieval model;
[0315] A third training module 1104, configured to perform text and image alignment training on the intermediate multi-modal retrieval model by using the training data set to generate a target multi-modal retrieval model, where the target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result.
[0316] Optionally, the multiple types of objects of interest include: vital sign objects, and the first acquisition module 1101 is further configured to: acquire first image data from a preset image storage area, where the display content of the first image data includes: vital sign objects; generate first annotation data and second annotation data based on the first image data, where the first annotation data is re-identification annotation data of the vital sign objects, and the second annotation data is description text annotation data of the vital sign objects.
[0317] Optionally, the first acquisition module 1101 is further configured to: perform spatio-temporal domain clustering processing on the vital sign objects based on the time information and geographical information of the first image data to obtain a clustering result; perform confidence backflow processing on the clustering result to obtain the first annotation data.
[0318] Optionally, the first acquisition module 1101 is further configured to: determine multiple local regions corresponding to the vital sign objects; perform split description on the multiple local regions by using a multi-modal language model to obtain a split result; perform combined description on the split result to obtain the second annotation data.
[0319] Optionally, the multiple types of objects of interest include: tool objects. The first acquisition module 1101 is further configured to: acquire second image data from a preset image storage area, where the display content of the second image data includes: tool objects; generate third annotation data based on the second image data, where the third annotation data is descriptive text annotation data of the tool objects.
[0320] Optionally, the first acquisition module 1101 is further configured to: acquire a basic description of the tool object; use a multimodal language model to perform a perspective-supplementary description on the basic description to obtain a supplementary result; and perform a combined description on the basic description and the supplementary result to obtain the third annotation data.
[0321] Optionally, the apparatus further includes: a processing module, configured to perform data augmentation processing on the second annotation data and the third annotation data to obtain a data augmentation result, where the data augmentation processing includes at least one of the following: randomly adjusting the word order, randomly erasing the text, and text back-translation.
[0322] Optionally, the initial multimodal retrieval model includes: an initial image encoder and an initial text encoder. The first training module 1102 is further configured to: train the initial image encoder using the first annotation data, and train the initial text encoder using the second annotation data to obtain a first loss, where the first loss is used to determine the contrast loss between the first annotation data and the second annotation data; in response to both the initial image encoder and the initial text encoder being in a frozen state, generate a learnable prompt based on the first loss.
[0323] Optionally, the initial multimodal retrieval model includes: an initial image encoder and an initial text encoder. The second training module 1103 is further configured to: train the initial image encoder using the first annotation data, and train the initial text encoder using the second annotation data and the learnable prompt to obtain a second loss, where the second loss is used to determine the relative distance loss between different vital sign objects and the identity classification loss of different vital sign objects; in response to the initial text encoder being in a frozen state, adjust the parameters of the initial image encoder based on the second loss to generate a target image encoder; and use the initial text encoder and the target image encoder to determine an intermediate multimodal retrieval model.
[0324] Optionally, the intermediate multi-modal retrieval model includes: a target image encoder and an initial text encoder. The above third training module 1104 is further configured to: train the target image encoder at least using first annotation data, and train the initial text encoder using the data augmentation result to obtain a third loss, where the third loss is used to determine the contrast loss between the first annotation data and the data augmentation result; in response to the target image encoder being in a frozen state, adjust the parameters of the initial text encoder based on the third loss to generate a target text encoder; and use the target text encoder and the target image encoder to determine a target multi-modal retrieval model.
[0325] By adopting the embodiment of the present application, by obtaining a training data set, where the training data set includes: image annotation data of various types of attention objects and description text annotation data associated with the image annotation data, then training an initial multi-modal retrieval model using the training data set to generate learnable prompts, and then performing image encoding training on the initial multi-modal retrieval model using the training data set and the learnable prompts to generate an intermediate multi-modal retrieval model, and finally performing text and image alignment training on the intermediate multi-modal retrieval model using the training data set to generate a target multi-modal retrieval model, where the target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result, thereby achieving the purpose of improving the performance of the multi-modal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multi-modal model to complex environments and improving the performance of the multi-modal model in an open scenario, and further solving the technical problem that the multi-modal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance.
[0326] It should be noted here that the above first acquisition module 1101, first training module 1102, second training module 1103, and third training module 1104 correspond to steps S21 to S24 in the embodiment. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also run in the server 10 provided in the embodiment.
[0327] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above data processing method. Figure 8 is a schematic structural diagram of another data processing device according to an embodiment of the present application, as Figure 8 shown, the device includes:
[0328] A second acquisition module 1201, configured to acquire a target retrieval request;
[0329] The first retrieval module 1202 is configured to perform multimodal retrieval on a target retrieval request by using a target multimodal retrieval model to obtain a target retrieval result; wherein, the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0330] Optionally, the information carried in the target retrieval request includes: a custom description text. The apparatus further includes: a rewriting module configured to rewrite the format of the custom description text by using a multimodal language model to obtain a target description text; the first retrieval module 1202 is further configured to: perform image-text retrieval on the target description text by using the target multimodal retrieval model to obtain a target retrieval image.
[0331] Optionally, the information carried in the target retrieval request includes: a custom image. The first retrieval module 1202 is further configured to: perform image retrieval on the custom image by using the target multimodal retrieval model to obtain a target retrieval image.
[0332] By applying the embodiments of the present application, by obtaining a target retrieval request, and then performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model to obtain a target retrieval result; wherein, the target multimodal retrieval model is generated according to the model training method described in any one of the above, thereby achieving the purpose of improving the performance of the multimodal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario, and further solving the technical problem that the multimodal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance.
[0333] It should be noted here that the above second acquisition module 1201 and the first retrieval module 1202 correspond to steps S51 and S52 in the embodiment. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also run in the server 10 provided in the embodiment.
[0334] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above data processing method. Figure 13 is a schematic structural diagram of another data processing device according to an embodiment of the present application, as Figure 13 shown. The device includes:
[0335] A third acquisition module 1301 is configured to acquire a target retrieval request, wherein the target retrieval request is used to query a specified participating object in an urban traffic scenario.
[0336] The second retrieval module 1302 is configured to perform multimodal retrieval on a target retrieval request by using a target multimodal retrieval model to obtain a specified participant retrieval result, where the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0337] By applying the embodiment of the present application, a target retrieval request is obtained, where the target retrieval request is used to query a specified participant in an urban traffic scenario, and then the target multimodal retrieval model is used to perform multimodal retrieval on the target retrieval request to obtain a specified participant retrieval result. The target multimodal retrieval model is generated according to the model training method described in any one of the above. Thus, the purpose of improving the performance of the multimodal model in a specific scenario is achieved, thereby realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario. Furthermore, the technical problem that the multimodal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance is solved.
[0338] It should be noted here that the above third acquisition module 1301 and second retrieval module 1302 correspond to steps S61 and S62 in the embodiment. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above module or unit may be a hardware component or a software component stored in a memory and processed by one or more processors, and the above module may also run in the server 10 provided in the embodiment.
[0339] According to an embodiment of the present application, another device embodiment for implementing the above data processing method is also provided. Figure 14 It is a schematic structural diagram of another data processing device according to an embodiment of the present application, as Figure 14 shown. The device includes:
[0340] A fourth acquisition module 1401 is configured to obtain a data processing request through a first application programming interface, where the request data carried in the data processing request includes: a target retrieval request;
[0341] A first return module 1402 is configured to return a data processing response through a second application programming interface, where the response data carried in the data processing response includes: a target retrieval result, which is obtained by performing multimodal retrieval on the target retrieval request by using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0342] By adopting the embodiment of the present application, a data processing request is obtained through a first application programming interface. Among them, the request data carried in the data processing request includes: a target retrieval request. Then, a data processing response is returned through a second application programming interface. Among them, the response data carried in the data processing response includes: a target retrieval result, which is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model. The target multimodal retrieval model is generated according to the model training method described in any one of the above, thereby achieving the purpose of improving the performance of the multimodal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario. Furthermore, the technical problem that the multimodal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance is solved.
[0343] It should be noted here that the above-mentioned fourth acquisition module 1401 and the first return module 1402 correspond to steps S71 and S72 in the embodiment. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The above modules can also run in the server 10 provided in the embodiment.
[0344] According to the embodiment of the present application, another device embodiment for implementing the above data processing method is also provided. Figure 15 It is a schematic structural diagram of another data processing device according to the embodiment of the present application, as Figure 15 shown. The device includes:
[0345] A fifth acquisition module 1501, configured to acquire a current input data processing dialogue request. Among them, the request data carried in the data processing dialogue request includes: a target retrieval request;
[0346] A second return module 1502, configured to respond to the data processing dialogue request and return a data processing dialogue reply. Among them, the information carried in the data processing dialogue reply includes: a target retrieval result, which is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model. The target multimodal retrieval model is generated according to the model training method described in any one of the above;
[0347] A display module 1503, configured to display the target retrieval result in a graphical user interface.
[0348] By adopting the embodiment of the present application, a data processing dialogue request is obtained based on the currently input data. The request data carried in the data processing dialogue request includes a target retrieval request. Then, in response to the data processing dialogue request, a data processing dialogue reply is returned. The information carried in the data processing dialogue reply includes a target retrieval result, which is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model. The target multimodal retrieval model is generated according to the model training method described in any one of the above. Finally, the target retrieval result is displayed in the graphical user interface, thereby achieving the purpose of improving the performance of the multimodal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario. Furthermore, the technical problem in the related art that the multimodal retrieval model cannot adapt to variability and large-scale data requirements and has poor model performance is solved.
[0349] It should be noted here that the above-mentioned fifth acquisition module 1501, second return module 1502, and display module 1503 correspond to steps S81 to S82 in the embodiment. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above-mentioned module or unit can be a hardware component or a software component stored in the memory and processed by one or more processors, and the above-mentioned module can also run in the server 10 provided in the embodiment.
[0350] According to the embodiment of the present application, another device embodiment for implementing the above data processing method is also provided. Figure 16 is a schematic structural diagram of another data processing device according to the embodiment of the present application, as Figure 16 shown. The device includes:
[0351] A first display module 1601, configured to display a target retrieval request on the operation interface in response to an input instruction acting on the operation interface;
[0352] A second display module 1602, configured to display a target retrieval result on the operation interface in response to a processing instruction acting on the operation interface; wherein, the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above.
[0353] By adopting the embodiment of the present application, by responding to the input instruction acting on the operation interface, the target retrieval request is displayed on the operation interface, and then by responding to the processing instruction acting on the operation interface, the target retrieval result is displayed on the operation interface; wherein, the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request by using the target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of the above, thereby achieving the purpose of improving the performance of the multimodal model in a specific scenario, thus realizing the technical effect of enhancing the adaptability of the multimodal model to complex environments and improving the performance of the multimodal model in an open scenario, and further solving the technical problem that the multimodal retrieval model in the related art cannot adapt to variability and large-scale data requirements and has poor model performance.
[0354] It should be noted here that the above first display module 1601 and second display module 1602 correspond to steps S91 and S92 in the embodiment. The functions of the two modules are the same as those of the corresponding steps in terms of implementation examples and application scenarios, but are not limited to the content disclosed in the above embodiment. It should be noted that the above module or unit can be a hardware component or a software component stored in the memory and processed by one or more processors, and the above module can also run in the server 10 provided in the embodiment.
[0355] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in the embodiments, but are not limited to the schemes provided in the embodiments.
[0356] The embodiment of the present application can provide a computing device. Figure 17 It is a structural block diagram of a computing device according to an embodiment of the present application. As Figure 17 shown, the computing device A may include: one or more ( Figure 17 only one is shown in the figure) processors 1702, a memory 1704, a storage controller, and a peripheral interface. The peripheral interface may be connected to a radio frequency module, an audio module, a display screen, etc., which are not limited here.
[0357] The above computing device A can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a PC (Personal Computer), a model all-in-one machine, etc. And the model described in the above embodiments of the present application may be pre-installed in the computing device.
[0358] Specifically, the computing device A can pre-set various types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi-modal task processing, etc., so as to provide diverse model selections. In different product forms, the computing device A can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference and application, etc. In some product forms, the computing device A also supports model management, including but not limited to multi-type model management (supporting the management of various types of models such as discriminative and generative models), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effect of the model based on model evaluation tools), etc. In other product forms, the computing device A can also create applications based on models, provide API invocation capabilities, and can call the model into the created application through the API interface, and at the same time provide application management tools to realize the management and monitoring of the application.
[0359] Furthermore, the computing device A can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master AI technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.
[0360] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the methods in the above embodiments. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include memories remotely set relative to the processor, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0361] The processor can call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0362] Those of ordinary skill in the art can understand that the structure shown Figure 17 is only schematic, and the computing device A can also be a terminal device such as a smart phone, a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD, etc. The Figure 17It does not limit the structure of the above computing device. For example, computing device A may also include more or fewer components (such as network interfaces, display devices, etc.) than those shown in the Figure 17 figure, or have a different configuration from that shown in the Figure 17 figure.
[0363] Embodiments of the present application may provide an electronic device. Figure 18 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 18 shown, the electronic device may include: an input / output device 182; a memory 184, and a processor 186, where the processor 186 is connected to the input / output device 182 and the memory 184 through a bus 188.
[0364] Among them, the memory may be used to store software programs and modules, such as program instructions / modules corresponding to the methods and apparatuses in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories may be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0365] The processor may call the executable program stored in the memory through a transmission device to execute the method described in any one of the above embodiments.
[0366] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0367] Embodiments of the present application also provide a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium may be used to store the program code executed by the model training method or data processing method provided in the above embodiments.
[0368] Optionally, in this embodiment, the above computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0369] Embodiments of the present application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements any one of the above model training methods or data processing methods.
[0370] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0371] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0372] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0373] In addition, the functional units in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0374] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0375] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A model training method, characterized in that, Including: Obtain a training data set, where the training data set includes: image annotation data of multiple types of objects of interest and descriptive text annotation data associated with the image annotation data; Use the training data set to train an initial multi-modal retrieval model to generate learnable prompts; Use the training data set and the learnable prompts to perform image encoding training on the initial multi-modal retrieval model to generate an intermediate multi-modal retrieval model; Use the training data set to perform text-image alignment training on the intermediate multi-modal retrieval model to generate a target multi-modal retrieval model, where the target multi-modal retrieval model is used to perform multi-modal retrieval on a target retrieval request to obtain a target retrieval result.
2. The model training method according to claim 1, wherein The multiple types of objects of interest include: vital sign objects. Obtaining the training data set includes: Obtain first image data from a preset image storage area, where the display content of the first image data includes: the vital sign objects; Generate first annotation data and second annotation data based on the first image data, where the first annotation data is re-identification annotation data of the vital sign objects, and the second annotation data is descriptive text annotation data of the vital sign objects.
3. The model training method according to claim 2, wherein Generating the first annotation data based on the first image data includes: Perform spatio-temporal domain clustering processing on the vital sign objects based on the time information and geographical information of the first image data to obtain a clustering result; Perform confidence backflow processing on the clustering result to obtain the first annotation data.
4. The model training method according to claim 2, wherein Generating the second annotation data based on the first image data includes: Determine multiple local regions corresponding to the vital sign objects; Use a multi-modal language model to perform split description on the multiple local regions to obtain a split result; Perform combined description on the split result to obtain the second annotation data.
5. The model training method according to claim 2, wherein The multiple types of objects of interest include: tool objects. Obtaining the training data set further includes: Obtain second image data from a preset image storage area, where the display content of the second image data includes: the tool objects; Generate third annotation data based on the second image data, where the third annotation data is descriptive text annotation data of the tool objects.
6. The model training method according to claim 5, wherein Generating the third annotation data based on the second image data includes: Obtain the basic description of the tool objects; Use a multi-modal language model to perform supplementary description from different perspectives on the basic description to obtain a supplementary result; Perform combined description on the basic description and the supplementary result to obtain the third annotation data.
7. The model training method according to claim 5, wherein The model training method further includes: Perform data augmentation processing on the second annotation data and the third annotation data to obtain a data augmentation result, where the data augmentation processing includes at least one of the following: randomly adjusting the word order, randomly erasing the text, and text back-translation.
8. The model training method according to claim 2, wherein The initial multi-modal retrieval model includes: an initial image encoder and an initial text encoder. Using the training data set to train the initial multi-modal retrieval model to generate the learnable prompts includes: Train the initial image encoder using the first labeled data, and train the initial text encoder using the second labeled data to obtain a first loss, where the first loss is used to determine the contrast loss between the first labeled data and the second labeled data; In response to both the initial image encoder and the initial text encoder being in a frozen state, generate the learnable prompt based on the first loss.
9. The model training method according to claim 2, wherein The initial multimodal retrieval model includes an initial image encoder and an initial text encoder. Use the training dataset and the learnable prompt to perform image encoding training on the initial multimodal retrieval model to generate the intermediate multimodal retrieval model, including: Train the initial image encoder using the first labeled data, and train the initial text encoder using the second labeled data and the learnable prompt to obtain a second loss, where the second loss is used to determine the relative distance loss between different vital sign class objects and the identity classification loss of different vital sign class objects; In response to the initial text encoder being in a frozen state, adjust the parameters of the initial image encoder based on the second loss to generate a target image encoder; Use the initial text encoder and the target image encoder to determine the intermediate multimodal retrieval model.
10. The model training method according to claim 7, wherein The intermediate multimodal retrieval model includes a target image encoder and an initial text encoder. Use the training dataset to perform text-image alignment training on the intermediate multimodal retrieval model to generate the target multimodal retrieval model, including: At least train the target image encoder using the first labeled data, and train the initial text encoder using the data augmentation result to obtain a third loss, where the third loss is used to determine the contrast loss between the first labeled data and the data augmentation result; In response to the target image encoder being in a frozen state, adjust the parameters of the initial text encoder based on the third loss to generate a target text encoder; Use the target text encoder and the target image encoder to determine the target multimodal retrieval model.
11. A data processing method, characterized in that, Include: Obtain a target retrieval request; Perform multimodal retrieval on the target retrieval request using the target multimodal retrieval model to obtain a target retrieval result; Wherein, the target multimodal retrieval model is generated according to the model training method described in any one of claims 1 to 10.
12. The data processing method according to claim 11, wherein The information carried in the target retrieval request includes: a custom description text, and the data processing method further includes: Use a multimodal language model to rewrite the format of the custom description text to obtain a target description text; Performing multimodal retrieval on the target retrieval request using the target multimodal retrieval model to obtain the target retrieval result includes: Perform text-image retrieval on the target description text using the target multimodal retrieval model to obtain a target retrieval image.
13. The data processing method according to claim 11, wherein The information carried in the target retrieval request includes: a custom image. Performing multimodal retrieval on the target retrieval request using the target multimodal retrieval model, the obtained target retrieval result includes: Performing image retrieval on the custom image using the target multimodal retrieval model to obtain a target retrieval image.
14. A data processing method, characterized in that Including: Obtaining a target retrieval request, where the target retrieval request is used to query a specified participating object in an urban traffic scenario; Performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model to obtain a retrieval result of the specified participating object; Wherein, the target multimodal retrieval model is generated according to the model training method described in any one of claims 1 to 10.
15. A data processing method, characterized in that, Including: Obtaining a data processing request through a first application programming interface, where the request data carried in the data processing request includes: a target retrieval request; Returning a data processing response through a second application programming interface, where the response data carried in the data processing response includes: a target retrieval result, and the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of claims 1 to 10.
16. A data processing method, characterized in that, Including: Obtaining a current input data processing dialogue request, where the request data carried in the data processing dialogue request includes: a target retrieval request; Responding to the data processing dialogue request and returning a data processing dialogue reply, where the information carried in the data processing dialogue reply includes: a target retrieval result, and the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of claims 1 to 10; Displaying the target retrieval result within a graphical user interface.
17. A data processing method, characterized in that Including: Responding to an input instruction on an operation interface and displaying a target retrieval request on the operation interface; Responding to a processing instruction on the operation interface and displaying a target retrieval result on the operation interface; Wherein, the target retrieval result is obtained by performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model, and the target multimodal retrieval model is generated according to the model training method described in any one of claims 1 to 10.
18. A data processing system, characterized in that, Including: A client for sending a target retrieval request; A server connected to the client for performing multimodal retrieval on the target retrieval request using a target multimodal retrieval model to obtain a target retrieval result; The client is further configured to output the target retrieval result; Wherein, the target multimodal retrieval model is generated according to the model training method described in any one of claims 1 to 10.
19. An electronic device, characterized in that, Including: A memory storing an executable program; A processor for running the program, where when the program runs, it executes the model training method described in any one of claims 1 to 10 or the data processing method described in any one of claims 11 to 17.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the model training method described in any one of claims 1 to 10 or the data processing method described in any one of claims 11 to 17.
Citation Information
Patent Citations
Method and device for acquiring semantic labels of digital images
CN105740402A
Pedestrian re-identification method based on contrast language image pre-training model CLIP
CN115393902A
Multi-view multi-person scene detection method under camera calibration-free condition
CN116824497A
Vehicle identification labeling method and fusion system for multi-view image data
CN117132859A
CLIP multi-mode fused dish identification method, device and equipment
CN119007189A
Cited By
On-orbit calculation-oriented remote sensing image annotation box intelligent generation method and device
CN120913096A
Multi-modal data analysis agent and method based on large model
CN120929505A
Model training method, task processing method and system, electronic equipment, storage medium and computer program product
CN121352074A