Data processing method, computing device and electronic device
通过筛选和增强初始多模态数据,构建高质量的目标多模态数据,解决了现有技术中多模态数据质量低下的问题,提高了模型的训练效率和性能。
Patent Information
- Application Number
- CN202510067425.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the multimodal data used in model training is of poor quality, resulting in severe redundancy and noise in the data set, affecting model performance.
By obtaining initial multimodal data, using the first network model to screen the first initial modal data, obtain the first target modal data and object recognition results, and input the object recognition results and the second initial modal data to the second network model for enhancement, obtain the second target modal data, and finally construct the target multimodal data based on the first target modal data and the second target modal data.
The multimodal data quality used in model training is improved, data redundancy and noise are reduced, information consistency between modes is enhanced, and the training efficiency and final performance of the model are improved.
Smart Images

Figure CN119988905A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to large model technology and data processing fields, and more specifically, to a data processing method, a computing device, and an electronic device. Background Art
[0002] In the field of general artificial intelligence, the Multimodal Large Language Model (MLLM) has promoted the development of the technological frontier. As the core part of MLLM, the Visual-Language Alignment Model (VLA) realizes complex cross-modal interactions by aligning visual and textual representations. The current e-commerce platform, with its text and image product retrieval and content-based recommendation functions, relies on VLA to achieve accurate cross-modal feature understanding. Given that the improvement of VLA performance is closely related to the scale and quality of the data set, expanding high-quality data sets to enhance model capabilities has become the current effect. However, in actual use, high-quality image-text matching data only accounts for a small proportion. In addition, the time-consuming and labor-intensive manual annotation and the inherent redundancy and noise of network data make it difficult to build a high-quality data set.
[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0004] The embodiments of the present application provide a data processing method, a computing device, and an electronic device to at least solve the technical problem of poor data quality used in model training in the related art.
[0005] According to one aspect of an embodiment of the present application, a data processing method is provided, comprising: obtaining initial multimodal data, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing an object; screening the first initial modal data using a first network model to obtain first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of a target type contained in the first initial modal data; inputting the object recognition result and the second initial modal data into a second network model, and enhancing the second initial modal data using the second network model to obtain second target modal data; and obtaining target multimodal data based on the first target modal data and the second target modal data.
[0006] According to one aspect of an embodiment of the present application, a data processing method is provided, comprising: in response to an input instruction acting on an operation interface, displaying initial multimodal data on the operation interface, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing an object; in response to a processing instruction acting on the operation interface, displaying target multimodal data on the operation interface, wherein the target multimodal data is determined based on the first target modal data and the second target modal data, the second target modal data is obtained by inputting an object recognition result and the second initial modal data into a second network model, and enhancing the second initial modal data using the second network model, the first target modal data and the object recognition result are obtained by screening the first initial modal data using the first network model, and the object recognition result is used to represent an object of a target type contained in the first initial modal data.
[0007] According to one aspect of an embodiment of the present application, a data processing method is provided, comprising: acquiring initial multimodal data by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter includes the initial multimodal data, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing an object; screening the first initial modal data using a first network model to obtain first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of a target type contained in the first initial modal data; inputting the object recognition result and the second initial modal data into a second network model, and enhancing the second initial modal data using the second network model to obtain second target modal data; obtaining target multimodal data based on the first target modal data and the second target modal data; and outputting the target multimodal data by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the target multimodal data.
[0008] According to another aspect of the embodiments of the present application, a computing device is further provided, including: a memory storing an executable program; and a processor for running the program, wherein the method in each embodiment of the present application is executed when the program is running.
[0009] According to another aspect of the embodiments of the present application, an electronic device is also provided, including: a memory storing an executable program; a processor connected to the memory via a bus and used to run the program, wherein the method of each embodiment of the present application is executed when the program is running.
[0010] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in each embodiment of the present application.
[0011] According to another aspect of the embodiments of the present application, a computer program product is also provided, including a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.
[0012] According to another aspect of an embodiment of the present application, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.
[0013] According to another aspect of the embodiments of the present application, a computer program is also provided, and when the computer program is executed by a processor, the methods in the various embodiments of the present application are implemented.
[0014] In an embodiment of the present application, initial multimodal data is obtained, wherein the initial multimodal data includes a first initial modal data and a second initial modal data for describing an object; the first initial modal data is screened using a first network model to obtain first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of a target type contained in the first initial modal data; the object recognition result and the second initial modal data are input into a second network model, and the second initial modal data is enhanced using the second network model to obtain second target modal data; based on the first target modal data and the second target modal data, target multimodal data is obtained, thereby improving the quality of multimodal data used for model training; it is easy to notice that by processing data of different modalities in a targeted manner, the semantic consistency and accuracy of data of different modalities can be improved, and the data quality can be improved, so that when integrating these processed modal data, the constructed target multimodal data set not only reduces redundancy and noise, but also enhances the information consistency between modalities. This enables subsequent models to learn purer and higher-quality cross-modal semantic relationships during training, thereby improving the training efficiency and final performance of the model, and further solving the technical problem of poor data quality used in model training in related technologies.
[0015] It is easy to notice that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0017] Figure 1 is a schematic diagram of an application scenario of data processing according to an embodiment of the present application;
[0018] Figure 2 is a flow chart of a data processing method according to an embodiment of the present application;
[0019] Figure 3 is a structural schematic diagram of a multimodal data processing method according to an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of a governance framework according to an embodiment of the present application;
[0021] Figure 5 It is a schematic diagram of a data governance framework structure of a single-modal basic model according to an embodiment of the present application;
[0022] Figure 6 is a flow chart of a data processing method according to an embodiment of the present application;
[0023] Figure 7 is a flow chart of a data processing method according to an embodiment of the present application;
[0024] Figure 8 is a schematic diagram of a data processing device according to an embodiment of the present application;
[0025] Fig. 9 is a schematic diagram of a data processing device according to an embodiment of the present application;
[0026] Fig.10 is a schematic diagram of a data processing device according to an embodiment of the present application;
[0027] Fig.11 is a structural block diagram of a computing device according to an embodiment of the present application;
[0028] Fig.12 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] The technical solution provided in this application is mainly implemented by large-scale model technology. The large model here refers to a deep learning model with large-scale model parameters, which can usually contain hundreds of millions, tens of billions, hundreds of billions, trillions or even more than 10 trillion model parameters. The large model can also be called a cornerstone model / foundation model (Foundation Model). The large model is pre-trained through large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as large-scale language model (Large Language Model, LLM), multi-modal pre-training model (multi-modal pre-training model), etc.
[0032] It should be noted that when the large model is actually applied, the pre-trained model can be fine-tuned through a small number of samples so that the large model can be applied to different tasks. For example, the large model can be widely used in natural language processing (NLP), computer vision, speech processing and other fields, specifically in computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, etc. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0033] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0034] Data Governance: Currently, model structures are gradually converging to a few solutions, and the importance of data to project results has greatly increased. Data governance refers to the method of improving the performance of deep learning models by improving the data set. On the basis of the initial training data, by improving the quality of the training data, a more powerful model can be achieved.
[0035] Cross-modal consistency: refers to the consistency and coordination of information and expression between data of different modalities in a sample. Stronger cross-modal consistency can improve the ability of the cross-modal alignment model.
[0036] Data noise / redundancy: Different samples in the data have different effects on model performance. Samples that hinder model performance are noise, and samples that do not help improve model capabilities are redundant.
[0037] Visual-language correlation is important for modern search recommendations and conversational question-answering, and its capabilities are closely related to the scale and quality of data. Currently, in relevant scenarios, only a small portion of data is a good image-text match. Manual data cleaning is time-consuming and labor-intensive, which hinders development. Although the current popular data governance framework in the industry is effective, it has obvious defects, mixing valid information and noise in the data, losing diversity, and requiring repeated iterations of models and data.
[0038] Existing data governance solutions have different focuses, but all have shortcomings. The model fitting ability is improved by screening low-noise text data through probabilistic prediction, but the data diversity is sacrificed, resulting in distribution shift. Although the teacher-student framework expands the training samples and increases the environmental adaptability of the model, it is limited by the distribution of labeled data and it is difficult to maintain data diversity. Although the cross-merging technology of the support vector machine improves the training efficiency, it affects the model performance due to the reduction in sample size. The data acquisition scheme improves data accuracy through clustering and screening networks, ignores the problems of redundancy and noise within the sample, and relies on high-quality seed data. In contrast, the present application innovatively uses a unimodal basic model to deduplicate and enhance images and text respectively, solving the problems of redundancy and noise between and within samples, and does not rely on high-quality seed data, achieving more efficient and comprehensive data quality improvement.
[0039] In general, current solutions improve data quality or model performance in specific scenarios, but are generally limited by data distribution, diversity, and computational efficiency. This application overcomes these limitations through a dual-branch design of a single-modal model and achieves a wider range of data governance effects.
[0040] This application proposes a data governance framework based on a single-model base model to improve data quality from the perspectives of redundancy removal and denoising. The visual branch uses the base model to evaluate the semantic contribution of each block and reduce redundancy. The text branch uses the base model to rewrite the product title. The framework designed by this application is easy to use and does not require repeated iterations. This application has conducted sufficient experiments in image-text retrieval and multimodal large model benchmarks to prove the improvement of data quality and the enhancement of model performance.
[0041] According to an embodiment of the present application, a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0042] Considering the huge amount of model parameters of the large model and the limited computing resources of the mobile terminal, the above method provided in the embodiment of the present application can be applied to Figure 1 The application scenarios shown are not limited to these. Figure 1 is a schematic diagram of an application scenario of data processing according to an embodiment of the present application. Figure 1 In the application scenario shown, the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client device 20 here may include but is not limited to: a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the large model, thereby implementing the method provided in the embodiment of the present application.
[0043] In an embodiment of the present application, a system composed of a client device and a server can perform the following steps: the client device generates initial multimodal data. The server uses the first network model to screen the first initial modal data to obtain first target modal data and object recognition results, wherein the object recognition results are used to represent the object of the target type contained in the first initial modal data; the object recognition results and the second initial modal data are input into the second network model, and the second initial modal data is enhanced using the second network model to obtain second target modal data; based on the first target modal data and the second target modal data, target multimodal data is obtained.
[0044] It should be noted that, with the rapid development of high-performance computing units, in other application scenarios, the above method provided in the embodiment of the present application can also be applied to the model all-in-one machine. In an optional embodiment, a plurality of models are built into the model all-in-one machine, and the user can choose to adjust a model as needed to obtain the user's own model, so that the high-performance computing unit built into the model all-in-one machine can directly call the adjusted model to execute the above method provided in the embodiment of the present application. In another optional embodiment, a trained model is built into the large model all-in-one machine, so that the high-performance computing unit built into the model all-in-one machine can directly call the model to execute the above method provided in the embodiment of the present application.
[0045] Furthermore, when users need to train their own models, they can also upload their own data sets through the client, which are sent by the client to the server, so that the server can adjust the pre-trained model with the data set to obtain the user's own model and then deploy it to the production environment. In order to facilitate users' needs for model adjustment, the server can provide complete adjustment tools, development frameworks and processes, and can support multiple adjustment strategies, so that the adjusted model can better adapt to applications in different fields and achieve a high degree of customization.
[0046] Under the above operating environment, this application provides Figure 2 The data processing method shown. Figure 2 is a flow chart of a data processing method according to an embodiment of the present application. Figure 2 As shown, the method may include the following steps:
[0047] Step S202, obtaining initial multimodal data;
[0048] The initial multimodal data includes first initial modal data and second initial modal data for describing the object.
[0049] In multimodal applications, it is usually necessary to collect data from two or more different types of modalities to jointly describe an object or scene. These modalities can be visual (images or videos), text (descriptions, titles, comments), audio (speech or background sound), etc.
[0050] The first initial modal data mentioned above can be the first or original visual data describing the object, such as a product picture. These image data directly reflect the appearance, features or usage scenarios of the product, and are the basis for the visual basic model to extract global and local features of the image. Through the processing of the visual basic model, the importance of image blocks can be evaluated, redundant parts can be identified, and category prediction can be performed to provide guidance for subsequent text processing.
[0051] The second initial modal data mentioned above can be the first or original text data describing the same object, such as a product title or description. The text data provides detailed information about the product, such as name, specifications, instructions for use, etc. The text-based model will process these original text data to remove noise and redundant information, enhance the clarity and accuracy of the text description, make it more consistent with the actual information reflected by the image, and improve the consistency and quality of the multimodal data.
[0052] In an optional embodiment, the initial multimodal data obtained can be collected from a related environment, including images of goods and related text descriptions, which may have problems such as redundancy, noise, inconsistency, or missing information. By collecting these initial image and text data, initial multimodal data is provided for the subsequent data governance process. Through the processing of the visual base model and the text base model, the quality of the data set is improved to make it more suitable for training a more efficient and accurate visual-language alignment model. This process is important for improving the performance of multimodal tasks (such as image and text retrieval, conversational question and answer, and image description generation), and is an indispensable step in the multimodal data governance framework.
[0053] Step S204, using the first network model to screen the first initial modality data to obtain first target modality data and an object recognition result;
[0054] The object recognition result is used to represent an object of a target type contained in the first initial modal data.
[0055] The above-mentioned first network model can be a visual base model, whose function is to process and understand image data. It can be any pre-trained visual model. It can be used to identify and extract features of target objects from images. The above-mentioned first network model can also be a text-based model, whose function can be to process text data, and it can be any pre-trained text model, which is not limited here. When the first network model is a visual base model, the above-mentioned first target modal data can be image data obtained after processing by the first network model, which is usually a de-redundant and denoised image feature representation, or a selected set of image blocks.
[0056] The above-mentioned object recognition result can be a classification result of the target type in the image generated by the first network model, including the category of the object and its confidence, to guide the correction and enhancement process of the text base model.
[0057] In an optional embodiment, the first network model can be used to perform recognition analysis on the image. This process can include image segmentation, feature extraction, object detection, etc., to obtain the first target modal data, that is, the processed image features, which are more focused on the key objects or information in the image, and the object recognition results, that is, the type of specific objects recognized from the image and their possible confidence, which is helpful for the correction and supplement of the subsequent text description.
[0058] Step S206, inputting the object recognition result and the second initial modality data into the second network model, and enhancing the second initial modality data using the second network model to obtain second target modality data;
[0059] The second network model can be a text-based model that can be used to process and understand text data. Based on the results of image recognition, the text description can be corrected and enhanced to make it more accurate and richer.
[0060] The above-mentioned second target modal data may be text data obtained after being processed by the second network model. By correcting and enhancing the text, the data describes the objects in the target scene more accurately and in detail.
[0061] In an optional embodiment, the second network model can be used to process text data based on the object recognition results obtained in the previous step. The second network model will use these object information as context clues to denoise, correct errors or enhance details of the text description to generate more accurate, richer and consistent "second target modal data" with the image information. For example, if the image recognizes that the product is a "tea set", the unclear description or typos that may exist in the text will be corrected, and redundant information in the text description that is not related to the "tea set" will also be removed.
[0062] Step S208, obtaining target multimodal data based on the first target modal data and the second target modal data.
[0063] The above-mentioned target multimodal data integrates the processed image features and enhanced text descriptions to form higher quality multimodal data pairs for subsequent model training or application.
[0064] The target multimodal data is constructed by combining the filtered image features (first target modal data) and the corrected text description (second target modal data). This is a set of purer and higher-quality image-text pairs. Compared with the original multimodal data, the target multimodal data is more semantically consistent, clearer in information, and redundancy and noise are significantly reduced. Such a dataset is more suitable as training material for training multimodal alignment models, thereby improving the performance and robustness of the model.
[0065] In an optional embodiment, starting from the original multimodal data, the information of the two modalities of image and text is processed separately through the processing of two basic models of vision and text, and finally integrated into a higher quality multimodal data set. This series of operations is aimed at improving the purity and semantic consistency of the data set, thereby improving the performance of the multimodal model trained based on the data set. For example, taking the e-commerce scenario as an example, through image de-redundancy and text enhancement, product images that are more in line with product information and more accurate product descriptions can be obtained, thereby improving the user experience and system effects of multimodal applications such as product recommendation and search.
[0066] This application achieves the goal of data governance, which is to improve the quality of multimodal data used in model training. Specifically, redundancy and noise are removed from the original data, data diversity is maintained, and data distribution deviation is avoided.
[0067] In e-commerce scenarios, product images and description texts can express product features more accurately and consistently, reducing misunderstandings caused by inaccurate descriptions or redundant images. This process not only improves the performance of multimodal models, such as the accuracy of visual-language alignment models in tasks such as product recommendation and product search, but also reduces the computational cost during training and reduces over-reliance on high-quality seed data, enabling the model to perform better on data sets closer to actual applications, thereby enhancing the robustness and generalization of the model. By fine-tuning product images (removing redundant image blocks and evaluating the probability of object existence) and denoising and enhancing the text of product descriptions (correcting text descriptions based on image information), the target multimodal data obtained is of higher quality and more suitable for training multimodal alignment models. This not only improves the performance of the model, but also improves the efficiency of model training, solving common data quality and technical problems in large-scale data governance.
[0068] Through the above steps, initial multimodal data is obtained, wherein the initial multimodal data includes a first initial modal data and a second initial modal data for describing an object; the first initial modal data is screened using a first network model to obtain first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of a target type contained in the first initial modal data; the object recognition result and the second initial modal data are input into a second network model, and the second initial modal data is enhanced using the second network model to obtain second target modal data; based on the first target modal data and the second target modal data, target multimodal data is obtained, thereby improving the quality of multimodal data used for model training; it is easy to notice that by processing data of different modalities in a targeted manner, the semantic consistency and accuracy of data of different modalities can be improved, and the data quality can be improved, so that when integrating these processed modal data, the constructed target multimodal data set not only reduces redundancy and noise, but also enhances the information consistency between modalities. This enables subsequent models to learn purer and higher-quality cross-modal semantic relationships during training, thereby improving the training efficiency and final performance of the model, and further solving the technical problem of poor data quality used in model training in related technologies.
[0069] In the above embodiment of the present application, the first initial modal data is image data, and the first network model is used to filter the first initial modal data to obtain first target modal data and object recognition results, including: slicing the image data to obtain multiple image blocks; extracting the global feature representation of the image data, and extracting the local feature representation of the multiple image blocks; filtering the multiple image blocks based on the global feature representation and the local feature representation to obtain the first target modal data; performing object recognition on the first target modal data to obtain an object recognition result.
[0070] The first network model mentioned above may be a visual base model, and the global feature representation of the visual base model mentioned above refers to the comprehensive features extracted from the entire image data, representing the overall semantic information of the image. In the present application, the global feature representation is extracted by the visual base model.
[0071] The above-mentioned local feature representation refers to the features extracted from a specific area or image block of the image, reflecting the semantic information of the local area of the image. In this application, the contribution of each image block to the global semantics is evaluated by comparing the global features with the local image block features.
[0072] The first target modality data mentioned above refers to the image data that has been screened and adjusted, retaining the image blocks that contribute most to describing the objects in the scene and removing the redundant and noisy parts. In this application, this is equivalent to a set of image blocks with high semantic contribution selected from the original image.
[0073] The above object recognition results are generated by the visual basic model about the type of target object in the image and its confidence, which are used to guide the correction and enhancement of text description. In this application, the object recognition results provide a discretized expression of image semantics, namely object category and existence probability distribution.
[0074] The image data can be divided into multiple image patches. This process can be regarded as decomposing the image into smaller components so that its semantic information can be analyzed more finely. Then, the global feature representation of the entire image and the local feature representation of each image patch are extracted through the visual basis model.
[0075] The similarity between the local feature representation and the global feature representation of the image block can be compared, and the top k image blocks with relatively high contribution to the global semantics can be screened out as the first target modal data. Finally, object recognition is performed on these screened image blocks to obtain the type of object in each image block and its corresponding confidence, that is, the object recognition result.
[0076] For example, in a picture containing a tea set, through the above steps, image blocks depicting key objects such as tea sets and tea tables can be retained, and clear recognition results about these objects can be obtained, providing a reference for the correction and enhancement of subsequent text descriptions.
[0077] By cutting images into pieces, extracting features, and filtering, we can achieve preliminary redundancy removal and denoising of image data. The key to this step is to use the visual basis model to evaluate the semantic importance of image blocks, thereby retaining the parts that contribute more to the overall scene description while removing redundant and irrelevant information. Through the object recognition results, we can obtain the category and existence probability of the target object in the image, which is an important basis for the discretization of image semantics and provides a clear visual semantic reference for subsequent text processing.
[0078] Exemplarily, the first initial modal data is a picture of a product, and the second initial modal data is a text description of the product. In the first step, the product picture can be processed using a visual base model, divided into multiple image blocks, and the local feature representations of these blocks and the global feature representation of the entire image are extracted respectively. In the second step, by calculating the cosine similarity between the local feature representation and the global feature representation, the image blocks with relatively high contribution are screened out, for example, the parts of the product picture that clearly display key objects such as teapots, tea trays, and tea cups. In the third step, object recognition is performed on the screened image blocks to obtain the recognition results of specific objects in the "tea set", such as identifying teapots, tea trays, tea cups, etc., as well as the confidence of each object, to form an object recognition result. Finally, these adjusted image blocks (first target modal data) and object recognition results will be used for subsequent text branch processing to generate more accurate and consistent multimodal data.
[0079] By performing the above-mentioned processing on the image data, the embodiment of the present application can improve the quality of multimodal data used for model training. First, by cutting and screening the image, the redundant information in the image can be reduced, and the image blocks with semantic value can be retained, which helps to improve the purity and efficiency of the data set. Secondly, by using the comparison of local and global feature representations, the importance of the image blocks can be more accurately evaluated to ensure that the key information that can describe the target scene is retained. Thirdly, the object recognition results not only provide a clear image semantic reference, but also provide guidance for text processing, ensure the consistency of image and text information, and avoid the problem of decreased data diversity in the data governance process. Finally, through the synergy of visual de-redundancy and text enhancement, the accuracy and quality of the data can be improved without sacrificing data diversity, thereby improving the training effect and final performance of the multimodal model, and solving the problem of low quality of model training data in related technologies. For example, in an e-commerce scenario, the adjusted image obtained by image processing and the effective object information identified from the image can guide the correction and enhancement of the text description, generate more accurate and consistent commodity image and text data pairs, and improve the user experience and system efficiency of commodity recommendation and search.
[0080] In the above embodiment of the present application, multiple image blocks are screened based on the global feature representation and the local feature representation to obtain the first target modal data, including: determining the similarity between the semantic information of the image data and the multiple image blocks based on the global feature representation and the local feature representation; screening the multiple image blocks based on the similarity to obtain the target image block, wherein the first similarity of the target image block is greater than the second similarity of other image blocks in the multiple image blocks except the target image block, the first similarity is the similarity between the local feature representation corresponding to the target image block and the global feature representation, and the second similarity is the similarity between the local feature representation corresponding to the other image blocks and the global feature representation; and determining the target image block to be the first target modal data.
[0081] The above-mentioned target image block is an image block that can represent the global semantics among multiple image blocks according to the similarity between its local feature representation and the global feature representation, that is, the “target image block”.
[0082] The first similarity mentioned above is the similarity between the local feature representation of the target image block and the global feature representation, while the second similarity is the similarity between the local feature representation of other image blocks except the target image block and the global feature representation. These two are used to quantify the contribution of the image block to the global semantics.
[0083] The above screening process involves a detailed analysis of the image data to determine which image blocks can reflect the overall semantics of the image. First, a global feature representation is extracted from the original image data, which is a comprehensive understanding of the semantics of the entire image. Then, the image is cut into blocks and a local feature representation is extracted for each image block. Next, the similarity between the local feature representation and the global feature representation of each image block can be calculated, which is usually achieved through cosine similarity. By comparing the similarities, those image blocks that are similar to the global feature representation can be found, namely the "target image blocks". The first similarity of the target image block is higher than the second similarity of other image blocks, which indicates that the target image block is more capable of representing the key information of the entire image. Then, these target image blocks are determined as the "first target modal data" for subsequent data governance and model training. This step ensures that those image blocks that contribute significantly to the overall semantics are retained, significantly reducing redundancy in the data while avoiding data diversity loss.
[0084] Obtain a global feature representation of the image, which is usually provided by the global average pooling layer of the model, describing the overall semantics and content of the image. Next, extract local feature representations of multiple image patches, each of which is generated by the corresponding local layer of the model, such as the convolutional layer or the local attention mechanism, reflecting the details and local semantics of different regions in the image. In order to determine the relationship between the semantic information of the image data and each image patch, the similarity between the global feature representation and each local feature representation can be calculated. This process can be implemented in many ways, but a common method is to calculate cosine similarity, that is, to measure the similarity between vectors by the cosine value of the angle between them. Through the above similarity calculation, the similarity score between each image patch and the global representation can be obtained. These scores reflect the contribution of each patch to the overall semantics of the image, or its importance. A patch with high importance means that the patch is crucial to understanding the content of the image, while a patch with low importance may contain redundant or non-critical information. Finally, based on the importance of the patch, it can be decided how to handle these patches in subsequent data governance and model training. For example, in a visual base model, patches can be selectively retained or deleted based on the similarity score to reduce redundancy, improve image quality and training efficiency. In the text-based model, similarity information can be used to guide the enhancement or correction of text descriptions, ensuring that the text content is consistent with the visual semantics of the image.
[0085] In the above screening step, the image data is adjusted by comparing the global and local features. The global feature representation provides a reference point for the overall semantics of the image, while the local feature representation describes the specific semantics of each block in the image. By calculating the similarity, those image blocks that contribute more to the global semantics, namely the target image blocks, can be identified as the first target modal data. This not only helps to reduce data redundancy, but also ensures that the model can learn the core semantic information from the image, thereby improving model performance.
[0086] For example, taking the e-commerce scenario as an example, suppose there is a product "New Chinese Style Tea Table Combination Solid Wood Office Kung Fu Tea Household Set", and its corresponding multimodal data contains a product picture and the corresponding product description text. In the first step, the pre-trained model can be used to process the product image, extract the global feature representation, and segment the image into multiple blocks, each of which produces a local feature representation. In the second step, the similarity between each image block and the global feature representation is calculated. For example, suppose there is a block in the image that focuses on showing the tea table, another block shows the details of the tea set, and other blocks may contain irrelevant background or repeated information. Through calculation, it is found that the tea table and tea set detail blocks have a high similarity with the global feature representation, so they are regarded as target image blocks. In the third step, these target image blocks are used as the first target modal data for subsequent data governance and model training, which can more effectively retain the key visual information of the product, reduce irrelevant or repeated visual information, and improve data quality and model efficiency.
[0087] Through the above screening steps, the quality of image data used for subsequent multimodal model training can be significantly improved. Specifically, the screening of target image blocks (i.e., the first target modal data) is based on its similarity to the global feature representation, which ensures that only those image blocks that can accurately reflect the core features of the product are retained, thereby reducing the redundancy and irrelevant information in the image data. At the same time, since the screening is performed in the original image data instead of relying on a high-quality cross-modal alignment model, this method breaks through the mutual dependence between high-quality data and high-quality models, lowers the threshold of data governance, and improves data governance efficiency. Without sacrificing data diversity, purer and higher-quality image data can be obtained, which is crucial to improving the performance of multimodal alignment models. When processing large-scale, high-noise real data sets, this data governance strategy can significantly improve the training effect and actual application performance of the model. For example, by removing redundant information in the image that is not related to the product description, the multimodal alignment model can more accurately learn the visual and language features of the product and improve the accuracy of the product retrieval and recommendation system.
[0088] In the above embodiment of the present application, object recognition is performed on the first target modal data to obtain an object recognition result, including: performing semantic recognition on the first target modal data to obtain multiple types of objects and confidence levels of the multiple types of objects; and screening multiple types of objects based on the confidence levels to obtain an object recognition result.
[0089] The above-mentioned semantic recognition refers to using a visual basic model to conduct an in-depth analysis of image data (first target modality data) to identify various object types and their semantic information contained in the image.
[0090] The above confidence level is the confidence level of the model that the object type identified in a certain image block is correctly identified, usually expressed as a value between 0 and 1. The larger the value, the higher the confidence level, indicating that the model believes that the recognition result is more reliable.
[0091] The above screening is based on the type of object and the corresponding confidence level, and selects the object type with high confidence level from the semantic recognition results to form an "object recognition result".
[0092] The first target modal data can be semantically recognized, which utilizes the deep learning capabilities of the visual base model. The visual base model can extract features from the image and identify different object types. During the recognition process, the model not only recognizes the object, but also assigns a confidence to each object to quantify the model's certainty of the recognition result. Then, the multiple types of objects recognized and their confidences can be screened, and only those objects with higher confidences are retained. The recognition results of these objects constitute the "object recognition results". For example, when processing product images, if the image block contains "tea cup" and "tea table", and the confidences given by the model are 0.9 and 0.65 respectively, then "tea cup" will be retained first due to its higher confidence and become part of the object recognition result. The values here are only for example and are not specifically limited.
[0093] Object recognition is the process of in-depth analysis of image data by the visual basic model. It aims to identify and extract objects with high semantic value and their confidence in the image. By screening high-confidence objects, a purer and more accurate set of image information, namely object recognition results, can be obtained, which provides key visual semantic support for subsequent text processing and data management.
[0094] For example, in an e-commerce scenario, for the image data of the product "New Chinese Style Tea Table Combination Solid Wood Office Kung Fu Tea Household Set", the target image blocks have been segmented and screened out through the model, and these blocks contain the core visual information of the product. Next, these image blocks will be semantically recognized, and the visual basic model will be used to identify the object type in each image block, such as "tea table", "tea set", "office decorations", etc., and a confidence level will be assigned to each identified object. Assume that in the recognition results, "tea table" and "tea set" have confidence levels of 0.95 and 0.8 respectively, while the confidence level of "office decorations" is 0.35. Based on the screening of confidence levels, only "tea table" and "tea set" are retained as object recognition results because they have higher confidence levels and are more likely to be the core objects in the product images. These object recognition results will be used to guide subsequent text description corrections and enhancements to ensure the consistency and accuracy of graphic and text information. The values here are only for example explanations and are not specifically limited.
[0095] Through the above-mentioned object recognition and screening steps, the quality of image data used for model training can be significantly improved, and irrelevant or low-confidence object information can be eliminated, thereby ensuring the purity and accuracy of image data. The object recognition results not only provide a clear visual information reference for the text processing branch, but also help the text description to more accurately reflect the characteristics of the product, which is crucial to improving the performance of the multimodal alignment model. For example, for the product search and recommendation functions of e-commerce, more accurate image information can help the model understand the product content more accurately, thereby improving the accuracy of product search and the personalization of recommendations to improve the user shopping experience. In this series of processing, not only the redundancy and noise in the image data are removed, but the data is also improved through the knowledge of the unimodal basic model, avoiding the problem of decreased data diversity in the traditional data governance process, and improving the efficiency of data governance and the training effect of the model.
[0096] In the above embodiment of the present application, the second initial modal data is text data, the object recognition result and the second initial modal data are input into the second network model, and the second initial modal data is enhanced using the second network model to obtain second target modal data, including: correcting the text data in the initial multimodal data based on preset prompt words to obtain corrected text data; enhancing the text data based on the object recognition result to obtain enhanced text data; and determining the second target modal data based on the corrected text data and the enhanced text data.
[0097] The above-mentioned preset prompt words can be a series of pre-set keywords or phrases in text processing in order to guide the text model to perform specific types of corrections or enhancements. They can be generated based on the object recognition results to prompt the model how to modify the text.
[0098] The above-mentioned modified text data can be to adjust the text data based on the preset prompt words, aiming to remove redundant information in the text, correct erroneous descriptions or fix grammatical problems, so as to obtain a more accurate and concise text description.
[0099] The enhanced text data mentioned above can supplement and enrich the text data based on the object recognition results, aiming to improve the details and clarity of the text description, ensure the consistency of the text and image information, and obtain a more comprehensive and richer text description.
[0100] The above-mentioned second target modal data can be the integration result of the corrected text data and the enhanced text data to form an adjusted text description, which can be the product of the processed text part in the initial multimodal data.
[0101] The text data can be processed using a text-based model to improve the quality of the text and its consistency with the image information. First, the text data is corrected based on preset prompt words, which usually involves removing duplicate information, correcting incorrect descriptions (such as typos, grammatical errors), and adjusting the text style to make the text description more accurate, clear, and concise. Next, the text data is enhanced based on the object recognition results, that is, based on the list of high-confidence objects identified from the image, the missing details in the text description are supplemented and improved to ensure that the text description fully reflects all key objects in the image. Finally, by integrating the corrected text data and the enhanced text data, the "second target modality data", that is, the adjusted text description, is obtained, which not only removes the redundancy and errors in the original text, but also adds detailed information about the image objects, improving the cross-modal consistency between text and image.
[0102] The second network model mentioned above, namely the text-based model, provides a powerful tool for denoising, correcting and enhancing text data. The introduction of preset prompt words, based on the object recognition results, provides clear guidance for the text-based model when correcting text, and can more accurately remove errors and redundancies in the text and improve the quality of the text. At the same time, text enhancement based on object recognition results ensures the integrity of text information and enhances the relevance of text and image by supplementing the description of high-confidence objects identified in the image, thereby improving the overall quality of multimodal data.
[0103] For example, in an e-commerce scenario, the product is "New Chinese Style Tea Table Combination Solid Wood Office Kung Fu Tea Household Set", and the initial text data may contain redundant information and description errors, such as "Tea Table Combination Solid Wood Office Kung Fu Tea Household Set Integrated Balcony Small Tea Brewing Table". Based on the preset prompt words, the text-based model is used to correct repeated or irrelevant information in the text. For example, words such as "integrated" and "small" that may appear repeatedly in the product title are deleted to obtain the corrected text data: "New Chinese Style Tea Table Combination Solid Wood Office Kung Fu Tea Household Set Balcony Tea Brewing Table". Next, based on the object recognition results (such as recognition of "tea set", "tea table", etc.), the text-based model can be used to enhance the text description, supplement missing details or correct inaccurate descriptions, and obtain enhanced text data: "New Chinese style tea table combination, including solid wood tea table, Kung Fu tea set, suitable for office and home, can be used for making tea on the balcony." Finally, the corrected text data and the enhanced text data are integrated to obtain the second target modal data, that is, the adjusted text description: "New Chinese style tea table combination solid wood office Kung Fu tea home set, including solid wood tea table, Kung Fu tea set, suitable for office and home, can be used for making tea on the balcony".
[0104] Through the steps of text processing, the quality of text data and the consistency of image information can be significantly improved without sacrificing data diversity. Specifically, the corrected text data eliminates redundancy and erroneous descriptions in the text, while the enhanced text data supplements the missing details provided in the image recognition results, making the text description more complete and accurate. This two-pronged strategy not only improves the purity of the text description, but also ensures the close connection between text and image information, thereby improving the training effect and actual application performance of the multimodal alignment model. For example, for applications such as image-text matching and product recommendations, the adjusted text description can more accurately reflect product features, reduce user confusion caused by inconsistent information or unclear descriptions, and improve the user's shopping experience and the accuracy of system recommendations.
[0105] The above-mentioned processes, from image segmentation, feature extraction, object recognition, to text correction and text enhancement, are finally integrated into adjusted multimodal data, aiming to solve the key problem of multimodal data governance in scenarios such as e-commerce, namely how to screen and adjust high-quality image-text pairings in massive data to improve the training efficiency and performance of the model. By integrating the processing results of the visual basic model and the text basic model, not only the redundancy and noise in the data are removed, but also the correlation between images and text is enhanced, achieving a comprehensive improvement in data quality. This is of great significance to promoting the application of multimodal models in actual scenarios and improving user experience and system efficiency.
[0106] In the above embodiments of the present application, text data is enhanced based on the object recognition result to obtain enhanced text data, including: determining object information corresponding to an object of a target type contained in the object recognition result; determining associated text data associated with the object information in the text data; and enhancing the associated text data in the text data to obtain enhanced text data.
[0107] The above-mentioned object information may be the specific information carried by each target type of object in the object recognition result, including the type of object, possible attributes (such as color, material), location information, etc. This information is an important basis for text enhancement.
[0108] The above-mentioned associated text data can be the part of the text data related to the object information, which can be the words that directly describe the object, or the text content that indirectly mentions the object's attributes or location. This part of the data will be used for enhancement to improve the accuracy and richness of the text description.
[0109] The enhanced text data mentioned above may be text data obtained by supplementing and correcting the associated text data through the object information, which generally involves adding missing details, correcting errors in the description, and enhancing the clarity and completeness of the text description associated with the object.
[0110] Based on the object recognition results, the corresponding object information of each identified target type object is first determined, which may include the object's category, color, material and other attributes. Next, the text content associated with these object information in the text data can be found, that is, the associated text data. For example, if the image recognizes "tea set", you can search whether the text description mentions information related to the tea set, such as "teapot", "tea cup" and other words. Then, the associated text data is enhanced using the text-based model, which may involve adding missing details (such as the material of the tea set), correcting errors in the description (such as typos), and enriching the text description associated with the object (such as adding "tea set consistent with the picture, including exquisite teapot, tea cup, tea tray") to ensure that the text description fully reflects the key information identified in the image. Finally, the enhanced associated text data is integrated to obtain the adjusted "enhanced text data", which not only contains the information in the original text, but also supplements the object details identified in the image, improving the accuracy and richness of the text description.
[0111] The above process aims to guide the enhancement of text descriptions through image recognition results to ensure the consistency and completeness of text and image information. By supplementing missing details related to objects in the image and correcting errors in text descriptions, enhanced text data can describe product features more accurately and in detail while maintaining the original information of the text, thereby improving the overall quality of multimodal data.
[0112] For example, in an e-commerce scenario, the product is "New Chinese style tea table combination solid wood office Kung Fu tea household set", and the initial text data may contain insufficient descriptions, such as: "Tea table combination solid wood office Kung Fu tea household set integrated balcony small tea brewing table". Based on the object recognition results (such as "tea table", "tea set", etc.), the object information related to these objects is determined. Then, the parts of the text data that are related to these object information are found, such as keywords such as "tea table", "tea set", etc. Next, using the text-based model, these associated text data can be enhanced. For example, the specific details of "tea table" and "tea set" can be supplemented to obtain: "New Chinese style tea table combination, including high-quality solid wood tea table, matched with Kung Fu tea set set, suitable for office and home use, and can be used for balcony tea making." Finally, the enhanced text description can be integrated to obtain the adjusted enhanced text data: "New Chinese style tea table combination, made of high-quality solid wood, matched with a complete Kung Fu tea set set, including teapots, teacups, tea trays, etc., suitable for office and home environments, can be easily placed on the balcony to make tea." This text enhancement makes the description more comprehensive and accurate, more consistent with the image information, improves the clarity and attractiveness of the product information, and is expected to enhance the user's search and purchasing experience.
[0113] By enhancing text data based on object recognition results, the accuracy and completeness of text descriptions can be significantly improved, which is crucial to improving the overall quality of multimodal data, improving model training effects and practical application performance. Specifically, enhancing text data not only removes redundancy and errors in text descriptions, but also increases the details and richness of text by supplementing the detailed information of high-confidence objects identified in images, thereby improving the cross-modal consistency between images and text.
[0114] In the field of e-commerce, this text enhancement can make product descriptions more accurate, so that users can get more detailed and consistent information when searching and browsing products, thereby improving the efficiency and satisfaction of purchasing decisions. At the same time, for multimodal models (such as visual-language alignment models), training based on enhanced text data can help the model understand the semantics of images and text more accurately, improve the performance and robustness of the model, and thus improve the effects of applications such as recommendation, search, and question-answering based on multimodal data. For example, for a product recommendation system, the adjusted image and text data pairs can more accurately reflect the user's interests and product features, thereby providing more personalized recommendation results, improving the user's shopping experience and the accuracy of system recommendations.
[0115] In the above embodiment of the present application, the method also includes: using the target multimodal data to train the initial multimodal alignment model to obtain a target multimodal alignment model, wherein the target multimodal alignment model is used to align data of multiple modalities in the target multimodal data, and the alignment processing is used to represent mapping data of multiple modalities in the target multimodal data to the same semantic space.
[0116] The pairing of image and text data obtained from the above target multimodal data has higher consistency and accuracy, and is suitable for training a multimodal alignment model.
[0117] The above-mentioned initial multimodal alignment model is a model used for preliminary processing of image and text data before data governance. It may contain certain deviations or insufficient performance and needs to be further trained and improved through adjusted data.
[0118] The target multimodal alignment model is an updated model obtained by training the initial multimodal alignment model using target multimodal data, and is more accurate and powerful in processing the consistency of images and texts and understanding product features.
[0119] After completing the data governance process, you can enter the model training phase. The goal of this phase is to use the governed high-quality target multimodal data to train and improve the initial multimodal alignment model, so as to obtain a target multimodal alignment model with better performance. First, ensure that the target multimodal data is ready. This includes the adjusted image patches (after de-redundancy and denoising), as well as the corrected and enhanced text descriptions, which constitute the paired image and text data. Input the target multimodal data into the initial multimodal alignment model for training. This process involves the model encoding the image and text data, and adjusting the model parameters through a loss function (such as contrastive loss) so that the feature representations of the image and text are more closely aligned in the multimodal space. By training on the target multimodal data, the model can more accurately understand and represent the image and text information, reduce the impact of data noise, and improve the performance of cross-modal matching.
[0120] The results of data governance can be applied to model training. By using high-quality and highly consistent target multimodal data, the performance of the initial multimodal alignment model can be improved, making it more accurate in processing semantic understanding of images and texts, cross-modal matching, etc., thereby obtaining the target multimodal alignment model.
[0121] The performance of the target multimodal alignment model trained with target multimodal data has been significantly improved. Since the data governance process effectively removes redundant information and noise and increases the consistency between images and text, the target multimodal alignment model can more accurately understand product features and improve the accuracy of product retrieval and recommendation. Specifically, the adjusted model can more accurately capture the core information of the product and reduce information bias when processing such as matching product images and titles. For example, for the product "New Chinese Style Tea Table Combination Solid Wood Office Kung Fu Tea Home Set", the adjusted model can accurately distinguish the characteristics of the tea table and tea sets, and provide high-quality search and recommendation results even when the image quality is not high or the text description is incomplete, significantly improving user experience and system efficiency.
[0122] The above-mentioned alignment processing means that the model can represent the semantic information in the image and text in a consistent way through learning, that is, make the corresponding features of the image and text as close as possible in the shared semantic space.
[0123] After training, the target multimodal alignment model can align new multimodal datasets, that is, map images and texts into a shared semantic space, which is crucial for the execution of various multimodal tasks, such as image and text retrieval, visual question answering, image description generation, etc. It breaks the boundaries of modality and enables the model to understand and process cross-modal information, thereby performing more complex tasks in the open world.
[0124] Through the governance of high-quality data, the performance of the multimodal alignment model can be effectively improved, breaking the interdependence between high-quality data and high-quality models, making model training more efficient, while ensuring data diversity and model robustness. It is of great value in promoting the application of multimodal technology in e-commerce, search, recommendation and other fields.
[0125] In the above embodiment of the present application, the target multimodal data is used to train the initial multimodal alignment model to obtain the target multimodal alignment model, including: using the initial multimodal alignment model to align the target multimodal data to obtain third modal data and fourth modal data corresponding to the third modal data, wherein the third modal data and the fourth modal data are used to describe the same object; constructing a contrast loss function of the third modal data and the fourth modal data; and adjusting the model parameters of the initial multimodal alignment model based on the contrast loss function to obtain the target multimodal alignment model.
[0126] The third modality data and the fourth modality data mentioned above respectively represent image and text modality data after the target multimodal data alignment process in multimodal data processing, but they have been mapped to the same semantic space, so that they can be compared and matched at the semantic level.
[0127] The above-mentioned alignment processing refers to the process of mapping data from different modalities (such as images and texts) into the same semantic space in multimodal alignment technology, so that data from different modalities can be effectively compared and aligned in the space.
[0128] The above contrast loss function is used to quantify whether the feature representations of the same sample in different modalities match in the semantic space when training the multimodal alignment model. This function usually encourages different modal representations of the same sample to be close, while modal representations of different samples are far apart.
[0129] After completing data governance and obtaining the target multimodal data, we can enter the stage of multimodal alignment model training. First, the target multimodal data (adjusted image and text pairs) can be input into the initial multimodal alignment model for alignment processing. In this process, the image and text are encoded separately and mapped to the same semantic space to generate third modal data and fourth modal data. Next, a contrastive loss function can be constructed to measure the semantic distance between the third modal data and the fourth modal data, as well as the distance between them and the modal representations of other samples, to ensure that the image and text feature representations from the same sample are closely aligned in the semantic space, while the features from different samples maintain a certain distance. Based on the contrastive loss function, the parameters of the initial multimodal alignment model can be adjusted through back propagation to minimize the loss, thereby obtaining a better target multimodal alignment model.
[0130] The adjusted multimodal dataset can be used to train and improve the initial multimodal alignment model by constructing and minimizing the contrastive loss function. The adjusted model can more accurately align image and text features in the same semantic space, improving the model's multimodal understanding and matching capabilities.
[0131] For example, in the e-commerce scenario, the target multimodal data, namely the adjusted image and text description of "New Chinese Style Tea Table Combination Solid Wood Office Kung Fu Tea Household Set", has been obtained through data governance. Next, the initial multimodal alignment model is used to align these target multimodal data, converting the image features into third-modal data and the text features into fourth-modal data, so that they can be compared in the same semantic space. By constructing a contrast loss function, the model parameters are adjusted to minimize the semantic distance between these feature representations, while ensuring that a certain distance is maintained from the feature representations of different samples. During the training process, the adjusted model can better understand and characterize the features of the tea table combination and Kung Fu tea set. Even in the case of poor image quality or incomplete text description, the model's alignment ability in the semantic space can provide accurate product retrieval and recommendation results.
[0132] By using the target multimodal data for training, the performance of the initial multimodal alignment model can be significantly improved, and the target multimodal alignment model can be obtained. The adjusted model can more accurately understand product features when processing image and text data, and improve the accuracy of product search and recommendation. Specifically, the construction of the contrast loss function ensures that the model can effectively align image and text features in the semantic space, reduces the impact of data noise and redundancy on model performance, and improves the robustness and generalization ability of the model.
[0133] For example, in a product recommendation system, the target multimodal alignment model can more accurately capture the semantics of user search and browsing behaviors, and provide product recommendations that better match user preferences, thereby improving user satisfaction and purchase conversion rates. In the search function, even if the keywords provided by the user do not completely match the product title, the model can provide closer product search results based on the semantic alignment of images and texts, thereby improving search efficiency and user experience. In this series of processing, not only the data quality is improved, but also the multimodal understanding and matching capabilities of the model are significantly enhanced.
[0134] Figure 3 is a structural diagram of a multimodal data processing method according to an embodiment of the present application, such as Figure 3 As shown, the original data can be simultaneously input into two different versions or stages of the multimodal data alignment model, which are marked as visual language alignment model 1 (VLA-v1) and visual language alignment model 0 (VLA-v0), representing different cycles or versions of model training. The training and filtering steps are designed to improve the matching quality of image-text pairs and the overall utility of the dataset. The specific steps are as follows:
[0135] Input raw data. The raw dataset contains images and corresponding text descriptions, which may contain poor image-text matching, redundancy and noise.
[0136] To train the VLA model (VLA-v0) using seed data, first, a small number of high-quality image-text matching pairs are needed as "seed data". The seed data is used to initially train the VLA model (VLA-v0) to establish the model's basic visual-language alignment capabilities.
[0137] Evaluation and screening of data, the VLA-v0 model after initial training is used to evaluate the quality of multiple image-text pairs in the dataset, and identify low-quality pairs by calculating the degree of match between the image and the text description. This process screens out a batch of higher-quality data pairs from the original data as an improved dataset.
[0138] Iterative training (VLA-v1), retraining the VLA model with the filtered high-quality dataset to generate the VLA-v1 version of the model. This process usually improves the performance of the model because it is trained on higher quality data. At this stage, "data v1" is the dataset obtained after evaluation and filtering by the VLA-v0 model, which is used to train the VLA-v1 model.
[0139] Data evaluation and screening (generating data v2),The VLA-v1 model is again used to evaluate each sample in the dataset, identify its matching quality and potential redundancy or noise, and further screen out higher quality data pairs. The result of this step is a new high-quality dataset, "data v2".
[0140] Repeat the training and filtering process. After the training process is completed, the VLA-v1 model performs data evaluation, and the filtered data pairs become "data v2" and are used to train the model again to generate the VLA-v2 model. This process can be iterated multiple times until the model performance reaches the expected level or the data quality is high enough.
[0141] After multiple rounds of data evaluation and screening, a high-quality dataset is finally generated, in which the image-text pairs have a high matching degree and reduce redundancy and noise. Through multiple rounds of training on the high-quality dataset, an adjusted VLA model (e.g., VLA-vn) is finally obtained, which has more accurate visual-language alignment capabilities and can better handle multimodal data.
[0142] The multimodal Bootstrap method in this application starts with a small amount of high-quality "seed data" and gradually improves the performance of the VLA model and adjusts the quality of the dataset through multiple iterations of training and screening. Each iterative cycle includes evaluating data quality using the current model version, screening high-quality data pairs, and then training the next model version using the screened data. This method is particularly effective when dealing with large-scale multimodal datasets because it allows the model to focus on data that is helpful for learning while reducing the impact of noise and redundancy, ultimately outputting a high-quality multimodal dataset and a VLA model with relatively good performance. However, the entire process needs to be carefully controlled to avoid data distribution deviation or over-screening to ensure the generalization ability of the model in various applications.
[0143] Visual Language Alignment Model 1 and Visual Language Alignment Model 0 each process the raw input data to evaluate the contribution of multiple image patches to the overall semantics, as well as the clarity and accuracy of the text description. Through such an evaluation, data that has a positive impact on model performance can be screened out, while redundant image patches and text noise can be reduced. There is a dependency between VLA-v1 and VLA-v0, that is, the output of one version can affect the input of another version, forming a cyclic iterative process that aims to continuously update the dataset and improve the overall quality of the training data. The final output data has been processed through multiple iterations, which should theoretically reduce redundancy and noise, improve the matching quality of images and texts, and thus be more suitable for training more powerful VLA models.
[0144] Figure 4 is a schematic diagram of a governance framework according to an embodiment of the present application, such as Figure 4 As shown in the figure, the unified governance framework is designed to improve data quality from the perspectives of redundancy removal and denoising. It does not rely on the iteration of the cross-modal alignment model, but processes image and text data using the visual base model and text base model respectively, thereby improving the performance of the cross-modal alignment model. The framework consists of two parts, the visual branch and the text branch.
[0145] The visual branch first uses a pre-trained model that can be trained on massive unsupervised visual data to extract the global semantic embedding of the image and the local embedding of different image blocks. By calculating the cosine similarity between the image block embedding and the semantic embedding, the contribution of different image blocks to the overall semantics is evaluated and their importance scores are determined. Image blocks with higher importance scores will be retained, while irrelevant or redundant blocks will be removed. Based on the importance scores, the visual manager retains more relevant image blocks, reduces redundancy, and improves image quality. Image blocks with low importance scores are evenly sampled to ensure data diversity. Using the streamlined image blocks, the category predictor extracts discretized semantic information, that is, the category of the object and its probability of existence, to provide guidance information for the text branch to ensure the consistency between the image semantics and the text description.
[0146] The text branch processes the text description through a large language model to correct information such as repetition, noise, and grammatical errors, thereby improving the clarity of the text. Based on the category prediction information provided by the visual branch, the text branch uses the knowledge in the large language model to supplement and enhance the text description and improve the cross-modal consistency between text and images.
[0147] The image-text pairs processed by the visual branch and the text branch are integrated together to generate a less redundant and lower-noise dataset for training the VLA model. This dataset has been processed to not only reduce the redundancy and noise in images and texts, but also maintain the diversity of the data, providing higher-quality training data for the VLA model.
[0148] The dataset generated by this application trains the VLA model from scratch to improve its ability to align images and texts. The model architecture can be a typical VLA model, and the training goal is to match the visual and text representations from the same sample, and enhance the model's cross-modal understanding ability by reducing inconsistent parts.
[0149] This application does not rely on high-quality seed data and the initial multimodal alignment model, which lowers the threshold for data governance and is more suitable for actual industrial scenarios. By processing images and text separately, it avoids repeated iterations of multimodal data, reduces data governance costs and computing resource requirements. While eliminating redundancy and noise, it maintains data diversity and avoids data distribution deviation through balanced sampling and text enhancement.
[0150] The above description provides an efficient data governance method that does not rely on high-quality seed data, which can effectively improve the quality of multimodal data and then adjust the performance of the VLA model trained based on these data.
[0151] Figure 5 is a schematic diagram of a data governance framework structure of a single-modal basic model according to an embodiment of the present application, such as Figure 5 As shown in the figure, it is divided into two stages. Stage 1 is mainly for data governance, and stage 2 is mainly for training. In stage 1, it includes a visual branch and a text branch. The visual branch processes the images in the original data to extract global semantic embedding and local embedding of different image blocks. The text branch uses a large language model to process text data and identify and correct noise information in the text. In stage 2, pre-training is mainly carried out. The processed image-text pairs can be used as input to train the VLA model. The typical VLA model architecture is adopted, including an image encoder and a text encoder for processing image and text data. Contrastive learning is used as the training goal, aiming to match the visual and text representations from the same data sample, reduce cross-modal inconsistencies, and improve the performance of the model through text rewriting and enhancement.
[0152] The visual branch is also used for image segmentation to remove redundancy and discretize semantics, and image block importance estimation. By calculating the cosine similarity between the image block embedding and the global semantic embedding, the contribution of different image blocks to the global semantics is evaluated. Image blocks with higher importance scores will be retained, while redundant or low-contribution image blocks will be removed. Image block selection can retain the first k image blocks with relatively high importance scores, while performing balanced sampling of image blocks with low importance scores to maintain data diversity. Visual category prediction can use category predictors to extract discretized semantic information, i.e., object categories and their existence probabilities, from the streamlined image representation to provide guidance for the text branch. The text branch is used for text denoising and enhancement, and can use the knowledge in the large language model to correct possible repetitions, noise, and typos in the text. Combined with the category prediction provided in the visual branch, the knowledge in the large language model is used to supplement and enhance the text description to ensure the consistency and completeness of image and text information.
[0153] Through the above process, a set of de-redundant and denoised image-text pairs are generated, which can improve data quality while maintaining data diversity, making the data more suitable for training the VLA model.
[0154] Figure 5 The entire process starts from inputting raw data, preprocessing images and texts through the visual branch and text branch respectively, and then integrating the results of the two branches to generate higher quality training data. Finally, the processed data is used to train the VLA model, aiming to improve the model's understanding and alignment capabilities of multimodal data through de-redundancy and denoising. The advantage of this application is that it can effectively utilize the knowledge of the single-modal basic model, avoiding the reliance on a large amount of high-quality seed data in traditional data governance methods, lowering the threshold for data governance, while maintaining data diversity when dealing with data redundancy and noise, and improving the generalization ability of the model.
[0155] This application designs a data governance framework based on a unimodal basic model. For a data governance framework, the key is to deal with redundancy and noise problems in the data, while breaking the interdependence of high-quality data and high-quality models, so as to achieve higher data quality improvement. The framework proposed in this application consists of two branches-visual network and text network. The visual branch reduces redundancy by removing semantically irrelevant image blocks and evaluates the probability of the existence of objects in the scene, effectively discretizing the image semantics. Subsequently, the text branch uses visual information and is guided by the existence of objects to improve the clarity and accuracy of the text description. Through these two branches, the image-text pairs generated by this application have less redundancy and lower noise, and are more suitable for pre-training of VLA. The pre-trained VLA model can then be applied to various downstream tasks, and its performance can be used as an effectiveness indicator of the data manager.
[0156] The visual branch reduces image patch redundancy by evaluating the contribution of multiple image patches to the overall semantics, and provides the probability of object existence through category prediction to guide the enhancement of text description.
[0157] In the data enrichment framework, we first evaluate the contribution of multiple image patches to the overall semantics. The visual manager is initialized with a pre-trained model to extract the global semantic embedding of the image and the local embedding of each image patch. The importance score of the image patch is determined by calculating the cosine similarity between the image patch embedding and the local embedding. The image patch with a higher importance score is considered to contribute more to the global semantics and is thus retained first. Figure 5 As shown in the first stage in Figure 1, for a given image, the model can first extract image patches that are more representative of the product's visual semantics. These important image patches are also areas with dense semantic information. This in-sample semantic importance modeling leverages the knowledge gained by the model pre-trained on massive unsupervised visual data, and has a more accurate estimate of the data distribution compared to data managers trained on limited seed data.
[0158] After the importance estimation is completed, the visual manager retains the top k image blocks with significant contributions and removes the remaining irrelevant or redundant blocks. The importance scores corresponding to multiple image blocks can be drawn using an importance graph. Optionally, the darker areas are image blocks with high importance. In this way, the streamlined image only contains the semantically critical parts, thereby reducing redundancy and improving image quality. After fewer image blocks are input to the visual encoder, the number of tokens required for visual information encoding is reduced, thereby reducing the required computational cost and video memory usage. At the same time, in order to take into account global semantics, some image blocks with low importance scores are also sampled evenly. By modeling the redundancy of image blocks within the sample, data redundancy is reduced while ensuring data diversity as much as possible. Avoid excessive de-redundancy, which will lead to a decrease in data diversity, thereby reducing model performance and robustness to unseen scenes.
[0159] During the training process, information redundancy is reduced by reducing the input of image blocks. During the reasoning process, for ease of use and to retain all semantics, all image blocks can be input into the visual encoder. Through the image block selection strategy in training, this different processing of images in training and reasoning still allows the pre-trained model to maintain relatively good performance on downstream tasks.
[0160] At the same time, due to the differences in the intrinsic information structure of image and text modalities, the redundancy of image semantics is much greater than that of text, and there are many visual semantics that do not have a good correspondence with the corresponding text semantics. Based on the training goal of contrastive learning, the removal of irrelevant visual semantics such as unimportant image blocks can also help enhance the cross-modal consistency of image and text. Therefore, by removing irrelevant image blocks, the parts with poor consistency in visual semantics are reduced, and the data quality is improved. Due to the redundant characteristics of image information, visual semantics can provide very rich and precise semantic details and have stronger robustness to noise. Therefore, the overall semantics of the sample can be extracted through visual signals, and text with insufficient details and noise can be enhanced.
[0161] The category predictor extracts discretized semantic information from the streamlined image representation, namely the object category and its probability distribution. These category predictions not only effectively represent the image semantics, but also provide guidance for the text branch to ensure that the image semantics are consistent with the text description. In this way, the visual manager converts the original visual semantics into clear category insights, providing a more refined input for multimodal pre-training. For example, the category information and confidence values provided by the visual category prediction can be reconfirmed and corrected. In addition, for product descriptions that lack details, the corresponding text descriptions can be supplemented by visual categories.
[0162] The above-mentioned visual branch aims to use the visual basic model to process the data redundancy and noise existing in the image information. In addition, the text branch of this application uses the knowledge in the text basic model to enhance the text description in the sample through cross-validation between the semantics provided by the visual category and the initial text. This process can be summarized as text enhancement. In addition, in order to reduce the impact of noise in the input text, the text manager performs denoising based on the context of the text with the assistance of visual semantics. The specific processing of text data includes text denoising, text enhancement, and pre-training.
[0163] For text denoising, such as Figure 5 As shown in the figure, the design of the text branch uses a universal large language model to correct text noise through the massive knowledge in the large language model. Since the large language model has achieved outstanding results in a variety of open language tasks, the language model can correct repetition, noise, typos and other information through the context of the initial text, and realize text denoising at the language expression level. In addition, text denoising has a certain degree of flexibility, and the large language model can adjust the semantic granularity, style and other characteristics of the text description by constructing different prompt words.
[0164] For text enhancement, in the visual branch, the category prediction process converts redundant visual semantics into discrete object categories, providing clear semantic insights. Prompt words can be constructed, containing the category name with high confidence values provided in the category prediction, and the corresponding confidence. The category prediction given by the visual branch is used as the global semantic reference of the sample to supplement and enhance the details in the text. Using the knowledge in the large language model, a reasonable text description can be inferred based on the overall semantics of the sample. When the category prediction provides less information, only some object details are supplemented. When the visual branch provides more category prediction information, combined with the knowledge of the large model, more accurate scene inference can be performed, thereby enhancing more complex text descriptions.
[0165] For pre-training, by improving the dataset, we can remove redundancy in image blocks and enhance text, thus achieving the goal of removing data redundancy and noise. We can train the visual-language alignment model from scratch on new data.
[0166] In terms of model architecture, image-text training data can be given to train a model to uniformly encode image and text data. The basic components of the model are image encoder and text encoder. The method of this application is also applicable to data management of other image-text alignment models. Specifically, for a certain image-text, it can be encoded with image / text encoders respectively to obtain visual / text representations.
[0167] In terms of training objectives, Information Noise-Contrastive Estimation (Info NCE) can be used as the image-text contrast learning loss, whose goal is to match the visual and text representations from a sample. Specifically, the cross-modal dot product similarity of the current visual / text embedding and multiple samples in the batch can be calculated to obtain the probability distribution of the activation function (Soft Max). Thus, the mutual information contrast loss (ITC) is obtained.
[0168] The basic model-based data governance framework involved in this application is a general idea in the multimodal pre-training scenario. The visual and text basic models themselves and the specific image and text processing details can be modified and replaced. For example, the basic model in the network can also be replaced with other types of basic models. The image and text processing method is not limited to image block sampling and text enhancement, and can be other modifications made using the knowledge provided by the basic model.
[0169] This application abandons the cross-modal basic model and turns to the use of a single-modal basic model to build a data governance framework based on a single-modal basic model. The sample can be rewritten through the basic model to remove the noise and redundancy inside the sample. Deleting image blocks and rewriting and enhancing text content effectively avoids the data offset problem caused by excessive deletion of samples. And the goal of data de-redundancy and denoising is achieved.
[0170] This application does not require a large amount of high-quality graphic data as a prerequisite, which lowers the threshold for data governance and is closer to actual scenarios. However, both text search and photo search rely on representation models to more accurately model product descriptions and product images. In this scenario, a data governance solution that does not rely on human annotation construction is required, otherwise the cost will be unbearable. In this application, the knowledge of the unimodal basic model can be introduced, and two branches, image and text, are designed to process the redundancy in the image slices and the noise in the text, respectively, so as to achieve effective data de-redundancy and denoising. In other words, for the framework of this application, there is a dependence on the unimodal basic model, and the data required for the training of the unimodal basic model is easier to obtain, and does not require complex data pipelines for repeated iterative processing. While utilizing the knowledge of the basic model, the entire framework is easier to implement.
[0171] In ultra-large-scale data scenarios, data governance usually has implementation cost issues and brings additional overhead. When designing this framework, the model's multiple iteration loop is opened and divided into two branches, visual de-duplication and text enhancement. Visual de-duplication mainly considers how to de-duplication the information structure of the image and use its high semantic accuracy to extract the overall semantics of the sample and provide category prediction as reference information for text enhancement. In the text enhancement branch, for samples whose confidence in category prediction is not very high, this application uses the context of the text to rewrite and enhance the text, including grammatical errors, repeated expressions and other issues. At the same time, for samples that provide sufficient confidence values for category prediction, the knowledge in the text basic model can be combined to infer and enhance the text details.
[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0173] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0174] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0175] According to an embodiment of the present application, a data processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here. Figure 6 is a flow chart of a data processing method according to an embodiment of the present application, such as Figure 6 As shown, the method includes:
[0176] Step S602, in response to an input instruction acting on the operation interface, displaying initial multimodal data of the target scene on the operation interface;
[0177] The initial multimodal data is used to describe the objects contained in the target scene through data of multiple modalities.
[0178] When the user executes an input command on the operation interface, the system responds to this operation and presents the initial multimodal data of the target scene to the user. The initial multimodal data, that is, the combination of raw image and text data without any processing, contains a preliminary description of the objects in the target scene, which may contain redundant, inaccurate or incomplete information. For example, on an e-commerce platform, a user may enter a command through a search or recommendation interface to request to view the details page of a certain product. In response to this command, the system displays the original image of the product (for example, the main image and detailed image of the product) and the description text (product title, details page description) on the operation interface so that the user can have a preliminary understanding of the product information.
[0179] Step S604, in response to the processing instruction acting on the operation interface, the target multimodal data of the target scene is displayed on the operation interface.
[0180] Among them, the target multimodal data is determined according to the first target modal data and the second target modal data, the second target modal data is obtained by inputting the object recognition result into the second network model, and using the second network model to recognize the initial multimodal data, the first target modal data and the object recognition result are obtained by using the first network model to recognize the initial multimodal data, and the object recognition result is used to represent the object of the target type contained in the multimodal data.
[0181] After the user further interacts with the operation interface and executes the processing instruction, the system will display the adjusted multimodal data of the target scene, i.e., the target multimodal data. The target multimodal data is the result of processing by the data governance framework of the embodiment of the present application, which combines the first target modal data (adjusted image data) and the second target modal data (adjusted text data). The acquisition of the first target modal data and the second target modal data is based on the processing of the initial multimodal data, including using the first network model for object recognition, and using the second network model for text recognition and enhancement based on the object recognition result. The target multimodal data is more accurate and pure in semantics, and can more comprehensively reflect the information of the objects in the target scene, providing users with clearer and more detailed product descriptions. For example, when viewing product details, the user may select an "update view" option as a processing instruction, and the system will respond to this instruction to display adjusted pictures and text descriptions, which are improved in the accuracy of object information, the richness of details, and the cross-modal consistency between images and texts.
[0182] Through the data governance framework of the embodiment of the present application, the target multimodal data has been adjusted in multiple dimensions, including but not limited to the elimination of image redundancy, the removal of text noise, the enhancement of text description, and the improvement of cross-modal consistency between image and text. For users, this means that more accurate, detailed and consistent descriptions can be obtained when viewing product information, thereby improving search efficiency, reducing ambiguity in information understanding, and enhancing the shopping experience. For example, for the visual redundancy that may exist in product images, the system can identify and delete or blur the images through the visual basic model, so that the images seen by users highlight the key features of the products. For deficiencies or errors in the text description, the system corrects and supplements them through the text basic model to provide more complete and accurate product information, such as correcting typos in product names, supplementing product details, etc.
[0183] Through the above steps, in response to the input instruction on the operation interface, the initial multimodal data is displayed on the operation interface, wherein the initial multimodal data includes the first initial modal data and the second initial modal data for describing the object; in response to the processing instruction on the operation interface, the target multimodal data is displayed on the operation interface, wherein the target multimodal data is determined according to the first target modal data and the second target modal data, the second target modal data is obtained by inputting the object recognition result and the second initial modal data into the second network model, and enhancing the second initial modal data using the second network model, the first target modal data and the object recognition result are obtained by screening the first initial modal data using the first network model, and the object recognition result is used to represent the object of the target type contained in the first initial modal data, thereby improving the quality of the multimodal data used for model training; it is easy to notice that by processing data of different modalities in a targeted manner, the semantic consistency and accuracy of data of different modalities can be improved, and the data quality can be improved, so that when integrating these processed modal data, the constructed target multimodal data set not only reduces redundancy and noise, but also enhances the information consistency between modalities. This enables subsequent models to learn purer and higher-quality cross-modal semantic relationships during training, thereby improving the training efficiency and final performance of the model, and further solving the technical problem of poor data quality used in model training in related technologies.
[0184] According to an embodiment of the present application, a data processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here. Figure 7 is a flow chart of a data processing method according to an embodiment of the present application, such as Figure 7 As shown, the method includes:
[0185] Step S702, acquiring initial multimodal data of the target scene by calling the first interface;
[0186] The first interface includes a first parameter, a parameter value of the first parameter includes initial multimodal data, and the initial multimodal data is used to describe objects included in the target scene through data of multiple modalities.
[0187] The above-mentioned first interface can be an interface for data interaction between the cloud server and the client, and the initial multimodal data can be passed into the interface function as the first parameter of the interface function to achieve the purpose of uploading the initial multimodal data to the cloud server.
[0188] Step S704, using the first network model to identify the initial multimodal data to obtain first target modal data and an object recognition result;
[0189] The object recognition result is used to represent the object of the target type contained in the multimodal data.
[0190] Step S706, inputting the object recognition result into the second network model, using the second network model to recognize the initial multimodal data, and obtaining second target modal data;
[0191] Step S708, obtaining target multimodal data of the target scene based on the first target modal data and the second target modal data;
[0192] Step S710: output target multimodal data by calling the second interface.
[0193] The second interface includes a second parameter, and a parameter value of the second parameter includes target multimodal data.
[0194] The above-mentioned second interface can be an interface for data interaction between the cloud server and the client. The cloud server can pass the target multimodal data into the interface function as the second parameter of the interface function to achieve the purpose of sending the target multimodal data to the client.
[0195] Through the above steps, initial multimodal data is obtained by calling the first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes initial multimodal data, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing the object; the first initial modal data is screened using the first network model to obtain first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of the target type contained in the first initial modal data; the object recognition result and the second initial modal data are input into the second network model, and the second initial modal data is enhanced using the second network model to obtain to the second target modal data; based on the first target modal data and the second target modal data, the target multimodal data is obtained; the target multimodal data is output by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target multimodal data, thereby improving the quality of the multimodal data used for model training; it is easy to notice that by processing data of different modalities in a targeted manner, the semantic consistency and accuracy of data of different modalities can be improved, and the data quality can be improved, so that when integrating these processed modal data, the constructed target multimodal data set not only reduces redundancy and noise, but also enhances the consistency of information between modalities. This enables subsequent models to learn purer and higher-quality cross-modal semantic relationships during training, thereby improving the training efficiency and final performance of the model, and thus solving the technical problem of poor data quality used for model training in related technologies.
[0196] According to an embodiment of the present application, a data processing device for implementing the above data processing method is also provided. Figure 8 is a schematic diagram of a data processing device according to an embodiment of the present application, such as Figure 8 As shown, the device 800 includes: an acquisition module 802 , an identification module 804 , an input module 806 , and a determination module 808 .
[0197] Among them, the acquisition module is used to acquire initial multimodal data, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing the object; the recognition module is used to screen the first initial modal data using the first network model to obtain first target modal data and object recognition results, wherein the object recognition results are used to represent objects of the target type contained in the first initial modal data; the input module is used to input the object recognition results and the second initial modal data into the second network model, and enhance the second initial modal data using the second network model to obtain second target modal data; the determination module is used to obtain target multimodal data based on the first target modal data and the second target modal data.
[0198] It should be noted that the acquisition module 802, the identification module 804, the input module 806, and the determination module 808 correspond to steps S202 to S208 in the above embodiment, and the four modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also be part of the device and can be run in the server 10 provided in the above embodiment.
[0199] In the above embodiments of the present application, the recognition module is used to slice the image data into blocks to obtain multiple image blocks; extract the global feature representation of the image data, and extract the local feature representation of the multiple image blocks, wherein the global feature representation is used to describe the overall semantics of the image data, and the local feature representation is used for the local semantics of different image blocks in the image data; based on the global feature representation and the local feature representation, the multiple image blocks are screened to obtain the first target modal data; and object recognition is performed on the first target modal data to obtain an object recognition result.
[0200] In the above embodiment of the present application, the recognition module is used to determine the similarity between the semantic information of the image data and multiple image blocks based on the global feature representation and the local feature representation; the multiple image blocks are screened based on the similarity to obtain a target image block, wherein the first similarity of the target image block is greater than the second similarity of other image blocks except the target image block in the multiple image blocks, the first similarity is the similarity between the local feature representation corresponding to the target image block and the global feature representation, and the second similarity is the similarity between the local feature representation corresponding to the other image blocks and the global feature representation; and the target image block is determined to be the first target modal data.
[0201] In the above embodiments of the present application, the recognition module is used to perform semantic recognition on the first target modal data to obtain multiple types of objects and confidence levels of multiple types of objects; and to screen multiple types of objects based on the confidence levels to obtain object recognition results.
[0202] In the above embodiments of the present application, the recognition module is used to correct the text data based on the preset prompt words to obtain corrected text data; enhance the text data based on the object recognition result to obtain enhanced text data; and determine the second target modal data based on the corrected text data and the enhanced text data.
[0203] In the above embodiments of the present application, the recognition module is used to determine the object information corresponding to the object of the target type contained in the object recognition result; determine the associated text data associated with the object information in the text data; and perform data enhancement on the associated text data in the text data to obtain enhanced text data.
[0204] In the above embodiments of the present application, the device further includes: a training module.
[0205] Among them, the training module is used to train the initial multimodal alignment model using the target multimodal data to obtain the target multimodal alignment model, wherein the target multimodal alignment model is used to align the data of multiple modes in the target multimodal data, and the alignment processing is used to represent mapping the data of multiple modes in the target multimodal data to the same semantic space.
[0206] In the above embodiment of the present application, the training module is used to use the initial multimodal alignment model to perform alignment processing on the target multimodal data to obtain third modal data and fourth modal data corresponding to the third modal data, wherein the third modal data and the fourth modal data are used to describe the same object; construct a contrast loss function of the third modal data and the fourth modal data; and adjust the model parameters of the initial multimodal alignment model based on the contrast loss function to obtain the target multimodal alignment model.
[0207] According to an embodiment of the present application, a data processing device for implementing the above data processing method is also provided. Fig. 9is a schematic diagram of a data processing device according to an embodiment of the present application, such as Fig. 9 As shown, the device 900 includes: a first display module 902 and a second display module 904 .
[0208] Among them, the first display module is used to respond to the input instruction acting on the operation interface, and display the initial multimodal data on the operation interface, wherein the initial multimodal data includes the first initial modal data and the second initial modal data for describing the object; the second display module is used to respond to the processing instruction acting on the operation interface, and display the target multimodal data on the operation interface, wherein the target multimodal data is determined according to the first target modal data and the second target modal data, the second target modal data is input into the second network model by inputting the object recognition result and the second initial modal data, and the second initial modal data is enhanced by using the second network model, the first target modal data and the object recognition result are obtained by screening the first initial modal data by using the first network model, and the object recognition result is used to represent the object of the target type contained in the first initial modal data.
[0209] It should be noted that the first display module 902 and the second display module 904 correspond to steps S602 to S604 in the above embodiment, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also be run as part of the device in the server 10 provided in the above embodiment.
[0210] According to an embodiment of the present application, a data processing device for implementing the above data processing method is also provided. Fig.10 is a schematic diagram of a data processing device according to an embodiment of the present application, such as Fig.10 As shown, the device 1000 includes: an acquisition module 1002 , an identification module 1004 , an input module 1006 , a determination module 1008 , and an output module 1010 .
[0211] Among them, the acquisition module is used to obtain initial multimodal data by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the initial multimodal data, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing the object; the recognition module is used to screen the first initial modal data using the first network model to obtain first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of the target type contained in the first initial modal data; the input module is used to input the object recognition result and the second initial modal data into the second network model, and enhance the second initial modal data using the second network model to obtain second target modal data; the determination module is used to obtain target multimodal data based on the first target modal data and the second target modal data; the output module is used to output the target multimodal data by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target multimodal data.
[0212] It should be noted that the acquisition module 1002, identification module 1004, input module 1006, determination module 1008, and output module 1010 correspond to steps S702 to S710 in the above embodiment, and the five modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above modules can also be run in the server 10 provided in the above embodiment as part of the device.
[0213] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in the above embodiments, as well as the application scenario and implementation process, but is not limited to the scheme provided in the above embodiments.
[0214] An embodiment of the present application may provide a computing device. Fig.11 is a structural block diagram of a computing device according to an embodiment of the present application. Fig.11 As shown, the computing device 100 may include: one or more ( Fig.11 (only one is shown) processor 102, memory 104, storage controller, and peripheral interfaces.
[0215] The above-mentioned computing device can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a PC (Personal Computer), a model all-in-one machine, etc., and the computing device can be pre-installed with the model in the above-mentioned embodiment of the present application.
[0216] Specifically, the computing device can pre-set multiple types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multimodal task processing, etc., so as to provide a variety of model choices. In different product forms, the computing device can support one or more model usage methods, including but not limited to model training, model calling, model fine-tuning, model deployment, model reasoning and application, etc. In some product forms, the computing device also supports model management, including but not limited to multi-type model management (supporting the management of multiple types of models such as discriminants and generative models), model version control (supporting the control of different model versions), model evaluation (based on model evaluation tools, evaluating the performance and effect of the model), etc. In other product forms, the computing device can also create applications based on the model, provide API calling capabilities, and can call the model to the created application through the API interface, while providing application management tools to achieve management and monitoring of the application.
[0217] Furthermore, the computing device can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master AI technology), and basic management and control capabilities (providing enterprise-level basic management and control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive, integrated AI development, training, deployment and application device is provided.
[0218] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0219] The processor may call the executable program stored in the memory through the transmission device to execute any method in the above embodiments.
[0220] An embodiment of the present application may provide an electronic device. Fig.12 is a structural block diagram of an electronic device according to an embodiment of the present application. Fig.12 As shown, the electronic device may include: an input / output device 112 ; a memory 114 and a processor 116 , wherein the processor 116 is connected to the input / output device 112 and the memory 114 via a bus 118 .
[0221] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0222] The processor may call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0223] Those skilled in the art will understand that Fig.12 The structure shown is for illustration only, and the computing device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. The figure does not limit the structure of the above computing device. For example, the computing device 100 may also include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a configuration different from that shown in the figure.
[0224] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0225] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the method provided in the above embodiment.
[0226] Optionally, in this embodiment, the above storage medium may be located in a computing device.
[0227] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in any one of the above embodiments.
[0228] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and the computer program implements the method provided in the embodiment when executed by a processor.
[0229] The embodiments of the present application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be used to store a computer program, and when the computer program is executed by a processor, the method provided in the embodiments is implemented.
[0230] The embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the method provided in the above embodiment is implemented.
[0231] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0232] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0233] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0234] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0235] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0236] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data processing method, characterized in that: include: Acquire initial multimodal data, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing an object; Using a first network model to screen the first initial modal data, obtaining first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of a target type contained in the first initial modal data; Inputting the object recognition result and the second initial modality data into a second network model, and enhancing the second initial modality data using the second network model to obtain second target modality data; Based on the first target modality data and the second target modality data, target multimodal data is obtained.
2. The data processing method according to claim 1, characterized in that: The first initial modality data is image data. The first initial modality data is screened using a first network model to obtain first target modality data and an object recognition result, including: Slicing the image data to obtain a plurality of image blocks; Extracting a global feature representation of the image data, and extracting local feature representations of the multiple image blocks, wherein the global feature representation is used to describe the overall semantics of the image data, and the local feature representation is used for local semantics of different image blocks in the image data; Filtering the multiple image blocks based on the global feature representation and the local feature representation to obtain the first target modality data; Perform object recognition on the first target modality data to obtain the object recognition result.
3. The data processing method according to claim 2, characterized in that: The method further comprises: filtering the plurality of image blocks based on the global feature representation and the local feature representation to obtain the first target modality data, including: Determining similarities between the semantic information of the image data and the plurality of image blocks based on the global feature representation and the local feature representation; The plurality of image blocks are screened based on the similarities to obtain a target image block, wherein a first similarity of the target image block is greater than a second similarity of other image blocks except the target image block among the plurality of image blocks, the first similarity is a similarity between a local feature representation corresponding to the target image block and the global feature representation, and the second similarity is a similarity between a local feature representation corresponding to the other image blocks and the global feature representation; Determine the target image block as the first target modality data.
4. The data processing method according to claim 2, characterized in that: Performing object recognition on the first target modality data to obtain the object recognition result includes: Performing semantic recognition on the first target modality data to obtain multiple types of objects and confidence levels of the multiple types of objects; The multiple types of objects are screened based on the confidence level to obtain the object recognition result.
5. The data processing method according to claim 1, characterized in that: The second initial modality data is text data, the object recognition result and the second initial modality data are input into a second network model, and the second initial modality data is enhanced by using the second network model to obtain second target modality data, including: Modifying the text data based on the preset prompt word to obtain modified text data; Enhance the text data based on the object recognition result to obtain enhanced text data; The second target modality data is determined based on the revised text data and the enhanced text data.
6. The data processing method according to claim 5, characterized in that: The text data is enhanced based on the object recognition result to obtain enhanced text data, including: Determining object information corresponding to the object of the target type contained in the object recognition result; Determining associated text data in the text data that is associated with the object information; Data enhancement is performed on the associated text data in the text data to obtain the enhanced text data.
7. The data processing method according to any one of claims 1 to 6, characterized in that: The method further comprises: The target multimodal data is used to train an initial multimodal alignment model to obtain a target multimodal alignment model, wherein the target multimodal alignment model is used to align data of multiple modalities in the target multimodal data, and the alignment processing is used to represent mapping data of multiple modalities in the target multimodal data to the same semantic space.
8. The data processing method according to claim 7, characterized in that: The target multimodal data is used to train the initial multimodal alignment model to obtain the target multimodal alignment model, including: Performing alignment processing on the target multimodal data using the initial multimodal alignment model to obtain third modal data and fourth modal data corresponding to the third modal data, wherein the third modal data and the fourth modal data are used to describe the same object; Constructing a contrast loss function of the third modality data and the fourth modality data; The model parameters of the initial multimodal alignment model are adjusted based on the contrast loss function to obtain the target multimodal alignment model.
9. A data processing method, characterized in that: include: In response to an input instruction acting on an operation interface, initial multimodal data is displayed on the operation interface, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing an object; In response to a processing instruction acting on the operation interface, target multimodal data is displayed on the operation interface, wherein the target multimodal data is determined based on first target modal data and second target modal data, the second target modal data is obtained by inputting an object recognition result and the second initial modal data into a second network model, and enhancing the second initial modal data using the second network model, the first target modal data and the object recognition result are obtained by screening the first initial modal data using the first network model, and the object recognition result is used to represent an object of a target type contained in the first initial modal data.
10. A data processing method, characterized in that: include: Acquiring initial multimodal data by calling a first interface, wherein the first interface includes a first parameter, a parameter value of the first parameter includes the initial multimodal data, wherein the initial multimodal data includes first initial modal data and second initial modal data for describing an object; Using a first network model to screen the first initial modal data, obtaining first target modal data and an object recognition result, wherein the object recognition result is used to represent an object of a target type contained in the first initial modal data; Inputting the object recognition result and the second initial modality data into a second network model, and enhancing the second initial modality data using the second network model to obtain second target modality data; Obtaining target multimodal data based on the first target modal data and the second target modal data; The target multimodal data is output by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the target multimodal data.
11. A computing device, characterized in that: include: A memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 10 when running.
12. An electronic device, characterized in that: include: A memory storing an executable program; A processor, connected to the memory via a bus, and configured to run the program, wherein the program executes the method described in any one of claims 1 to 10 when running.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Cited By
Complex equipment fault diagnosis method and system based on multi-modal knowledge graph
CN120217264A
Material mechanical property prediction method and equipment based on transfer learning and ensemble learning, and medium
CN120706185A