A data processing method, device and equipment

By introducing an early withdrawal network mechanism into the network model and dynamically adjusting the model processing method, the problem of resource waste in the existing technology is solved, and efficient processing of multimodal data is achieved.

CN119475043BActive Publication Date: 2025-05-09ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510059324.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-09
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

When existing network models deal with multimodal data, especially long-tail problems, there is a problem of resource waste, because the model needs to complete forward propagation regardless of whether the data is simple or complex.

Method used

By introducing an early retreat network mechanism into the network model, the processing method of the model is dynamically adjusted according to the difficulty of the data. Specifically, the target data of the multimodal data processing model is obtained and processed through the subnet of the encoded network and the early exit network. If the output result satisfies the output condition of the early exit network, the processing is terminated; otherwise, the result is input to the next subnet for further processing.

Benefits of technology

It realizes dynamic adjustment of model processing methods based on the difficulty of data, saves computing resources, and optimizes overall computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119475043B_ABST
    Figure CN119475043B_ABST
Patent Text Reader

Abstract

The embodiments of this specification disclose a data processing method, apparatus and equipment, which method includes: obtaining target data belonging to a first modality, the risk identification model includes a coding network and a data processing network corresponding to each modality, the coding network includes multiple sub-networks and corresponding early-leaving networks connected in series, the output end of the last sub-network is connected to the data processing network, and each sub-network is respectively connected to the corresponding early-leaving network and the next sub-network; based on the target data, using the first sub-network of the coding network corresponding to the first modality, determining a first intermediate result output by the first sub-network, and inputting the first intermediate result into the first early-leaving network connected to the first sub-network to obtain a first output result; if the first output result meets the output condition, the first output result is used as the output result of the risk identification model; otherwise, the first output result is input into the next sub-network for processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of computer technology, and in particular to a data processing method, device and equipment. Background Art

[0002] Network models have increasingly wide applications, such as content security and privacy protection in risk prevention and control or risk identification. Taking the CLIP model as an example, the CLIP model performs comparative learning on a large number of image-text pair datasets, so that the CLIP model can map images and related descriptive texts into the same vector space, so that similar images and texts are close to each other in this space. This powerful cross-modal understanding capability makes the CLIP model widely applicable to various tasks.

[0003] Although the network model has shown excellent performance in processing multimodal data content, especially long-tail problems, it faces some challenges in actual deployment. For a more complex network model, whether it is processing simple data or complex data (or difficult data), the execution cost is usually the same, because the network model needs to complete the full forward propagation to generate prediction results, which may lead to a waste of resources, especially when processing a large amount of simple data. For this reason, it is necessary to provide a better way for the model to process data, so that the model can adjust the way it processes data according to the difficulty of the data, thereby saving computing resources and optimizing the overall computing efficiency. Summary of the invention

[0004] The purpose of the embodiments of this specification is to provide a more optimal way for a model to process data, so that the model's processing method for data can be adjusted according to the difficulty of the data, thereby saving computing resources and optimizing overall computing efficiency.

[0005] In order to implement the above technical solution, the embodiments of this specification are implemented as follows:

[0006] A data processing method provided in an embodiment of the present specification includes: obtaining target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each modality of the multiple modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the risk identification model is used to determine the first intermediate result output by the first subnetwork, and the first intermediate result is input into the first early-leaving network connected to the first subnetwork to obtain a first output result. If the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result through the risk identification model is terminated, and the first output result is used as the output result of the risk identification model. If the first output result does not meet the output condition, the first output result is input into the next subnetwork as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0007] A data processing method provided in an embodiment of the present specification includes: obtaining target data belonging to a first mode applied to a multimodal data processing model, wherein the multimodal data processing model includes a coding network and a data processing network corresponding to each mode in multiple modes, wherein the coding network corresponding to each mode is composed of multiple network layers in series, wherein the coding network corresponding to each mode includes multiple sub-networks in series and an early-leaving network corresponding to each sub-network except the last sub-network in the multiple sub-networks in series, wherein the output end of the last sub-network in the multiple sub-networks in series is connected to the input end of the data processing network, wherein each sub-network includes at least one network layer, wherein each sub-network except the last sub-network in the multiple sub-networks in series is respectively connected to the corresponding early-leaving network and the next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first sub-network of the coding network corresponding to the first mode in the multimodal data processing model is used to determine the first intermediate result output by the first sub-network, and the first intermediate result is input into the first early-leaving network connected to the first sub-network to obtain the first output result. If the first output result satisfies the output condition corresponding to the first early-leaving network, the multimodal data processing model is terminated from continuing to process the first output result, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is used as the target data belonging to the first modality and input into the next sub-network for processing until the output result of the multimodal data processing model is obtained.

[0008] A data processing device provided in an embodiment of the present specification includes: a data acquisition module, which acquires target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each modality of the multiple modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. The early exit processing module, based on the target data, uses the first sub-network of the encoding network corresponding to the first modality in the risk identification model to determine the first intermediate result output by the first sub-network, and inputs the first intermediate result into the first early exit network connected to the first sub-network to obtain the first output result. The early exit module, if the first output result meets the output condition corresponding to the first early exit network, terminates the processing of the first output result through the risk identification model, and uses the first output result as the output result of the risk identification model. The early exit decision module, if the first output result does not meet the output condition, inputs the first output result as the target data belonging to the first modality into the next sub-network for processing until the output result of the risk identification model is obtained.

[0009] A data processing device provided in an embodiment of the present specification comprises: a data acquisition module, which acquires target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model comprises a coding network and a data processing network corresponding to each modality in a plurality of modalities, wherein the coding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the coding network corresponding to each modality comprises a plurality of subnetworks connected in series and an early-leaving network corresponding to each subnetwork except the last subnetwork in the plurality of subnetworks connected in series, wherein the output end of the last subnetwork in the plurality of subnetworks connected in series is connected to the input end of the data processing network, wherein each subnetwork comprises at least one network layer, wherein each subnetwork except the last subnetwork in the plurality of subnetworks connected in series is respectively connected to a corresponding early-leaving network and a next subnetwork, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. A processing module, based on the target data, uses the first subnetwork of the coding network corresponding to the first modality in the multimodal data processing model, determines a first intermediate result output by the first subnetwork, and inputs the first intermediate result into a first early-leaving network connected to the first subnetwork, thereby obtaining a first output result. Early exit module, if the first output result meets the output condition corresponding to the first early exit network, then stop processing the first output result through the multimodal data processing model, and use the first output result as the output result of the multimodal data processing model. Decision module, if the first output result does not meet the output condition, then use the first output result as the target data belonging to the first modality and input it into the next sub-network for processing until the output result of the multimodal data processing model is obtained.

[0010] A data processing device provided in an embodiment of the present specification comprises: a processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor is configured to: obtain target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model comprises an encoding network and a data processing network corresponding to each modality of the multiple modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality comprises a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network comprises at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the risk identification model is used to determine the first intermediate result output by the first subnetwork, and the first intermediate result is input into the first early-leaving network connected to the first subnetwork to obtain a first output result. If the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result through the risk identification model is terminated, and the first output result is used as the output result of the risk identification model. If the first output result does not meet the output condition, the first output result is input into the next subnetwork as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0011] A data processing device provided in an embodiment of the present specification comprises: a processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor is configured to: obtain target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model comprises an encoding network and a data processing network corresponding to each modality of the plurality of modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality comprises a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network comprises at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first sub-network of the encoding network corresponding to the first modality in the multimodal data processing model is used to determine the first intermediate result output by the first sub-network, and the first intermediate result is input into the first early-leaving network connected to the first sub-network to obtain a first output result. If the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result by the multimodal data processing model is terminated, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

[0012] An embodiment of the present specification also provides a storage medium, which is used to store computer-executable instructions. When the executable instructions are executed by a processor, the following process is implemented: obtaining target data belonging to a first modality of a risk identification model applied to multimodal data, the target data of the first modality being text data of a text modality or image data of an image modality, the risk identification model including an encoding network and a data processing network corresponding to each modality of the multiple modalities, the encoding network corresponding to each modality being composed of a plurality of network layers connected in series, the encoding network corresponding to each modality including a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, each sub-network includes at least one of the network layers, and each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, and the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the risk identification model is used to determine the first intermediate result output by the first subnetwork, and the first intermediate result is input into the first early-leaving network connected to the first subnetwork to obtain a first output result. If the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result through the risk identification model is terminated, and the first output result is used as the output result of the risk identification model. If the first output result does not meet the output condition, the first output result is input into the next subnetwork as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0013] The embodiment of the present specification also provides a storage medium, which is used to store computer-executable instructions. When the executable instructions are executed by a processor, the following process is implemented: obtaining target data belonging to a first modality applied to a multimodal data processing model, the multimodal data processing model includes a coding network and a data processing network corresponding to each modality of a plurality of modalities, the coding network corresponding to each modality is composed of a plurality of network layers connected in series, the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, each sub-network includes at least one of the network layers, and each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, and the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first sub-network of the encoding network corresponding to the first modality in the multimodal data processing model is used to determine the first intermediate result output by the first sub-network, and the first intermediate result is input into the first early-leaving network connected to the first sub-network to obtain a first output result. If the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result by the multimodal data processing model is terminated, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

[0014] The embodiment of the present specification also provides a computer program product, including a computer program, which implements the following process when executed by a processor: obtaining target data belonging to a first modality of a risk identification model applied to multimodal data, the target data of the first modality being text data of a text modality or image data of an image modality, the risk identification model including an encoding network and a data processing network corresponding to each modality of the multiple modalities, the encoding network corresponding to each modality being composed of a plurality of network layers connected in series, the encoding network corresponding to each modality including a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, the output end of the last sub-network in the plurality of sub-networks connected in series being connected to the input end of the data processing network, each sub-network including at least one of the network layers, each sub-network except the last sub-network in the plurality of sub-networks connected in series being respectively connected to a corresponding early-leaving network and a next sub-network, the early-leaving network and the data processing network being used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the risk identification model is used to determine the first intermediate result output by the first subnetwork, and the first intermediate result is input into the first early-leaving network connected to the first subnetwork to obtain a first output result. If the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result through the risk identification model is terminated, and the first output result is used as the output result of the risk identification model. If the first output result does not meet the output condition, the first output result is input into the next subnetwork as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0015] The embodiment of the present specification also provides a computer program product, including a computer program, which implements the following process when executed by a processor: obtaining target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes a coding network and a data processing network corresponding to each modality of the multiple modalities, wherein the coding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements. Based on the target data, the first sub-network of the encoding network corresponding to the first modality in the multimodal data processing model is used to determine the first intermediate result output by the first sub-network, and the first intermediate result is input into the first early-leaving network connected to the first sub-network to obtain a first output result. If the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result by the multimodal data processing model is terminated, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor.

[0017] Figure 1 This is an embodiment of a data processing method of this specification;

[0018] Figure 2 A schematic diagram of a data processing process of this specification;

[0019] Figure 3 This is a schematic diagram of a model training process in this manual;

[0020] Figure 4Another data processing method embodiment of this specification;

[0021] Figure 5 A schematic diagram of another data processing process of this specification;

[0022] Figure 6 A schematic diagram of a transfer learning process in this specification;

[0023] Figure 7 This is another data processing method embodiment of the present specification;

[0024] Figure 8 This is another data processing method embodiment of the present specification;

[0025] Fig. 9 This is an embodiment of a data processing device of the present specification;

[0026] Fig.10 Another data processing device embodiment of the present specification;

[0027] Fig.11 This is an embodiment of a data processing device in this specification. DETAILED DESCRIPTION

[0028] The embodiments of this specification provide a data processing method, device and equipment.

[0029] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0030] The embodiments of this specification provide a mechanism for dynamically adjusting the execution of a network model. Large models have increasingly wide applications, such as content security applications in risk prevention and control or risk identification, and are particularly suitable for dealing with the long-tail problems therein. Taking the CLIP model as an example, the CLIP model performs comparative learning on a large number of image-text pair data sets, so that the CLIP model has the ability to map images and related descriptive texts into the same vector space, so that similar images and texts are close to each other in the space. This powerful cross-modal understanding ability enables the CLIP model to be widely used in various tasks (including content security in risk prevention and control or risk identification). In the field of content security, the CLIP model can be used to identify and filter non-compliant images and texts. This long-tail problem usually refers to content that appears less frequently but has a larger variety. This content is often difficult to accurately identify and process. The emergence of the CLIP model provides a new processing approach. First, the CLIP model can perform multimodal understanding. Specifically, the CLIP model can process image data and text data, so that the CLIP model can understand the content more comprehensively. For example, the CLIP model can identify an image with offensive text, or the CLIP model can identify the non-compliant information implicit in the image. Second, zero-shot learning can be achieved. Specifically, the training method of the CLIP model can perform zero-shot learning. Learning), that is, even if there is no specific category of data directly trained, the CLIP model can classify or identify these categories of data, which is very useful for the processing of long-tail problems; secondly, the CLIP model has good scalability and adaptability. Specifically, the training of the CLIP model is based on a large amount of data, so it has good generalization ability and adaptability. Moreover, the CLIP model can easily adapt to new content and changing risk prevention and control standards; thirdly, transfer learning can be performed. The data representation generated in the CLIP model can be used for transfer learning. Even when there is less labeled data for a specific task, it can be fine-tuned to adapt to specific risk prevention and control needs; finally, the CLIP model can improve data processing efficiency. Specifically, the CLIP model can evaluate the similarity between image data and text data with only one forward propagation, so that the CLIP model can quickly process a large amount of data. In addition, although the CLIP model has shown excellent performance in processing the content of multimodal data, especially long-tail problems, it faces some challenges in actual deployment, mainly including high computational cost and the need to adapt to diverse risk prevention and control scenarios. The CLIP model is a large deep learning model with a large number of parameters, which means that without proper optimization, a forward inference of the CLIP model may require high computing resources, which may become a bottleneck for risk prevention and control scenarios that require rapid processing of large amounts of data.At the same time, there are many scenarios for risk prevention and control, and deploying an adaptive large model for each scenario brings great deployment efficiency problems. In addition, for large models such as the CLIP model, the execution cost is usually the same whether processing simple data or complex data, because the CLIP model needs to complete the full forward propagation to generate prediction results, which may lead to a waste of resources, especially when processing a large amount of simple data, which could have been processed by a smaller and faster model.

[0031] In the case of more scenarios mentioned above, model compression can be used to reduce the size of the CLIP model. For example, some unimportant parameters can be removed through model pruning to reduce the size of the CLIP model and reduce the consumption of computing resources while maintaining the performance of the CLIP model as much as possible. Alternatively, the parameters and activation values ​​of the CLIP model can be converted from floating point numbers to low-precision representations (such as INT8) through quantization to reduce the size of the CLIP model and accelerate calculations. Alternatively, the knowledge of the CLIP model can be transferred to a smaller and faster student model through model distillation to reduce the computational burden while maintaining good performance. However, the above methods do not fundamentally alleviate the situation where multiple models need to be deployed multiple times for multiple scenarios of risk prevention and control. Moreover, they will affect the deployment efficiency and significantly reduce the effectiveness of the model.

[0032] In addition, for the above-mentioned difficult and easy data, it can be processed through the cascade model. Specifically, a model cascade system (Cascade) from small to large is established. In this system, one or more smaller and faster models are first used to process the data. These models can quickly identify most of the simple data and only forward the difficult-to-judge data to the next level of more complex and computationally expensive models (such as CLIP models, etc.). Through the above method, a lot of resources can be saved, because only a few difficult data need to go through the entire model cascade. Alternatively, it can be processed through dynamic difficulty judgment. Specifically, a mechanism is developed to dynamically evaluate the processing difficulty of data. For data evaluated as simple, it can be Simple models or rules are directly applied for processing, and for difficult data, they are provided to high-performance models such as CLIP models for processing. The difficulty assessment itself needs to be fast and resource-efficient, which can be implemented based on some simple features or heuristic rules, or it can also be processed by data screening and preprocessing. Specifically, before the data is input into the CLIP model, some preprocessing mechanisms are used to screen and simplify the data. For example, a simple method of text or image analysis can be used to eliminate data that is obviously irrelevant or meets specific conditions. For some problems that can be quickly solved by simple rules (such as filtering out text data containing specific keywords, etc.), the above processing can be directly applied, thereby reducing the dependence on the CLIP model. However, the above method cannot effectively utilize the attribute of the CLIP model with pre-trained parameters, and the model needs to be retrained, which loses the ability of Zero-Shot. Routing data through rules and other methods has a large demand for operating costs, and the effect cannot be guaranteed. The embodiment of this specification provides a feasible processing method, by reasonably designing the early exit network, dynamically adjusting the network that performs data processing in the model according to the difficulty of the data, thereby improving the overall resource throughput. For specific processing, please refer to the specific content in the following embodiment.

[0033] like Figure 1 As shown, an embodiment of this specification provides a data processing method, and the execution subject of the method can be a terminal device or a server, etc., wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, or a computer device such as a laptop or a desktop computer, or an IoT device (specifically such as a smart watch, a car-mounted device, etc.), etc., wherein the server can be an independent server, or a server cluster composed of multiple servers, etc., and the server can be a background server for financial services or online shopping services, or a background server for an application, etc. In this embodiment, the execution subject is taken as an example for detailed description. For the case where the execution subject is a terminal device, please refer to the following server situation processing, which will not be repeated here. The method can specifically include the following steps:

[0034] In step S102, target data belonging to a first modality is obtained and applied to a multimodal data processing model. The multimodal data processing model includes an encoding network and a data processing network corresponding to each modality of the multiple modalities. The encoding network corresponding to each modality is composed of multiple network layers connected in series. The encoding network corresponding to each modality includes multiple sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the multiple sub-networks connected in series. The output end of the last sub-network in the multiple sub-networks connected in series is connected to the input end of the data processing network. Each sub-network includes at least one network layer. Each sub-network except the last sub-network in the multiple sub-networks connected in series is respectively connected to the corresponding early-leaving network and the next sub-network. The early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements.

[0035] Among them, the first modality can be any modality, and the modality can be the type or attribute described in the data, for example, the modality corresponding to the text data can be the text modality, the modality corresponding to the image data can be the image modality, the modality corresponding to the audio data can be the audio modality, etc., which can be set according to the actual situation. The multimodal data processing model can be a model for processing data of multiple different modalities. Specifically, the multimodal data processing model can be a model for performing image and text retrieval (i.e., a model for retrieving corresponding text data through image data or retrieving corresponding image data through text data), the multimodal data processing model can be a model for risk prevention and control or risk identification in the financial field, etc. The multimodal data processing model can be constructed by a variety of algorithms and / or networks. In actual applications, the multimodal data processing model can be constructed by a deep neural network, or it can be a specified large model, specifically, it can be a large language model, such as ChatGLM, GPT-4, etc., which can be set according to the actual situation.

[0036] The multimodal data processing model may include a modal data processing network for each of the multiple modalities. For example, the multimodal data processing model can be used to process image data and text data. In this case, a modal data processing network for the image modality and a modal data processing network for the text modality are provided in the multimodal data processing model. The modal data processing networks of different modalities in the multimodal data processing model may be independent of each other, or they may share the same partial network with each other, etc. The specific settings may be based on actual conditions. The modal data processing network of each modality may include an encoding network and a data processing network. The encoding network can be used to encode the input modal data. The encoding networks corresponding to data of different modalities may be different. For example, the encoding network corresponding to the text data of the text modality may be a network for encoding the text data, and the encoding network corresponding to the image data of the image modality may be a network for encoding the image data. The networks for encoding the text data and the image data are usually different. The encoding network can set a corresponding structure according to the modal information of the corresponding data. For example, the encoding network can be constructed by a convolutional neural network, or the encoding network can be constructed by a recurrent neural network, or the encoding network can also be constructed by multiple Transformer modules, etc., wherein the structure of the encoding network can include multiple identical network layers (or units). For example, the encoding network can be constructed by a convolutional neural network including multiple specified convolutional layers, or the encoding network can be constructed by multiple network layers including one or more Transformer modules, etc., which can be set according to actual conditions. The data processing network can be used to further process the data output by the encoding network. The data processing network can be set according to the purpose or requirements to be achieved by the multimodal data processing model. For example, the multimodal data processing model is used for risk prevention and control processing. The data processing network can identify whether there is a specified risk based on the data output by the encoding network, and perform corresponding risk prevention and control processing when it is determined that there is a specified risk. For another example, the multimodal data processing model is used for image and text retrieval processing. The data processing network can perform similarity calculation based on the data output by the encoding network and the data in the retrieval library to retrieve retrieval data that matches the data output by the encoding network, etc. The specific setting can be based on actual conditions, and the embodiments of this specification are not limited to this.

[0037] like Figure 2As shown, the multiple network layers included in the encoding network can be divided into multiple parts, each part can include one or more network layers, each part can be used as a sub-network, and the output end of each sub-network except the last sub-network in the encoding network can be connected to the input end of the next sub-network and the early-exit network respectively, and the output end of the last sub-network in the encoding network is only connected to the input end of the data processing network. The early-leaving network is used to process the data output by the network layer (i.e., the last network layer in the subnetwork to which the input end of the early-leaving network is connected) according to data processing requirements (i.e., the purpose or requirements to be achieved by the multimodal data processing model, such as risk prevention and control processing or image and text retrieval processing, etc.), and the data processing network is used to process the data output by the network layer (i.e., the last network layer in the last subnetwork in the encoding network) according to the above-mentioned data processing requirements (i.e., the purpose or requirements to be achieved by the multimodal data processing model, etc.). In practical applications, the early-leaving network and the data processing network may have the same network structure, for example, both the early-leaving network and the data processing network are constructed by linear classifiers, or both the early-leaving network and the data processing network are constructed by multilayer perceptrons MLP, etc. The early-leaving network and the data processing network may also have different network structures, which may be set specifically according to actual conditions.

[0038] In implementation, see Figure 2 The structure of the multimodal data processing model shown in the figure, when it is necessary to execute a certain business through the multimodal data processing model, the data to be processed (that is, the target data belonging to the first modality applied to the multimodal data processing model) can be obtained, wherein it should be noted that the target data belonging to the first modality can be data that has not been processed by the multimodal data processing model. At this time, the data input by the user (such as image data, text data or video data, etc.) or the specified data generated based on the data input by the user (such as payment data or transfer data, etc.) can be obtained, and the above-mentioned acquired data can be used as the target data belonging to the first modality applied to the multimodal data processing model, or the target data belonging to the first modality can also be processed by the multimodal data processing model. Data processed by some sub-networks (excluding the last two sub-networks in the encoding network) in the encoding network in the modal data processing model. Specifically, the encoding network includes 5 sub-networks, and the 5 sub-networks (which can be recorded as sub-network 1, sub-network 2, sub-network 3, sub-network 4 and sub-network 5) are connected in series based on a specified order (such as sub-network 1-sub-network 2-sub-network 3-sub-network 4-sub-network 5). The data output by sub-network 1 can be used as the target data belonging to the first modality applied to the multi-modal data processing model, and the data output by sub-network 3 can be used as the target data belonging to the first modality applied to the multi-modal data processing model, etc. The specific setting can be based on actual conditions, and the embodiments of this specification are not limited to this.

[0039] In step S104, based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the multimodal data processing model is used to determine the first intermediate result output by the first subnetwork, and the first intermediate result is input into the first early-retirement network connected to the first subnetwork to obtain the first output result.

[0040] Among them, the first subnetwork can be a subnetwork specified in the encoding network corresponding to the first modality, and specifically can be a subnetwork that the target data needs to input. For example, if the target data is data that has not been processed by the multimodal data processing model, the first subnetwork can be the first subnetwork in the encoding network corresponding to the first modality. If the target data is data processed by partial subnetworks in the encoding network in the multimodal data processing model (excluding the last two subnetworks in the encoding network), the first subnetwork can be a subnetwork connected after the above-mentioned partial subnetwork (that is, the first subnetwork connected after the above-mentioned partial subnetwork).

[0041] In implementation, the target data can be directly input into the first sub-network of the encoding network corresponding to the first mode in the multimodal data processing model, and the target data can be processed by the first sub-network to obtain a corresponding output result, which is the first intermediate result output by the first sub-network. Considering that the data processed by the multimodal data processing model may be difficult data or simple data, in order to save computing resources and improve data processing efficiency, data of different difficulty levels can be output at different depth levels in the encoding network in the multimodal data processing model. For this purpose, one or more early exit points can be set. As mentioned above, each early exit point is located between two adjacent sub-networks, and each early exit point is provided with an early exit network. After the data is processed and output at different depth levels in the encoding network in the multimodal data processing model, it can be first input into the corresponding early exit network for processing to obtain a corresponding output result. Specifically, the first intermediate result can be first input into the first early exit network connected to the first sub-network to obtain a first output result. Then, it can be determined whether it is necessary to exit from the multimodal data processing model based on the first output result. For details, see the processing of step S106 or step S108 below.

[0042] In step S106, if the first output result satisfies the output condition corresponding to the first early-leaving network, the processing of the first output result by the multimodal data processing model is terminated, and the first output result is used as the output result of the multimodal data processing model.

[0043] Among them, the output condition corresponding to the first early exit network may include multiple types. For example, the output condition corresponding to the first early exit network may be that the first output result exceeds a preset threshold. At this time, the first output result may be compared with the preset threshold to determine whether the first output result satisfies the output condition corresponding to the first early exit network. The output condition corresponding to the first early exit network may also be that the first output result includes specified keywords. At this time, it may be detected whether the first output result includes specified keywords to determine whether the first output result satisfies the output condition corresponding to the first early exit network, etc. In addition, different early exit networks may have different corresponding output conditions, which may be set specifically according to actual conditions, and the embodiments of this specification do not limit this.

[0044] In implementation, if it is determined through calculation or detection that the first output result satisfies the output condition corresponding to the first early-leaving network, it indicates that the target data is relatively simple data, and the target data does not need to be processed using all the networks in the multimodal data processing model, that is, the output requirements corresponding to the complete multimodal data processing model can be fully met (such as the error between the result output by the early-leaving network and the result output by the complete multimodal data processing model is within a preset range, etc.). At this time, the multimodal data processing model can be used to terminate the processing of the first output result, and the first output result can be directly used as the output result of the multimodal data processing model, and the first output result can be output to the user or management party as the output result of the multimodal data processing model.

[0045] In step S108, if the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

[0046] In implementation, if it is determined through calculation or detection that the first output result does not meet the output condition corresponding to the first early-leaving network, it indicates that the first output result cannot be used as the output result of the multimodal data processing model. At this time, the subsequent sub-network can be used to continue processing the first output result, that is, the first output result can be used as the target data belonging to the first modality in the above-mentioned step S102, and the processing of steps S104 and S106 can be continued or the processing of steps S104 and S108 can be performed, that is, the target data belonging to the first modality (that is, the first output result) can be input into the encoding network corresponding to the first modality in the multimodal data processing model. The first sub-network (at this time, actually the next sub-network connected to the above-mentioned first sub-network) is inputted, and the output intermediate result is inputted into the early-retirement network connected to the above-mentioned sub-network to obtain the corresponding output result, and the processing as in the above-mentioned step S106 or S108 is continued. In order to clearly explain the above-mentioned processing process, the above-mentioned processing is further explained below. Specifically, the first output result can be inputted into the next sub-network connected to the first sub-network in the encoding network corresponding to the first modality in the multimodal data processing model to obtain the second intermediate result outputted by the next sub-network (for the convenience of explanation, it is temporarily named as the second intermediate result. If there are the same names in the future, they are the same as those in the following steps unless otherwise specified). The second intermediate result may be different), and the second intermediate result is input into a second early-leaving network (for the sake of convenience, temporarily named as the second early-leaving network, and if there is a same name later, it may be different from the second early-leaving network unless otherwise specified) connected to the next sub-network to obtain a second output result (for the sake of convenience, temporarily named as the second output result, and if there is a same name later, it may be different from the second output result unless otherwise specified). If the second output result meets the output condition corresponding to the second early-leaving network, the multimodal data processing model is terminated to continue processing the second output result, and the second output result is used as the output of the multimodal data processing model As a result, if the second output result does not meet the output condition corresponding to the second early-leaving network, the second output result can be input into the next sub-network connected to the above-mentioned next sub-network in the encoding network corresponding to the first modality in the multimodal data processing model, and then the above-mentioned processing process can be repeated. Finally, data output from a certain early-leaving network can be obtained as the output result of the multimodal data processing model, or the data output by the last sub-network can be input into the data processing network (at this time, the complete data processing process of the complete multimodal data processing model is passed) to obtain the corresponding output result, which is the output result of the multimodal data processing model.In the above manner, data of different difficulty levels can be processed at different levels in the network of the multimodal data processing model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the multimodal data processing model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the conditions for early-leaving can be met in a deeper network layer, or even need to be processed by a complete multimodal data processing model. In this way, the amount of computation required to process simple data can be significantly reduced, while ensuring that sufficient computing resources are still given to difficult data for accurate prediction, so that the multimodal data processing model can adaptively process data of different difficulty levels, which can not only save computing resources but also optimize overall computing efficiency.

[0047] The embodiment of the present specification provides a data processing method, by acquiring target data belonging to a first modality applied to a multimodal data processing model, the multimodal data processing model includes a coding network and a data processing network corresponding to each modality in a plurality of modalities, the coding network corresponding to each modality is composed of a plurality of network layers connected in series, the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, each sub-network includes at least one network layer, each sub-network in the plurality of sub-networks connected in series except the last sub-network is respectively connected to the corresponding early-leaving network and the next sub-network, the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, using the first sub-network of the coding network corresponding to the first modality in the multimodal data processing model, determine a first intermediate result output by the first sub-network, and input the first intermediate result into the first early-leaving network connected to the first sub-network to obtain a first output result, wherein if the first input If the result meets the output condition corresponding to the first early-leaving network, the multimodal data processing model is terminated to continue processing the first output result, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the multimodal data processing model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the multimodal data processing model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the multimodal data processing model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete multimodal data processing model, which can not only significantly reduce the amount of calculation required to process simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the multimodal data processing model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0048] In practical applications, each network layer includes one or more serially connected encoding modules, the encoding modules included in different network layers have the same structure, and the encoding modules included in each network layer have the same structure.

[0049] Among them, the encoding module may include multiple types. For example, the encoding module may be composed of one or more convolutional layers, or the encoding module may be composed of one or more specified encoders (such as encoders constructed by recurrent neural networks, etc.). The encoding module can be used as the smallest unit constituting a network layer. The network layer may be composed of one encoding module, or it may be composed of multiple (such as 2 or 5, etc.) encoding modules connected in series, etc. Different network layers may include different numbers of encoding modules. For example, a network layer includes 2 encoding modules connected in series, and the next network layer connected to the network layer may include 4 encoding modules, etc. The specific setting can be based on actual conditions.

[0050] In practical applications, based on the above structure, the multimodal data processing model can be a model built based on the CLIP model, the encoding module can be a Transformer module (i.e., Transformer Block), the target data belonging to the first modality is text data belonging to the text modality or image data belonging to the image modality, the multimodal data processing model can be a model for risk prevention and control or risk identification, and the early exit network is a network of linear classifiers.

[0051] In implementation, one or more Transformer modules can constitute different network layers, one or more different network layers can constitute a sub-network, multiple sub-networks constitute an encoding network of a certain modality, the sub-networks except the last sub-network and the corresponding early-leaving networks, and the last sub-network and the data processing network constitute a multimodal data processing model, and the modalities corresponding to the data that can be processed by the multimodal data processing model include text modality and image modality.

[0052] Based on the multimodal data processing model of the above structure, the complete multimodal data processing model can be trained as a whole through training samples, or the part of the multimodal data processing model belonging to the CLIP model can be pre-trained through training samples, and then the early-leaving network is fine-tuned to finally obtain the trained multimodal data processing model. Based on the above two methods, the specific processing may include the following method one and method two.

[0053] Method 1: The complete multimodal data processing model is trained as a whole through training samples. For details, please refer to the processing of steps A2 to A8 below.

[0054] In step A2, sample data belonging to the second modality for training the multimodal data processing model is obtained.

[0055] The second mode may be the same as or different from the first mode.

[0056] In step A4, based on the sample data belonging to the second modality, the second subnetwork of the encoding network corresponding to the second modality in the multimodal data processing model is used to determine the second intermediate result sample output by the second subnetwork, and the second intermediate result sample is input into the second early-retirement network connected to the second subnetwork to obtain a second output result sample.

[0057] In step A6, if the second output result sample satisfies the output condition corresponding to the second early-leaving network, the processing of the second output result sample by the multimodal data processing model is terminated, and the second output result sample is used as the output result sample of the multimodal data processing model; if the second output result sample does not meet the output condition corresponding to the second early-leaving network, the second output result sample is input into the next sub-network as sample data belonging to the second modality for processing until the output result sample of the multimodal data processing model is obtained.

[0058] The specific processing process of the above steps A2 to A6 can be found in the above related content and will not be repeated here.

[0059] In step A8, the multimodal data processing model is trained based on the output result samples of the multimodal data processing model to obtain a trained multimodal data processing model.

[0060] In implementation, the multimodal data processing model can be trained in a variety of different ways. For example, supervised model training can be performed. In this case, based on the output result samples of the multimodal data processing model, a preset loss function can be used to calculate the loss information between the output result samples and the corresponding label information, and the model parameters in the multimodal data processing model can be adjusted based on the loss information. Then, the above processing process is repeated until the loss function converges, thereby obtaining a trained multimodal data processing model. Alternatively, unsupervised model training can be performed. In this case, the multimodal data processing model can be trained by contrastive learning. Specifically, if the multimodal data processing model is a model constructed based on the CLIP model, the multimodal data processing model can be trained by contrastive learning, etc. In addition to training the multimodal data processing model by contrastive learning, the multimodal data processing model can also be trained in a variety of different ways, which can be set specifically according to actual conditions, and the embodiments of this specification do not limit this.

[0061] Method 2: If Figure 3 As shown, part of the network in the multimodal data processing model is pre-trained through training samples, and then the early-falling network is fine-tuned to finally obtain the trained multimodal data processing model. For details, please refer to the processing of the following steps B2 to B10.

[0062] In step B2, sample data belonging to the third modality for training the multimodal data processing model is obtained.

[0063] The third mode may be the same as or different from the first mode.

[0064] In step B4, the early-falling network in the multimodal data processing model is shielded, and the shielded multimodal data processing model is pre-trained based on sample data belonging to the third modality to obtain a pre-trained shielded multimodal data processing model.

[0065] In implementation, the early-falling network included in the multimodal data processing model can be temporarily shielded (or removed) so that the early-falling network included in the multimodal data processing model cannot play any role or is invalid. Then, the remaining network (i.e., the shielded multimodal data processing model) can be pre-trained using sample data belonging to the third modality to obtain a pre-trained partial network (i.e., the pre-trained shielded multimodal data processing model).

[0066] In step B6, sample data belonging to the fourth modality is obtained.

[0067] The fourth mode may be the same as or different from the first mode. Furthermore, the fourth mode may be the same as or different from the third mode.

[0068] In step B8, the sample data belonging to the fourth modality is input into the pre-trained masked multimodal data processing model to obtain the encoded data output by the encoding network corresponding to the fourth modality in the pre-trained masked multimodal data processing model.

[0069] In step B10, the model parameters in the pre-trained masked multimodal data processing model are fixed, and knowledge distillation training is performed on the early-falling network in the multimodal data processing model based on the encoded data to obtain a trained multimodal data processing model.

[0070] In practical applications, in the above step S104, based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the multimodal data processing model is used to determine the specific processing method of the first intermediate result output by the first subnetwork. The following is an optional processing method, such as Figure 4 As shown, please refer to the processing of steps S1042 to S1048 below for details.

[0071] In step S1042, scenario information corresponding to the target data is obtained, where the scenario information includes one or more of fraud, privacy data leakage, illegal financial activities, and discrimination.

[0072] In implementation, some businesses (such as risk prevention and control business or risk identification business) often include many different scenarios. Usually, an adapted model needs to be deployed for each scenario, resulting in low deployment efficiency. For this reason, Figure 5 As shown, the prompt information can be used to replace the differentiated deployment of different scenarios in the above business. The base model can also be a multimodal data processing model. Different scenarios only need to train a different prompt information Prompt, and it can be superimposed with the input data. In this way, the multimodal data processing model only needs to be deployed once, and different prompt information Prompts only need to be set when different scenarios are accessed. The prompt information Promt can be deployed in a lightweight manner, and only needs to be superimposed with the input data. Based on this, the scene information corresponding to the target data can be obtained.

[0073] In step S1044, based on the scene information corresponding to the target data, a prompt information construction rule matching the scene information is determined.

[0074] The prompt information construction rule may be a rule for constructing the prompt information Prompt.

[0075] In implementation, the correspondence between different scenarios and prompt information construction rules can be pre-set according to actual conditions. After the scenario information corresponding to the target data is obtained, the prompt information construction rules corresponding to the scenario information can be found from the above correspondence.

[0076] In step S1046, a rule and target data are constructed based on the determined prompt information to generate prompt information corresponding to the target data.

[0077] Among them, the prompt information corresponding to the generated target data can be information that can effectively guide the multimodal data processing model to solve specific tasks. The prompt information is usually a descriptive text message, which will be combined with the input data (image data or text data, etc.) to help the multimodal data processing model better understand the tasks to be performed. For example, in the risk prevention and control scenario of image content, if the multimodal data processing model is required to detect the non-compliant content contained therein, the prompt information can be such as "Does this image contain non-compliant content?"; in the risk prevention and control scenario of text content, the prompt information can be such as "Does this text contain discriminatory remarks?" The multimodal data processing model will adjust its output results according to the above prompt information to adapt to different task requirements.

[0078] In step S1048, the prompt information and the target data are input into the first sub-network of the encoding network corresponding to the first modality in the multimodal data processing model to obtain a first intermediate result output by the first sub-network.

[0079] A significant advantage of using the prompt method is the lightweight deployment. Compared with training a specific model or fine-tuning the model, designing an effective prompt only requires a small amount of text data and does not require changing the architecture of the multimodal data processing model. Deploying a new prompt usually only requires combining it with the input data and processing it through the base model, thereby achieving rapid online and adaptable deployment.

[0080] In practical applications, the above prompt information construction rules can be constructed in the following manner, and for details, please refer to the processing of the following steps C2 to C6.

[0081] In step C2, a prompt information construction rule is constructed, and the constructed prompt information construction rule is initialized to obtain an initial prompt information construction rule.

[0082] In implementation, blank prompt information construction rules or prompt information construction rules that lack key information (such as key parameter information, etc.) to be supplemented can be constructed based on actual conditions or expert experience. Then, the constructed prompt information construction rules can be initialized using preset parameters to obtain initial prompt information construction rules.

[0083] In step C4, based on the initial prompt information construction rule and the sample data of the scene information corresponding to the initial prompt information construction rule, a first prompt information sample is generated, and label information corresponding to the first prompt information sample is determined.

[0084] In step C6, based on the first prompt information sample and the above-mentioned label information, the prompt information construction rule is fine-tuned using the multimodal data processing model to obtain a fine-tuned prompt information construction rule.

[0085] In practical applications, the prompt information construction rules constructed based on the above method can also be transferred to the following methods: Figure 6 As shown, please refer to the processing of steps D2 and D4 below for details.

[0086] In step D2, second prompt information is generated based on the fine-tuned prompt information construction rules and target sample data for model training of the multimodal data processing sub-model, where the multimodal data processing sub-model is a sub-model of the multimodal data processing model.

[0087] Among them, the multimodal data processing sub-model is a sub-model of the multimodal data processing model. The multimodal data processing sub-model can be generated in a variety of different ways. For example, the multimodal data processing model can be trained by knowledge distillation to obtain a modal data processing sub-model, or the multimodal data processing sub-model can be generated by federated learning, etc. The specific setting can be based on actual conditions.

[0088] In step D4, the multimodal data processing sub-model is trained based on the second prompt information and the target sample data to obtain a trained multimodal data processing sub-model.

[0089] In implementation, Figure 6 As shown, the prompt information applied to the larger model can be used as the prompt information for initialization of the smaller model by means of transfer learning, that is, the prompt information Prompt trained or optimized on a large pre-trained model can be transferred to other smaller models as the initialization of the prompt information Prompt of these models. Specifically, the second prompt information is generated based on the fine-tuned prompt information construction rules and the target sample data for model training of the multimodal data processing sub-model. The second prompt information is the prompt information Prompt trained or optimized on the large pre-trained model (i.e., the multimodal data processing model). The second prompt information can be used as the prompt information Prompt for initialization of the multimodal data processing sub-model, and the multimodal data processing sub-model can be trained based on the second prompt information and the target sample data to obtain the trained multimodal data processing sub-model. In theory, the above method can use the rich knowledge learned by the multimodal data processing model to help the smaller model (i.e., the multimodal data processing sub-model) better adapt to specific tasks, especially when there is less available training data.

[0090] The processing of the above steps D2 and D4 inherits the prompt information Prompt of the larger model (i.e., the multimodal data processing model) as an initialization prompt information, which can accelerate the training of the smaller model (i.e., the multimodal data processing sub-model). The smaller model can directly benefit from the knowledge learned by the larger model, reducing the exploration time and resource consumption during its own training process. In addition, the performance of the smaller model can also be improved. The larger model is usually pre-trained on a wider range of and diverse data, so it has better generalization ability. Passing the knowledge therein to the smaller model through the prompt information Prompt can enhance the processing ability of the smaller model for specific tasks. In addition, the dependence on a large amount of labeled data can be reduced. The prompt information Prompt inherited from the larger model can be used to reduce the smaller model's dependence on a large amount of labeled data, because the larger model has learned a wide range of language patterns and knowledge on rich data.

[0091] The embodiment of the present specification provides a data processing method, by acquiring target data belonging to a first modality applied to a multimodal data processing model, the multimodal data processing model includes a coding network and a data processing network corresponding to each modality in a plurality of modalities, the coding network corresponding to each modality is composed of a plurality of network layers connected in series, the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, each sub-network includes at least one network layer, each sub-network in the plurality of sub-networks connected in series except the last sub-network is respectively connected to the corresponding early-leaving network and the next sub-network, the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, using the first sub-network of the coding network corresponding to the first modality in the multimodal data processing model, determine a first intermediate result output by the first sub-network, and input the first intermediate result into the first early-leaving network connected to the first sub-network to obtain a first output result, wherein if the first input If the result meets the output condition corresponding to the first early-leaving network, the multimodal data processing model is terminated to continue processing the first output result, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the multimodal data processing model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the multimodal data processing model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the multimodal data processing model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete multimodal data processing model, which can not only significantly reduce the amount of calculation required to process simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the multimodal data processing model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0092] The following is a detailed description of a data processing method provided in an embodiment of this specification in conjunction with a specific application scenario, wherein a multimodal data processing model may be a risk identification model applied to multimodal data, and the multimodal data includes text data in a text modality and image data in an image modality.

[0093] like Figure 7As shown, an embodiment of this specification provides a data processing method, and the execution subject of the method can be a terminal device or a server, etc., wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, or a computer device such as a laptop or a desktop computer, or an IoT device (specifically such as a smart watch, a car-mounted device, etc.), etc., wherein the server can be an independent server, or a server cluster composed of multiple servers, etc., and the server can be a background server for financial services or online shopping services, or a background server for an application, etc. In this embodiment, the execution subject is taken as an example for detailed description. For the case where the execution subject is a terminal device, please refer to the following server situation processing, which will not be repeated here. The method can specifically include the following steps:

[0094] In step S702, target data belonging to a first modality of a risk identification model applied to multimodal data is obtained, the target data of the first modality is text data of a text modality or image data of an image modality, the risk identification model includes an encoding network and a data processing network corresponding to each of the multiple modalities, the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, the encoding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each of the plurality of sub-networks connected in series except the last sub-network, the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, each sub-network includes at least one network layer, each of the plurality of sub-networks connected in series except the last sub-network is respectively connected to a corresponding early-leaving network and a next sub-network, and the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements.

[0095] In implementation, the risk identification model applied to multimodal data is a model constructed based on the CLIP model. Specifically, an encoding network and a data processing network can be set in the CLIP model. The encoding network can be constructed by multiple network layers, and each network layer can be composed of one or more encoding modules composed of Transformer modules. Multiple early exit points can be set on the basis of the above-mentioned original CLIP model structure, that is, in the above-mentioned original CLIP model structure, multiple early exit points can be set after different network layers of the Transformer module, and each early exit point can include an early exit network. The task of the early exit network can be to predict the correlation between image data and text data (such as the value of logit) based on the representation output by the current network layer. Each early exit network can set a corresponding output condition to determine whether to stop the calculation of the subsequent network layer in advance, and directly output the output result of the current early exit network as the final result. Based on the structure of the above-mentioned risk identification model, the target data belonging to the first modality of the risk identification model applied to multimodal data can be obtained. The specific processing method can refer to the above-mentioned related content, which will not be repeated here.

[0096] In step S704, based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the risk identification model is used to determine the first intermediate result output by the first subnetwork, and the first intermediate result is input into the first early-exit network connected to the first subnetwork to obtain the first output result.

[0097] In step S706, if the first output result satisfies the output condition corresponding to the first early-leaving network, the processing of the first output result by the risk identification model is terminated, and the first output result is used as the output result of the risk identification model.

[0098] In step S708, if the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0099] The specific processing process of the above steps S704 to S708 can be found in the above related content and will not be repeated here.

[0100] In practical applications, each network layer includes one or more serially connected encoding modules, the encoding modules included in different network layers have the same structure, and the encoding modules included in each network layer have the same structure.

[0101] In practical applications, the risk identification model applied to multimodal data can be a model built based on the CLIP model, in which the encoding module is a Transformer module, and the early-retirement network is a network of linear classifiers or a network of machine learning models.

[0102] In practical applications, the early-falling network can be trained in a variety of different ways. The early-falling network can be trained together with the CLIP model during the training process, or it can be trained independently after the CLIP model pre-training is completed. Their goal is to classify or predict the data at each layer as accurately as possible. The specific method of training together with the CLIP model can refer to the processing of steps E2 to E8 below, and the specific method of training independently after the CLIP model pre-training is completed can refer to the processing of steps F2 to F10 below.

[0103] In step E2, sample data belonging to the second modality for training the risk identification model is obtained.

[0104] In step E4, based on the sample data belonging to the second modality, the second subnetwork of the encoding network corresponding to the second modality in the risk identification model is used to determine the second intermediate result sample output by the second subnetwork, and the second intermediate result sample is input into the second early-exit network connected to the second subnetwork to obtain a second output result sample.

[0105] In step E6, if the second output result sample meets the output condition corresponding to the second early-leaving network, the processing of the second output result sample through the risk identification model is terminated, and the second output result sample is used as the output result sample of the risk identification model; if the second output result sample does not meet the output condition corresponding to the second early-leaving network, the second output result sample is input into the next sub-network as sample data belonging to the second modality for processing until the output result sample of the risk identification model is obtained.

[0106] In step E8, the risk identification model is trained based on the output result samples of the risk identification model to obtain a trained risk identification model.

[0107] The specific processing process of the above steps E2 to E8 can be found in the above related content and will not be repeated here.

[0108] In step F2, sample data belonging to the third modality for training the risk identification model is obtained.

[0109] In step F4, the early-falling network in the risk identification model is shielded, and the shielded risk identification model is pre-trained based on sample data belonging to the third modality to obtain a pre-trained shielded risk identification model.

[0110] Among them, the shielded risk identification model obtained by shielding the early-retirement network in the risk identification model can be the CLIP model, and the pre-trained shielded risk identification model is the pre-trained CLIP model.

[0111] In step F6, sample data belonging to the fourth modality is obtained.

[0112] In step F8, the sample data belonging to the fourth modality is input into the pre-trained masked risk identification model to obtain the encoded data output by the encoding network corresponding to the fourth modality in the pre-trained masked risk identification model.

[0113] In step F10, the model parameters in the pre-trained masked risk identification model are fixed, and knowledge distillation training is performed on the early-retirement network in the risk identification model based on the encoded data to obtain a trained multimodal data processing model.

[0114] The specific processing procedures of the above steps F2 to F10 can be found in the above related contents and will not be repeated here.

[0115] In practical applications, in the above step S704, based on the target data, the first subnetwork of the encoding network corresponding to the first modality in the risk identification model is used to determine the specific processing method of the first intermediate result output by the first subnetwork. The following is an optional processing method, such as Figure 8 As shown, the processing may specifically include the following steps S7042 to S7048.

[0116] In step S7042, scenario information corresponding to the target data is obtained, where the scenario information includes one or more of fraud, privacy data leakage, illegal financial activities, and discrimination.

[0117] In implementation, risk prevention and control business or risk identification business often includes many different scenarios. Usually, an adapted model needs to be deployed for each scenario, resulting in low deployment efficiency. For this reason, the prompt information can be used to replace the differentiated deployment of different scenarios in the above business. The base model can also be a CLIP model (or risk identification model). Different scenarios only need to train a different prompt information Prompt, and it can be superimposed with the input data. In this way, the CLIP model (or risk identification model) only needs to be deployed once, and different prompt information Prompts only need to be set when different scenarios are connected. The prompt information Promt can be deployed in a lightweight manner, and it only needs to be superimposed with the input data. Based on this, the scene information corresponding to the target data can be obtained.

[0118] In step S7044, based on the scene information corresponding to the target data, a prompt information construction rule matching the scene information is determined.

[0119] In step S7046, rules and target data are constructed based on the determined prompt information to generate prompt information corresponding to the target data.

[0120] In step S7048, the prompt information and the target data are input into the first sub-network of the encoding network corresponding to the first modality in the risk identification model to obtain a first intermediate result output by the first sub-network.

[0121] The specific processing process of the above steps S7042 to S7048 can be found in the above related content and will not be repeated here.

[0122] In practical applications, the prompt information construction rules can be determined by fine-tuning, and details can be found in the processing of the following steps G2 to G6.

[0123] In step G2, a prompt information construction rule is constructed, and the constructed prompt information construction rule is initialized to obtain an initial prompt information construction rule.

[0124] In step G4, based on the initial prompt information construction rule and the sample data of the scene information corresponding to the initial prompt information construction rule, a first prompt information sample is generated, and label information corresponding to the first prompt information sample is determined.

[0125] In step G6, based on the first prompt information sample and the label information, the risk identification model is used to fine-tune the prompt information construction rule to obtain a fine-tuned prompt information construction rule.

[0126] The specific processing process of the above steps G2 to G6 can be found in the above related content and will not be repeated here.

[0127] In practical applications, based on the prompt information construction rules constructed in the above manner, transfer learning can also be performed in the following manner, and specific reference is made to the processing of the following steps H2 and H4.

[0128] In step H2, second prompt information is generated based on the fine-tuned prompt information construction rules and the target sample data for model training of the risk identification sub-model, where the risk identification sub-model is a sub-model of the risk identification model.

[0129] Among them, the risk identification sub-model is a sub-model of the risk identification model. The risk identification sub-model can be generated in a variety of different ways. For example, the risk identification model can be trained through knowledge distillation to obtain the risk identification sub-model, or the risk identification sub-model can be generated through federated learning, etc. The specific settings can be based on actual conditions.

[0130] In step H4, the risk identification sub-model is trained based on the second prompt information and the target sample data to obtain a trained risk identification sub-model.

[0131] The specific processing procedures of the above steps H2 and H4 can be found in the above-mentioned related contents and will not be repeated here.

[0132] The embodiment of the present specification provides a data processing method, by acquiring target data belonging to a first modality of a risk identification model applied to multimodal data, the target data of the first modality is text data of a text modality or image data of an image modality, the risk identification model includes a coding network and a data processing network corresponding to each modality of the multiple modalities, the coding network corresponding to each modality is composed of a plurality of network layers connected in series, the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, each sub-network includes at least one network layer, each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to the corresponding early-leaving network and the next sub-network, the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, the first sub-network of the coding network corresponding to the first modality in the risk identification model is used to determine a first intermediate result output by the first sub-network, and the first intermediate result is input to the first sub-network connected to the first sub-network. In an early-leaving network, a first output result is obtained, wherein if the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result through the risk identification model is terminated, and the first output result is used as the output result of the risk identification model; if the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the risk identification model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the risk identification model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the risk identification model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete risk identification model, which can not only significantly reduce the amount of calculation required for processing simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the risk identification model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0133] The above is a data processing method provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Fig. 9 shown.

[0134] The data processing device comprises: a data acquisition module 901, an early leaving processing module 902, an early leaving module 903 and an early leaving decision module 904, wherein:

[0135] A data acquisition module 901 is used to acquire target data belonging to a first modality of a risk identification model applied to multimodal data. The target data of the first modality is text data of a text modality or image data of an image modality. The risk identification model includes a coding network and a data processing network corresponding to each modality of the multiple modalities. The coding network corresponding to each modality is composed of a plurality of network layers connected in series. The coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series. The output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network. Each sub-network includes at least one of the network layers. Each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network. The early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements.

[0136] The early-leaving processing module 902 determines a first intermediate result output by the first subnetwork of the encoding network corresponding to the first modality in the risk identification model based on the target data, and inputs the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result;

[0137] The early exit module 903 stops processing the first output result through the risk identification model if the first output result satisfies the output condition corresponding to the first early exit network, and uses the first output result as the output result of the risk identification model;

[0138] Early exit decision module 904, if the first output result does not meet the output condition, then the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the risk identification model is obtained.

[0139] In the embodiments of the present specification, each network layer includes one or more serially connected encoding modules, the encoding modules included in different network layers have the same structure, and the encoding modules included in each network layer have the same structure.

[0140] In the embodiments of this specification, the risk identification model is a model constructed based on the CLIP model, the encoding module is a Transformer module, and the early-retirement network is a linear classifier network.

[0141] In the embodiment of this specification, the device further includes:

[0142] A first sample acquisition module, which acquires sample data belonging to the second modality for training the risk identification model;

[0143] A sample processing module, based on the sample data belonging to the second modality, uses the second subnetwork of the encoding network corresponding to the second modality in the risk identification model to determine a second intermediate result sample output by the second subnetwork, and inputs the second intermediate result sample into a second early-leaving network connected to the second subnetwork to obtain a second output result sample;

[0144] The model processing module terminates the processing of the second output result sample through the risk identification model if the second output result sample satisfies the output condition corresponding to the second early-leaving network, and uses the second output result sample as the output result sample of the risk identification model; if the second output result sample does not satisfy the output condition corresponding to the second early-leaving network, the second output result sample is input into the next sub-network as sample data belonging to the second modality for processing until the output result sample of the risk identification model is obtained;

[0145] The first training module trains the risk identification model based on the output result samples of the risk identification model to obtain a trained risk identification model.

[0146] In the embodiment of this specification, the device further includes:

[0147] A second sample acquisition module, which acquires sample data belonging to a third modality for training the risk identification model;

[0148] A second training module is used to shield the early-falling network in the risk identification model, and pre-train the shielded risk identification model based on the sample data belonging to the third modality to obtain a pre-trained shielded risk identification model;

[0149] A third sample acquisition module, which acquires sample data belonging to a fourth modality;

[0150] A sample encoding module, inputting the sample data belonging to the fourth modality into the pre-trained shielded risk identification model, and obtaining the encoded data output by the encoding network corresponding to the fourth modality in the pre-trained shielded risk identification model;

[0151] The third training module fixes the model parameters in the pre-trained shielded risk identification model, and performs knowledge distillation training on the early-retirement network in the risk identification model based on the encoded data to obtain a trained risk identification model.

[0152] In the embodiment of this specification, the early leave processing module 902 includes:

[0153] A scenario determination unit, which obtains scenario information corresponding to the target data, wherein the scenario information includes one or more of fraud, privacy data leakage, illegal financial activities, and discrimination;

[0154] a prompt information construction rule determination unit, which determines, based on the scene information corresponding to the target data, a prompt information construction rule that matches the scene information;

[0155] a prompt information generating unit, which generates prompt information corresponding to the target data based on the determined prompt information building rule and the target data;

[0156] The result determination unit inputs the prompt information and the target data into a first sub-network of an encoding network corresponding to the first modality in a risk identification model to obtain a first intermediate result output by the first sub-network.

[0157] In the embodiment of this specification, the device further includes:

[0158] An initialization module constructs the prompt information construction rule and performs initialization processing on the constructed prompt information construction rule to obtain an initial prompt information construction rule;

[0159] A prompt sample generating module, which generates a first prompt information sample based on the initial prompt information building rule and the sample data of the scene information corresponding to the initial prompt information building rule, and determines the label information corresponding to the first prompt information sample;

[0160] A fine-tuning module, based on the first prompt information sample and the label information, uses the risk identification model to fine-tune the prompt information construction rule to obtain a fine-tuned prompt information construction rule.

[0161] In the embodiment of this specification, the device further includes:

[0162] a prompt information generating module, which generates second prompt information based on the fine-tuned prompt information building rule and target sample data for model training of a risk identification sub-model, wherein the risk identification sub-model is a sub-model of the risk identification model;

[0163] The transfer learning module trains the risk identification sub-model based on the second prompt information and the target sample data to obtain a trained risk identification sub-model.

[0164] The embodiment of the present specification provides a data processing device, which obtains target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes a coding network and a data processing network corresponding to each modality of the multiple modalities, wherein the coding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one network layer, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to the corresponding early-leaving network and the next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, using the first sub-network of the coding network corresponding to the first modality in the multimodal data processing model, a first intermediate result output by the first sub-network is determined, and the first intermediate result is input into the first early-leaving network connected to the first sub-network to obtain a first output result, wherein if the first input If the result meets the output condition corresponding to the first early-leaving network, the multimodal data processing model is terminated to continue processing the first output result, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the multimodal data processing model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the multimodal data processing model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the multimodal data processing model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete multimodal data processing model, which can not only significantly reduce the amount of calculation required to process simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the multimodal data processing model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0165] Based on the same idea, the embodiment of this specification also provides a data processing device, such as Fig.10 shown.

[0166] The data processing device comprises: a data acquisition module 1001, a processing module 1002, an early leaving module 1003 and a decision module 1004, wherein:

[0167] A data acquisition module 1001 is used to acquire target data belonging to a first mode applied to a multimodal data processing model, wherein the multimodal data processing model includes a coding network and a data processing network corresponding to each mode of the multiple modes, wherein the coding network corresponding to each mode is composed of a plurality of network layers connected in series, wherein the coding network corresponding to each mode includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements;

[0168] The processing module 1002 determines a first intermediate result output by the first subnetwork of the encoding network corresponding to the first modality in the multimodal data processing model based on the target data, and inputs the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result;

[0169] The early exit module 1003 stops processing the first output result by the multimodal data processing model if the first output result satisfies the output condition corresponding to the first early exit network, and uses the first output result as the output result of the multimodal data processing model;

[0170] Decision module 1004: if the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

[0171] In the embodiments of the present specification, each network layer includes one or more serially connected encoding modules, the encoding modules included in different network layers have the same structure, and the encoding modules included in each network layer have the same structure.

[0172] In the embodiments of the present specification, the multimodal data processing model is a model constructed based on the CLIP model, the encoding module is a Transformer module, the target data belonging to the first modality is text data belonging to the text modality or image data belonging to the image modality, the multimodal data processing model is a model for risk prevention and control, and the early-leaving network is a linear classifier network.

[0173] In the embodiment of this specification, the device further includes:

[0174] A first sample acquisition module, which acquires sample data belonging to the second modality for training the multimodal data processing model;

[0175] A sample processing module, based on the sample data belonging to the second modality, uses a second sub-network of the encoding network corresponding to the second modality in the multimodal data processing model to determine a second intermediate result sample output by the second sub-network, and inputs the second intermediate result sample into a second early-leaving network connected to the second sub-network to obtain a second output result sample;

[0176] A model processing module, if the second output result sample satisfies the output condition corresponding to the second early-leaving network, then the processing of the second output result sample by the multimodal data processing model is terminated, and the second output result sample is used as the output result sample of the multimodal data processing model; if the second output result sample does not satisfy the output condition corresponding to the second early-leaving network, then the second output result sample is input into the next sub-network as sample data belonging to the second modality for processing, until the output result sample of the multimodal data processing model is obtained;

[0177] The first training module trains the multimodal data processing model based on the output result samples of the multimodal data processing model to obtain a trained multimodal data processing model.

[0178] In the embodiment of this specification, the device further includes:

[0179] A second sample acquisition module, acquiring sample data belonging to a third modality for training the multimodal data processing model;

[0180] A second training module is configured to shield the early-falling network in the multimodal data processing model, and pre-train the shielded multimodal data processing model based on the sample data belonging to the third modality to obtain a pre-trained shielded multimodal data processing model;

[0181] A third sample acquisition module, which acquires sample data belonging to a fourth modality;

[0182] A sample encoding module, inputting the sample data belonging to the fourth modality into a pre-trained shielded multimodal data processing model, and obtaining encoded data output by an encoding network corresponding to the fourth modality in the pre-trained shielded multimodal data processing model;

[0183] The third training module fixes the model parameters in the pre-trained shielded multimodal data processing model, and performs knowledge distillation training on the early-falling network in the multimodal data processing model based on the encoded data to obtain a trained multimodal data processing model.

[0184] In the embodiment of this specification, the processing module 1002 includes:

[0185] A scenario determination unit, which obtains scenario information corresponding to the target data, wherein the scenario information includes one or more of fraud, privacy data leakage, illegal financial activities, and discrimination;

[0186] a prompt information construction rule determination unit, which determines, based on the scene information corresponding to the target data, a prompt information construction rule that matches the scene information;

[0187] a prompt information generating unit, which generates prompt information corresponding to the target data based on the determined prompt information building rule and the target data;

[0188] The result determination unit inputs the prompt information and the target data into a first subnetwork of an encoding network corresponding to the first modality in a multimodal data processing model to obtain a first intermediate result output by the first subnetwork.

[0189] In the embodiment of this specification, the device further includes:

[0190] An initialization module constructs the prompt information construction rule and performs initialization processing on the constructed prompt information construction rule to obtain an initial prompt information construction rule;

[0191] A prompt sample generating module, which generates a first prompt information sample based on the initial prompt information building rule and the sample data of the scene information corresponding to the initial prompt information building rule, and determines the label information corresponding to the first prompt information sample;

[0192] A fine-tuning module, based on the first prompt information sample and the label information, uses the multimodal data processing model to fine-tune the prompt information construction rule to obtain a fine-tuned prompt information construction rule.

[0193] In the embodiment of this specification, the device further includes:

[0194] a prompt information generating module, which generates second prompt information based on the fine-tuned prompt information construction rule and target sample data for model training of a multimodal data processing sub-model, wherein the multimodal data processing sub-model is a sub-model of the multimodal data processing model;

[0195] The transfer learning module trains the multimodal data processing sub-model based on the second prompt information and the target sample data to obtain a trained multimodal data processing sub-model.

[0196] The embodiment of the present specification provides a data processing device, which obtains target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, and the risk identification model includes a coding network and a data processing network corresponding to each modality of the multiple modalities, wherein the coding network corresponding to each modality is composed of a plurality of network layers connected in series, and the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one network layer, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to the corresponding early-leaving network and the next sub-network, and the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, the first sub-network of the coding network corresponding to the first modality in the risk identification model is used to determine a first intermediate result output by the first sub-network, and the first intermediate result is input to the first sub-network connected to the first sub-network. In an early-leaving network, a first output result is obtained, wherein if the first output result meets the output condition corresponding to the first early-leaving network, the processing of the first output result through the risk identification model is terminated, and the first output result is used as the output result of the risk identification model; if the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the risk identification model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the risk identification model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the risk identification model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete risk identification model, which can not only significantly reduce the amount of calculation required for processing simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the risk identification model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0197] The above is a data processing device provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Fig.11 shown.

[0198] The data processing device may provide a terminal device or a server, etc. for the above-mentioned embodiments.

[0199] The data processing device may have relatively large differences due to different configurations or performances, and may include one or more processors 1101 and memory 1102, and the memory 1102 may store one or more storage applications or data. Among them, the memory 1102 may be a short-term storage or a persistent storage. The application stored in the memory 1102 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the data processing device. Furthermore, the processor 1101 may be configured to communicate with the memory 1102 and execute a series of computer executable instructions in the memory 1102 on the data processing device. The data processing device may also include one or more power supplies 1103, one or more wired or wireless network interfaces 1104, one or more input and output interfaces 1105, and one or more keyboards 1106.

[0200] Specifically in this embodiment, the data processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions in the data processing device, and the one or more programs are configured to be executed by one or more processors, including computer executable instructions for performing the following:

[0201] Obtain target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each of the multiple modalities, wherein the encoding network corresponding to each of the multiple modalities is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each of the multiple modalities includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each of the multiple sub-networks connected in series except the last sub-network, wherein the output end of the last sub-network in the multiple sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each of the multiple sub-networks connected in series except the last sub-network is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements;

[0202] Based on the target data, using a first subnetwork of the encoding network corresponding to the first modality in the risk identification model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result;

[0203] If the first output result satisfies the output condition corresponding to the first early-leaving network, then terminating the processing of the first output result by the risk identification model, and using the first output result as the output result of the risk identification model;

[0204] If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0205] In addition, specifically in this embodiment, the data processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions in the data processing device, and the one or more programs configured to be executed by one or more processors include computer executable instructions for performing the following:

[0206] Obtain target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes an encoding network and a data processing network corresponding to each modality of the multiple modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements;

[0207] Based on the target data, using a first subnetwork of an encoding network corresponding to the first modality in a multimodal data processing model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-retirement network connected to the first subnetwork to obtain a first output result;

[0208] If the first output result satisfies the output condition corresponding to the first early-leaving network, terminating the processing of the first output result by the multimodal data processing model, and using the first output result as the output result of the multimodal data processing model;

[0209] If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

[0210] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0211] The embodiment of the present specification provides a data processing device, which obtains target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes a coding network and a data processing network corresponding to each modality of the multiple modalities, wherein the coding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one network layer, wherein each sub-network in the plurality of sub-networks connected in series except the last sub-network is respectively connected to the corresponding early-leaving network and the next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, using the first sub-network of the coding network corresponding to the first modality in the multimodal data processing model, a first intermediate result output by the first sub-network is determined, and the first intermediate result is input into the first early-leaving network connected to the first sub-network to obtain a first output result, wherein if the first input If the result meets the output condition corresponding to the first early-leaving network, the multimodal data processing model is terminated to continue processing the first output result, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the multimodal data processing model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the multimodal data processing model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the multimodal data processing model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete multimodal data processing model, which can not only significantly reduce the amount of calculation required to process simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the multimodal data processing model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0212] Furthermore, based on the above Figures 1 to 8 In one embodiment, the present specification further provides a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be implemented:

[0213] Obtain target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each of the multiple modalities, wherein the encoding network corresponding to each of the multiple modalities is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each of the multiple modalities includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each of the multiple sub-networks connected in series except the last sub-network, wherein the output end of the last sub-network in the multiple sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each of the multiple sub-networks connected in series except the last sub-network is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements;

[0214] Based on the target data, using a first subnetwork of the encoding network corresponding to the first modality in the risk identification model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result;

[0215] If the first output result satisfies the output condition corresponding to the first early-leaving network, then terminating the processing of the first output result by the risk identification model, and using the first output result as the output result of the risk identification model;

[0216] If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0217] In addition, in another specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc., and the computer executable instruction information stored in the storage medium can implement the following process when executed by the processor:

[0218] Obtain target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes an encoding network and a data processing network corresponding to each modality of the multiple modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements;

[0219] Based on the target data, using a first subnetwork of an encoding network corresponding to the first modality in a multimodal data processing model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-retirement network connected to the first subnetwork to obtain a first output result;

[0220] If the first output result satisfies the output condition corresponding to the first early-leaving network, terminating the processing of the first output result by the multimodal data processing model, and using the first output result as the output result of the multimodal data processing model;

[0221] If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

[0222] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0223] The embodiment of the present specification provides a storage medium, which obtains target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes a coding network and a data processing network corresponding to each modality of the multiple modalities, wherein the coding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one network layer, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, using the first sub-network of the coding network corresponding to the first modality in the multimodal data processing model, a first intermediate result output by the first sub-network is determined, and the first intermediate result is input into a first early-leaving network connected to the first sub-network to obtain a first output result, wherein if the first output If the result meets the output condition corresponding to the first early-leaving network, the multimodal data processing model is terminated to continue processing the first output result, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the multimodal data processing model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the multimodal data processing model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the multimodal data processing model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete multimodal data processing model, which can not only significantly reduce the amount of calculation required to process simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the multimodal data processing model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0224] Furthermore, based on the above Figures 1 to 8 In one or more embodiments of the present specification, a computer program product is provided, including a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented:

[0225] Obtain target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each of the multiple modalities, wherein the encoding network corresponding to each of the multiple modalities is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each of the multiple modalities includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each of the multiple sub-networks connected in series except the last sub-network, wherein the output end of the last sub-network in the multiple sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each of the multiple sub-networks connected in series except the last sub-network is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements;

[0226] Based on the target data, using a first subnetwork of the encoding network corresponding to the first modality in the risk identification model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result;

[0227] If the first output result satisfies the output condition corresponding to the first early-leaving network, then terminating the processing of the first output result by the risk identification model, and using the first output result as the output result of the risk identification model;

[0228] If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

[0229] In addition, in another specific embodiment, the computer program product includes a computer program, and when the computer program in the computer program product is executed by a processor, it can implement the following process:

[0230] Obtain target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes an encoding network and a data processing network corresponding to each modality of the multiple modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements;

[0231] Based on the target data, using a first subnetwork of an encoding network corresponding to the first modality in a multimodal data processing model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-retirement network connected to the first subnetwork to obtain a first output result;

[0232] If the first output result satisfies the output condition corresponding to the first early-leaving network, terminating the processing of the first output result by the multimodal data processing model, and using the first output result as the output result of the multimodal data processing model;

[0233] If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

[0234] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0235] The embodiment of the present specification provides a computer program product, which obtains target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes a coding network and a data processing network corresponding to each modality of the multiple modalities, wherein the coding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the coding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one network layer, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to the corresponding early-leaving network and the next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements, and then, based on the target data, using the first sub-network of the coding network corresponding to the first modality in the multimodal data processing model, a first intermediate result output by the first sub-network is determined, and the first intermediate result is input into the first early-leaving network connected to the first sub-network to obtain a first output result, wherein if the first input If the result meets the output condition corresponding to the first early-leaving network, the multimodal data processing model is terminated to continue processing the first output result, and the first output result is used as the output result of the multimodal data processing model. If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first mode for processing until the output result of the multimodal data processing model is obtained. In this way, data of different difficulty levels can be processed at different levels in the network of the multimodal data processing model, that is, simple data can output the corresponding results in advance from the shallower early-leaving network in the multimodal data processing model, and difficult data or complex data may require a deeper network structure to provide sufficient representation capabilities, so that the early-leaving conditions can be met in the deeper network layer, and even need to be processed by a complete multimodal data processing model, which can not only significantly reduce the amount of calculation required to process simple data, but also ensure that sufficient computing resources are still given to difficult data for accurate prediction, so that the multimodal data processing model can adaptively process data of different difficulty levels, which can save computing resources and optimize the overall computing efficiency.

[0236] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0237] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0238] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0239] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0240] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0241] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0242] The embodiments of this specification are described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable fraud case serial and parallel device to produce a machine, so that the instructions executed by the processor of the computer or other programmable fraud case serial and parallel device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0243] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable fraud case serial and parallel device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0244] These computer program instructions may also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0245] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0246] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0247] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0248] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0249] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0250] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0251] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0252] The above description is only an embodiment of this specification and is not intended to limit this document. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this specification.

Claims

1. A data processing method, the method comprising: Obtain target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each of the multiple modalities, wherein the encoding network corresponding to each of the multiple modalities is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each of the multiple modalities includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each of the multiple sub-networks connected in series except the last sub-network, wherein the output end of the last sub-network in the multiple sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each of the multiple sub-networks connected in series except the last sub-network is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements; Based on the target data, using a first subnetwork of the encoding network corresponding to the first modality in the risk identification model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result; If the first output result satisfies the output condition corresponding to the first early-leaving network, then terminating the processing of the first output result by the risk identification model, and using the first output result as the output result of the risk identification model; If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

2. According to the method of claim 1, each network layer includes one or more serially connected encoding modules, the encoding modules contained in different network layers have the same structure, and the encoding modules contained in each network layer have the same structure.

3. According to the method of claim 2, the risk identification model is a model built based on the CLIP model, the encoding module is a Transformer module, and the early-retirement network is a linear classifier network.

4. The method according to claim 3, further comprising: Acquiring sample data belonging to a second modality for training the risk identification model; Based on the sample data belonging to the second modality, using the second subnetwork of the encoding network corresponding to the second modality in the risk identification model, determining a second intermediate result sample output by the second subnetwork, and inputting the second intermediate result sample into a second early-leaving network connected to the second subnetwork to obtain a second output result sample; If the second output result sample satisfies the output condition corresponding to the second early-leaving network, then the processing of the second output result sample by the risk identification model is terminated, and the second output result sample is used as the output result sample of the risk identification model; If the second output result sample does not meet the output condition corresponding to the second early-leaving network, the second output result sample is input into the next sub-network as sample data belonging to the second modality for processing until the output result sample of the risk identification model is obtained; Based on the output result samples of the risk identification model, the risk identification model is trained to obtain a trained risk identification model.

5. The method according to claim 3, further comprising: Acquiring sample data belonging to a third modality for training the risk identification model; Shielding the early-falling network in the risk identification model, and pre-training the shielded risk identification model based on the sample data belonging to the third modality to obtain a pre-trained shielded risk identification model; Obtaining sample data belonging to the fourth mode; Inputting the sample data belonging to the fourth modality into the pre-trained shielded risk identification model, and obtaining the encoded data output by the encoding network corresponding to the fourth modality in the pre-trained shielded risk identification model; The model parameters in the pre-trained shielded risk identification model are fixed, and knowledge distillation training is performed on the early-retirement network in the risk identification model based on the encoded data to obtain a trained risk identification model.

6. The method according to claim 3 or 4, wherein based on the target data, using the first subnetwork of the encoding network corresponding to the first modality in the risk identification model, determining the first intermediate result output by the first subnetwork comprises: Acquiring scenario information corresponding to the target data, wherein the scenario information includes one or more of fraud, privacy data leakage, illegal financial activities, and discrimination; Based on the scene information corresponding to the target data, determining a prompt information construction rule matching the scene information; Based on the determined prompt information construction rule and the target data, generating prompt information corresponding to the target data; The prompt information and the target data are input into a first subnetwork of an encoding network corresponding to the first modality in a risk identification model to obtain a first intermediate result output by the first subnetwork.

7. The method according to claim 6, further comprising: Constructing the prompt information construction rule, and initializing the constructed prompt information construction rule to obtain an initial prompt information construction rule; Generate a first prompt information sample based on the initial prompt information construction rule and sample data of the scene information corresponding to the initial prompt information construction rule, and determine label information corresponding to the first prompt information sample; Based on the first prompt information sample and the label information, the prompt information construction rule is fine-tuned using the risk identification model to obtain a fine-tuned prompt information construction rule.

8. The method according to claim 7, further comprising: generating second prompt information based on the fine-tuned prompt information construction rule and target sample data for model training of a risk identification sub-model, the risk identification sub-model being a sub-model of the risk identification model; The risk identification sub-model is trained based on the second prompt information and the target sample data to obtain a trained risk identification sub-model.

9. A data processing method, the method comprising: Obtain target data belonging to a first modality applied to a multimodal data processing model, wherein the multimodal data processing model includes an encoding network and a data processing network corresponding to each modality of the multiple modalities, wherein the encoding network corresponding to each modality is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each modality includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each sub-network except the last sub-network in the plurality of sub-networks connected in series, wherein the output end of the last sub-network in the plurality of sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each sub-network except the last sub-network in the plurality of sub-networks connected in series is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements; Based on the target data, using a first subnetwork of an encoding network corresponding to the first modality in a multimodal data processing model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-retirement network connected to the first subnetwork to obtain a first output result; If the first output result satisfies the output condition corresponding to the first early-leaving network, terminating the processing of the first output result by the multimodal data processing model, and using the first output result as the output result of the multimodal data processing model; If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the multimodal data processing model is obtained.

10. A data processing device, comprising: A data acquisition module, which acquires target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each of the multiple modalities, wherein the encoding network corresponding to each of the multiple modalities is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each of the multiple modalities includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each of the multiple sub-networks connected in series except the last sub-network, wherein the output end of the last sub-network in the multiple sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each of the multiple sub-networks connected in series except the last sub-network is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements; An early-leaving processing module, based on the target data, uses a first subnetwork of the encoding network corresponding to the first modality in the risk identification model to determine a first intermediate result output by the first subnetwork, and inputs the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result; an early exit module, which stops processing the first output result through the risk identification model if the first output result satisfies the output condition corresponding to the first early exit network, and uses the first output result as the output result of the risk identification model; The early exit decision module inputs the first output result as target data belonging to the first mode into the next sub-network for processing if the first output result does not meet the output condition, until the output result of the risk identification model is obtained.

11. A data processing device, comprising: processor; as well as a memory arranged to store computer executable instructions which, when executed, cause the processor to: Obtain target data belonging to a first modality of a risk identification model applied to multimodal data, wherein the target data of the first modality is text data of a text modality or image data of an image modality, wherein the risk identification model includes an encoding network and a data processing network corresponding to each of the multiple modalities, wherein the encoding network corresponding to each of the multiple modalities is composed of a plurality of network layers connected in series, wherein the encoding network corresponding to each of the multiple modalities includes a plurality of sub-networks connected in series and an early-leaving network corresponding to each of the multiple sub-networks connected in series except the last sub-network, wherein the output end of the last sub-network in the multiple sub-networks connected in series is connected to the input end of the data processing network, wherein each sub-network includes at least one of the network layers, wherein each of the multiple sub-networks connected in series except the last sub-network is respectively connected to a corresponding early-leaving network and a next sub-network, wherein the early-leaving network and the data processing network are used to process the data output by the network layer according to the same data processing requirements; Based on the target data, using a first subnetwork of the encoding network corresponding to the first modality in the risk identification model, determining a first intermediate result output by the first subnetwork, and inputting the first intermediate result into a first early-leaving network connected to the first subnetwork to obtain a first output result; If the first output result satisfies the output condition corresponding to the first early-leaving network, then terminating the processing of the first output result by the risk identification model, and using the first output result as the output result of the risk identification model; If the first output result does not meet the output condition, the first output result is input into the next sub-network as the target data belonging to the first modality for processing until the output result of the risk identification model is obtained.

Citation Information

Patent Citations

  • Power distribution equipment anomaly detection method and device, electronic equipment and storage medium

    CN116385963A

  • Illusion relieving method and device for multi-modal large model, electronic equipment and medium

    CN119128061A