Object Processing Method, Apparatus, Readable Medium, and Electronic Device

By introducing dynamic configurations of feature extraction module, object segmentation module and task processing module in the target model, the problem of high deployment cost of different categories of media material models in the prior art is solved, and more efficient model deployment and training is achieved.

CN116339868BActive Publication Date: 2025-07-25BEIJING YOUZHUJU NETWORK TECH CO LTD

Patent Information

Application Number
CN202310317575.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2025-07-25
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

In the prior art, different categories of media materials need to be trained and iterated separately, resulting in high deployment costs.

Method used

A unified target model is adopted, including feature extraction module, multiple object segmentation modules and multiple task processing modules, corresponding to different objects and task types, and dynamically switches through the configuration module to adapt to different types of target objects and tasks.

Benefits of technology

Through a unified target model, the cost of model deployment is reduced, and the efficiency of model deployment and training flexibility is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116339868B_ABST
    Figure CN116339868B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to an object processing method, apparatus, readable medium, and electronic device. The method includes: obtaining a target object to be processed, determining the object type of the target object, determining the task type corresponding to the target task for processing the target object, and inputting the target object, object type, and task type into a pre-generated target model to obtain a target result output by the target model; wherein, the target model may include a feature extraction module, multiple object segmentation modules, and multiple task processing modules, different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types. In this way, different types of target objects and different types of target tasks can be processed through a unified target model, which facilitates model training and iteration, reduces the cost of model deployment, and improves the efficiency of model deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to an object processing method, apparatus, readable medium, and electronic device. Background Art

[0002] With the progress of information technology, various types of media materials such as texts, voices, images, and videos emerge in an endless stream, and artificial intelligence technologies can perform processing such as category recognition, content recommendation, and intelligent creation on different types of media materials.

[0003] In related technologies, for different types of media materials, respective corresponding models can be used for training and iteration, and the deployment cost is relatively high. Summary of the Invention

[0004] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the subsequent Detailed Implementation section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0005] According to a first aspect of embodiments of the present disclosure, an object processing method is provided. The method includes:

[0006] Obtaining a target object to be processed;

[0007] Determining the object type of the target object;

[0008] Determining the task type corresponding to a target task for processing the target object;

[0009] Inputting the target object, the object type, and the task type into a pre-generated target model to obtain a target result output by the target model;

[0010] Wherein, the target model includes a feature extraction module, a plurality of object segmentation modules, and a plurality of task processing modules. Different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types.

[0011] According to a second aspect of embodiments of the present disclosure, an object processing apparatus is provided. The apparatus includes:

[0012] An object acquisition module, configured to obtain a target object to be processed;

[0013] A first determination module, configured to determine the object type of the target object;

[0014] A second determination module, configured to determine the task type corresponding to a target task for processing the target object;

[0015] An object processing module, configured to input the target object, the object type, and the task type into a pre-generated target model, and obtain a target result output by the target model;

[0016] Wherein, the target model includes a feature extraction module, a plurality of object segmentation modules, and a plurality of task processing modules. Different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types.

[0017] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processing device, the steps of the method according to the first aspect of the present disclosure are implemented.

[0018] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0019] A storage device, on which a computer program is stored;

[0020] A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to the first aspect of the present disclosure.

[0021] By adopting the above technical solution, a target object to be processed is obtained, the object type of the target object is determined, the task type corresponding to the target task for processing the target object is determined, and the target object, the object type, and the task type are input into a pre-generated target model, and a target result output by the target model is obtained; wherein, the target model may include a feature extraction module, a plurality of object segmentation modules, and a plurality of task processing modules. Different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types. In this way, different types of target objects and different types of target tasks can be processed through a unified target model, which is convenient for model training and iteration, reduces the cost of model deployment, and improves the efficiency of model deployment.

[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn to scale. In the drawings:

[0024] Figure 1 is a flowchart of an object processing method shown according to an exemplary embodiment.

[0025] Figure 2 It is a schematic diagram of a target model shown according to an exemplary embodiment.

[0026] Figure 3 It is a schematic diagram of another target model shown according to an exemplary embodiment.

[0027] Figure 4 It is according to Figure 1 The flowchart of step S104 shown according to the illustrated embodiment.

[0028] Figure 5 It is a flowchart of a method for generating a target model shown according to an exemplary embodiment.

[0029] Figure 6 It is a block diagram of an object processing device shown according to an exemplary embodiment.

[0030] Figure 7 It is a block diagram of an object processing device shown according to an exemplary embodiment.

[0031] Figure 8 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0032] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0033] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0034] The term "including" and its variations used in the present disclosure are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0035] It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0036] It should be noted that the modifications of "one" and "plural" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more". In the description of this disclosure, unless otherwise stated, "plural" means two or more, and other quantifiers are similar; "at least one (item)", "one (item) or more (items)" or similar expressions refer to any combination of these items (items), including any combination of single item (item) or plural items (items). For example, at least one (item) a can represent any number of a; for another example, one (item) or more (items) of a, b and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or plural; "and / or" is a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. The character " / " indicates that the associated objects before and after are in an "or" relationship. The singular forms of "a", "one kind", "one item", "the" and "this" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0037] In the embodiments of this disclosure, although operations or steps are described in a specific order in the drawings, it should not be understood that these operations or steps are required to be performed in the specific order shown or in a serial order, nor that all the operations or steps shown are required to be performed to obtain the desired result. In the embodiments of this disclosure, these operations or steps can be performed serially; they can also be performed in parallel; or a part of these operations or steps can be performed.

[0038] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0039] It can be understood that before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to users and the authorization of users should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0040] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.

[0041] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, a pop-up window manner, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0042] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0043] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0044] The present disclosure will be described below with reference to specific embodiments.

[0045] Figure 1 is a flowchart of an object processing method shown according to an exemplary embodiment. The method can be applied to an electronic device, and the electronic device may include a terminal device, such as a smart phone, a smart wearable device, a smart speaker, a smart tablet, a PDA (Personal Digital Assistant), a CPE (Customer Premise Equipment), a personal computer, a vehicle-mounted terminal, etc.; the electronic device may also include a server, such as a local server or a cloud server. As Figure 1 shown, the method may include:

[0046] S101. Obtain a target object to be processed.

[0047] Wherein, the target object may include at least one of a target text, a target voice, a target image, and a target video.

[0048] S102. Determine the object type of the target object.

[0049] Wherein, the object type of the target object may be one or more.

[0050] S103. Determine the task type corresponding to the target task for processing the target object.

[0051] Similarly, the target task can be one or multiple. Each target task corresponds to one task type.

[0052] S104. Input the target object, object type, and task type into a pre-generated target model to obtain the target result output by the target model.

[0053] Among them, the target model can include a feature extraction module, multiple object segmentation modules, and multiple task processing modules. Different object segmentation modules correspond to different object types respectively, and different task processing modules correspond to different task types.

[0054] Figure 2 is a schematic diagram of a target model shown according to an exemplary embodiment. As Figure 2 shown, the target model can include a feature extraction module 201, an object segmentation module 202, and a task processing module 203. Among them, there can be multiple object segmentation modules, and there can also be multiple task processing modules.

[0055] In some embodiments, the target model can include multiple object segmentation modules. Each object segmentation module corresponds to one object type. Exemplarily, the multiple object segmentation modules can include at least one of the following:

[0056] A text segmentation module 2021 for segmenting text-type objects; this text segmentation module can be used to perform word segmentation on text-type objects, and this text segmentation module can also be called a Text Tokenizer.

[0057] An audio segmentation module 2022 for segmenting audio-type objects; this audio segmentation module can be used to perform segmentation processing on audio-type objects, and this audio segmentation module can also be called an AudioTokenizer.

[0058] An image segmentation module 2023 for segmenting image-type objects; this image segmentation module can be used to perform segmentation processing on image-type objects, and this image segmentation module can also be called an ImageTokenizer.

[0059] A video segmentation module 2024 for segmenting video-type objects. This video segmentation module can be used to perform segmentation processing on video-type objects, and this video segmentation module can also be called a VideoTokenizer.

[0060] In some embodiments, the target model may include a plurality of task processing modules. For example, the plurality of task processing modules may include at least one of the following: a first task processing module 2031, a second task processing module 2032, ……, an Nth task processing module 203N. Different task processing modules may be used to execute different types of tasks.

[0061] For example, the plurality of task processing modules may include task modules for text processing such as text segmentation, text recognition, text classification, or other text processing, task modules for image processing such as image classification, image detection, semantic segmentation, face recognition, or other image processing, task modules for video processing such as action recognition, scene content classification, face recognition, or other video processing, and task modules for voice processing such as speech recognition, voice control, or other voice processing.

[0062] It should be noted that the plurality of processing task modules may include one or more of the above task modules.

[0063] In this way, through Figure 2 the shown target model, target objects of different object types and target tasks of different task types can be processed to obtain a target result output by the target model.

[0064] The target result may be an object recognition result (such as text recognition, speech recognition, face recognition, action recognition, object recognition, etc.), an object classification result, or other expected results, which are not limited in this disclosure.

[0065] In some embodiments, the feature extraction module 201 may include a plurality of Transformer layers. For example, it may include K Transformer layers: Transformer layer 1, Transformer layer 2, ……, Transformer layer K.

[0066] It should be noted that the feature extraction module uses a plurality of Transformer layers as the backbone network of the entire target model. The input of the feature extraction module may be the output after the object segmentation module segments the target object, and the output of the feature extraction module may be used as the output of different task processing modules, so as to connect the object segmentation module and the task processing module. Through the self-attention mechanism of the Transformer layer, target objects of different object types can be processed.

[0067] Furthermore, the feature extraction module may include a shared layer and a candidate adaptation layer. Among them, the shared layer may include at least one Transformer layer, and the candidate adaptation layer may also include at least one Transformer layer. The shared layer may be a Transformer layer shared by multiple different task types; the candidate adaptation layer may be a Transformer layer corresponding to a specific task type. If the target task of the specific task type is to be executed, this Transformer layer can be used for feature extraction; conversely, if the target task of the specific task type is not to be executed, this Transformer layer may not be used.

[0068] In this way, through the candidate adaptation layer, the capacity of the model can be improved in large-scale tasks, and the processing performance of the model can be enhanced. Moreover, in the case of inputting sample objects of multiple different object types, information complementarity between different sample objects can be achieved through the self-attention mechanism of the Transformer.

[0069] Figure 3 It is a schematic diagram of another target model shown according to an exemplary embodiment. As Figure 3 shown, the target model may further include a configuration module 204. The configuration module 204 may be used to determine a target segmentation module for processing the target object from multiple object segmentation modules according to the object type of the target object; the configuration module 204 may also be used to determine a target processing module for processing the target object from multiple task processing modules according to the task type of the target task; the configuration module 204 may also be used to determine a target adaptation layer for processing the target object from multiple candidate adaptation layers of the feature extraction module according to the task type of the target task.

[0070] In some embodiments, the configuration module 204 may include an automated switch and a configuration management system. The configuration module may be used in the model training stage and also in the model inference application stage to configure the feature extraction module 201, the object segmentation module 202, and the task processing module 203 according to the object type and the task type.

[0071] In some embodiments, the configuration module 204 may configure the target segmentation module and the task processing module through a first switching unit 205 and a second switch 206. The above switching unit may adopt gating logic to control data diversion based on the input object type and task type. The gating logic may be predefined in the form of a configuration file. Exemplarily, the object segmentation module corresponding to each object type, the task processing module corresponding to each task type, or at least one candidate adaptation layer corresponding to each task type may be predefined.

[0072] In this way, when the model is deployed, multiple modalities and multiple tasks can share one model framework, while also taking into account the inference speed of a single model.

[0073] In some embodiments of the present disclosure, the target object may be a unimodal object, that is, the target object may only include one of target text, target speech, target image, or target video.

[0074] The object type of the target object is also one. The target tasks for processing the target object may be one or more, and the task types of different target tasks may be different.

[0075] For example, if the target object includes target text, the object type of the target object may be text type, and the task types of the target tasks for processing the target object may include at least one of text segmentation, text recognition, text classification, or other text processing task types.

[0076] For another example, if the target object includes a target image, the object type of the target object may be image type, and the task types of the target tasks for processing the target object may include at least one of image classification, image detection, semantic segmentation, face recognition, or other image processing task types.

[0077] For yet another example, if the target object includes a target video, the object type of the target object may be video type, and the task types of the target tasks for processing the target object may include at least one of action recognition, scene content classification, face recognition, or other video processing task types.

[0078] For yet another example, if the target object includes target speech, the object type of the target object may be speech type, and the task types of the target tasks for processing the target object may include at least one of speech recognition, speech control, or other speech processing task types.

[0079] In some other embodiments of the present disclosure, the target object may be a multimodal object, that is, the target object may include multiple of target text, target speech, target image, and target video. Similarly, the object type of the target object may also be multiple.

[0080] For example, the target object may include target text and target speech, the object type of the target object may include text type and speech type, and the task types of the target tasks for processing the target object may include multimodal processing tasks, such as correlation analysis of text and speech, scene recognition combining text and speech, and other tasks.

[0081] For another example, the target object may include target text, target speech, and target images. The object types of the target object may include text type, speech type, and image type. The task types of the target tasks for processing the target object may include multimodal processing tasks, such as multimodal scene recognition and multimodal object recognition that combine text, speech, and video.

[0082] Similarly, the target object, object type, and task type in the embodiments of the present disclosure may also be in other combinations, which will not be elaborated here.

[0083] Using the above method, obtain the target object to be processed, determine the object type of the target object, determine the task type corresponding to the target task for processing the target object, and input the target object, object type, and task type into a pre-generated target model to obtain the target result output by the target model. Among them, the target model may include a feature extraction module, multiple object segmentation modules, and multiple task processing modules. Different object segmentation modules correspond to different object types respectively, and different task processing modules correspond to different task types. In this way, different types of target objects and different types of target tasks can be processed through a unified target model, which is convenient for model training and iteration, reduces the cost of model deployment, and improves the efficiency of model deployment.

[0084] Figure 4 is based on Figure 1 shown in the flowchart of step S104 in the illustrated embodiment. As Figure 4 shown, the step S104 may include the following sub-steps:

[0085] S1041. Input the target object into the target segmentation module to obtain multiple first object features after segmentation.

[0086] Among them, the target segmentation module may be the object segmentation module corresponding to the object type.

[0087] Exemplarily, the target segmentation module may be determined according to the object type, and the target object is input into the target segmentation module to obtain multiple first object features after segmentation.

[0088] Taking the target object as the target text as an example, through this segmentation process, a long text (such as a sentence or a paragraph) can be segmented into multiple short texts (such as words or phrases), and first object features are generated according to the segmented short texts.

[0089] In some embodiments, for the target text, methods such as pre-trained multilingual Sentence-bert can be used for segmentation. The Sentence-bert method can include classtoken, embeddingtoken, positiontoken, etc., which respectively correspond to the category, feature, and position information of the token.

[0090] It should be noted that in the case where the target object includes target speech, target image, or target video, similar segmentation processing can also be performed. Exemplarily, the image can be segmented using pixbert.

[0091] In some embodiments, when the object type of the target object is an image type or a voice type, multiple sub-target objects can be segmented, and multiple first target object features can be determined based on the multiple sub-target objects. Among them, there may be an overlapping area between multiple adjacent sub-target objects in terms of position.

[0092] For example, when slicing the target image, a certain degree of overlap between the slices can be retained, that is, there is partial overlap between adjacent sub-target images obtained by slicing. The specific degree of overlap can be adjusted according to the image size, the number of slices, the repetition rate ratio, etc. In this way, the image correlation between the sub-target objects obtained by slicing can be better retained, thereby improving the accuracy of model processing.

[0093] For another example, when slicing the target voice, a certain degree of overlap between the slices can also be retained, that is, there is partial overlap between adjacent sub-target voices in terms of time position. The specific degree of overlap can be adjusted according to the voice length, the number of slices, the repetition rate ratio, etc. In this way, the voice correlation between the sub-target objects obtained by slicing can be better retained, thereby improving the accuracy of model processing.

[0094] S1042. Input the multiple first object features into the feature extraction module to obtain the second object features output by the feature extraction module.

[0095] In some embodiments, the feature extraction module can include at least one Transformer layer, and the first object features can be feature-extracted through at least one Transformer layer to obtain the second object features.

[0096] In other embodiments, the feature extraction module can include a shared layer and a candidate adaptation layer, and the first object features can be feature-extracted according to the shared layer and the target adaptation layer to obtain the second object features. Among them, the target adaptation layer can be a candidate adaptation layer corresponding to the task type.

[0097] It should be noted that the shared layer may include at least one Transformer layer, and the candidate adaptation layer may also include at least one Transformer layer. The shared layer may be a Transformer layer shared by multiple different task types. For example, the target tasks of each task type need to use this shared layer; the candidate adaptation layer may be a Transformer layer corresponding to a specific task type. If the target task of this specific task type is executed, this Transformer layer can be used for feature extraction; conversely, if the target task of this specific task type is not executed, this Transformer layer may not be used.

[0098] In this way, through the candidate adaptation layer, the capacity of the model can be improved in large-scale tasks, and the processing performance of the model can be enhanced. Moreover, in the case of inputting sample objects of multiple different object types, information complementarity between different sample objects can be achieved through the self-attention mechanism of the Transformer.

[0099] S1043. Input the second object feature into the target processing module to obtain the target result output by the target processing module.

[0100] Among them, the target processing module may be a task processing module corresponding to the task type.

[0101] In this way, different types of target objects and different types of target tasks can be processed through this unified target model, improving the model deployment efficiency.

[0102] Figure 5 is a flowchart of a method for generating a target model shown according to an exemplary embodiment. As Figure 5 shown, the method may include:

[0103] S501. Obtain multiple first sample sets.

[0104] Among them, each first sample set includes multiple sample objects and the sample results corresponding to each sample object. Different first sample sets correspond to different task types.

[0105] In some embodiments, if the target model includes N task processing modules, then in this step, N first sample sets can be obtained, and each first sample set corresponds to a task processing module.

[0106] It should be noted that in this step, the obtained sample data can also be preprocessed, and the specific preprocessing method can be user-defined. For example, the preprocessing may include at least one of data augmentation processing methods such as random cropping, flipping, mirroring, and adding noise.

[0107] S502. Determine the second sample set according to the first sample set.

[0108] In some embodiments, the first sample set may be used as the second sample set.

[0109] In other embodiments, the first sample set may be sampled according to the task type to obtain the second sample set.

[0110] Exemplarily, the sampling weight may be determined according to the task type; each first sample set is sampled according to the sampling weight to obtain a third sample set corresponding to each first sample set; and the second sample set is determined according to the third sample set.

[0111] Among them, the above sampling weight may be any value between 0 and 1, and the sampling weight may also be any value between 0% and 100%. The sampling weight corresponding to each task type may be preset. For example, the sampling weight corresponding to the first task type may be e1, the sampling weight corresponding to the first task type may be e2,..., and the sampling weight corresponding to the Nth task type may be e N , e1, e2,..., e N The sum value of may be equal to 1 or 100%.

[0112] In some embodiments, multiple third sample sets may be obtained by sampling through the following formula (1):

[0113]

[0114] Among them, sample_data represents multiple third sample sets, represents the sampling sample size of one of the third sample sets, t1, t2,..., t n represents the sample sizes of the first sample sets corresponding to each task type. For example, t1 represents the sample size of the first sample set corresponding to the first task type, t2 represents the sample size of the first sample set corresponding to the second task type, t n represents the sample size of the first sample set corresponding to the Nth task type; e1, e2,..., e n represents the sampling weights of the first sample sets corresponding to each task type. For example, e1 represents the sampling weight of the first sample set corresponding to the first task type, e2 represents the sampling weight of the first sample set corresponding to the second task type, e n represents the sampling weight of the first sample set corresponding to the Nth task type, and M represents the overall sampling sample size preset for training. According to this formula, each first sample set corresponding to each task type can be sampled to obtain multiple third sample sets.

[0115] In this way, through this method, the first sample set can be randomly sampled according to the sampling weights to obtain a third sample set. Each third sample set is a subset or the entire set of the first sample set, and the task type corresponding to the third sample set is the same as that of the first sample set.

[0116] There can also be one or more of the third sample sets, and the number of the third sample sets is the same as the number of the first sample set.

[0117] In some embodiments, the third sample set can be used as the second sample set.

[0118] In this way, the second sample set for training can include various types of samples, thereby increasing the number of samples.

[0119] In other embodiments, the third sample set with the same task type as the task processing module of the target model can be used as the second sample set.

[0120] S503. Train the multimodal model according to the second sample set to obtain the target model.

[0121] It should be noted that the structure of the multimodal model can be the same as that of the target model. For example, the multimodal model can also include a feature extraction module, multiple object segmentation modules, and multiple task processing modules. For another example, the multimodal model can further include a configuration module.

[0122] In some embodiments, the multimodal model can further include a sampling module, and the sampling module can be used to execute the steps of S501 and S502 above.

[0123] In some embodiments, according to the second sample set, the model training steps can be repeatedly executed until it is determined that the trained multimodal model meets the preset stop iteration condition, and the trained multimodal model is used as the target model.

[0124] Among them, the model training steps can include:

[0125] S11. Obtain the task loss value of each task processing module of the multimodal model according to the second sample set.

[0126] Among them, the task loss value can be used to characterize the difference degree between the prediction result output by the task processing module and the sample result.

[0127] In some embodiments, the task loss value can be obtained in the following manner:

[0128] First, input the sample object into the sample segmentation module to obtain multiple first sample object features after segmentation.

[0129] Among them, the sample splitting module includes object splitting modules corresponding to the object types of sample objects.

[0130] Exemplarily, in the case where the object type of the sample object is an image or speech, multiple sub-sample objects can be split, and multiple first sample object features can be determined based on the multiple sub-sample objects. Among them, there is an overlapping area between two adjacent sub-sample objects in terms of position.

[0131] For example, when slicing a sample image, a certain degree of overlap between the slices can be retained, that is, there is partial overlap between adjacent sub-sample images obtained by slicing. The specific degree of overlap can be adjusted according to the image size, the number of slices, the proportion of the repetition rate, etc. In this way, the image correlation between the sub-sample objects obtained by slicing can be better retained, thereby improving the accuracy of model processing.

[0132] For another example, when slicing a sample speech, a certain degree of overlap between the slices can also be retained, that is, there is partial overlap between adjacent sub-sample speeches in terms of time position. The specific degree of overlap can be adjusted according to the speech length, the number of slices, the proportion of the repetition rate, etc. In this way, the speech correlation between the sub-sample objects obtained by slicing can be better retained, thereby improving the accuracy of model processing.

[0133] Secondly, input the multiple first sample object features into the feature extraction module to obtain the second sample object features output by the feature extraction module.

[0134] Thirdly, input the second sample object features into the sample processing module to obtain the prediction result output by the sample processing module.

[0135] Among them, the sample processing module can be the task processing module corresponding to the task type.

[0136] Finally, obtain the task loss value of each task processing module according to the prediction result and the sample result.

[0137] Exemplarily, the difference between the prediction result and the sample result can be calculated according to the task loss function as the task loss value corresponding to the task processing module.

[0138] The task loss functions corresponding to each task processing module can be the same or different, and the present disclosure does not make any limitation thereto.

[0139] S12. Calculate the comprehensive loss value according to the task weights and task loss values of multiple task processing modules.

[0140] Among them, the task weights of each task processing module can be the same or different.

[0141] Exemplarily, the comprehensive loss value can be calculated through the following formula (2):

[0142]

[0143] Where N represents the total number of tasks, Loss_i represents the task loss value corresponding to the i-th task processing module, and P_i represents the task weight corresponding to the i-th task processing module. In this way, the comprehensive loss value can be calculated through formula (2).

[0144] It should be noted that the above task weight can be any value between 0 and 1, or the above task weight can be any value between 0% and 100%. The sum of the task weights corresponding to multiple task processing modules can be 1 or 100%. This task weight can set the loss weights of different task modules in joint training according to business needs, or can also be initialized to 1 / N.

[0145] In this way, the task losses of different task modules can be normalized, which can prevent the loss weight of a certain task module from being too large and causing the training of other task modules to fail. Moreover, by weighting the task losses of different modalities and different tasks as the comprehensive loss of overall joint training, it can promote the gradient backpropagation and model parameter update during training.

[0146] It should be noted that the loss function corresponding to each task processing module can also be obtained by combining multi-task loss methods such as Uncertainty Weighting and GradNorm. The present disclosure does not limit this.

[0147] S13. In the case where it is determined according to the comprehensive loss value that the multi-modal model does not meet the preset stop iteration condition, update the parameters of the multi-modal model to obtain the trained modal model, and use the trained modal model as the new modal model.

[0148] The preset stop iteration condition may include that the comprehensive loss value is less than or equal to a preset loss threshold, or the change value of the comprehensive loss value within a certain number of iterations is less than a preset change threshold, or it may also be the stop iteration conditions commonly used in related technologies. The present disclosure does not limit this either. The above preset loss threshold or preset change threshold can be any preset value.

[0149] In addition, if it is determined according to the comprehensive loss value that the target neural network model meets the preset stop iteration condition, the model training step can be stopped, and the trained multi-modal model is used as the target model.

[0150] It should be noted that the training method for this multi-modal model can also refer to the training methods in related technologies. The present disclosure does not limit this.

[0151] By adopting the above method, different object types and task types are taken into account within a training framework, achieving joint training of multiple modalities.

[0152] Figure 6 FIG. 4 is a block diagram of an object processing device 1100 shown according to an exemplary embodiment, as Figure 6 shown, the device 1100 may include:

[0153] An object acquisition module 1101, configured to acquire a target object to be processed;

[0154] A first determination module 1102, configured to determine the object type of the target object;

[0155] A second determination module 1103, configured to determine the task type corresponding to the target task for processing the target object;

[0156] An object processing module 1104, configured to input the target object, the object type, and the task type into a pre-generated target model, and obtain a target result output by the target model;

[0157] Wherein, the target model includes a feature extraction module, a plurality of object segmentation modules, and a plurality of task processing modules. Different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types.

[0158] According to one or more embodiments of the present disclosure, the object processing module 1104 is configured to input the target object into a target segmentation module to obtain a plurality of first object features after segmentation; wherein, the target segmentation module includes the object segmentation module corresponding to the object type; input the plurality of first object features into the feature extraction module to obtain second object features output by the feature extraction module; input the second object features into a target processing module to obtain the target result output by the target processing module; wherein, the target processing module includes the task processing module corresponding to the task type.

[0159] According to one or more embodiments of the present disclosure, the feature extraction module includes a shared layer and candidate adaptation layers; each candidate adaptation layer corresponds to a different task type; the object processing module 1104 is configured to perform feature extraction on the first object features according to the shared layer and a target adaptation layer to obtain the second object features; the target adaptation layer is the candidate adaptation layer corresponding to the task type.

[0160] Figure 7 FIG. 5 is a block diagram of another object processing device 1100 shown according to an exemplary embodiment, as Figure 7 shown, the device 1100 may further include:

[0161] A model generation module 1105, configured to obtain a plurality of first sample sets; wherein each of the first sample sets includes a plurality of sample objects and a sample result corresponding to each sample object, and different first sample sets correspond to different task types; determine a second sample set according to the first sample sets; and train a multimodal model according to the second sample set to obtain the target model.

[0162] According to one or more embodiments of the present disclosure, the model generation module 1105 is configured to use the first sample set as the second sample set; or sample the first sample set according to the task type to obtain the second sample set.

[0163] According to one or more embodiments of the present disclosure, the model generation module 1105 is configured to determine sampling weights according to the task type; sample each of the first sample sets according to the sampling weights to obtain a third sample set corresponding to each first sample set; and determine the second sample set according to the third sample sets.

[0164] According to one or more embodiments of the present disclosure, the model generation module 1105 is configured to use the third sample set as the second sample set; or use the third sample set with the same task type as the task processing module of the target model as the second sample set.

[0165] According to one or more embodiments of the present disclosure, the model generation module 1105 is configured to, according to the second sample set, repeatedly execute the model training step until it is determined that the trained multimodal model meets a preset stop iteration condition, and use the trained multimodal model as the target model;

[0166] The model training step includes:

[0167] Obtain a task loss value of each task processing module of the multimodal model according to the second sample set; the task loss value is used to characterize the difference between the prediction result output by the task processing module and the sample result;

[0168] Calculate a comprehensive loss value according to the task weights of a plurality of the task processing modules and the task loss values;

[0169] In the case where it is determined according to the comprehensive loss value that the multimodal model does not meet the preset stop iteration condition, update the parameters of the multimodal model to obtain a trained modal model, and use the trained modal model as a new modal model.

[0170] According to one or more embodiments of the present disclosure, the model generation module 1105 is configured to input the sample object into the sample segmentation module to obtain a plurality of first sample object features after segmentation; wherein, the sample segmentation module includes the object segmentation module corresponding to the object type of the sample object; input the plurality of first sample object features into the feature extraction module to obtain second sample object features output by the feature extraction module; input the second sample object features into the sample processing module to obtain a prediction result output by the sample processing module; wherein, the sample processing module includes the task processing module corresponding to the task type; obtain the task loss value of each task processing module according to the prediction result and the sample result.

[0171] According to one or more embodiments of the present disclosure, the model generation module 1105 is configured to, when the object type of the sample object is an image or voice, segment to obtain a plurality of sub-sample objects; wherein, there is an overlapping area between adjacent sub-sample objects in terms of position; determine a plurality of first sample object features according to the plurality of sub-sample objects.

[0172] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0173] Next, refer to Figure 8 , which shows a schematic structural diagram of an electronic device 2000 (such as a terminal device or a server) suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The server in the embodiments of the present disclosure may include, but is not limited to, local servers, cloud servers, single servers, distributed servers, etc. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0174] As Figure 8As shown, the electronic device 2000 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 2001, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2002 or a program loaded from a storage device 2008 into a random access memory (RAM) 2003. In the RAM 2003, various programs and data required for the operation of the electronic device 2000 are also stored. The processing device 2001, the ROM 2002, and the RAM 2003 are connected to each other through a bus 2004. An input / output (I / O) interface 2005 is also connected to the bus 2004.

[0175] Generally, the following devices may be connected to the input / output interface 2005: an input device 2006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 2007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 2008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 2009. The communication device 2009 may allow the electronic device 2000 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 2000 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0176] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 2009, or installed from the storage device 2008, or installed from the ROM 2002. When the computer program is executed by the processing device 2001, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0177] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0178] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0179] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0180] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a target object to be processed; determine the object type of the target object; determine the task type corresponding to the target task for processing the target object; input the target object, the object type, and the task type into a pre-generated target model to obtain a target result output by the target model; wherein the target model includes a feature extraction module, a plurality of object segmentation modules, and a plurality of task processing modules, different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types.

[0181] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0182] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0183] The modules involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the object acquisition module can also be described as "the module for acquiring the target object to be processed".

[0184] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0185] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0186] According to one or more embodiments of the present disclosure, an object processing method is provided. The method includes:

[0187] Acquire a target object to be processed;

[0188] Determine the object type of the target object;

[0189] Determine the task type corresponding to the target task for processing the target object;

[0190] Input the target object, the object type, and the task type into a pre-generated target model to obtain a target result output by the target model;

[0191] Wherein, the target model includes a feature extraction module, a plurality of object segmentation modules, and a plurality of task processing modules. Different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types.

[0192] According to one or more embodiments of the present disclosure, inputting the target object, the object type, and the task type into a pre-generated target model to obtain the target result output by the target model includes:

[0193] Inputting the target object into a target segmentation module to obtain a plurality of first object features after segmentation; wherein, the target segmentation module includes the object segmentation module corresponding to the object type;

[0194] Inputting the plurality of first object features into the feature extraction module to obtain second object features output by the feature extraction module;

[0195] Inputting the second object features into a target processing module to obtain the target result output by the target processing module; wherein, the target processing module includes the task processing module corresponding to the task type.

[0196] According to one or more embodiments of the present disclosure, the feature extraction module includes a shared layer and candidate adaptation layers; each candidate adaptation layer corresponds to a different task type; the inputting the plurality of first object features into the feature extraction module to obtain the second object features output by the feature extraction module includes:

[0197] Performing feature extraction on the first object features according to the shared layer and a target adaptation layer to obtain the second object features; the target adaptation layer is the candidate adaptation layer corresponding to the task type.

[0198] According to one or more embodiments of the present disclosure, the target model is generated by the following method:

[0199] Obtaining a plurality of first sample sets; wherein, each first sample set includes a plurality of sample objects and the sample results corresponding to each sample object, and different first sample sets correspond to different task types;

[0200] Determining a second sample set according to the first sample set;

[0201] Training a multimodal model according to the second sample set to obtain the target model.

[0202] According to one or more embodiments of the present disclosure, the determining the second sample set according to the first sample set includes:

[0203] Taking the first sample set as the second sample set; or,

[0204] Sampling the first sample set according to the task type to obtain the second sample set.

[0205] According to one or more embodiments of the present disclosure, sampling the first sample set according to the task type to obtain the second sample set includes:

[0206] Determine a sampling weight according to the task type;

[0207] Sample each of the first sample sets according to the sampling weight to obtain a third sample set corresponding to each of the first sample sets;

[0208] Determine the second sample set according to the third sample set.

[0209] According to one or more embodiments of the present disclosure, determining the second sample set according to the third sample set includes:

[0210] Taking the third sample set as the second sample set; or,

[0211] Taking the third sample set with the same task type as the task processing module of the target model as the second sample set.

[0212] According to one or more embodiments of the present disclosure, training the multimodal model according to the second sample set to obtain the target model includes:

[0213] According to the second sample set, repeatedly execute the model training step until it is determined that the trained multimodal model meets the preset stop iteration condition, and take the trained multimodal model as the target model;

[0214] The model training step includes:

[0215] Obtain the task loss value of each task processing module of the multimodal model according to the second sample set; the task loss value is used to characterize the difference degree between the prediction result output by the task processing module and the sample result;

[0216] Calculate a comprehensive loss value according to the task weights of multiple task processing modules and the task loss values;

[0217] In the case where it is determined according to the comprehensive loss value that the multimodal model does not meet the preset stop iteration condition, update the parameters of the multimodal model to obtain a trained modal model, and take the trained modal model as a new modal model.

[0218] According to one or more embodiments of the present disclosure, obtaining the task loss value of each task processing module of the multimodal model according to the second sample set includes:

[0219] Input the sample object into the sample splitting module to obtain multiple first sample object features after splitting; wherein, the sample splitting module includes the object splitting module corresponding to the object type of the sample object.

[0220] Input multiple first sample object features into the feature extraction module to obtain second sample object features output by the feature extraction module.

[0221] Input the second sample object features into the sample processing module to obtain a prediction result output by the sample processing module; wherein, the sample processing module includes the task processing module corresponding to the task type.

[0222] Obtain the task loss value of each task processing module according to the prediction result and the sample result.

[0223] According to one or more embodiments of the present disclosure, the inputting the sample object into the sample splitting module to obtain multiple first sample object features after splitting includes:

[0224] In the case where the object type of the sample object is an image or voice, split to obtain multiple sub-sample objects; wherein, there is an overlapping area between adjacent sub-sample objects in terms of position.

[0225] Determine multiple first sample object features according to the multiple sub-sample objects.

[0226] According to one or more embodiments of the present disclosure, an object processing device is provided, and the device includes:

[0227] An object acquisition module for acquiring a target object to be processed.

[0228] A first determination module for determining the object type of the target object.

[0229] A second determination module for determining the task type corresponding to the target task for processing the target object.

[0230] An object processing module for inputting the target object, the object type, and the task type into a pre-generated target model to obtain a target result output by the target model.

[0231] Wherein, the target model includes a feature extraction module, multiple object splitting modules, and multiple task processing modules, different object splitting modules respectively correspond to different object types, and different task processing modules correspond to different task types.

[0232] According to one or more embodiments of the present disclosure, the object processing module is configured to input the target object into the target segmentation module to obtain a plurality of first object features after segmentation; wherein, the target segmentation module includes the object segmentation module corresponding to the object type; input the plurality of first object features into the feature extraction module to obtain second object features output by the feature extraction module; input the second object features into the target processing module to obtain the target result output by the target processing module; wherein, the target processing module includes the task processing module corresponding to the task type.

[0233] According to one or more embodiments of the present disclosure, the feature extraction module includes a shared layer and candidate adaptation layers; each candidate adaptation layer corresponds to a different task type; the object processing module is configured to perform feature extraction on the first object features according to the shared layer and the target adaptation layer to obtain the second object features; the target adaptation layer is the candidate adaptation layer corresponding to the task type.

[0234] According to one or more embodiments of the present disclosure, the apparatus further includes:

[0235] A model generation module, configured to obtain a plurality of first sample sets; wherein, each first sample set includes a plurality of sample objects and a sample result corresponding to each sample object, and different first sample sets correspond to different task types; determine a second sample set according to the first sample sets; train a multimodal model according to the second sample sets to obtain the target model.

[0236] According to one or more embodiments of the present disclosure, the model generation module is configured to use the first sample set as the second sample set; or sample the first sample set according to the task type to obtain the second sample set.

[0237] According to one or more embodiments of the present disclosure, the model generation module is configured to determine a sampling weight according to the task type; sample each first sample set according to the sampling weight to obtain a third sample set corresponding to each first sample set; determine the second sample set according to the third sample sets.

[0238] According to one or more embodiments of the present disclosure, the model generation module is configured to use the third sample set as the second sample set; or use the third sample set with the same task type as the task processing module of the target model as the second sample set.

[0239] According to one or more embodiments of the present disclosure, the model generation module is configured to cyclically execute model training steps according to the second sample set until it is determined that the trained multi-modal model meets a preset stop iteration condition, and use the trained multi-modal model as the target model;

[0240] The model training steps include:

[0241] Obtain the task loss value of each task processing module of the multi-modal model according to the second sample set; the task loss value is used to characterize the difference between the prediction result output by the task processing module and the sample result;

[0242] Calculate a comprehensive loss value according to the task weights of multiple task processing modules and the task loss values;

[0243] In the case where it is determined according to the comprehensive loss value that the multi-modal model does not meet the preset stop iteration condition, update the parameters of the multi-modal model to obtain a trained modal model, and use the trained modal model as a new modal model.

[0244] According to one or more embodiments of the present disclosure, the model generation module is configured to input the sample object into the sample splitting module to obtain multiple first sample object features after splitting; wherein, the sample splitting module includes the object splitting module corresponding to the object type of the sample object; input the multiple first sample object features into the feature extraction module to obtain second sample object features output by the feature extraction module; input the second sample object features into the sample processing module to obtain a prediction result output by the sample processing module; wherein, the sample processing module includes the task processing module corresponding to the task type; obtain the task loss value of each task processing module according to the prediction result and the sample result.

[0245] According to one or more embodiments of the present disclosure, the model generation module is configured to, when the object type of the sample object is an image or speech, split to obtain multiple sub-sample objects; there is an overlapping area between adjacent sub-sample objects in terms of position; determine multiple first sample object features according to the multiple sub-sample objects.

[0246] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0247] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0248] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. With regard to the apparatus in the above embodiments, the specific manner in which each module performs an operation has been described in detail in the embodiments related to the method, and will not be elaborated herein.

Claims

1. An object processing method, characterized in that, The method includes: Obtaining a target object to be processed, where the target object is a single-modal object or a multi-modal object; Determining the object type of the target object; Determining the task type corresponding to the target task for processing the target object; Inputting the target object, the object type, and the task type into a pre-generated target model to obtain a target result output by the target model, where the target result is an object recognition result and / or an object classification result; Wherein, the target model includes a feature extraction module, multiple object segmentation modules, and multiple task processing modules, different object segmentation modules respectively correspond to different object types, and different task processing modules correspond to different task types; The inputting the target object, the object type, and the task type into a pre-generated target model to obtain a target result output by the target model includes: Inputting the target object into a target segmentation module to obtain multiple first object features after segmentation; wherein, the target segmentation module includes the object segmentation module corresponding to the object type; Inputting the multiple first object features into the feature extraction module to obtain second object features output by the feature extraction module, and the feature extraction module includes a shared layer and a candidate adaptation layer; Determining the task processing module corresponding to the task type from the multiple task processing modules as the target processing module; and Inputting the second object features into the target processing module to obtain the target result output by the target processing module.

2. The method according to claim 1, characterized in that, Each candidate adaptation layer corresponds to a different task type; the inputting the multiple first object features into the feature extraction module to obtain second object features output by the feature extraction module includes: Performing feature extraction on the first object features according to the shared layer and the target adaptation layer to obtain the second object features; the target adaptation layer is the candidate adaptation layer corresponding to the task type.

3. The method according to any one of claims 1 or 2, characterized in that, The target model is generated in the following manner: Obtaining multiple first sample sets; wherein, each first sample set includes multiple sample objects and a sample result corresponding to each sample object, and different first sample sets correspond to different task types; Determining a second sample set according to the first sample set; Training a multi-modal model according to the second sample set to obtain the target model.

4. The method according to claim 3, wherein The determining a second sample set according to the first sample set includes: Using the first sample set as the second sample set; or, Sampling the first sample set according to the task type to obtain the second sample set.

5. The method according to claim 4, characterized in that The sampling the first sample set according to the task type to obtain the second sample set includes: Determining a sampling weight according to the task type; Sampling each first sample set according to the sampling weight to obtain a third sample set corresponding to each first sample set; Determining the second sample set according to the third sample set.

6. The method according to claim 5, characterized in that, The determining the second sample set according to the third sample set includes: Using the third sample set as the second sample set; or, Use the third sample set with the same task type as the task processing module of the target model as the second sample set.

7. The method according to claim 3, characterized in that, The training the multimodal model according to the second sample set to obtain the target model includes: According to the second sample set, repeatedly execute the model training step until it is determined that the trained multimodal model meets the preset stop iteration condition, and use the trained multimodal model as the target model; The model training step includes: Obtain the task loss value of each task processing module of the multimodal model according to the second sample set; the task loss value is used to characterize the difference between the prediction result output by the task processing module and the sample result; Calculate the comprehensive loss value according to the task weights of multiple task processing modules and the task loss value; In the case where it is determined according to the comprehensive loss value that the multimodal model does not meet the preset stop iteration condition, update the parameters of the multimodal model to obtain the trained modal model, and use the trained modal model as the new modal model.

8. The method according to claim 7, wherein The obtaining the task loss value of each task processing module of the multimodal model according to the second sample set includes: Input the sample object into the sample splitting module to obtain multiple first sample object features after splitting; wherein, the sample splitting module includes the object splitting module corresponding to the object type of the sample object; Input the multiple first sample object features into the feature extraction module to obtain the second sample object features output by the feature extraction module; Input the second sample object features into the sample processing module to obtain the prediction result output by the sample processing module; wherein, the sample processing module includes the task processing module corresponding to the task type; Obtain the task loss value of each task processing module according to the prediction result and the sample result.

9. The method according to claim 8, wherein The inputting the sample object into the sample splitting module to obtain multiple first sample object features after splitting includes: In the case where the object type of the sample object is an image or speech, split to obtain multiple sub-sample objects; wherein, there is an overlapping area between adjacent sub-sample objects in terms of position; Determine multiple first sample object features according to the multiple sub-sample objects.

10. An object processing device, characterized in that, The apparatus includes: An object acquisition module, configured to acquire a target object to be processed, where the target object is a unimodal object or a multimodal object; A first determination module, configured to determine the object type of the target object; A second determination module, configured to determine the task type corresponding to the target task for processing the target object; An object processing module, configured to input the target object, the object type, and the task type into a pre-generated target model to obtain a target result output by the target model, where the target result is an object recognition result and / or an object classification result; wherein, the target model includes a feature extraction module, multiple object splitting modules, and multiple task processing modules, different object splitting modules respectively correspond to different object types, and different task processing modules correspond to different task types; The object processing module is further configured to input the target object into the target segmentation module to obtain a plurality of first object features after segmentation; wherein, the target segmentation module includes the object segmentation module corresponding to the object type; input the plurality of first object features into the feature extraction module to obtain second object features output by the feature extraction module, and the feature extraction module includes a shared layer and a candidate adaptation layer; determine the task processing module corresponding to the task type from the plurality of task processing modules as the target processing module; and input the second object features into the target processing module to obtain the target result output by the target processing module.

11. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processing device, the steps of the method according to any one of claims 1 to 9 are implemented.

12. An electronic device, characterized in that, Comprising: A storage device having a computer program stored thereon; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-modal classification model generation method and device

    CN113762321A

Cited By

  • Object processing method, device, readable medium and electronic device

    EP4443397A1

  • Object processing method, device, readable medium and electronic device

    EP4443397B1