Training method of modal classification model and classification method based on multi-modal data

By collaboratively training the weight values of the modal classification model and multiple modalities, adjusting the weight values of each modal state, the dependence problem of the modal classification model on simple modal states during the training process is solved, and balanced learning and accurate prediction of multimodal data are realized.

CN120296482APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410042255.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing modal classification models tend to prioritize learning simple modal data during training, while ignoring difficult modal data, resulting in performance degradation and classification accuracy.

Method used

By collaboratively training the weight values of the modal classification model and multiple modalities, adjust the weight values of each modality to balance the learning of different modalities, and adjust the model parameters using the loss value to ensure that the model performs balanced learning of multiple modalities.

Benefits of technology

It improves the robustness of the modal classification model and the accuracy of event category prediction, enhances the processing ability and robustness of multimodal data, reduces dependence on simple modes, and improves training efficiency and prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296482A_ABST
    Figure CN120296482A_ABST
Patent Text Reader

Abstract

The invention provides a modal classification model training method and a classification method based on multi-modal data, and belongs to the technical field of artificial intelligence. Comprising the following steps: for each piece of sample scene data, obtaining a plurality of single-mode classification results of the sample scene data through a mode classification model; fusing the plurality of single-mode classification results based on the respective weight values of the plurality of modes to obtain a first multi-mode classification result; based on the plurality of single-mode classification results, the first multi-mode classification result and the annotation event category, adjusting respective weight values of the plurality of modes; fusing the plurality of single-mode classification results based on the respective adjusted weight values of the plurality of modes to obtain a second multi-mode classification result; and determining a loss value based on the respective labeled event categories of the plurality of sample scene data and the second multi-modal classification result, and adjusting model parameters of the modal classification model based on the loss value, that is, improving the robustness of the modal classification model through cooperative training of the modal classification model and the weight values of the plurality of modals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a training method for a modality classification model and a classification method based on multi-modal data. Background Art

[0002] With the development of autonomous driving technology, cars are no longer just mechanical devices. They have become more intelligent, such as being able to perceive the environment and make corresponding decisions and then execute them. Among them, cars collect data of multiple modalities such as images and lidar point clouds through sensor devices such as cameras and lidar to obtain detailed information about the vehicle's surrounding environment. Accordingly, the car determines the event category of the event generated at the current moment (such as encountering a red light) based on the data of multiple modalities, and then makes a decision (such as stopping). Therefore, how to determine the event category based on the data of multiple modalities is very important for autonomous driving.

[0003] In related technologies, generally, a modality classification model is trained through data of multiple modalities. Through the modality classification model, the event category is determined based on the data of multiple modalities collected. However, the training difficulties of the data of multiple modalities are different. The modality classification model often tends to preferentially learn the data of modalities that are easy to train, while ignoring the data of modalities that are difficult to train, resulting in a decline in the performance of the modality classification model, and thus reducing the classification accuracy. Summary of the Invention

[0004] Embodiments of this application provide a training method for a modality classification model and a classification method based on multi-modal data. This method improves the robustness of the modality classification model through the collaborative training of the modality classification model and the weight values of multiple modalities. The technical solutions are as follows:

[0005] On the one hand, a training method for a modality classification model is provided. The method includes:

[0006] Obtain multiple sample scenario data of a target scenario and the respective labeled event categories of the multiple sample scenario data. The sample scenario data includes scenario sub-data of multiple modalities collected under the target scenario;

[0007] For each sample scenario data, through the modality classification model, obtain multiple single-modal classification results of the sample scenario data. Each of the single-modal classification results corresponds to scenario sub-data of one modality, and each of the single-modal classification results is used to indicate the event category corresponding to the corresponding scenario sub-data;

[0008] Based on the weight values of the multiple modalities, fuse the multiple single-modal classification results of the sample scenario data to obtain a first multi-modal classification result of the sample scenario data. The first multi-modal classification result is used to indicate the event category corresponding to the scenario sub-data of the multiple modalities;

[0009] Adjust the weight values of the multiple modalities respectively based on the multiple unimodal classification results of the sample scenario data, the first multimodal classification result of the sample scenario data, and the labeled event category of the sample scenario data;

[0010] Fuse the multiple unimodal classification results of the sample scenario data based on the adjusted weight values of the multiple modalities respectively to obtain the second multimodal classification result of the sample scenario data;

[0011] Determine a loss value based on the labeled event categories of the multiple sample scenario data respectively and the second multimodal classification results of the multiple sample scenario data respectively, and adjust the model parameters of the modality classification model based on the loss value.

[0012] On the other hand, a classification method based on multimodal data is provided, and the method includes:

[0013] Obtain the scenario data of the target scenario, where the scenario data includes scenario sub-data of multiple modalities collected under the target scenario;

[0014] Input the scenario data into the modality classification model to obtain multiple unimodal classification results of the scenario data, each unimodal classification result corresponding to the scenario sub-data of one modality, the modality classification model being trained by the method described in any one of the above, and each unimodal classification result being used to indicate the event category corresponding to the corresponding scenario sub-data;

[0015] Fuse the multiple unimodal classification results of the scenario data based on the first weight values of the multiple modalities respectively to obtain the target classification result of the scenario data, the target classification result being used to indicate the event category corresponding to the scenario sub-data of the multiple modalities.

[0016] On the other hand, a training device for a modality classification model is provided, and the device includes:

[0017] An acquisition module, configured to acquire multiple sample scenario data of the target scenario and the labeled event categories of the multiple sample scenario data respectively, where the sample scenario data includes scenario sub-data of multiple modalities collected under the target scenario;

[0018] A classification module, configured to, for each sample scenario data, obtain multiple unimodal classification results of the sample scenario data through the modality classification model, each unimodal classification result corresponding to the scenario sub-data of one modality, and each unimodal classification result being used to indicate the event category corresponding to the corresponding scenario sub-data;

[0019] A fusion module, configured to fuse multiple unimodal classification results of the sample scenario data based on the respective weight values of the multiple modalities, to obtain a first multimodal classification result of the sample scenario data, where the first multimodal classification result is used to indicate an event category corresponding to the scenario sub-data of the multiple modalities;

[0020] An adjustment module, configured to adjust the respective weight values of the multiple modalities based on the multiple unimodal classification results of the sample scenario data, the first multimodal classification result of the sample scenario data, and the labeled event category of the sample scenario data;

[0021] The fusion module is further configured to fuse multiple unimodal classification results of the sample scenario data based on the adjusted respective weight values of the multiple modalities, to obtain a second multimodal classification result of the sample scenario data;

[0022] A determination module, configured to determine a loss value based on the labeled event categories of multiple sample scenario data and the second multimodal classification results of the multiple sample scenario data, and adjust model parameters of the modality classification model based on the loss value.

[0023] In some embodiments, the adjustment module is configured to:

[0024] If the first multimodal classification result of the sample scenario data matches the labeled event category of the sample scenario data, decrease the weight value of a first modality among the multiple modalities, and increase the weight value of a second modality among the multiple modalities, where the unimodal classification result corresponding to the first modality is the same as the labeled event category of the sample scenario data, and the unimodal classification result corresponding to the second modality is different from the labeled event category of the sample scenario data;

[0025] If the first multimodal classification result of the sample scenario data does not match the labeled event category of the sample scenario data, no longer adjust the weight value of the first modality, and increase the weight value of the second modality.

[0026] In some embodiments, the adjustment module is configured to:

[0027] Determine a product of the weight value of the first modality and a preset change rate, and use a difference between the weight value of the first modality and the product as the decreased weight value of the first modality;

[0028] Determine a product of the weight value of the second modality and the preset change rate, and use a sum value of the weight value of the second modality and the corresponding product as the increased weight value of the second modality.

[0029] In some embodiments, the determination module is configured to:

[0030] Determine the product of the weight value of the second modality and a preset multiple of the preset change rate, and use the sum value of the weight value of the second modality and the product as the increased weight value of the second modality.

[0031] In some embodiments, the determining module is configured to:

[0032] For each sample scenario data, determine the sub-loss value corresponding to the sample scenario data based on the labeled event category of the sample scenario data and the second multi-modal classification result of the sample scenario data;

[0033] Determine the sum value of the weight values of the multiple modalities adjusted corresponding to the sample scenario data;

[0034] Determine the product of the sub-loss value corresponding to the sample scenario data and the sum value;

[0035] Use the sum value of the products corresponding to the multiple sample scenario data respectively as the loss value.

[0036] In some embodiments, the fusion module is configured to:

[0037] Based on the adjusted weight values of the multiple modalities respectively, perform weighted summation on the multiple single-modal classification results of the sample scenario data to obtain a weighted sum value;

[0038] Based on the sum value of the adjusted weight values of the multiple modalities and the weighted sum value, determine the second multi-modal classification result, where the second multi-modal classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple weight values.

[0039] On the other hand, a classification device based on multi-modal data is provided, and the device includes:

[0040] An acquisition module, configured to acquire scenario data of a target scenario, where the scenario data includes scenario sub-data of multiple modalities collected in the target scenario;

[0041] A classification module, configured to input the scenario data into a modality classification model to obtain multiple single-modal classification results of the scenario data, each single-modal classification result corresponding to scenario sub-data of a modality, the modality classification model being trained by the method described in any one of the above, and each single-modal classification result being used to indicate the event category corresponding to the corresponding scenario sub-data;

[0042] A fusion module, configured to fuse multiple unimodal classification results of the scenario data based on respective first weight values of the multiple modalities, to obtain a target classification result of the scenario data, where the target classification result is used to indicate an event category corresponding to the scenario data, and the target classification result is used to indicate an event category corresponding to scenario sub-data of the multiple modalities.

[0043] In some embodiments, the fusion module is configured to:

[0044] For each modality, determine a second weight value of the modality based on the first weight value of the modality, where the second weight value is negatively correlated with the first weight value;

[0045] Based on the respective second weight values of the multiple modalities, perform a weighted sum on the multiple unimodal classification results of the scenario data to obtain a weighted sum value;

[0046] Based on the sum value of the respective first weight values of the multiple modalities and the weighted sum value, determine the target classification result, where the target classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple first weight values.

[0047] In some embodiments, the fusion module is configured to:

[0048] Determine the sum value of the respective first weight values corresponding to the multiple modalities;

[0049] For each modality, use the difference between the first weight value of the modality and the sum value as the second weight value of the modality.

[0050] In some embodiments, the acquisition module is further configured to:

[0051] For each modality, acquire a weight set of the modality, where the weight set includes adjusted weight values of the modality corresponding to multiple sample scenario data;

[0052] Determine the sum value of the multiple weight values in the weight set to obtain the first weight value of the modality.

[0053] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and the memory is used to store at least one program, and the at least one program is loaded and executed by the processor to implement the training method of the modality classification model or the classification method based on multi-modal data in the embodiments of the present application.

[0054] On the other hand, a computer-readable storage medium is provided, in which at least one program is stored, and the at least one program is loaded and executed by a processor to implement the training method of the modality classification model or the classification method based on multimodal data in the embodiments of the present application.

[0055] On the other hand, a computer program product is provided, the computer program product includes at least one program, the at least one program is stored in a computer-readable storage medium, a processor of a computer device reads the at least one program from the computer-readable storage medium, and the processor executes the at least one program, so that the computer device executes the training method of the modality classification model or the classification method based on multimodal data described in any of the above implementation manners.

[0056] The embodiments of the present application provide a training method for a modality classification model. This method adjusts the weight values of multiple modalities based on multiple unimodal classification results, multimodal classification results, and labeled event categories, and then adjusts the model parameters based on the classification results determined by the adjusted weight values. In this way, not only is the model training achieved, but also the weight values of each modality are adjusted during the model training process, realizing the balance of the weights of different modalities. Moreover, since the loss value is determined based on the second multimodal classification result, and the second multimodal classification result is determined based on the adjusted weight values, that is, the weight values of multiple modalities are indirectly introduced into the loss value. In this way, the model parameters are adjusted based on the loss value introducing the weight values, so that the weight values affect the adjustment of the model parameters, and then based on the weight values, the modality classification model can perform balanced learning on multiple modalities, rather than only focusing on the modalities that are easy to learn, thereby improving the processing ability of the modality classification model for any modality data and enhancing the robustness of the modality classification model. That is, this method improves the training efficiency through the co-training of the modality classification model and the weight values of multiple modalities. Based on the trained modality classification model, accurate unimodal classification results can be obtained. On this basis, combined with the trained weight values, accurate multimodal classification results can be obtained. That is, on the basis of improving the event category prediction efficiency through the modality classification model, the accuracy of event category prediction is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 It is a schematic diagram of an implementation environment provided by the embodiments of the present application;

[0059] Figure 2 It is a flowchart of a method for training a modality classification model provided by an embodiment of the present application;

[0060] Figure 3 It is a flowchart of another method for training a modality classification model provided by an embodiment of the present application;

[0061] Figure 4 It is a flowchart of a classification method based on multi-modal data provided by an embodiment of the present application;

[0062] Figure 5 It is a flowchart of another classification method based on multi-modal data provided by an embodiment of the present application;

[0063] Figure 6 It is a block diagram of a device for training a modality classification model provided by an embodiment of the present application;

[0064] Figure 7 It is a block diagram of a classification device based on multi-modal data provided by an embodiment of the present application;

[0065] Figure 8 It is a block diagram of a terminal provided by an embodiment of the present application;

[0066] Figure 9 It is a block diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0067] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0068] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited.

[0069] In the present application, the term "at least one" means one or more, and the meaning of "multiple" is two or more.

[0070] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the sample scenario data involved in the present application is obtained under full authorization.

[0071] The following introduces the professional terms involved in this application:

[0072] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0073] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0074] Autonomous driving technology refers to the vehicle's ability to drive itself without the operation of a driver. It usually includes technologies such as high-precision maps, environmental perception, computer vision, behavior decision-making, path planning, and motion control. Autonomous driving includes multiple development paths such as single-vehicle intelligence, vehicle-road cooperation, and networked cloud control. Autonomous driving technology has a wide range of application prospects. Currently, in addition to the fields of logistics, public transportation, taxis, and intelligent transportation, it will be further developed in the future.

[0075] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0076] A modality is a form of manifestation of things and a description of things from a specific angle. Multimodality usually contains two or more modality forms and refers to the description of things from multiple perspectives. When people perceive the world, multiple senses always receive external information simultaneously, such as seeing images, hearing sounds, smelling odors, and tactile perception. With the development of multimedia technology, the types and magnitudes of available media data have increased significantly. For example, sensors can not only generate images or videos but also contain matching depth, temperature information, etc.

[0077] Multimodal perception technology is a technology that realizes the comprehensive perception and understanding of multiple modality information in the real world by fusing data from multiple sensors. It can be applied to identify, classify, and analyze various perception objects, such as humans, objects, environments, etc. Multimodal perception technology combines different sensor technologies, such as cameras, microphones, radars, infrared sensors, accelerometers, etc., to obtain diverse information about the target. By fusing and analyzing this information, more comprehensive and accurate perception results can be obtained.

[0078] The following introduces the implementation environment involved in this application:

[0079] The training method of the modality classification model or the classification method based on multimodal data provided by the embodiments of this application can be executed by a computer device, which can be provided as a server or a terminal. The following introduces the schematic diagram of the implementation environment of the method provided by the embodiments of this application.

[0080] See Figure 1 , Figure 1A schematic diagram of an implementation environment provided by an embodiment of this application. The implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. In some embodiments, the server 102 is used to train a modality classification model, and the trained modality classification model is used to determine the event category of the event occurring in the target scenario based on the scenario data. A target application is installed on the terminal 101, and this target application is used to predict the event category. In some embodiments, the trained modality classification model is embedded on the terminal 101. The terminal 101 determines the event category of the event occurring in the target scenario based on the scenario data collected in the target scenario through the modality classification model. In other embodiments, the terminal 101 sends the scenario data collected in the target scenario to the server 102, and the server 102 determines the event category of the event occurring in the target scenario through the modality classification model on it.

[0081] In the embodiment of this application, the target scenario can be an autonomous driving scenario, an environmental detection scenario, a medical diagnosis scenario, etc. For example, if the target scenario is an autonomous driving scenario, the terminal 101 is a vehicle-mounted terminal. The vehicle-mounted terminal obtains data of multiple modalities collected by devices such as cameras, lidars, and speed sensors on the autonomous driving vehicle. Correspondingly, the data of multiple modalities are respectively data of the image modality collected by the camera, data of the point cloud modality collected by the lidar, data of the speed modality collected by the sensor, etc. Then, based on the data of multiple modalities, through the modality classification model, the event category of the event occurring in the current driving scenario is determined. For example, the event category can be an event of encountering a red light, an event of encountering a zebra crossing, etc. Correspondingly, the vehicle-mounted terminal can make decisions based on the determined event category, such as controlling the autonomous driving vehicle to stop, decelerate, etc.

[0082] For example, the target scenario is an environmental detection scenario. The terminal 101 collects data of multiple modalities through devices such as a soil detector, an air temperature and humidity sensor, and a camera; correspondingly, the data of multiple modalities are respectively data of the soil modality, data of the air modality, and data of the image modality. Then, based on the data of multiple modalities, through the modality classification model, the event category of environmental pollution in the current environment is determined. For example, the event category can be air pollution, water pollution, etc.

[0083] For example, the target scenario is a medical diagnosis scenario. The terminal 101 collects data of multiple modalities through devices such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), and PET (Positron Emission Tomography). Accordingly, the data of multiple modalities includes data of CT image modality, MRI image modality, and PET image modality. Alternatively, the terminal 101 obtains data of multiple modalities such as imaging data, biochemical data, and clinical data. Then, based on the data of multiple modalities, the body state category of the patient is determined through a modality classification model.

[0084] In some embodiments, the terminal 101 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a VR (Virtual Reality) device, an AR (Augmented Reality) device, etc., but is not limited thereto. In some embodiments, the server 102 is an independent server or can also be a server cluster or a distributed system composed of multiple servers, and can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 mainly undertakes computing work, and the terminal 101 undertakes secondary computing work; or, the server 102 undertakes secondary computing services, and the terminal 101 undertakes primary computing work; or, a distributed computing architecture is adopted between the server 102 and the terminal 101 for collaborative computing.

[0085] See Figure 2 , Figure 2 is a flowchart of a training method for a modality classification model provided by an embodiment of the present application. The method includes the following steps.

[0086] 201. The computer device obtains multiple sample scenario data of the target scenario and the labeled event category of each of the multiple sample scenario data. The sample scenario data includes multiple modality scenario sub-data collected in the target scenario.

[0087] In the embodiments of the present application, the target scenario can be set and changed as needed. For example, the target scenario can be an autonomous driving scenario, an environmental detection scenario, a medical detection scenario, a financial scenario, etc. If the target scenario is an autonomous driving scenario, the scene sub-data of multiple modalities includes at least two of the image data collected by the autonomous driving vehicle through a camera, the lidar point cloud data collected by the lidar, and the vehicle state data collected by the sensor, etc.

[0088] The data of multiple modalities refers to the data of multiple perception modalities. Common perception modalities include images, voices, texts, etc. Images are the most commonly used perception modality. In the embodiments of the present application, the multiple modalities can be different according to different target scenarios, and no specific limitation is made here. The scene sub-data of multiple modalities can be data from different sensors. For example, the scene sub-data of multiple modalities includes the image data collected by the camera, the audio data collected by the recording device, etc.

[0089] In the embodiments of the present application, the scene sub-data of multiple modalities included in each sample scene data corresponds to the state at the same moment. For example, the scene sub-data of multiple modalities is collected at the same moment, or within the same time range. Further, the scene sub-data of multiple modalities is collected within the same regional range. For example, in the autonomous driving scenario, the collection area of the scene sub-data of multiple modalities is the area near the autonomous driving vehicle. Or, the scene sub-data of multiple modalities is collected for the same target object. For example, in the medical diagnosis scenario, the scene sub-data of multiple modalities is collected from the same patient. Again, in the environmental detection scenario, the scene sub-data of multiple modalities is collected from the same area.

[0090] The labeled event category is the event category of the event that occurs in the target scenario determined based on the scene sub-data of multiple modalities included in the sample scene data. If the target scenario is an autonomous driving scenario, the labeled event category can be an event of encountering a red light, an event of encountering a car accident, an event of encountering rainy weather, etc. Further, based on the data of multiple modalities, the speed, position, and direction of the vehicle can also be determined, and then it can be determined whether it is speeding, illegally parked, or going in the wrong direction, etc. Correspondingly, the labeled event category can also be an event of speeding, an event of illegal parking, an event of going in the wrong direction, etc.

[0091] In the embodiments of the present application, the labeled event category is used to indicate making a decision in the target scenario. For example, in the autonomous driving vehicle scenario, if it is determined based on the scene data that the current event is encountering a red light, then a decision to stop is made.

[0092] 202. For each sample scene data, the computer device obtains multiple single-modal classification results of the sample scene data through the modal classification model. Each single-modal classification result corresponds to the scene sub-data of one modality, and each single-modal classification result is used to indicate the event category corresponding to the corresponding scene sub-data.

[0093] In an embodiment of the present application, for the scene sub-data of each modality in the sample scene data, the computer device inputs the scene sub-data of the modality into the modality classification model to obtain the single-modality classification result corresponding to the scene sub-data of the modality.

[0094] In an embodiment of the present application, the single-modality classification result is a probability distribution vector. The multiple element values in the probability distribution vector are used to indicate the respective prediction probabilities on multiple event categories. Among them, generally, the event category corresponding to the maximum prediction probability is used as the predicted event category, and at the same time, the event category corresponding to the maximum prediction probability is also the event category indicated by the single-modality classification result.

[0095] In an embodiment of the present application, multiple event categories can correspond to each target scene, and the multiple event categories in any target scene can be set and changed as needed, and no specific limitation is made here.

[0096] 203. The computer device fuses the single-modality classification results of the sample scene data based on the respective weight values of multiple modalities to obtain the first multi-modality classification result of the sample scene data. The first multi-modality classification result is used to indicate the event category corresponding to the scene sub-data of multiple modalities.

[0097] In an embodiment of the present application, the first multi-modality classification result is also the fusion result of multiple single-modality classification results and is used to indicate the final event category.

[0098] In an embodiment of the present application, the first multi-modality classification result is a probability distribution vector. The multiple element values in the probability distribution vector are used to indicate the respective prediction probabilities on multiple event categories. Among them, generally, the event category corresponding to the maximum prediction probability is used as the predicted event category, and at the same time, the event category corresponding to the maximum prediction probability is also the event category indicated by the first multi-modality classification result.

[0099] In an embodiment of the present application, the multiple single-modality classification results can be weighted and summed or weighted averaged based on multiple weight values to obtain the first multi-modality classification result.

[0100] 204. The computer device adjusts the respective weight values of multiple modalities based on the multiple single-modality classification results of the sample scene data, the first multi-modality classification result of the sample scene data, and the labeled event category of the sample scene data.

[0101] In an embodiment of the present application, if the single-modality classification result of any modality matches the labeled event category, then based on the first multi-modality classification result, the weight value of this modality is reduced or not adjusted. If the single-modality classification result of any modality does not match the labeled event category, then based on the first multi-modality classification result, the weight value of this modality is increased.

[0102] 205. The computer device fuses the multiple single-modal classification results of the sample scenario data based on the weight values adjusted for each of the multiple modalities, to obtain the second multi-modal classification result of the sample scenario data.

[0103] In the embodiment of the present application, the second multi-modal classification result is used to indicate the event category corresponding to the scenario sub-data of the multiple modalities. The second multi-modal classification result is a probability distribution vector.

[0104] In the embodiment of the present application, the multiple single-modal classification results can be weighted and summed or weighted averaged, etc. based on the multiple weight values, to obtain the second multi-modal classification result.

[0105] 206. The computer device determines a loss value based on the labeled event categories of the multiple sample scenario data and the second multi-modal classification results of the multiple sample scenario data, and adjusts the model parameters of the modal classification model based on the loss value.

[0106] In the embodiment of the present application, the computer device iteratively executes steps 202-206 based on the multiple sample scenario data, to train the modal classification model. The multiple sample scenario data used in each iteration process are the same, that is, the multiple sample scenario data and the labeled event categories obtained in step 201 are used in each iteration process.

[0107] In the embodiment of the present application, each sample scenario data is used in multiple iteration processes respectively. Each sample scenario data corresponds to weight values of multiple modalities respectively. The weight values of the multiple modalities used by each sample scenario data in the current iteration process are the adjusted weight values obtained by adjusting the weight values based on this sample scenario data in the previous iteration process. That is, after the training is completed, not only the trained modal classification model is obtained, but also the weight values of the multiple modalities corresponding to the multiple sample scenario data are obtained respectively. Among them, one sample scenario data corresponds to a weight set, and the weight set includes the weight values obtained by adjusting the multiple modalities multiple times.

[0108] In the embodiment of the present application, the loss value is used to indicate the gap between the labeled event categories of the multiple sample scenario data and the second multi-modal classification results. The computer device iteratively trains the modal classification model based on the multiple sample scenario data and the labeled event categories of the multiple sample scenario data until the preset requirements are met, to obtain the target modal classification model.

[0109] Among them, meeting the preset requirements can be that the gap between the labeled event categories of the multiple sample scenario data and the second multi-modal classification results is less than the preset gap. Further, meeting the preset requirements means that the loss value converges, or the loss value reaches the preset threshold, or the number of iterations reaches the preset number, which is not specifically limited herein.

[0110] An embodiment of the present application provides a training method for a modal classification model. This method adjusts the weight values of multiple modalities based on multiple single-modal classification results, multi-modal classification results, and labeled event categories, and then adjusts the model parameters based on the classification results determined by the adjusted weight values. In this way, not only is the model training achieved, but also the weight values of each modality are adjusted during the model training process, realizing the balance of the weights of different modalities. Moreover, since the loss value is determined based on the second multi-modal classification result, and the second multi-modal classification result is determined based on the adjusted weight values, that is, the weight values of multiple modalities are indirectly introduced into the loss value. Adjusting the model parameters based on the loss value introduced with the weight values enables the weight values to affect the adjustment of the model parameters, and then enables the modal classification model to perform balanced learning on multiple modalities based on the weight values, rather than only focusing on the modalities that are easy to learn, thereby improving the processing ability of the modal classification model for any modal data and enhancing the robustness of the modal classification model. That is, this method improves the training efficiency through the collaborative training of the modal classification model and the weight values of multiple modalities. Based on the trained modal classification model, accurate single-modal classification results can be obtained. On this basis, combined with the trained weight values, accurate multi-modal classification results can be obtained, that is, on the basis of improving the event category prediction efficiency through the modal classification model, the accuracy of event category prediction is improved.

[0111] The above Figure 2 is the basic process of the training method for the modal classification model. Next, based on Figure 3 the training method for the modal classification model will be further introduced. See Figure 3 , Figure 3 which is a flowchart of a training method for a modal classification model provided by an embodiment of the present application. This method is executed by a computer device and includes the following steps.

[0112] 301. The computer device obtains multiple sample scenario data of the target scenario and the labeled event category of each of the multiple sample scenario data. The sample scenario data includes scenario sub-data of multiple modalities collected under the target scenario.

[0113] In an embodiment of the present application, the training data set composed of multiple sample scenario data and the labeled event category of each of the multiple sample scenario data can be expressed as where M represents the number of modalities, N represents the number of sample scenario data, represents the M-th modality of the i-th sample scenario data. y i ∈{1,…,K} represents the corresponding labeled event category, and K is an integer greater than 1, representing the event category.

[0114] In the j-th iteration process, the weight set of multiple modalities of any sample scenario data is expressed as Among them, represents the weight value of the i-th sample scenario data in the M-th modality during the j-th iteration, and the total number of iterations is T times. Before training, the initial value of is 1, that is

[0115] 302. For each sample scenario data, the computer device obtains multiple single-modal classification results of the sample scenario data through the modality classification model. Each single-modal classification result corresponds to the scenario sub-data of one modality, and each single-modal classification result is used to indicate the event category corresponding to the scenario sub-data it corresponds to.

[0116] In the embodiment of the present application, the goal of the modality classification model is to learn a parameterized function: f(x 1 ,…,x M ,θ)→z. Among them, θ represents the model parameters of the modality classification model, and x 1 represents the scenario sub-data of the sample scenario data x in the first modality, and x M represents the scenario sub-data of the sample scenario data x in the M-th modality. z represents the output of the modality classification model, and z is a vector containing K values, called the logits vector. Then, the logits vector is transformed through the softmax layer of the modality classification model to obtain the probability distribution vector Among them, the probability distribution of the sample scenario data x in the M-th modality is defined as Among them, P represents probability, and y represents the event category. Correspondingly, the event category predicted by the model is the event category with the highest corresponding probability, denoted as

[0117] In the embodiment of the present application, the modality classification model can process the scenario sub-data of any modality. After inputting the scenario sub-data of any modality into the modality classification model, the single-modal classification result corresponding to the scenario sub-data of this modality is output.

[0118] 303. The computer device fuses the multiple single-modal classification results of the sample scenario data based on the weight values of multiple modalities respectively to obtain the first multi-modal classification result of the sample scenario data. The first multi-modal classification result is used to indicate the event category corresponding to the scenario sub-data of multiple modalities.

[0119] In some embodiments, the process by which the above computer device fuses multiple single-modal classification results of sample scenario data based on the respective weight values of multiple modalities to obtain a first multi-modal classification result of the sample scenario data includes the following steps: The computer device performs a weighted sum of multiple single-modal classification results of the sample scenario data based on the respective weight values of multiple modalities to obtain a weighted sum value; Based on the sum value of the respective weight values of multiple modalities and this weighted sum value, a first multi-modal classification result is determined, where the first multi-modal classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of multiple weight values.

[0120] Among them, the process by which the above computer device determines the first multi-modal classification result based on the sum value of the respective weight values of multiple modalities and this weighted sum value includes the following steps: The computer device takes the quotient of the weighted sum value and the sum value of multiple weight values as the first multi-modal classification result.

[0121] Optionally, the computer device obtains the first multi-modal classification result through the following formula (1).

[0122]

[0123] Among them, z ij represents the first multi-modal classification result of the i-th sample scenario data in the j-th iteration process. represents the single-modal classification result of the i-th sample scenario data in the M-th modality in the j-th iteration process. represents the weight value of the i-th sample scenario data in the M-th modality in the j-th iteration process.

[0124] In the embodiments of the present application, the weight set of any weight value in each sample scenario data is initialized to the same value, for example, the initialization value of any weight value is 1.

[0125] 304. If the first multi-modal classification result of the sample scenario data matches the labeled event category of the sample scenario data, the computer device reduces the weight value of the first modality among multiple modalities and increases the weight value of the second modality among multiple modalities. The single-modal classification result corresponding to the first modality is the same as the labeled event category of the sample scenario data, and the single-modal classification result corresponding to the second modality is different from the labeled event category of the sample scenario data.

[0126] In some embodiments, the above computer device reduces the weight value of the first modality among multiple modalities and increases the weight value of the second modality among multiple modalities. The process includes the following steps: The computer device determines the product of the weight value of the first modality and a preset change rate, and takes the difference between the weight value of the first modality and the product as the reduced weight value of the first modality; The computer device determines the product of the weight value of the second modality and the preset change rate, and takes the sum of the weight value of the second modality and the corresponding product as the increased weight value of the second modality.

[0127] In the embodiments of the present application, the preset change rate is a preset hyperparameter, and the preset change rate can be set and changed as needed. In this embodiment, an example where the preset change rate is less than 0.5 is used for illustration.

[0128] Among them, if the first multi-modal classification result of the sample scenario data matches the labeled event category of the sample scenario data, that is Then the adjustment formula for the weight value can be seen in the following formula (2).

[0129]

[0130] Among them, represents the event category indicated by the first multi-modal classification result of the i-th sample scenario data in the j-th iteration process. y i represents the labeled event category of the i-th sample scenario data in the j-th iteration process. represents the event category indicated by the single-modal classification result of the i-th sample scenario data in the M-th modality in the j-th iteration process. represents the weight value of the i-th sample scenario data in the M-th modality in the j-th iteration process. represents the weight value of the i-th sample scenario data in the M-th modality in the j-th iteration process, that is, the adjusted weight value. α represents the preset change rate.

[0131] 305. If the first multi-modal classification result of the sample scenario data does not match the labeled event category of the sample scenario data, the computer device no longer adjusts the weight value of the first modality and increases the weight value of the second modality.

[0132] In some embodiments, the process of the above computer device increasing the weight value of the second modality includes the following steps: The computer device determines the product of the weight value of the second modality and a preset multiple of the preset change rate, and takes the sum of the weight value of the second modality and the product as the increased weight value of the second modality.

[0133] Among them, the preset multiple can be set and changed as needed. In this embodiment, an example where the preset multiple is 2 is used for illustration.

[0134] Among them, if the first multi-modal classification result of the sample scenario data does not match the labeled event category of the sample scenario data, that is Then the adjustment formula for the weight value is shown in the following formula (3).

[0135]

[0136] In the embodiment of the present application, through the above steps 304-305, the process of adjusting the weight values of multiple modalities respectively is realized based on the multiple single-modal classification results of the sample scenario data, the first multi-modal classification result of the sample scenario data, and the labeled event category of the sample scenario data.

[0137] In this embodiment, when the first multi-modal classification result matches the labeled event category, that is, the overall event category prediction is correct, the weight value of the first modality is decreased at a single change rate, and the weight value of the second modality is increased at a single change rate. When the overall event category prediction is incorrect, the weight value of the first modality is not adjusted, and the weight value of the second modality is increased at a double change rate. Since when the overall event category prediction is correct and a single modality prediction is correct, it indicates that the model has learned the data of this modality well, and thus the weight value of this modality can be decreased without having too much impact on the performance of this modality. When the overall event category prediction is correct and a single modality prediction is incorrect, it indicates that the model has difficulty learning this modality and has not learned it well enough. Therefore, the weight value of this modality is increased so that the model can focus more on learning this modality. When the overall event category prediction is incorrect and a single modality prediction is correct, it indicates that the model has learned this modality, but the degree of learning is not enough to offset the interference of the modality with incorrect prediction, that is, it has not been fully learned. Therefore, at this time, the weight value of the modality with correct prediction remains unchanged to avoid reducing the prediction effect of this modality after adjusting its weight value. When the overall event category prediction is incorrect and a single modality prediction is incorrect, then the incorrect prediction of the overall event category has a large weight due to this modality with incorrect prediction, that is, the degree of incorrect prediction of this modality is large, and the model has far from learned this modality. Therefore, the weight value of this modality is increased to a large extent so that the model can focus more on learning this modality. That is, this embodiment improves the flexibility and accuracy of weight value adjustment.

[0138] It should be noted that the above steps 304-305 are only an optional implementation manner to implement this process, and the computer device can also implement this process through other optional implementation manners. For example, if the first multi-modal classification result of the sample scenario data matches the labeled event category of the sample scenario data, increase the weight value of the first modality and decrease the weight value of the second modality among the multiple modalities. If the first multi-modal classification result of the sample scenario data does not match the labeled event category of the sample scenario data, increase the weight value of the first modality and do not adjust the weight value of the second modality.

[0139] It should be noted that the scene sub-data of any modality can be easily and correctly classified by the modality classification model, indicating that this modality is easy to be learned by the modality classification model. If the scene sub-data of any modality is always misclassified by the modality classification model, it indicates that this modality is not easy to be learned by the modality classification model. Correspondingly, the modality classification model will pay more attention to learning the easy modalities and ignore the difficult modalities. In this embodiment, for the misclassified modalities, their weight values are increased, and for the correctly classified modalities, their weight values are decreased, that is, weight amplification for difficult samples and weight attenuation for easy samples are realized, thereby balancing the weights of multiple modalities, enabling the modality classification model to focus more attention on learning difficult modality samples, reducing the over-reliance of the model on easy modalities, and improving the robustness of the modality classification model.

[0140] It should be noted that during the training process of the modality classification model, it is very important to distinguish easy samples and difficult samples and quantify the difficulty level of the samples. An intuitive approach is to use the frequency of correct sample classification as a quantization index during the model training process, that is, samples that are easily classified correctly are easy samples, and samples that are always misclassified are difficult samples. In multi-modal learning, the model decision is divided into two parts. First, the single-modal results are obtained, and then the results of multi-modal fusion are obtained. Therefore, in this embodiment of the present application, the relationship between the multi-modal fusion results and the single-modal results is considered when considering the difficulty level of the samples, that is, the dynamic weighting algorithm in this embodiment is performed.

[0141] In this embodiment, during the model training process, the dynamically adjusted modality weights will affect the decision fusion of the sample scene data in multiple modalities and the training of the sample scene data. In order to fully exploit the knowledge of difficult samples and prevent overfitting of the knowledge of easy samples, this embodiment balances the weights of different modalities by amplifying the weights of difficult samples and attenuating the weights of easy samples. This method enables the modality classification model not to be biased towards learning a certain modality, reduces the dependence of the model on easy modalities, and improves the robustness of the model.

[0142] It should be noted that the steps of steps 304-305 above are only for convenience of description and do not limit the execution order of the two. In any iteration process of the computer device, for any sample scene data, one of the steps 304-305 is selected to execute.

[0143] In some embodiments, during the model training process, the computer device monitors the value of the weight value. If the weight value exceeds the preset threshold, a request is sent to relevant personnel to request manual intervention to check whether the sample scene data is damaged, thereby avoiding overfitting of a certain modality caused by too large or too small weight values.

[0144] 306. The computer device fuses the multiple single-modal classification results of the sample scenario data based on the weight values adjusted for each of the multiple modalities, to obtain a second multi-modal classification result of the sample scenario data.

[0145] In the embodiment of the present application, the process by which the computer device fuses the multiple single-modal classification results of the sample scenario data based on the weight values adjusted for each of the multiple modalities to obtain a second multi-modal classification result of the sample scenario data includes the following steps: The computer device performs weighted summation on the multiple single-modal classification results of the sample scenario data based on the weight values adjusted for each of the multiple modalities, to obtain a weighted sum value; Based on the sum value of the weight values adjusted for each of the multiple modalities and the weighted sum value, a second multi-modal classification result is determined, and the second multi-modal classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple weight values.

[0146] Among them, the process by which the computer device determines the second multi-modal classification result based on the sum value of the weight values adjusted for each of the multiple modalities and the weighted sum value includes the following steps: The computer device uses the quotient of the weighted sum value and the sum value of the multiple weight values as the second multi-modal classification result. This process is shown in formula (1) and will not be elaborated here.

[0147] In this embodiment, the single-modal classification results of multiple modalities are subjected to weighted summation to obtain an overall event category. In this way, the classification results of the data of multiple modalities are fused, making the classification results more accurate. And it is normalized by the sum value of the weight values to obtain a second multi-modal classification result, which is convenient for obtaining the event category based on the second multi-modal classification result.

[0148] 307. For each sample scenario data, the computer device determines a sub-loss value corresponding to the sample scenario data based on the labeled event category of the sample scenario data and the second multi-modal classification result of the sample scenario data.

[0149] In the embodiment of the present application, the computer device obtains a sub-loss value corresponding to the sample scenario data through any one of the cross-entropy loss function, mean squared error loss function, KL divergence loss function, negative log-likelihood loss, etc., based on the labeled event category of the sample scenario data and the second multi-modal classification result.

[0150] 308. The computer device determines the sum value of the weight values adjusted for the multiple modalities corresponding to the sample scenario data, determines the product of the sub-loss value corresponding to the sample scenario data and the sum value, and uses the sum value of the products respectively corresponding to the multiple sample scenario data as the loss value.

[0151] Optionally, the computer device obtains the loss value through the following formula (4).

[0152]

[0153] Among them, L j represents the loss value in the j-th iteration process. L(z ij , y1) represents the second multi-modal classification result z based on the i-th sample scenario data ij and the determined sub-loss value with the labeled event category y i . represents the sum value of the weight values of multiple modalities.

[0154] In the embodiment of the present application, through the above steps 307-308, the process of determining the loss value is realized based on the labeled event categories of each of the multiple sample scenario data and the second multi-modal classification results of each of the multiple sample scenario data. In this embodiment, the loss value for adjusting the model parameters is determined based on the losses respectively corresponding to the multiple sample scenario data. Since multiple sample scenario data are considered, the accuracy of the loss value is improved. Moreover, the second multi-modal pre-training result is obtained based on the adjusted weight values, that is, the weight value factor is introduced into the sub-loss value, and the sum value of the adjusted weight values of multiple modalities is also introduced into the loss value. In this way, the model parameters are adjusted based on the loss value introduced with the weight value, so that when the multi-modal classification model is learning, the factor of the weight value is considered. And because the more difficult the modality to learn, the greater its weight value, that is, the greater the impact on the loss value. Then, based on the loss value, the model parameters are adjusted, so that the multi-modal classification model can pay more attention to the data of the difficult modality to learn, thereby improving the ability of the multi-modal classification model to process the data of the difficult modality, and thus improving the robustness of the model.

[0155] 309. The computer device adjusts the model parameters of the multi-modal classification model based on the loss value.

[0156] In the implementation of the present application, the computer device can adopt the gradient descent method or the gradient backpropagation method to adjust the model parameters based on the loss value, which is not specifically limited herein.

[0157] It should be noted that, in the embodiment of the present application, each iteration process includes the steps of adjusting the weight value and adjusting the model parameters, that is, the computer device iteratively executes steps 302-309.

[0158] The embodiment of the present application provides a weighted training method based on the decision result. This method dynamically learns the weight of each modality corresponding to each sample scenario data to better balance the information mining difficulty between different modalities. And, through the weight value for training early warning and adding manual intervention, the stability and accuracy of the weight value adjustment are ensured.

[0159] The method provided by the embodiments of the present application can be applied to multiple fields, such as autonomous driving, medicine, environmental monitoring, finance and other fields. The embodiments of the present application focus on multimodal data fusion and model optimization, improve the performance, reliability and applicability of multimodal data processing systems, and improve the accuracy of decision-making and predictive tasks, which are of great significance for the safety of autonomous driving, the accuracy of medical diagnosis, the reliability of environmental monitoring, the accuracy of financial decision-making, etc.

[0160] In the embodiments of the present application, applying this method to the autonomous driving scenario, the modality classification model, that is, the autonomous driving classification model, can improve the credibility of the autonomous driving classification model through weighted multimodal sample data. By means of weighted multimodal sample data, the learning difficulty of different modalities and different samples in autonomous driving can be understood, and users can have an overall and individual understanding of the quality of multimodal data in autonomous driving. In particular, during the training process of the autonomous driving classification model, the dynamics of weighted multimodal sample data can enable users to more conveniently understand the model learning situation. For example, which noise data may exist in the training dataset, which data has been learned sufficiently, and which data needs to be re-learned, etc.

[0161] In some fields of autonomous driving adaptive learning, online learning, etc., data supervision can be carried out by following the weight changes of multimodal sample data to give early warnings for abnormal weight values. For example, if the weight of some sample data is too large due to continuous misclassification, it is possible that the sample data is noise data. At this time, the computer device can seek expert help to correct the sample data and achieve effective human-computer interaction.

[0162] It should be noted that in addition to supervising individual sample data, the embodiments of the present application can also supervise different modalities and different event categories to effectively identify faults in different sensor acquisition devices and data label noise. That is, the embodiments of the present application can not only improve the accuracy, reliability and safety of the autonomous driving classification model, but also provide users with higher efficiency and decision-making support for human-computer interaction, thereby improving the user experience.

[0163] An embodiment of the present application provides a method for training a modality classification model. This method adjusts the weight values of multiple modalities based on multiple unimodal classification results, multimodal classification results, and labeled event categories, and then adjusts the model parameters based on the classification results determined by the adjusted weight values. In this way, not only is model training achieved, but also the weight values of each modality are adjusted during the model training process, realizing the balance of the weights of different modalities. Moreover, since the loss value is determined based on the second multimodal classification result, and the second multimodal classification result is determined based on the adjusted weight values, that is, the weight values of multiple modalities are indirectly introduced into the loss value. In this way, the model parameters are adjusted based on the loss value introduced with the weight values, enabling the weight values to affect the adjustment of the model parameters, and further enabling the modality classification model to perform balanced learning on multiple modalities based on the weight values, rather than only focusing on the modalities that are easy to learn, thereby improving the processing ability of the modality classification model for any modality data and enhancing the robustness of the modality classification model. That is, this method co-trains the modality classification model and the weight values of multiple modalities, improving the training efficiency while being able to obtain accurate unimodal classification results based on the trained modality classification model. On this basis, combined with the trained weight values, accurate multimodal classification results can be obtained, that is, on the basis of improving the event category prediction efficiency through the modality classification model, the accuracy of event category prediction is improved.

[0164] Through the above Figure 2 and Figure 3 embodiments, the modality classification model and the weight values of multiple modalities are trained respectively. Based on the following Figure 2 or Figure 3 trained modality classification model and the weight values of multiple modalities, the scenario data is classified to obtain the event category. Refer to Figure 4 , Figure 4 which is a classification method based on multimodal data provided by an embodiment of the present application. This method includes the following steps.

[0165] 401. The computer device acquires the scenario data of the target scenario, and the scenario data includes scenario sub-data of multiple modalities collected under the target scenario.

[0166] In the embodiment of the present application, the target scenario and the scenario data are the same as the explanations in step 201, and will not be elaborated here.

[0167] 402. The computer device inputs the scenario data into the modality classification model to obtain multiple unimodal classification results of the scenario data. Each unimodal classification result corresponds to the scenario sub-data of one modality, and each unimodal classification result is used to indicate the event category corresponding to the corresponding scenario sub-data.

[0168] In the embodiment of the present application, for each modality of the scene sub-data in the scene data, the computer device inputs it into the modality classification model respectively to obtain the single-modality classification result corresponding to the scene sub-data of this modality.

[0169] In the embodiment of the present application, the single-modality classification result is a probability distribution vector.

[0170] 403. The computer device fuses the multiple single-modality classification results of the scene data based on the first weight values of the multiple modalities respectively to obtain the target classification result of the scene data. The target classification result is used to indicate the event category corresponding to the scene sub-data of the multiple modalities.

[0171] In the embodiment of the present application, the target classification result is a probability distribution vector. The multiple element values in the probability distribution vector are used to indicate the respective prediction probabilities on the multiple event categories. The event category corresponding to the maximum prediction probability among the multiple prediction probabilities is also the event category indicated by the target classification result. In the embodiment of the present application, the first weight values of the multiple modalities are based on Figure 2 and Figure 3 the adjusted weight values obtained in the embodiment. The step of fusing the multiple single-modality classification results in step 403 is the same as that in step 203 and will not be elaborated here.

[0172] The embodiment of the present application provides a classification method based on multi-modal data. This method obtains the event category based on the modality classification model, and this modality classification model is co-trained with the weight values of the multiple modalities. Based on the weight values, the modality classification model can perform balanced learning on the multiple modalities, rather than only focusing on the modalities that are easy to learn, thereby improving the processing ability of the modality classification model for any modality data and improving the robustness of the modality classification model. That is, when obtaining the event category through the modality classification model and the weight values of the multiple modalities, accurate single-modality classification results can be obtained based on the modality classification model. On this basis, combined with the trained weight values, accurate multi-modal classification results can be obtained. That is, on the basis of improving the prediction efficiency of the event category through the modality classification model, the accuracy of the event category prediction is improved.

[0173] Figure 4 The embodiment of... is the basic process of the classification method based on multi-modal data. Next, based on Figure 5 the embodiment of... will further introduce the classification method based on multi-modal data. Refer to Figure 5 , Figure 5 which is the flowchart of a classification method based on multi-modal data provided by the embodiment of the present application. This method includes the following steps.

[0174] 501. The computer device acquires the scene data of the target scene. The scene data includes the scene sub-data of multiple modalities collected under the target scene.

[0175] In the embodiments of the present application, step 501 is the same as step 201, and will not be elaborated here.

[0176] 502. The computer device inputs the scenario data into the modality classification model to obtain multiple unimodal classification results of the scenario data. Each unimodal classification result corresponds to scenario sub-data of one modality, and each unimodal classification result is used to indicate the event category corresponding to the scenario sub-data it corresponds to.

[0177] In the embodiments of the present application, the computer device inputs the scenario sub-data of each modality in the scenario data into the modality classification model respectively, and the modality classification model outputs the unimodal classification results corresponding to the scenario sub-data of each modality respectively.

[0178] 503. For each modality, the computer device determines a second weight value of the modality based on the first weight value of the modality, and the second weight value is negatively correlated with the first weight value.

[0179] In some embodiments, the process of obtaining the first weight values corresponding to multiple modalities includes the following steps: For each modality, the computer device obtains a weight set of the modality, and the weight set includes the weight values adjusted for the modality corresponding to multiple sample scenario data; determines the sum value of the multiple weight values in the weight set to obtain the first weight value of the modality.

[0180] It should be noted that after the modality classification model is trained, that is, after T iterative trainings, for each sample scenario data, after T adjustments of the weight values, the weight values corresponding to the multiple modalities of the sample scenario data are obtained, and each modality corresponds to one weight value. Correspondingly, for each modality, the weight values corresponding to the multiple sample scenario data of the modality are obtained, and each sample scenario data corresponds to one weight value. This weight value is also the weight value obtained after the last adjustment.

[0181] Therefore, for each modality, there is a corresponding weight set, and the weight set includes the weight values adjusted for the modality corresponding to multiple sample scenario data. For each sample data, there is a corresponding weight set, and the weight set includes the weight values adjusted for the multiple modalities corresponding to the sample scenario data.

[0182] For example, the computer device determines the first weight value through the following formula (5).

[0183]

[0184] where T represents the total number of iterations, represents the adjusted weight value of the i-th sample scenario data corresponding to the M-th modality after T iterations, It represents the first weight value of the M-th modality after T iterations.

[0185] In this embodiment, for each modality, the first weight value of the modality is determined based on the weight values obtained by separately training on multiple sample scenario data, which improves the accuracy of the first weight value.

[0186] In some embodiments, the process by which the above computer device determines the second weight value of a modality based on the first weight value of the modality includes the following steps: The computer device determines the sum value of the first weight values corresponding to the multiple modalities; for each modality, the difference between the first weight value of the modality and the sum value is used as the second weight value of the modality. In this embodiment, using the difference between the first weight value and the sum value as the second weight value improves the efficiency of determining the second weight value; and this makes it so that the larger the first weight value, the smaller the second weight value, ensuring that when determining the overall event category, more attention is paid to the influence of simple modalities, thereby improving the accuracy of event category judgment.

[0187] In some other embodiments, for each modality, the computer device uses the quotient of the sum value of the modality and the first weight value as the second weight value of the modality. Alternatively, for each modality, the computer device directly uses the reciprocal of the modality as the second weight value of the modality.

[0188] 504. The computer device weights and sums the multiple single-modal classification results of the scenario data based on the second weight values of the multiple modalities respectively to obtain a weighted sum value.

[0189] In the embodiments of the present application, the multiple single-modal classification results are probability distribution vectors. Correspondingly, the weighted sum value is a vector.

[0190] 505. The computer device determines the target classification result based on the sum value of the first weight values of the multiple modalities and the weighted sum value. The target classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple first weight values.

[0191] Optionally, the process by which the above computer device determines the target classification result based on the sum value of the second weight values of the multiple modalities and the weighted sum value includes the following steps: The computer device determines the sum value of the multiple second weight values, the computer device determines the product of the sum value and the number of multiple modalities minus 1, and uses the weighted sum value and this product as the target classification result.

[0192] For example, the computer device obtains the target classification result through the following formula (6).

[0193]

[0194] where z i represents the target classification result of the i-th scenario data. Represents the sum value of the first weight values corresponding to multiple modalities. Represents the second weight value of the a-th modality. Represents the unimodal classification result of the i-th scenario data in the m-th modality.

[0195] In the embodiment of the present application, through the above steps 503-505, the process of fusing the unimodal classification results of scenario data based on the first weight values of multiple modalities to obtain the target classification result of the scenario data is realized. In this embodiment, the unimodal classification results of multiple modalities are weighted and summed with the second weight value. Since the second weight value is negatively correlated with the first weight value, and the first weight value of the more difficult-to-learn modality is larger, and the first weight value of the easier-to-learn modality is smaller, this makes it more concerned about the influence of the easier-to-learn modality when determining the overall event category, improving the accuracy of event category judgment. And normalization is performed based on the sum of the weight values of multiple modalities, making the obtained target classification result convenient for processing.

[0196] It should be noted that steps 503-505 are only an optional implementation method for realizing the above process, and the computer device can also realize this process through other optional implementation methods, which will not be elaborated here.

[0197] During the use of the modality classification model, we cannot know in advance the difficulty level of each modality's decision. In order to make a more credible and accurate decision, we often need to trust the modalities that are easier to learn during training. Therefore, a larger weight should be given to this simple modality. During the training process of the modality classification model, a higher weight is given to the difficult-to-learn modalities, and a lower weight is given to the easy-to-learn modalities. Therefore, in any modality, the sum of the weight values of all sample scenario data in this modality is used as the weight of this modality. The larger the weight, the more difficult it is to learn this modality, and vice versa. Therefore, the second weight value that is negatively correlated with the first weight value is used as the weight value during the actual use process, so that the simple modality has a higher weight during decision-making.

[0198] The multi-modal data weighted training and weight dynamic learning methods provided in the embodiments of the present application can bring various effects, including but not limited to the following aspects. On the one hand, it can improve the performance of autonomous vehicles. In the field of autonomous driving, through reasonable weighted training of multi-modal data, the modal classification model can better utilize the information provided by each sensor, improve the perception accuracy of the surrounding environment, which will help reduce the occurrence of traffic accidents and improve driving safety. On the other hand, it can optimize the integration of multi-modal data. The method provided in the embodiments of the present application can ensure that data of different modalities (such as cameras, lidars, sensors, etc.) are fully explored without bias towards learning a certain modality, thereby improving the integration quality of multi-modal data and enabling the autonomous driving system to understand the surrounding environment more comprehensively. On the other hand, it can reduce the model's dependence on simple modalities. In most cases, the model tends to overly rely on data of simple modalities while ignoring data of difficult modalities. The dynamic weight learning method in the embodiments of the present application can balance the weights of different modalities, thus reducing the over-reliance on simple modalities and improving the robustness of the model. On the other hand, it can improve the interpretability of model predictions. By dynamically adjusting the weights during the result fusion process, the prediction results of the model become more interpretable, which is very important for users and regulatory agencies of autonomous driving systems as they need to understand why the autonomous driving system makes certain decisions.

[0199] The embodiments of the present application provide a classification method based on multi-modal data. This method obtains event categories based on a modal classification model, and this modal classification model is co-trained with weight values of multiple modalities. Based on the weight values, the modal classification model can perform balanced learning on multiple modalities, rather than just focusing on easily learnable modalities, thereby improving the processing ability of the modal classification model for any modal data and enhancing the robustness of the modal classification model. That is, when obtaining event categories through the modal classification model and the weight values of multiple modalities, accurate single-modal classification results can be obtained based on the modal classification model. On this basis, combined with the trained weight values, accurate multi-modal classification results can be obtained. That is, on the basis of improving the prediction efficiency of event categories through the modal classification model, the accuracy of event category prediction is improved.

[0200] Figure 6 It is a block diagram of a training device for a modal classification model provided according to an embodiment of the present application. Refer to Figure 6 , the device includes:

[0201] An acquisition module 601, configured to acquire multiple sample scenario data of a target scenario and the respective labeled event categories of the multiple sample scenario data, where the sample scenario data includes scenario sub-data of multiple modalities collected under the target scenario;

[0202] A classification module 602, which is used to obtain multiple single-modal classification results of the sample scenario data for each sample scenario data through a modal classification model. Each single-modal classification result corresponds to the scenario sub-data of one modality, and each single-modal classification result is used to indicate the event category corresponding to the scenario sub-data it corresponds to;

[0203] A fusion module 603, which is used to fuse multiple single-modal classification results of the sample scenario data based on the respective weight values of multiple modalities to obtain the first multi-modal classification result of the sample scenario data. The first multi-modal classification result is used to indicate the event category corresponding to the scenario sub-data of multiple modalities;

[0204] An adjustment module 604, which is used to adjust the respective weight values of multiple modalities based on multiple single-modal classification results of the sample scenario data, the first multi-modal classification result of the sample scenario data, and the labeled event category of the sample scenario data;

[0205] The fusion module 603 is also used to fuse multiple single-modal classification results of the sample scenario data based on the respective adjusted weight values of multiple modalities to obtain the second multi-modal classification result of the sample scenario data;

[0206] A determination module 605, which is used to determine a loss value based on the respective labeled event categories of multiple sample scenario data and the respective second multi-modal classification results of multiple sample scenario data, and adjust the model parameters of the modal classification model based on the loss value.

[0207] In some embodiments, the adjustment module 604 is used to:

[0208] If the first multi-modal classification result of the sample scenario data matches the labeled event category of the sample scenario data, reduce the weight value of the first modality among multiple modalities and increase the weight value of the second modality among multiple modalities. The single-modal classification result corresponding to the first modality is the same as the labeled event category of the sample scenario data, and the single-modal classification result corresponding to the second modality is different from the labeled event category of the sample scenario data;

[0209] If the first multi-modal classification result of the sample scenario data does not match the labeled event category of the sample scenario data, no longer adjust the weight value of the first modality and increase the weight value of the second modality.

[0210] In some embodiments, the adjustment module 604 is used to:

[0211] Determine the product of the weight value of the first modality and a preset change rate, and use the difference between the weight value of the first modality and the product as the reduced weight value of the first modality;

[0212] Determine the product of the weight value of the second modality and a preset change rate, and use the sum of the weight value of the second modality and the corresponding product as the increased weight value of the second modality.

[0213] In some embodiments, the determining module 605 is configured to:

[0214] Determine the product of the weight value of the second modality and a preset multiple of the preset change rate, and use the sum value of the weight value of the second modality and the product as the increased weight value of the second modality.

[0215] In some embodiments, the determining module 605 is configured to:

[0216] For each sample scenario data, determine the sub-loss value corresponding to the sample scenario data based on the labeled event category of the sample scenario data and the second multi-modal classification result of the sample scenario data;

[0217] Determine the sum value of the weight values adjusted for multiple modalities corresponding to the sample scenario data;

[0218] Determine the product of the sub-loss value corresponding to the sample scenario data and the sum value;

[0219] Use the sum value of the products corresponding to multiple sample scenario data respectively as the loss value.

[0220] In some embodiments, the fusion module 603 is configured to:

[0221] Based on the weight values adjusted for multiple modalities respectively, perform weighted summation on the single-modal classification results of the sample scenario data to obtain a weighted sum value;

[0222] Based on the sum value of the weight values adjusted for multiple modalities respectively and the weighted sum value, determine the second multi-modal classification result, where the second multi-modal classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple weight values.

[0223] The embodiment of the present application provides a training of a modal classification model. Based on multiple unimodal classification results, multimodal classification results, and labeled event categories, the weight values of multiple modalities are adjusted, and then the model parameters are adjusted based on the classification results determined by the adjusted weight values. In this way, not only the model training is realized, but also the weight values of each modality are adjusted during the model training process, achieving the balance of the weights of different modalities. Moreover, since the loss value is determined based on the second multimodal classification result, and the second multimodal classification result is determined based on the adjusted weight values, that is, the weight values of multiple modalities are indirectly introduced into the loss value. In this way, the model parameters are adjusted based on the loss value introducing the weight values, so that the weight values have an impact on the adjustment of the model parameters. Furthermore, based on the weight values, the modal classification model can perform balanced learning on multiple modalities, rather than only focusing on the modalities that are easy to learn, thereby improving the processing ability of the modal classification model for any modal data and enhancing the robustness of the modal classification model. That is, through the collaborative training of the modal classification model and the weight values of multiple modalities, while improving the training efficiency, based on the trained modal classification model, accurate unimodal classification results can be obtained. On this basis, combined with the trained weight values, accurate multimodal classification results can be obtained. That is, on the basis of improving the event category prediction efficiency through the modal classification model, the accuracy of the event category prediction is improved.

[0224] Figure 7 is a block diagram of a classification device based on multimodal data according to an embodiment of the present application.

[0225] See Figure 7 , the device includes:

[0226] An acquisition module 701, configured to acquire scene data of a target scene, where the scene data includes scene sub-data of multiple modalities collected in the target scene;

[0227] A classification module 702, configured to input the scene data into a modal classification model to obtain multiple unimodal classification results of the scene data, each unimodal classification result corresponding to scene sub-data of one modality, and the modal classification model is trained by the method of any one of the above, and each unimodal classification result is used to indicate the event category corresponding to the corresponding scene sub-data;

[0228] A fusion module 703, configured to fuse multiple unimodal classification results of the scene data based on the first weight values of multiple modalities respectively to obtain a target classification result of the scene data, where the target classification result is used to indicate the event category corresponding to the scene data, and the target classification result is used to indicate the event category corresponding to the scene sub-data of multiple modalities.

[0229] In some embodiments, the fusion module 703 is configured to:

[0230] For each modality, based on the first weight value of the modality, determine the second weight value of the modality, where the second weight value is negatively correlated with the first weight value;

[0231] Based on the second weight values of the multiple modalities respectively, perform a weighted sum of the multiple single-modal classification results of the scenario data to obtain a weighted sum value;

[0232] Based on the sum value of the first weight values of the multiple modalities and the weighted sum value, determine the target classification result, where the target classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple first weight values.

[0233] In some embodiments, the fusion module 703 is configured to:

[0234] Determine the sum value of the first weight values respectively corresponding to the multiple modalities;

[0235] For each modality, use the difference between the first weight value of the modality and the sum value as the second weight value of the modality.

[0236] In some embodiments, the acquisition module 701 is further configured to:

[0237] For each modality, acquire the weight set of the modality, where the weight set includes the weight values of the modality adjusted respectively corresponding to multiple sample scenario data;

[0238] Determine the sum value of the multiple weight values in the weight set to obtain the first weight value of the modality.

[0239] The embodiments of the present application provide a classification device based on multi-modal data. The device obtains an event category based on a modality classification model, and the modality classification model is jointly trained with the weight values of multiple modalities. Based on the weight values, the modality classification model can perform balanced learning on multiple modalities, rather than only focusing on the modalities that are easy to learn, thereby improving the processing ability of the modality classification model for any modality data and improving the robustness of the modality classification model. That is, when the device obtains the event category through the modality classification model and the weight values of multiple modalities, accurate single-modal classification results can be obtained based on the modality classification model. On this basis, combined with the trained weight values, accurate multi-modal classification results can be obtained. That is, on the basis of improving the prediction efficiency of the event category through the modality classification model, the accuracy of the event category prediction is improved.

[0240] In the embodiments of the present application, the computer device can be a terminal or a server. When the computer device is a terminal, the terminal is used as the execution subject to implement the technical solution provided by the embodiments of the present application; when the computer device is a server, the server is used as the execution subject to implement the technical solution provided by the embodiments of the present application; or, the technical solution provided by the present application is implemented through the interaction between the terminal and the server. The embodiments of the present application do not limit this.

[0241] Figure 8 The block diagram of the terminal 800 provided by an exemplary embodiment of the present application is shown.

[0242] Generally, the terminal 800 includes: a processor 801 and a memory 802.

[0243] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0244] The memory 802 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 is used to store at least one program code, and the at least one program code is used to be executed by the processor 801 to implement the training method of the modality classification model or the classification method based on multi-modal data provided by the method embodiments in the present application.

[0245] In some embodiments, the terminal 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 803 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.

[0246] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0247] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 804 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0248] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on or above the surface of the display screen 805. The touch signals can be input to the processor 801 as control signals for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 805, which is provided on the front panel of the terminal 800; in other embodiments, there may be at least two display screens 805, which are respectively provided on different surfaces of the terminal 800 or in a foldable design; in other embodiments, the display screen 805 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 800. Even, the display screen 805 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 805 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0249] The camera module 806 is used to capture images or videos. Optionally, the camera module 806 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of the main camera, the depth-of-field camera, the wide-angle camera, and the telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 806 may also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. The two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0250] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 801 for processing, or input to the radio frequency circuit 804 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into audible sound waves for humans, but also convert the electrical signal into inaudible sound waves for humans for uses such as ranging. In some embodiments, the audio circuit 807 may further include a headphone jack.

[0251] The power supply 808 is used to supply power to each component in the terminal 800. The power supply 808 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 808 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0252] In some embodiments, the terminal 800 further includes one or more sensors 809. The one or more sensors 809 include but are not limited to: an acceleration sensor 810, a gyroscope sensor 811, a pressure sensor 812, an optical sensor 813, and a proximity sensor 814.

[0253] The acceleration sensor 810 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 800. For example, the acceleration sensor 810 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 810. The acceleration sensor 810 can also be used for collecting game or user's motion data.

[0254] The gyroscope sensor 811 can detect the body direction and rotation angle of the terminal 800. The gyroscope sensor 811 can cooperate with the acceleration sensor 810 to collect the 3D actions of the user on the terminal 800. According to the data collected by the gyroscope sensor 811, the processor 801 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0255] The pressure sensor 812 can be disposed on the side frame of the terminal 800 and / or the lower layer of the display screen 805. When the pressure sensor 812 is disposed on the side frame of the terminal 800, it can detect the holding signal of the user on the terminal 800, and the processor 801 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 812. When the pressure sensor 812 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 805. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0256] The optical sensor 813 is used to collect the ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 according to the ambient light intensity collected by the optical sensor 813. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera module 806 according to the ambient light intensity collected by the optical sensor 813.

[0257] The proximity sensor 814, also known as the distance sensor, is usually disposed on the front panel of the terminal 800. The proximity sensor 814 is used to collect the distance between the user and the front of the terminal 800. In one embodiment, when the proximity sensor 814 detects that the distance between the user and the front of the terminal 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from the lit state to the off state; when the proximity sensor 814 detects that the distance between the user and the front of the terminal 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from the off state to the lit state.

[0258] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the terminal 800, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0259] Figure 9It is a schematic structural diagram of a server provided according to an embodiment of the present application. The server 900 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 901 and one or more memories 902. Among them, the memory 902 is used to store executable program codes, and the processor 901 is configured to execute the above-mentioned executable program codes to implement the training method of the modality classification model or the classification method based on multimodal data provided by each of the above method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0260] An embodiment of the present application also provides a computer-readable storage medium. At least one segment of program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the training method of the modality classification model or the classification method based on multimodal data in any of the above implementation manners.

[0261] An embodiment of the present application also provides a computer program product. The computer program product includes at least one segment of program. The at least one segment of program is stored in a computer-readable storage medium. The processor of the computer device reads the at least one segment of program from the computer-readable storage medium, and the processor executes the at least one segment of program, so that the computer device executes the training method of the modality classification model or the classification method based on multimodal data in any of the above implementation manners.

[0262] In some embodiments, the computer program product involved in the embodiments of the present application may be deployed to be executed on a computer device, or on multiple computer devices located at one place. Or, it may be executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network may form a blockchain system.

[0263] All of the above optional technical solutions may be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here. The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for a modal classification model, characterized in that The method includes: Obtaining multiple sample scenario data of a target scenario and the respective labeled event categories of the multiple sample scenario data, where the sample scenario data includes multiple modal sub-scenario data collected under the target scenario; For each sample scenario data, through a modal classification model, obtaining multiple single-modal classification results of the sample scenario data, each of the single-modal classification results corresponding to a modal sub-scenario data, and each of the single-modal classification results being used to indicate the event category corresponding to the corresponding sub-scenario data; Based on the respective weight values of the multiple modalities, fusing the multiple single-modal classification results of the sample scenario data to obtain a first multi-modal classification result of the sample scenario data, where the first multi-modal classification result is used to indicate the event category corresponding to the multiple modal sub-scenario data; Based on the multiple single-modal classification results of the sample scenario data, the first multi-modal classification result of the sample scenario data, and the labeled event category of the sample scenario data, adjusting the respective weight values of the multiple modalities; Based on the respective adjusted weight values of the multiple modalities, fusing the multiple single-modal classification results of the sample scenario data to obtain a second multi-modal classification result of the sample scenario data; Based on the respective labeled event categories of the multiple sample scenario data and the second multi-modal classification results of the multiple sample scenario data, determining a loss value, and adjusting the model parameters of the modal classification model based on the loss value.

2. The method according to claim 1, characterized in that, The adjusting the respective weight values of the multiple modalities based on the multiple single-modal classification results of the sample scenario data, the first multi-modal classification result of the sample scenario data, and the labeled event category of the sample scenario data includes: If the first multi-modal classification result of the sample scenario data matches the labeled event category of the sample scenario data, reducing the weight value of a first modality among the multiple modalities and increasing the weight value of a second modality among the multiple modalities, where the single-modal classification result corresponding to the first modality is the same as the labeled event category of the sample scenario data, and the single-modal classification result corresponding to the second modality is different from the labeled event category of the sample scenario data; If the first multi-modal classification result of the sample scenario data does not match the labeled event category of the sample scenario data, no longer adjusting the weight value of the first modality and increasing the weight value of the second modality.

3. The method according to claim 2, characterized in that, The reducing the weight value of the first modality among the multiple modalities and increasing the weight value of the second modality among the multiple modalities includes: Determining the product of the weight value of the first modality and a preset change rate, and using the difference between the weight value of the first modality and the product as the reduced weight value of the first modality; Determining the product of the weight value of the second modality and the preset change rate, and using the sum of the weight value of the second modality and the corresponding product as the increased weight value of the second modality.

4. The method according to claim 2, characterized in that The increasing the weight value of the second modality includes: Determining the product of the weight value of the second modality and a preset multiple of the preset change rate, and using the sum of the weight value of the second modality and the product as the increased weight value of the second modality.

5. The method according to claim 1, characterized in that Determining a loss value based on the labeled event categories of the respective multiple sample scenario data and the respective second multi-modal classification results of the multiple sample scenario data includes: For each sample scenario data, determining a sub-loss value corresponding to the sample scenario data based on the labeled event category of the sample scenario data and the second multi-modal classification result of the sample scenario data; Determining the sum value of the weighted values adjusted for the multiple modalities corresponding to the sample scenario data; Determining the product of the sub-loss value corresponding to the sample scenario data and the sum value; Taking the sum value of the products corresponding to the multiple sample scenario data respectively as the loss value.

6. The method according to claim 1, wherein The method of fusing the multiple single-modal classification results of the sample scenario data based on the weighted values adjusted for the respective multiple modalities to obtain the second multi-modal classification result of the sample scenario data includes: Based on the weighted values adjusted for the respective multiple modalities, performing weighted summation on the multiple single-modal classification results of the sample scenario data to obtain a weighted sum value; Based on the sum value of the weighted values adjusted for the respective multiple modalities and the weighted sum value, determining the second multi-modal classification result, where the second multi-modal classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple weighted values.

7. A classification method based on multimodal data, characterized in that, The method includes: Obtaining scenario data of a target scenario, where the scenario data includes scenario sub-data of multiple modalities collected in the target scenario; Inputting the scenario data into a modality classification model to obtain multiple single-modal classification results of the scenario data, where each single-modal classification result corresponds to scenario sub-data of one modality, the modality classification model is trained by the method according to any one of claims 1-6, and each single-modal classification result is used to indicate the event category corresponding to the corresponding scenario sub-data; Based on the first weighted values of the respective multiple modalities, fusing the multiple single-modal classification results of the scenario data to obtain a target classification result of the scenario data, where the target classification result is used to indicate the event category corresponding to the scenario sub-data of the multiple modalities.

8. The classification method according to claim 7, wherein The method of fusing the multiple single-modal classification results of the scenario data based on the first weighted values of the respective multiple modalities to obtain a target classification result of the scenario data includes: For each modality, determining a second weighted value of the modality based on the first weighted value of the modality, where the second weighted value is negatively correlated with the first weighted value; Based on the second weighted values of the respective multiple modalities, performing weighted summation on the multiple single-modal classification results of the scenario data to obtain a weighted sum value; Based on the sum value of the first weighted values of the respective multiple modalities and the weighted sum value, determining the target classification result, where the target classification result is positively correlated with the weighted sum value and negatively correlated with the sum value of the multiple first weighted values.

9. The classification method according to claim 8, characterized in that, The method of determining the second weighted value of the modality based on the first weighted value of the modality includes: Determining the sum value of the first weighted values corresponding to the respective multiple modalities; For each modality, taking the difference between the first weighted value of the modality and the sum value as the second weighted value of the modality.

10. The classification method according to claim 7, characterized in that The process of obtaining the first weight values corresponding to the multiple modalities includes: For each modality, obtain the weight set of the modality, where the weight set includes the adjusted weight values of the modality corresponding to multiple sample scenario data; Determine the sum value of the multiple weight values in the weight set to obtain the first weight value of the modality.

11. A training device for a modal classification model, characterized in that, The device includes: An acquisition module, configured to acquire multiple sample scenario data of a target scenario and the labeled event categories of the multiple sample scenario data, where the sample scenario data includes scenario sub-data of multiple modalities collected in the target scenario; A classification module, configured to, for each sample scenario data, obtain multiple single-modal classification results of the sample scenario data through a modality classification model, each of the single-modal classification results corresponding to scenario sub-data of a modality, and each of the single-modal classification results being used to indicate the event category corresponding to the corresponding scenario sub-data; A fusion module, configured to fuse the multiple single-modal classification results of the sample scenario data based on the weight values of the multiple modalities respectively, to obtain a first multi-modal classification result of the sample scenario data, where the first multi-modal classification result is used to indicate the event category corresponding to the scenario sub-data of the multiple modalities; An adjustment module, configured to adjust the weight values of the multiple modalities respectively based on the multiple single-modal classification results of the sample scenario data, the first multi-modal classification result of the sample scenario data, and the labeled event category of the sample scenario data; The fusion module is further configured to fuse the multiple single-modal classification results of the sample scenario data based on the adjusted weight values of the multiple modalities respectively, to obtain a second multi-modal classification result of the sample scenario data; A determination module, configured to determine a loss value based on the labeled event categories of the multiple sample scenario data respectively and the second multi-modal classification results of the multiple sample scenario data respectively, and adjust the model parameters of the modality classification model based on the loss value.

12. A classification device based on multi-modal data, characterized in that, The device includes: An acquisition module, configured to acquire scenario data of a target scenario, where the scenario data includes scenario sub-data of multiple modalities collected in the target scenario; A classification module, configured to input the scenario data into a modality classification model to obtain multiple single-modal classification results of the scenario data, each of the single-modal classification results corresponding to scenario sub-data of a modality, where the modality classification model is trained by the method according to any one of claims 1-6, and each of the single-modal classification results is used to indicate the event category corresponding to the corresponding scenario sub-data; A fusion module, configured to fuse the multiple single-modal classification results of the scenario data based on the first weight values of the multiple modalities respectively, to obtain a target classification result of the scenario data, where the target classification result is used to indicate the event category corresponding to the scenario sub-data of the multiple modalities.

13. A computer device, characterized in that, The computer device includes a processor and a memory. The memory is used to store at least one program, and the at least one program is loaded and executed by the processor to perform the training method of the modality classification model according to any one of claims 1-6 or the classification method based on multimodal data according to any one of claims 7-10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one program, and the at least one program is used to perform the training method of the modality classification model according to any one of claims 1-6 or the classification method based on multimodal data according to any one of claims 7-10.

15. A computer program product, characterized in that, The computer program product includes at least one program. The at least one program is stored in a computer-readable storage medium. The processor of the computer device reads the at least one program from the computer-readable storage medium, and the processor executes the at least one program, so that the computer device performs the training method of the modality classification model according to any one of claims 1-6 or the classification method based on multimodal data according to any one of claims 7-10.