Visual decision model training method and related methods, devices, equipment and media

Through the visual decision model training method, combining image data and visual event data, fusion features are extracted and iteratively trained, the problem of inaccurate execution of visual tasks in complex light scenes is solved, and the accuracy of visual tasks is achieved is achieved.

CN119672499BActive Publication Date: 2025-05-06PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510180760.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-06
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

In the prior art, vision technology based on traditional graphic data is difficult to accurately complete visual tasks in complex light scenes, resulting in the inability of artificial intelligence models to effectively perform corresponding actions.

Method used

A visual decision model training method is proposed. By obtaining image data and visual event data, inputting it into the coding submodel to output fusion features, and inputting the fusion features into the decision submodel to output target actions and evaluation values. By adjusting the model parameters for iterative training until the loss meets the preset conditions, the trained visual decision model is obtained.

Benefits of technology

It is possible to accurately determine the target execution actions corresponding to the visual task in complex light scenarios, and improve the execution accuracy of the artificial intelligence model in visual tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672499B_ABST
    Figure CN119672499B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a visual decision model training method and related methods, devices, equipment and media. The visual decision model includes a coding sub-model and a decision sub-model. By acquiring image data and visual event data, the image data and visual event data are input into the coding sub-model, and fusion features are output; the fusion features are input into the decision sub-model, and the target action and the evaluation value corresponding to the target action are output; the first loss corresponding to the coding sub-model is determined according to the image data, visual event data and fusion features; the second loss corresponding to the decision sub-model is determined according to the target action and the evaluation value; the coding sub-model is iteratively trained according to the first loss to obtain the trained coding sub-model; the decision sub-model is iteratively trained according to the second loss to obtain the trained decision sub-model; the trained visual decision model includes the trained coding sub-model and the trained decision sub-model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to a visual decision model training method and related methods, devices, equipment and media. Background Art

[0002] With the development of artificial intelligence technology, many branches of artificial intelligence technology have emerged, such as visual technology. Currently, artificial intelligence based on visual technology is widely used in different scenarios, such as autonomous driving, navigation and other scenarios.

[0003] In related technologies, visual technology often relies on traditional graphic data, such as obtaining image data by shooting the environment with a camera, and then extracting features from the image data to obtain image noise features used to characterize the image data. Finally, the artificial intelligence model performs corresponding visual tasks based on the image noise features. However, due to the low dynamic range of the camera, in complex light scenes, the visual information that can be provided by the image data is limited, which results in the image noise features being able to represent limited information in the actual environment, and ultimately the artificial intelligence model cannot accurately complete the corresponding visual tasks based on the image noise features. Summary of the invention

[0004] The embodiments of the present application provide a visual decision model training method and related methods, devices, equipment and media. The trained visual decision model can accurately determine the target execution action corresponding to the visual task by combining the target image data and target visual event data corresponding to the visual task.

[0005] In order to achieve the above-mentioned object, an embodiment of the present application provides a visual decision model training method on the one hand, wherein the visual decision model includes an encoding sub-model and a decision sub-model, and the method includes:

[0006] Acquire image data and visual event data, input the image data and the visual event data into an encoding sub-model, and output fusion features;

[0007] Inputting the fusion feature into the decision sub-model, and outputting the target action and the evaluation value corresponding to the target action;

[0008] Determine a first loss corresponding to the encoding sub-model according to the image data, the visual event data and the fusion feature;

[0009] Determine a second loss corresponding to the decision sub-model according to the target action and the evaluation value;

[0010] When the first loss does not satisfy a first preset loss condition, adjusting a model parameter corresponding to the encoding sub-model according to the first loss, and returning to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model;

[0011] When the second loss does not satisfy the second preset loss condition, adjusting the model parameters corresponding to the decision sub-model according to the second loss, returning to execute inputting the fusion feature into the decision sub-model, so as to iteratively train the decision sub-model until the second loss satisfies the second preset loss condition, thereby obtaining a trained decision sub-model;

[0012] The trained visual decision model includes the trained encoding sub-model and the trained decision sub-model.

[0013] In order to achieve the above-mentioned object, an embodiment of the present application provides a visual decision model training device on the one hand, wherein the visual decision model includes an encoding sub-model and a decision sub-model, and the device includes:

[0014] A first input module, used to obtain image data and visual event data, input the image data and the visual event data into an encoding sub-model, and output fusion features;

[0015] A second input module is used to input the fusion feature into the decision sub-model, and output the target action and the evaluation value corresponding to the target action;

[0016] A first determination module, configured to determine a first loss corresponding to the encoding sub-model according to the image data, the visual event data and the fusion feature;

[0017] A second determination module, used to determine a second loss corresponding to the decision sub-model according to the target action and the evaluation value;

[0018] a first iteration module, configured to adjust a model parameter corresponding to the encoding sub-model according to the first loss when the first loss does not satisfy a first preset loss condition, and return to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model;

[0019] a second iteration module, configured to adjust the model parameters corresponding to the decision sub-model according to the second loss when the second loss does not satisfy the second preset loss condition, and return to execute inputting the fusion feature into the decision sub-model to iteratively train the decision sub-model until the second loss satisfies the second preset loss condition, thereby obtaining a trained decision sub-model;

[0020] The trained visual decision model includes the trained encoding sub-model and the trained decision sub-model.

[0021] In some implementations, the encoding sub-model includes a first encoder, a second encoder, and a third encoder; the first input module is configured to:

[0022] Inputting the image data into the first encoder and outputting image noise characteristics;

[0023] Inputting the visual event data into a second encoder and outputting event noise features;

[0024] The image data and the visual event data are input into a third encoder, and a fusion feature is output.

[0025] In some embodiments, the first determination module includes a first loss submodule, a second loss submodule, and a third loss submodule;

[0026] A first loss submodule, configured to determine a first sub-loss value corresponding to the first encoder according to the image data, the fusion feature and the image noise feature;

[0027] A second loss submodule, configured to determine a second sub-loss value corresponding to the second encoder according to the visual event data, the fusion feature and the event noise feature;

[0028] The third loss sub-module is used to determine the third sub-loss value corresponding to the third encoder according to the image noise feature, the event noise feature and the fusion feature, and the first loss includes a first sub-loss value, a second sub-loss value and a third sub-loss value.

[0029] In some implementations, the first loss submodule is configured to:

[0030] Inputting the fusion feature and the image noise feature into a first pre-trained decoder, and outputting decoded image data;

[0031] Determine a first reconstruction loss value corresponding to the first encoder according to a difference between the decoded image data and the image data;

[0032] Determining a first similarity between the image noise feature and the fusion feature, and determining a first classification loss value corresponding to the first encoder according to the first similarity;

[0033] The first reconstruction loss value and the first classification loss value are added to obtain a first sub-loss value corresponding to the first encoder.

[0034] In some embodiments, the second loss submodule is used to:

[0035] Inputting the fusion feature and the event noise feature into a second pre-trained decoder, and outputting decoded visual event data;

[0036] determining a second reconstruction loss value corresponding to the second encoder according to a difference between the decoded visual event data and the visual event data;

[0037] Determine a second similarity between the event noise feature and the fusion feature, and determine a second classification loss value corresponding to the second encoder according to the second similarity;

[0038] The second reconstruction loss value and the second classification loss value are added to obtain a second sub-loss value corresponding to the second encoder.

[0039] In some implementations, the third loss submodule is configured to:

[0040] Determine a first similarity between the event noise feature and the fusion feature, and determine a second similarity between the event noise feature and the fusion feature;

[0041] Determine a third classification loss value corresponding to the third encoder according to the first similarity and the second similarity;

[0042] A third sub-loss value corresponding to the third encoder is determined according to the third classification loss value.

[0043] In some implementations, the third loss submodule is configured to:

[0044] Inputting the fused features and the target action into a third pre-trained decoder, and outputting the decoded fused features at the next moment;

[0045] Determine a first prediction loss value corresponding to the third encoder according to a difference between the decoded fusion feature and the fusion feature at the next moment;

[0046] Inputting the fused feature and the target action into a fourth pre-trained decoder, and outputting a decoding reward value at the next moment;

[0047] Determining a second prediction loss value corresponding to the third encoder according to a difference between the decoded reward value and the reward value corresponding to the target action;

[0048] The first prediction loss value, the second prediction loss value and the third classification loss value are added to obtain a third sub-loss value corresponding to the third encoder.

[0049] In some embodiments, the first iteration module is configured to determine a sub-loss value that does not satisfy the first preset loss condition when any one of the first sub-loss value, the second sub-loss value, and the third sub-loss value does not satisfy the first preset loss condition;

[0050] Determining a target encoder corresponding to the sub-loss value among the first encoder, the second encoder, and the third encoder;

[0051] The model parameters corresponding to the target encoder are adjusted according to the sub-loss value, and the image data and the visual event data are input into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model, wherein the trained encoding sub-model includes a trained third encoder.

[0052] In some embodiments, the decision sub-model includes a policy network and a value network, and a second input module for:

[0053] Input the fused features into the strategy network and output the target action;

[0054] The target action and the fusion feature are input into the value network, and the evaluation value corresponding to the target action is output.

[0055] In some implementations, the second determining module is configured to:

[0056] Determining an action value corresponding to the fusion feature and the target action;

[0057] Determine a target first sub-loss value corresponding to the strategy network according to the action value;

[0058] A target second sub-loss value corresponding to the value network is determined according to the action value and the evaluation value, and the second loss includes the target first sub-loss value and the target second sub-loss value.

[0059] In some embodiments, the second iteration module is used to:

[0060] When either the target first sub-loss value or the target second sub-loss value does not satisfy a second preset loss condition, determining a target sub-loss value that does not satisfy the second preset loss condition;

[0061] Determining a target network corresponding to the target sub-loss value in the strategy network and the value network;

[0062] The model parameters corresponding to the target network are adjusted according to the target sub-loss value, and the fusion feature is input into the decision sub-model to perform iterative training on the decision sub-model until the second loss satisfies the second preset loss condition, thereby obtaining the trained decision sub-model.

[0063] In order to achieve the above-mentioned purpose, an embodiment of the present application provides a visual data processing method, including:

[0064] Obtain target image data and target visual event data corresponding to the visual task;

[0065] Inputting the target image data and the target visual event data into the trained encoding sub-model, and outputting target fusion features;

[0066] Inputting the target fusion feature into the trained decision sub-model, and outputting the target execution action corresponding to the visual task;

[0067] Among them, the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model, and the trained visual decision model is trained according to the visual decision model training method provided in the embodiment of the present application.

[0068] In order to achieve the above-mentioned purpose, an embodiment of the present application provides a visual data processing device, including:

[0069] An acquisition module is used to acquire target image data and target visual event data corresponding to a visual task;

[0070] A fusion module, used for inputting the target image data and the target visual event data into the trained encoding sub-model, and outputting target fusion features;

[0071] A decision module, used for inputting the target fusion feature into the trained decision sub-model, and outputting the target execution action corresponding to the visual task;

[0072] Among them, the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model, and the trained visual decision model is trained according to the visual decision model training method provided in the embodiment of the present application.

[0073] In order to achieve the above-mentioned objectives, on the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the visual decision model training method provided by the embodiment of the present application or the visual data processing method provided by the embodiment of the present application.

[0074] In order to achieve the above-mentioned objectives, on the one hand, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the visual decision model training method provided in the embodiment of the present application or the visual data processing method provided in the embodiment of the present application is implemented.

[0075] In an embodiment of the present application, the visual decision model includes a coding sub-model and a decision sub-model. By acquiring image data and visual event data, the image data and the visual event data are input into the coding sub-model, and a fusion feature is output; the fusion feature is input into the decision sub-model, and a target action and an evaluation value corresponding to the target action are output; a first loss corresponding to the coding sub-model is determined according to the image data, the visual event data and the fusion feature; a second loss corresponding to the decision sub-model is determined according to the target action and the evaluation value; when the first loss does not meet a first preset loss condition, the model parameters corresponding to the coding sub-model are adjusted according to the first loss, and the image data and the visual event data are input into the coding sub-model to iteratively train the coding sub-model until the first loss meets the first preset loss condition, thereby obtaining a trained coding sub-model; when the second loss does not meet a second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the fusion feature is input into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining a trained decision sub-model; wherein the trained visual decision model includes a trained coding sub-model and a trained decision sub-model.

[0076] In this way, the encoding sub-model in the visual decision model can extract visual-related features from the image data and the visual event data, thereby obtaining fused features. The fused features have richer visual information than the image features extracted from the image data alone. Then, the fused features are input into the decision sub-model in the visual decision model, thereby obtaining the target action output by the decision sub-model. In the process of the visual decision model processing the image data and the visual event data, the first loss corresponding to the encoding sub-model can be determined according to the image data, the visual event data and the fused features, and the second loss corresponding to the decision sub-model can be determined according to the target action and the evaluation value. When the first loss does not meet the first preset loss condition, the model parameters corresponding to the encoding sub-model are adjusted according to the first loss, and the image data and the visual event data are input into the encoding sub-model to iteratively train the encoding sub-model until the first loss meets the first preset loss condition, thereby obtaining the trained encoding sub-model. When the second loss does not meet the second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the fused features are input into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining the trained decision sub-model. In the process of visual decision model training, the encoding sub-model and the decision sub-model are trained synchronously. Through continuous iteration of the encoding sub-model, the trained encoding sub-model can more accurately obtain the fusion features between the visual event data and the image data. Through continuous iteration of the decision sub-model, the decision sub-model can more accurately determine the target action to be performed based on the fusion features. Compared with the related art that only uses image data to obtain visual information, the visual decision model finally trained in this application can output the target fusion features that can more accurately represent the visual information of the environment based on the target image data and the target visual event data in the environment when performing the corresponding visual task, and accurately determine the target execution action corresponding to the visual task based on the target fusion features.

[0077] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0079] Figure 1 It is a schematic diagram of the system framework corresponding to the visual decision model training method and the visual data processing method provided in the embodiment of the present application;

[0080] Figure 2 It is a flowchart of the visual decision model training method provided in the embodiment of the present application;

[0081] Figure 3 is a schematic diagram of the structure of the visual decision model provided in the embodiment of the present application;

[0082] Figure 4 is a flowchart of step 230 provided in an embodiment of the present application;

[0083] Figure 5 is another flowchart of the visual decision model training method provided in an embodiment of the present application;

[0084] Figure 6 is a flowchart of a visual data processing method provided in an embodiment of the present application;

[0085] Figure 7 is a schematic diagram of images and visual events provided by embodiments of the present application;

[0086] Figure 8 is a schematic diagram of a trained visual decision model provided in an embodiment of the present application;

[0087] Fig. 9 It is a structural schematic diagram of a visual decision model training device provided in an embodiment of the present application;

[0088] Fig.10 is a schematic diagram of the structure of a visual data processing device provided in an embodiment of the present application;

[0089] Fig.11 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0090] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0091] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on image data and visual event data, the user's permission or consent will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or jump to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0092] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, multiple steps appearing in a specific order are included, but it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second" or "target" in this document are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0093] The visual decision model training method and visual data processing method provided in the embodiments of the present application relate to the field of artificial intelligence technology. The visual decision model training method and visual data processing method provided in the embodiments of the present application can be applied to a terminal, can also be applied to a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms and other basic cloud computing services. Cloud servers; software can be applications that implement visual decision model training methods and visual data processing methods, etc., but are not limited to the above forms.

[0094] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0095] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:

[0096] Visual event data: Based on the data obtained by the dynamic vision sensor (DVS), a new type of visual sensor that is very different from traditional frame-based image sensors (such as CCD and CMOS sensors). Traditional image sensors capture images at a fixed frame rate (such as 30 frames per second). They take a whole image at each time interval, regardless of whether there are changes in the scene. The dynamic vision sensor is an event-driven sensor that only responds to dynamic changes in the scene, such as the movement of objects, changes in light, etc. Each pixel in the dynamic vision sensor works independently and can sense changes in brightness in the scene (such as the movement of objects). When the pixel detects that the brightness change exceeds a certain threshold, an "event" is generated. This event contains information such as the pixel position, the time when the event occurred, and the polarity of the brightness change (whether it becomes brighter or darker). The data corresponding to this event is the visual event data.

[0097] Image data: Image data refers to the grayscale value or color value of each pixel expressed in numerical values. It is obtained through traditional image sensors (such as CMOS).

[0098] Voxelization: Voxelization is a process of converting a 3D geometric model (such as a CAD model, 3D scan data, etc.) or a spatial region into a voxel representation. A voxel is a volume element in a 3D space, similar to a pixel in a 2D image. It is the smallest unit after the 3D space is discretized.

[0099] Event Polarity: In event-based vision systems (such as DVS - Dynamic Vision Sensor), event polarity refers to the direction of the brightness change represented by the event. It is used to distinguish whether the brightness at the pixel is increasing or decreasing. When a pixel detects a brightness change exceeding a certain threshold, an event is generated. If the brightness increases, the event polarity is usually marked as "Positive"; if the brightness decreases, the event polarity is marked as "Negative". For example, in a simple scene there is a white object moving on a black background. When the leading edge of the object enters the field of view of a pixel, the brightness of the pixel will increase due to receiving more light, and the event polarity generated is positive; when the trailing edge of the object leaves the field of view of the pixel, the brightness of the pixel will decrease, and the event polarity generated is negative.

[0100] First, let’s describe the technical problems existing in the relevant technologies:

[0101] With the development of artificial intelligence technology, many branches of artificial intelligence technology have emerged, such as visual technology. Currently, artificial intelligence based on visual technology is widely used in different scenarios, such as autonomous driving, navigation and other scenarios.

[0102] In related technologies, visual technology often relies on traditional graphic data, such as obtaining image data by shooting the environment with a camera, and then extracting features from the image data to obtain image noise features used to characterize the image data. Finally, the artificial intelligence model performs corresponding visual tasks based on the image noise features. However, due to the low dynamic range of the camera, in complex light scenes, the visual information that can be provided by the image data is limited, which results in the image noise features being able to represent limited information in the actual environment, and ultimately the artificial intelligence model cannot accurately complete the corresponding visual tasks based on the image noise features.

[0103] For example, the artificial intelligence model cannot accurately determine the action corresponding to the visual task based on the image noise characteristics, resulting in the inability to complete the visual task after the predicted action is executed.

[0104] In order to solve the above problems, the embodiment of the present application proposes a visual decision model and a training method thereof. The encoding sub-model in the visual decision model can extract visual-related features from image data and visual event data, thereby obtaining fused features. The fused features have richer visual information than the image features extracted from the image data alone. Then, the fused features are input into the decision sub-model in the visual decision model, thereby obtaining the target action output by the decision sub-model. In the process of the visual decision model processing the image data and the visual event data, the first loss corresponding to the encoding sub-model can be determined according to the image data, the visual event data and the fused features, and the second loss corresponding to the decision sub-model can be determined according to the target action and the evaluation value. When the first loss does not meet the first preset loss condition, the model parameters corresponding to the encoding sub-model are adjusted according to the first loss, and the image data and the visual event data are input into the encoding sub-model to perform iterative training on the encoding sub-model until the first loss meets the first preset loss condition, thereby obtaining the trained encoding sub-model. When the second loss does not meet the second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the execution is returned to input the fusion features into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining the trained decision sub-model. In the process of visual decision model training, the encoding sub-model and the decision sub-model are trained synchronously, and through continuous iteration of the encoding sub-model, the trained encoding sub-model can more accurately obtain the fusion features between the visual event data and the image data, and through continuous iteration of the decision sub-model, the decision sub-model can more accurately determine the target action to be executed based on the fusion features. Compared with the related art that only uses image data to obtain visual information, the visual decision model finally trained in the present application can output target fusion features that can more accurately characterize the visual information of the environment based on the target image data and the target visual event data in the environment when performing the corresponding visual task, and accurately determine the target execution action corresponding to the visual task based on the target fusion features.

[0105] From the above, it can be seen that the trained visual decision model in the present application can accurately determine the target execution action corresponding to the visual task by combining the target image data and the target visual event data corresponding to the visual task, and accurately complete the visual task through the target execution action.

[0106] See also Figure 1 , Figure 1 1 is a schematic diagram of a system framework corresponding to the visual decision model training method and the visual data processing method provided in the embodiment of the present application. The visual decision model training method and the visual data processing method provided in the embodiment of the present application can be applied to the system framework.

[0107] The terminal 140 or the server 110 may be a device for executing a visual decision model training method or a visual data processing method.

[0108] The terminal 140 includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud office, enterprise management, etc. In addition, it can be a single device or a collection of multiple devices. For example, multiple desktop computers are interconnected through a local area network, share a display, etc. to work together, and together constitute a terminal 140. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0109] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with the ordinary terminal 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0110] The gateway 120 is also called an internetwork connector or a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. The gateway is a translator between two systems that use different communication protocols, data formats or languages, or even completely different architectures. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 must be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 must also be sent to the corresponding terminal 140 through the gateway 120.

[0111] In the embodiment of the present application, the visual decision model trained by the visual decision model training method can be applied to various scenarios such as autonomous driving, navigation, and environmental detection, and there is no limitation on the scenarios in which the trained visual decision model in the present application is applied.

[0112] See also Figure 2 , Figure 2 : is a flow chart of a visual decision model training method provided in an embodiment of the present application. The visual decision model includes an encoding sub-model and a decision sub-model, and the visual decision model training method may include the following steps:

[0113] Step 210: Obtain image data and visual event data, input the image data and visual event data into the encoding sub-model, and output fusion features;

[0114] Step 220: Input the fusion feature into the decision sub-model, and output the target action and the evaluation value corresponding to the target action;

[0115] Step 230: determining a first loss corresponding to the encoding sub-model according to the image data, the visual event data and the fusion feature;

[0116] Step 240: Determine a second loss corresponding to the decision sub-model according to the target action and the evaluation value;

[0117] Step 250: when the first loss does not satisfy the first preset loss condition, adjust the model parameters corresponding to the encoding sub-model according to the first loss, and return to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model;

[0118] Step 260: When the second loss does not meet the second preset loss condition, adjust the model parameters corresponding to the decision sub-model according to the second loss, return to execute and input the fusion feature into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, and obtain the trained decision sub-model, wherein the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model.

[0119] Steps 210 to 260 will be described in detail below.

[0120] In step 210, image data and visual event data are acquired, the image data and visual event data are input into the encoding sub-model, and fusion features are output.

[0121] The image data can be acquired by a camera shooting environment, and the visual event data can be acquired by a dynamic visual sensor sensing environment. The image data and the visual event data are training data for the visual decision model.

[0122] Please combine Figure 3 , Figure 3 It is a structural diagram of the visual decision model provided in the embodiment of the present application.

[0123] The encoding sub-model includes a first encoder, a second encoder and a third encoder, and the first encoder, the second encoder and the third encoder respectively correspond to different input data. The input of the first encoder is image data, the input of the second encoder is visual event data, and the input of the third encoder is visual event data and image data.

[0124] In some embodiments, image data and visual event data are input into an encoding sub-model, and fusion features are output, including:

[0125] (1.1) Inputting image data into the first encoder and outputting image noise features;

[0126] (1.2) Inputting the visual event data into the second encoder and outputting the event noise features;

[0127] (1.3) The image data and visual event data are input into the third encoder, and the fused features are output.

[0128] Among them, the first encoder, the second encoder and the third encoder can be constructed based on a convolutional neural network, and the network structures corresponding to the first encoder, the second encoder and the third encoder can be the same, but the training tasks of the first encoder, the second encoder and the third encoder are different.

[0129] The training task corresponding to the first encoder is to extract image noise features from image data, the training task corresponding to the second encoder is to extract event noise features from visual event data, and the training task corresponding to the third encoder is to extract fusion features corresponding to both image data and visual event data.

[0130] Among them, the first encoder input dimension is ( ), is the number of channels, namely RGB three channels, and are the height and width of the image, respectively. Since the decision sub-model needs historical time series information to make decisions, c is set to the number of channels in a single image multiplied by 3. 3 represents three moments, and the three RGB channels are multiplied by 3, so the number of channels , and Can be set to 128, 128, or other higher resolutions. ( ) These three variables can be set according to the scene, and there is no restriction here. Finally, the first encoder outputs the image noise feature, which is represented as the feature vector , , ), where t is time, is the first encoder, , , They are image data at three different times. The spatial dimension of this feature vector can be set according to actual needs, for example, the spatial dimension is 50.

[0131] For the second encoder, the visual event data is first voxelized and represented as , It contains many events, namely , this definition expresses Contains events gather, belong During this time period, It can be 0.05 seconds. At the same time, each event is a four-tuple consisting of These four types of quantities represent the pixel position of the event ( ), trigger time and event polarity Then Voxelization to Tensor , the dimension of this tensor is ( ),in and are the height and width of the image, respectively. , and the first encoder input and The dimensions are the same, is the dimension of the voxel, can be set to 5.

[0132] Among them, the tensor It can be expressed as:

[0133] .in, is the horizontal coordinate of a pixel, is the vertical coordinate of a pixel, is the nth time interval. Represents the normalized event time, which is calculated as follows: , express The moment of the first event in express The moment of the i-th event in .

[0134] Similarly, since the decision sub-model needs historical time series information to make decisions, the final input of the second encoder is the tensor under three time series as input, and the final output is the event noise feature, which is the feature vector in the latent space, expressed as , ),in For the third encoder, , They are the tensors corresponding to the events in three time series respectively.

[0135] For the third encoder, the input is visual event data and image data, and the final output is the fusion feature corresponding to the visual event data and the image data. The fusion feature is used to characterize the visual information in the environment, which can be specifically expressed as: , ),in For the third encoder.

[0136] By combining image data and visual event data, the fusion features obtained can more accurately represent the visual information in the environment, provide more accurate training data for the subsequent training of the decision sub-model, and improve the training efficiency of the decision sub-model.

[0137] In step 220, the fused features are input into the decision sub-model, and the target action and the evaluation value corresponding to the target action are output.

[0138] Please combine Figure 3 The visual decision model includes a decision sub-model, which includes a policy network and a value network. The policy network is mainly used to generate actions. It is a mapping function from state space to action space, which determines what action the agent should take according to the current state of the environment.

[0139] The value network is used to evaluate the value of the state. It receives the state of the environment as input and outputs a numerical value to represent the potential value of the state for the agent to achieve its goal. This value is measured based on the long-term cumulative reward.

[0140] The fused features can be understood as "states". By inputting the fused features into the policy network, the action corresponding to the "state" can be output, that is, the target action corresponding to the fused features.

[0141] The target action and fusion features are input into the value network. The value network can evaluate the input "state" and "action" and output the evaluation value corresponding to the target action.

[0142] In step 230, a first loss corresponding to the encoding sub-model is determined based on the image data, the visual event data and the fusion features.

[0143] Among them, the training of the encoding sub-model is actually the training of the first encoder, the second encoder and the third encoder, so that each encoder is trained according to the corresponding training task to accurately extract features of the input data. For example, the first encoder can accurately extract the image noise features in the image data, the second encoder can accurately extract the event noise features in the visual event data, and the third encoder can accurately extract the fusion features corresponding to the visual event data and the image data. The fusion features can also be understood as features shared by the visual event data and the image data. The above features can all be understood as vectors. Each of the first encoder, the second encoder and the third encoder has a corresponding loss value, and the first loss includes the loss values ​​corresponding to the first encoder, the second encoder and the third encoder respectively.

[0144] See also Figure 4 , Figure 4It is a flowchart of step 230 provided in an embodiment of the present application.

[0145] In some embodiments, determining the first loss corresponding to the encoding sub-model according to the image data, the visual event data, and the fusion feature may include the following steps:

[0146] Step 301: Determine a first sub-loss value corresponding to a first encoder according to image data, fusion features and image noise features;

[0147] Step 302: Determine a second sub-loss value corresponding to a second encoder according to the visual event data, the fusion feature and the event noise feature;

[0148] Step 303: Determine a third sub-loss value corresponding to the third encoder according to the image noise feature, the event noise feature and the fusion feature, where the first loss includes a first sub-loss value, a second sub-loss value and a third sub-loss value.

[0149] Steps 301 to 303 are described in detail below.

[0150] In step 301, a first sub-loss value corresponding to a first encoder is determined based on image data, fusion features, and image noise features.

[0151] Among them, the fusion feature corresponds to the real visual information in the environment of the image data, and the image noise feature corresponds to the noise in the environment of the image data. Then the fusion feature and the image noise feature are combined and restored into image-type data, which should be close to the real image data. The training task of the first encoder is to accurately extract the image noise features in the image data.

[0152] Based on this, in the present application, the loss value between the data and the image data can be determined according to the first reconstruction loss value corresponding to the first encoder, that is, the fusion feature and the image noise feature are combined together and restored to the image type data.

[0153] In addition, the fused features can be considered as positive samples, and the image noise features can be considered as negative samples. In theory, the farther the distance between the positive samples and the negative samples is, the better. Based on this, the distance between the fused features and the image noise features can be distanced, and the distance operation can be converted into a classification problem of the image noise features and the fused features, so that the first classification loss value corresponding to the first encoder can be determined.

[0154] The first classification loss value and the first reconstruction loss value are added together to form a first sub-loss value corresponding to the first encoder.

[0155] In some implementations, determining a first sub-loss value corresponding to the first encoder according to the image data, the fusion feature, and the image noise feature includes:

[0156] (1.1) Inputting the fusion features and the image noise features into the first pre-trained decoder, and outputting the decoded image data;

[0157] (1.2) determining a first reconstruction loss value corresponding to the first encoder according to a difference between the decoded image data and the image data;

[0158] (1.3) determining a first similarity between the image noise feature and the fusion feature, and determining a first classification loss value corresponding to the first encoder according to the first similarity;

[0159] (1.4) The first reconstruction loss value and the first classification loss value are added to obtain a first sub-loss value corresponding to the first encoder.

[0160] The specific calculation method of the first reconstruction loss value for the first encoder is as follows:

[0161] .in, is the first pre-trained decoder, K represents a set of moments of a batch of randomly sampled training samples (image data), represents the sum of the elements at the corresponding positions of the two features, and A represents the image data, that is, . is the first reconstruction loss value.

[0162] By inputting the fusion features combined with the image noise features into the first pre-trained decoder, the decoded image data is restored. When the difference between the decoded image data and the image data is smaller, that is, the first reconstruction loss value is smaller, it means that the first encoder extracts the image noise features of the image data more accurately. In a batch of training samples, by adding the difference between each image data and the corresponding decoded image data, the first reconstruction loss when training the first encoder for the batch of training samples is obtained. Subsequently, the first reconstruction loss can be used to guide the adjustment of the parameters of the first encoder to achieve iteration of the first encoder.

[0163] Among them, the specific calculation method for the first classification loss value of the first encoder is as follows:

[0164] Among them, the image noise feature is used as a negative sample, the fusion feature is used as a positive sample, and f is a function for calculating the similarity of two feature vectors. is the first similarity between the image noise feature and the fusion feature, is the similarity between the corresponding fused features. The fused features can be considered as positive samples, and the image noise features can be considered as negative samples. In theory, the farther the distance between the positive and negative samples, the better. Based on this, the distance between the fused features and the image noise features can be distanced, and the distance operation can be converted into a classification problem of the image noise features and the fused features. In this way, the first classification loss value corresponding to the first encoder can be determined, that is, .

[0165] Finally, the first reconstruction loss value and the first classification loss value may be added to obtain a first sub-loss value corresponding to the first encoder. The first loss of the encoding sub-model includes the first sub-loss value.

[0166] In step 302, a second sub-loss value corresponding to a second encoder is determined based on the visual event data, the fusion feature and the event noise feature.

[0167] Among them, the fusion feature corresponds to the real visual information in the environment of the visual event data, and the visual event noise feature corresponds to the noise in the environment of the visual event data. Then the fusion feature and the event noise feature are combined and restored into the data of the visual event type, which should be close to the real visual event data. The training task of the second encoder is to accurately extract the event noise features in the visual event data.

[0168] Based on this, in the present application, the loss value between the data and the visual event data can be determined based on the second reconstruction loss value corresponding to the second encoder, that is, the fusion feature and the event noise feature are combined and restored to the data of the visual event type.

[0169] In addition, the fused features can be considered as positive samples, and the event noise features can be considered as negative samples. In theory, the farther the distance between the positive samples and the negative samples is, the better. Based on this, the distance between the fused features and the event noise features can be distanced, and the distance operation can be converted into a classification problem of the event noise features and the fused features, so that the second classification loss value corresponding to the second encoder can be determined.

[0170] The second classification loss value and the second reconstruction loss value are added together to form a second sub-loss value corresponding to the second encoder.

[0171] In some implementations, determining a second sub-loss value corresponding to a second encoder according to the visual event data, the fusion feature, and the event noise feature includes:

[0172] (2.1) Inputting the fusion features and the event noise features into the second pre-trained decoder, and outputting the decoded visual event data;

[0173] (2.2) determining a second reconstruction loss value corresponding to the second encoder according to a difference between the decoded visual event data and the visual event data;

[0174] (2.3) determining a second similarity between the event noise feature and the fusion feature, and determining a second classification loss value corresponding to the second encoder according to the second similarity;

[0175] (2.4) The second reconstruction loss value and the second classification loss value are added to obtain a second sub-loss value corresponding to the second encoder.

[0176] The specific calculation method of the second reconstruction loss value for the second encoder is as follows:

[0177] .in, is the second pre-trained decoder, K represents a set of moments of a batch of randomly sampled training samples (visual event data), represents the sum of the elements at the corresponding positions of the two features, and B represents the visual event data, that is, , . is the second reconstruction loss value.

[0178] By inputting the fusion features combined with the event noise features into the second pre-trained decoder, the decoded visual event data is restored. When the difference between the decoded visual event data and the visual event data is smaller, that is, the second reconstruction loss value is smaller, it means that the second encoder extracts the event noise features of the visual event data more accurately. In a batch of training samples, by adding the difference between each visual event data and the corresponding decoded visual event data, the second reconstruction loss when training the second encoder for the batch of training samples is obtained. Subsequently, the second reconstruction loss can be used to guide the parameter adjustment of the second encoder to achieve iteration of the second encoder.

[0179] Among them, the second classification loss value of the second encoder is specifically calculated as follows:

[0180] Among them, the event noise feature is used as a negative sample, the fusion feature is used as a positive sample, and f is a function for calculating the similarity of the two feature vectors. is the second similarity between the event noise feature and the fusion feature, is the similarity between the corresponding fused features. The fused features can be considered as positive samples, and the event noise features can be considered as negative samples. In theory, the farther the distance between the positive and negative samples, the better. Based on this, the distance between the fused features and the event noise features can be distanced, and the distance operation can be converted into a classification problem of the event noise features and the fused features. In this way, the second classification loss value corresponding to the second encoder can be determined, that is, .

[0181] Finally, the second reconstruction loss value and the second classification loss value can be added to obtain a second sub-loss value corresponding to the second encoder. The first loss of the encoding sub-model includes the second sub-loss value.

[0182] In step 303, a third sub-loss value corresponding to the third encoder is determined according to the image noise feature, the event noise feature and the fusion feature, and the first loss includes a first sub-loss value, a second sub-loss value and a third sub-loss value.

[0183] Among them, the fusion feature can be considered as a positive sample, and the event noise feature and the image noise feature can be considered as a negative sample. In theory, the farther the distance between the positive sample and the negative sample is, the better. Based on this, the distance between the fusion feature and the event noise feature can be widened, and the distance between the fusion feature and the image noise feature can be widened. The widening operation can be converted into a classification problem of the image noise feature, the event noise feature and the fusion feature, so that the third classification loss value corresponding to the third encoder can be determined. The third sub-loss value includes the third classification loss value.

[0184] In some implementations, determining a third sub-loss value corresponding to the third encoder according to the image noise feature, the event noise feature, and the fusion feature includes:

[0185] (3.1) determining a first similarity between the event noise feature and the fusion feature, and determining a second similarity between the event noise feature and the fusion feature;

[0186] (3.2) determining a third classification loss value corresponding to the third encoder according to the first similarity and the second similarity;

[0187] (3.3) Determine a third sub-loss value corresponding to the third encoder according to the third classification loss value.

[0188] Among them, the third classification loss value corresponding to the third encoder can be calculated as follows:

[0189] Among them, the event noise feature is used as a negative sample, the image noise feature is used as a negative sample, the fusion feature is used as a positive sample, and f is a function for calculating the similarity of two feature vectors. is the first similarity between the image noise feature and the fusion feature, is the second similarity between the event noise feature and the fusion feature, is the similarity between the corresponding fusion features. Theoretically, the farther the distance between the positive sample and the negative sample is, the better. Based on this, the distance between the fusion feature and the event noise feature can be distanced, and the distance between the fusion feature and the image noise feature can be distanced. The distance operation can be converted into a classification problem of image noise features, event noise features and fusion features. In this way, the third classification loss value corresponding to the third encoder can be determined, that is, .

[0190] Finally, the third sub-loss value corresponding to the third encoder can be determined according to the third classification loss value.

[0191] In some implementations, determining a third sub-loss value corresponding to a third encoder according to the third classification loss value includes:

[0192] (3.3.1) Input the fused features and the target action into the third pre-trained decoder, and output the decoded fused features at the next moment;

[0193] (3.3.2) determining a first prediction loss value corresponding to the third encoder according to a difference between the decoded fusion feature and the fusion feature at the next moment;

[0194] (3.3.3) Input the fused features and target action into the fourth pre-trained decoder and output the decoding reward value at the next moment;

[0195] (3.3.4) determining a second prediction loss value corresponding to the third encoder according to a difference between the decoded reward value and the reward value corresponding to the target action;

[0196] (3.3.5) The first prediction loss value, the second prediction loss value, and the third classification loss value are added to obtain a third sub-loss value corresponding to the third encoder.

[0197] Among them, the specific calculation method of the first prediction loss value corresponding to the third encoder is as follows:

[0198] .in, is the first prediction loss value, is the third pre-trained decoder, which is used to determine the decoding fusion features at the next moment according to the fusion features and the target action, and is the fusion feature of the next moment, which is specifically obtained from the experience pool and can be considered as the real label. Since the third pre-trained decoder determines the decoding fusion feature based on the fusion feature, by determining the difference between the decoding fusion feature and the fusion feature of the next moment, it can be reflected whether the fusion feature predicted by the third encoder based on the image data and the visual event data is accurate, thereby guiding the third encoder to adjust the corresponding parameters through the first prediction loss value to achieve the iteration of the third encoder.

[0199] The specific calculation method of the second prediction loss value corresponding to the third encoder is as follows:

[0200] .in is the second prediction loss value, is the fourth pre-trained decoder, which is used to determine the decoding reward value at the next moment according to the fusion feature and the target action, and The reward value after executing the target action is obtained from the experience pool and can be considered as the real label. Since the fourth pre-trained decoder determines the decoding reward value based on the fusion feature, the difference between the decoding reward value and the reward value corresponding to the target action can reflect whether the fusion feature predicted by the third encoder based on the image data and the visual event data is accurate, thereby guiding the third encoder to adjust the corresponding parameters through the second prediction loss value to achieve the iteration of the third encoder.

[0201] Finally, the first prediction loss value, the second prediction loss value and the third classification loss value may be added to obtain a third sub-loss value corresponding to the third encoder. The third sub-loss value is used to guide the third encoder to adjust the corresponding parameters to implement iteration of the third encoder.

[0202] From step 301 to step 303, it can be seen that through the joint training of the first encoder, the second encoder and the third encoder, since the calculation of the third sub-loss value of the third encoder requires the use of image noise features and event noise features, the synchronous training of the first encoder, the second encoder and the third encoder can promote the third encoder to accurately extract the fusion features corresponding to the visual event data and the image data. In this way, the trained third encoder can accurately extract the target fusion features corresponding to the target visual event data and the target image data.

[0203] In step 240, a second loss corresponding to the decision sub-model is determined according to the target action and the evaluation value.

[0204] Among them, the decision sub-model includes a policy network and a value network. During the training process of the decision sub-model, the policy network and the value network are trained synchronously. The policy network corresponds to the target first sub-loss value, and the value network corresponds to the target second sub-loss value. The second loss includes the target first sub-loss value and the target second sub-loss value.

[0205] In some implementations, determining a second loss corresponding to the decision sub-model according to the target action and the evaluation value includes:

[0206] (1.1) Determine the action value corresponding to the fusion feature and the target action;

[0207] (1.2) Determine the target first sub-loss value corresponding to the strategy network based on the action value;

[0208] (1.3) The target second sub-loss value corresponding to the value network is determined according to the action value and the evaluation value. The second loss includes the target first sub-loss value and the target second sub-loss value.

[0209] Among them, the action value corresponding to the fusion feature and the target action can be determined based on the Bellman equation, which can be specifically expressed as ,in is the target action output by the policy network. Then, the target first sub-loss value corresponding to the policy network is determined according to the action value, which can be specifically expressed as:

[0210] .in, is the target first sub-loss value, represents the policy network, α is a hyperparameter used to adjust the weight or proportion of each item in the loss function, Indicates that the condition is the fusion feature The action taken at time step t+1 when .

[0211] Then, the target second sub-loss value corresponding to the value network is determined according to the action value and the evaluation value. The second loss includes the target first sub-loss value and the target second sub-loss value. It can be specifically expressed as:

[0212] .in, represents the target second sub-loss value, where represents the reward value obtained at time step t, is the action value, which means Take the target action Expected return, is the evaluation value, indicating that from this state Start with an estimate of the cumulative rewards that may be received in the future, It is a discount factor used to balance the importance of immediate rewards and future rewards, and is usually between 0 and 1.

[0213] In the process of training the policy network, the corresponding parameters of the policy network need to be adjusted based on the target first sub-loss value as a guide. In the process of training the value network, the corresponding parameters of the value network need to be adjusted based on the target second sub-loss value as a guide.

[0214] During the training process, their updates are interrelated. The back-propagation algorithm is usually used to update the network parameters. When training the value network, the target value is calculated based on the action sequence generated by the policy network, and this target value is used to update the parameters of the value network. At the same time, the update of the policy network will also consider the state value evaluated by the value network. For example, in the improved algorithm of the policy gradient algorithm, the value information evaluated by the value network is used as an adjustment factor of the reward signal to update the policy network, so that the policy network is more inclined to choose actions that can bring higher value states.

[0215] In step 250, when the first loss does not satisfy the first preset loss condition, the model parameters corresponding to the encoding sub-model are adjusted according to the first loss, and the execution is returned to input the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining the trained encoding sub-model.

[0216] Among them, the encoding sub-model includes a first encoder, a second encoder and a third encoder. During the training process of the encoding sub-model, the first encoder, the second encoder and the third encoder need to be fully trained to obtain the trained encoding sub-model.

[0217] In some embodiments, when the first loss does not satisfy the first preset loss condition, adjusting the model parameters corresponding to the encoding sub-model according to the first loss, and returning to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, and obtaining the trained encoding sub-model, including:

[0218] (1.1) When any one of the first sub-loss value, the second sub-loss value, and the third sub-loss value does not satisfy the first preset loss condition, determining the sub-loss value that does not satisfy the first preset loss condition;

[0219] (1.2) determining a target encoder corresponding to the sub-loss value among the first encoder, the second encoder, and the third encoder;

[0220] (1.3) adjusting the model parameters corresponding to the target encoder according to the sub-loss value, and returning to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining the trained encoding sub-model, wherein the trained encoding sub-model includes the trained third encoder.

[0221] The first preset loss condition is that the first sub-loss value, the second sub-loss value and the third sub-loss value are all less than the corresponding preset loss thresholds. That is, if the first sub-loss value is less than the first preset loss threshold corresponding to the first encoder, and the second sub-loss value is less than the second preset loss threshold corresponding to the second encoder, and the third sub-loss value is less than the third preset loss threshold corresponding to the third encoder, then it is considered that the first loss meets the first preset loss condition.

[0222] When the first sub-loss value is greater than or equal to the first preset loss threshold corresponding to the first encoder, the first sub-loss value does not meet the first preset loss condition. A target encoder corresponding to the sub-loss value is determined among the first encoder, the second encoder, and the third encoder, that is, the first encoder corresponding to the first sub-loss value is determined as the target encoder.

[0223] Then, the model parameters corresponding to the target encoder are adjusted according to the sub-loss value, that is, the model parameters corresponding to the first encoder are adjusted according to the first sub-loss value to implement iteration of the first encoder. Then, the execution is returned to input the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining the trained encoding sub-model.

[0224] It should be noted that the trained encoding sub-model is mainly used to extract features from image data and visual event data to obtain fusion features, so the trained encoding sub-model may only include the third encoder. The first encoder and the second encoder actually play a role in assisting the third encoder in training.

[0225] In step 260, when the second loss does not meet the second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the execution is returned to input the fusion features into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining a trained decision sub-model, wherein the trained visual decision model includes a trained encoding sub-model and a trained decision sub-model.

[0226] Among them, the decision sub-model includes a policy network and a value network. During the training process of the decision sub-model, the policy network and the value network need to be fully trained before the trained decision sub-model can be obtained.

[0227] In some embodiments, when the second loss does not satisfy the second preset loss condition, adjusting the model parameters corresponding to the decision sub-model according to the second loss, returning to execute inputting the fusion feature into the decision sub-model to iteratively train the decision sub-model until the second loss satisfies the second preset loss condition, and obtaining the trained decision sub-model, including:

[0228] (1.1) when either the target first sub-loss value or the target second sub-loss value does not satisfy the second preset loss condition, determining the target sub-loss value that does not satisfy the second preset loss condition;

[0229] (1.2) Determine the target network corresponding to the target sub-loss value in the policy network and the value network;

[0230] (1.3) Adjust the model parameters corresponding to the target network according to the target sub-loss value, and return to execute to input the fused features into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining the trained decision sub-model.

[0231] Among them, the second preset loss condition is that the loss values ​​corresponding to the strategy network and the value network are both less than the target preset loss thresholds corresponding to the two networks. That is, the target first sub-loss value of the strategy network needs to be less than the target first preset loss threshold corresponding to the strategy network, and the target second sub-loss value of the value network needs to be less than the target second preset loss threshold corresponding to the value network.

[0232] When the target first sub-loss value of the policy network is not less than the target first preset loss threshold corresponding to the policy network, the policy network is determined as the target network, and the corresponding parameters of the policy network need to be adjusted according to the target first sub-loss value to implement the iteration of the policy network. Then return to execute and input the fusion features into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, and obtain the trained decision sub-model.

[0233] It should be noted that the above encoding sub-model and decision sub-model are synchronously trained and continuously iterated, and the trained encoding sub-model and the trained decision sub-model constitute the trained visual decision model. The trained visual decision model can accurately determine the target execution action corresponding to the visual task by combining the target image data and the target visual event data corresponding to the visual task.

[0234] In an embodiment of the present application, the visual decision model includes a coding sub-model and a decision sub-model. By acquiring image data and visual event data, the image data and the visual event data are input into the coding sub-model, and a fusion feature is output; the fusion feature is input into the decision sub-model, and a target action and an evaluation value corresponding to the target action are output; a first loss corresponding to the coding sub-model is determined according to the image data, the visual event data and the fusion feature; a second loss corresponding to the decision sub-model is determined according to the target action and the evaluation value; when the first loss does not meet a first preset loss condition, the model parameters corresponding to the coding sub-model are adjusted according to the first loss, and the image data and the visual event data are input into the coding sub-model to iteratively train the coding sub-model until the first loss meets the first preset loss condition, thereby obtaining a trained coding sub-model; when the second loss does not meet a second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the fusion feature is input into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining a trained decision sub-model; wherein the trained visual decision model includes a trained coding sub-model and a trained decision sub-model.

[0235] In this way, the encoding sub-model in the visual decision model can extract visual-related features from the image data and the visual event data, thereby obtaining fused features. The fused features have richer visual information than the image features extracted from the image data alone. Then, the fused features are input into the decision sub-model in the visual decision model, thereby obtaining the target action output by the decision sub-model. In the process of the visual decision model processing the image data and the visual event data, the first loss corresponding to the encoding sub-model can be determined according to the image data, the visual event data and the fused features, and the second loss corresponding to the decision sub-model can be determined according to the target action and the evaluation value. When the first loss does not meet the first preset loss condition, the model parameters corresponding to the encoding sub-model are adjusted according to the first loss, and the image data and the visual event data are input into the encoding sub-model to iteratively train the encoding sub-model until the first loss meets the first preset loss condition, thereby obtaining the trained encoding sub-model. When the second loss does not meet the second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the fused features are input into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining the trained decision sub-model. In the process of visual decision model training, the encoding sub-model and the decision sub-model are trained synchronously. Through continuous iteration of the encoding sub-model, the trained encoding sub-model can more accurately obtain the fusion features between the visual event data and the image data. Through continuous iteration of the decision sub-model, the decision sub-model can more accurately determine the target action to be performed based on the fusion features. Compared with the related art that only uses image data to obtain visual information, the visual decision model finally trained in this application can output the target fusion features that can more accurately represent the visual information of the environment based on the target image data and the target visual event data in the environment when performing the corresponding visual task, and accurately determine the target execution action corresponding to the visual task based on the target fusion features.

[0236] See also Figure 5 , Figure 5 1 is another flow chart of the visual decision model training method provided in an embodiment of the present application. The visual decision model training method may include the following steps:

[0237] Step 401, acquiring image data and visual event data, inputting the image data into a first encoder, and outputting image noise features;

[0238] Step 402: input the visual event data into a second encoder, and output event noise features;

[0239] Step 403: input the image data and the visual event data into a third encoder, and output a fusion feature;

[0240] Step 404: input the fused features into the policy network and output the target action;

[0241] Step 405: input the target action and the fusion feature into the value network, and output the evaluation value corresponding to the target action;

[0242] Step 406: Determine a first sub-loss value corresponding to the first encoder according to the image data, the fusion feature and the image noise feature;

[0243] Step 407: Determine a second sub-loss value corresponding to the second encoder according to the visual event data, the fusion feature and the event noise feature;

[0244] Step 408: Determine a third sub-loss value corresponding to the third encoder according to the image noise feature, the event noise feature and the fusion feature, where the first loss includes a first sub-loss value, a second sub-loss value and a third sub-loss value;

[0245] Step 409: when the first loss does not satisfy the first preset loss condition, adjust the model parameters corresponding to the encoding sub-model according to the first loss, and return to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining the trained encoding sub-model;

[0246] Step 410: determine the action value corresponding to the fusion feature and the target action, and determine the target first sub-loss value corresponding to the strategy network according to the action value;

[0247] Step 411: Determine the target second sub-loss value corresponding to the value network according to the action value and the evaluation value, where the second loss includes the target first sub-loss value and the target second sub-loss value;

[0248] Step 412: When the second loss does not meet the second preset loss condition, adjust the model parameters corresponding to the decision sub-model according to the second loss, return to execute and input the fusion feature into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, and obtain the trained decision sub-model, wherein the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model.

[0249] In the above embodiments, the description of each embodiment has its own focus. For the part that is not described in detail in a certain embodiment, please refer to the detailed description of the above-mentioned visual decision model training method, which will not be repeated here.

[0250] See also Figure 6 , Figure 6 : is a flow chart of a visual data processing method provided in an embodiment of the present application. The visual data processing method comprises the following steps:

[0251] Step 510: Obtain target image data and target visual event data corresponding to the visual task;

[0252] Step 520: input the target image data and the target visual event data into the trained encoding sub-model, and output the target fusion feature;

[0253] Step 530: input the target fusion feature into the trained decision sub-model, and output the target execution action corresponding to the visual task; wherein the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model, and the trained visual decision model is trained according to the visual decision model training method.

[0254] Steps 510 to 530 will be described in detail below.

[0255] In step 510, target image data and target visual event data corresponding to the visual task are obtained.

[0256] Among them, the target image data corresponding to the visual task can be obtained based on the shooting of a camera or a camera, and the target visual event data corresponding to the visual task can be obtained based on a dynamic visual sensor.

[0257] In some scenarios, in an autonomous driving scenario, the vehicle's camera captures target image data of the environment, and the vehicle's dynamic visual sensor captures target visual event data of the environment. Figure 7 As shown, Figure 7 A schematic diagram of images and visual events provided in the embodiments of the present application. Figure 7 The left side shows the target image data, including pixel points and pixel values, and the right side shows the target visual event data, including pixel position, trigger time, and event polarity.

[0258] In the autonomous driving scenario, the visual task can be to determine whether to stop the vehicle based on the visual information of the environment.

[0259] In step 520, the target image data and the target visual event data are input into the trained encoding sub-model, and the target fusion features are output.

[0260] The target image data and target visual event data are input into the trained encoding sub-model to output the target fusion feature. The target fusion feature is a feature that can characterize the visual information in the environment. Since it is generated based on the combination of the two modalities of image and event, the visual information contained in the target fusion feature is richer than the visual information contained in the single-modal image feature.

[0261] See also Figure 8 , Figure 8It is a schematic diagram of the trained visual decision model provided in an embodiment of the present application.

[0262] The trained encoding sub-model includes the trained third encoder, and the trained decision sub-model includes the trained strategy network and the trained value network. The target image data and the target visual event data are input into the trained third encoder, and the target fusion feature is output.

[0263] In step 530, the target fusion feature is input into the trained decision sub-model, and the target execution action corresponding to the visual task is output; wherein the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model, and the trained visual decision model is trained according to the visual decision model training method.

[0264] The target fusion features are input into the trained decision sub-model to output the target execution action corresponding to the visual task. Specifically, the target fusion features are input into the trained policy network to output the target execution action corresponding to the visual task. In the autonomous driving scenario, the target execution action can be to determine that there is a person on the road in front of the vehicle and perform operations such as braking and deceleration.

[0265] From the above, it can be seen that the trained visual decision model can accurately determine the target execution action corresponding to the visual task by combining the target image data and the target visual event data corresponding to the visual task.

[0266] See also Fig. 9 , Fig. 9 Schematic diagram of the structure of the visual decision model training device provided in the embodiment of the present application. The visual decision model training device is used to execute the above-mentioned visual decision model training method.

[0267] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0268] The visual decision model training device 600 includes:

[0269] A first input module 610 is used to obtain image data and visual event data, input the image data and visual event data into the encoding sub-model, and output fusion features;

[0270] The second input module 620 is used to input the fusion feature into the decision sub-model and output the target action and the evaluation value corresponding to the target action;

[0271] A first determination module 630, configured to determine a first loss corresponding to the encoding sub-model according to the image data, the visual event data and the fusion feature;

[0272] A second determination module 640, configured to determine a second loss corresponding to the decision sub-model according to the target action and the evaluation value;

[0273] A first iteration module 650 is used for adjusting the model parameters corresponding to the encoding sub-model according to the first loss when the first loss does not satisfy the first preset loss condition, and returning to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model;

[0274] A second iteration module 660 is used to adjust the model parameters corresponding to the decision sub-model according to the second loss when the second loss does not meet the second preset loss condition, and return to execute to input the fusion feature into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining a trained decision sub-model;

[0275] The trained visual decision model includes a trained encoding sub-model and a trained decision sub-model.

[0276] In some embodiments, the encoding sub-model includes a first encoder, a second encoder, and a third encoder; a first input module 610, configured to:

[0277] Inputting image data into a first encoder and outputting image noise characteristics;

[0278] Inputting visual event data into a second encoder and outputting event noise features;

[0279] The image data and visual event data are input into the third encoder, and the fused features are output.

[0280] In some embodiments, the first determination module 630 includes a first loss submodule, a second loss submodule, and a third loss submodule;

[0281] A first loss submodule, used to determine a first sub-loss value corresponding to the first encoder according to the image data, the fusion feature and the image noise feature;

[0282] A second loss submodule, used to determine a second sub-loss value corresponding to the second encoder according to the visual event data, the fusion feature and the event noise feature;

[0283] The third loss submodule is used to determine the third sub-loss value corresponding to the third encoder according to the image noise characteristics, event noise characteristics and fusion characteristics, and the first loss includes the first sub-loss value, the second sub-loss value and the third sub-loss value.

[0284] In some implementations, the first loss submodule is configured to:

[0285] Inputting the fusion features and the image noise features into a first pre-trained decoder, and outputting decoded image data;

[0286] Determine a first reconstruction loss value corresponding to the first encoder according to a difference between the decoded image data and the image data;

[0287] Determine a first similarity between the image noise feature and the fusion feature, and determine a first classification loss value corresponding to the first encoder according to the first similarity;

[0288] The first reconstruction loss value and the first classification loss value are added to obtain a first sub-loss value corresponding to the first encoder.

[0289] In some embodiments, the second loss submodule is used to:

[0290] Inputting the fused features and the event noise features into a second pre-trained decoder, and outputting decoded visual event data;

[0291] determining a second reconstruction loss value corresponding to the second encoder according to a difference between the decoded visual event data and the visual event data;

[0292] Determine a second similarity between the event noise feature and the fusion feature, and determine a second classification loss value corresponding to the second encoder according to the second similarity;

[0293] The second reconstruction loss value and the second classification loss value are added to obtain a second sub-loss value corresponding to the second encoder.

[0294] In some implementations, the third loss submodule is configured to:

[0295] Determine a first similarity between the event noise feature and the fusion feature, and determine a second similarity between the event noise feature and the fusion feature;

[0296] Determine a third classification loss value corresponding to a third encoder according to the first similarity and the second similarity;

[0297] A third sub-loss value corresponding to the third encoder is determined according to the third classification loss value.

[0298] In some implementations, the third loss submodule is configured to:

[0299] Input the fused features and the target action into the third pre-trained decoder, and output the decoded fused features at the next moment;

[0300] Determine a first prediction loss value corresponding to the third encoder according to a difference between the decoded fusion feature and the fusion feature at the next moment;

[0301] Input the fused features and the target action into the fourth pre-trained decoder, and output the decoding reward value at the next moment;

[0302] Determine a second prediction loss value corresponding to the third encoder according to a difference between the decoded reward value and the reward value corresponding to the target action;

[0303] The first prediction loss value, the second prediction loss value and the third classification loss value are added to obtain a third sub-loss value corresponding to the third encoder.

[0304] In some embodiments, the first iteration module 650 is configured to determine a sub-loss value that does not satisfy the first preset loss condition when any one of the first sub-loss value, the second sub-loss value, and the third sub-loss value does not satisfy the first preset loss condition;

[0305] Determining a target encoder corresponding to the sub-loss value among the first encoder, the second encoder, and the third encoder;

[0306] The model parameters corresponding to the target encoder are adjusted according to the sub-loss value, and the image data and the visual event data are input into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining the trained encoding sub-model, and the trained encoding sub-model includes the trained third encoder.

[0307] In some embodiments, the decision sub-model includes a policy network and a value network, and the second input module 620 is used to:

[0308] Input the fused features into the policy network and output the target action;

[0309] The target action and fusion features are input into the value network, and the evaluation value corresponding to the target action is output.

[0310] In some implementations, the second determination module 640 is configured to:

[0311] Determine the action value corresponding to the fusion feature and the target action;

[0312] Determine the target first sub-loss value corresponding to the strategy network according to the action value;

[0313] The target second sub-loss value corresponding to the value network is determined according to the action value and the evaluation value, and the second loss includes the target first sub-loss value and the target second sub-loss value.

[0314] In some embodiments, the second iteration module 660 is used to

[0315] When any one of the target first sub-loss value and the target second sub-loss value does not satisfy the second preset loss condition, determining the target sub-loss value that does not satisfy the second preset loss condition;

[0316] Determine the target network corresponding to the target sub-loss value in the policy network and the value network;

[0317] The model parameters corresponding to the target network are adjusted according to the target sub-loss value, and the fusion features are input into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining the trained decision sub-model.

[0318] In the above embodiments, the description of each embodiment has its own focus. For the part that is not described in detail in a certain embodiment, please refer to the detailed description of the above-mentioned visual decision model training method, which will not be repeated here.

[0319] See also Fig.10 , Fig.10 1 is a schematic diagram of the structure of a visual data processing device provided in an embodiment of the present application. The visual data processing device is used to execute the above-mentioned visual data processing method.

[0320] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0321] The visual data processing device 700 comprises:

[0322] An acquisition module 710 is used to acquire target image data and target visual event data corresponding to a visual task;

[0323] A fusion module 720, for inputting target image data and target visual event data into the trained encoding sub-model, and outputting target fusion features;

[0324] A decision module 730 is used to input the target fusion feature into the trained decision sub-model and output the target execution action corresponding to the visual task;

[0325] Among them, the trained visual decision model includes a trained encoding sub-model and a trained decision sub-model, and the trained visual decision model is trained according to the visual decision model training method provided in the embodiment of the present application.

[0326] The embodiment of the present application also provides a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned visual decision model training method or visual data processing method when executing the computer program. The computer device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0327] See also Fig.11 , Fig.11 The hardware structure of a computer device according to another embodiment is shown, and the computer device includes:

[0328] The processor 901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0329] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the visual decision model training method or visual data processing method of the embodiment of this application;

[0330] Input / output interface 903, used to implement information input and output;

[0331] Communication interface 904, used to realize communication interaction between the device and other devices, which can be realized through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0332] A bus 905 that transmits information between various components of the device (e.g., the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0333] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0334] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned visual decision model training method or visual data processing method.

[0335] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0336] In an embodiment of the present application, the visual decision model includes a coding sub-model and a decision sub-model. By acquiring image data and visual event data, the image data and the visual event data are input into the coding sub-model, and a fusion feature is output; the fusion feature is input into the decision sub-model, and a target action and an evaluation value corresponding to the target action are output; a first loss corresponding to the coding sub-model is determined according to the image data, the visual event data and the fusion feature; a second loss corresponding to the decision sub-model is determined according to the target action and the evaluation value; when the first loss does not meet a first preset loss condition, the model parameters corresponding to the coding sub-model are adjusted according to the first loss, and the image data and the visual event data are input into the coding sub-model to iteratively train the coding sub-model until the first loss meets the first preset loss condition, thereby obtaining a trained coding sub-model; when the second loss does not meet a second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the fusion feature is input into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining a trained decision sub-model; wherein the trained visual decision model includes a trained coding sub-model and a trained decision sub-model.

[0337] In this way, the encoding sub-model in the visual decision model can extract visual-related features from the image data and the visual event data, thereby obtaining fused features. The fused features have richer visual information than the image features extracted from the image data alone. Then, the fused features are input into the decision sub-model in the visual decision model, thereby obtaining the target action output by the decision sub-model. In the process of the visual decision model processing the image data and the visual event data, the first loss corresponding to the encoding sub-model can be determined according to the image data, the visual event data and the fused features, and the second loss corresponding to the decision sub-model can be determined according to the target action and the evaluation value. When the first loss does not meet the first preset loss condition, the model parameters corresponding to the encoding sub-model are adjusted according to the first loss, and the image data and the visual event data are input into the encoding sub-model to iteratively train the encoding sub-model until the first loss meets the first preset loss condition, thereby obtaining the trained encoding sub-model. When the second loss does not meet the second preset loss condition, the model parameters corresponding to the decision sub-model are adjusted according to the second loss, and the fused features are input into the decision sub-model to iteratively train the decision sub-model until the second loss meets the second preset loss condition, thereby obtaining the trained decision sub-model. In the process of visual decision model training, the encoding sub-model and the decision sub-model are trained synchronously. Through continuous iteration of the encoding sub-model, the trained encoding sub-model can more accurately obtain the fusion features between the visual event data and the image data. Through continuous iteration of the decision sub-model, the decision sub-model can more accurately determine the target action to be performed based on the fusion features. Compared with the related art that only uses image data to obtain visual information, the visual decision model finally trained in this application can output the target fusion features that can more accurately represent the visual information of the environment based on the target image data and the target visual event data in the environment when performing the corresponding visual task, and accurately determine the target execution action corresponding to the visual task based on the target fusion features.

[0338] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0339] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0340] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0341] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0342] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0343] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0344] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0345] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0346] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0347] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0348] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A visual decision model training method, characterized in that: The visual decision model includes an encoding sub-model and a decision sub-model, the encoding sub-model includes a first encoder, a second encoder and a third encoder, and the method includes: Acquire image data and visual event data, input the image data into the first encoder, and output image noise features; Inputting the visual event data into a second encoder and outputting event noise features; Inputting the image data and the visual event data into a third encoder and outputting fusion features; Inputting the fusion feature into the decision sub-model, and outputting the target action and the evaluation value corresponding to the target action; Determine a first loss corresponding to the encoding sub-model according to the image data, the visual event data and the fusion feature; Determine a second loss corresponding to the decision sub-model according to the target action and the evaluation value; When the first loss does not satisfy a first preset loss condition, adjusting a model parameter corresponding to the encoding sub-model according to the first loss, and returning to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model; When the second loss does not satisfy the second preset loss condition, adjusting the model parameters corresponding to the decision sub-model according to the second loss, returning to execute inputting the fusion feature into the decision sub-model, so as to iteratively train the decision sub-model until the second loss satisfies the second preset loss condition, thereby obtaining a trained decision sub-model; The trained visual decision model includes the trained encoding sub-model and the trained decision sub-model.

2. The visual decision model training method according to claim 1, characterized in that: The determining, according to the image data, the visual event data and the fusion feature, a first loss corresponding to the encoding sub-model comprises: Determine a first sub-loss value corresponding to the first encoder according to the image data, the fusion feature and the image noise feature; Determine a second sub-loss value corresponding to the second encoder according to the visual event data, the fusion feature and the event noise feature; A third sub-loss value corresponding to the third encoder is determined according to the image noise feature, the event noise feature and the fusion feature, and the first loss includes a first sub-loss value, a second sub-loss value and a third sub-loss value.

3. The visual decision model training method according to claim 2, characterized in that: The determining, according to the image data, the fusion feature and the image noise feature, a first sub-loss value corresponding to the first encoder comprises: Inputting the fusion feature and the image noise feature into a first pre-trained decoder, and outputting decoded image data; Determine a first reconstruction loss value corresponding to the first encoder according to a difference between the decoded image data and the image data; Determining a first similarity between the image noise feature and the fusion feature, and determining a first classification loss value corresponding to the first encoder according to the first similarity; The first reconstruction loss value and the first classification loss value are added to obtain a first sub-loss value corresponding to the first encoder.

4. The visual decision model training method according to claim 2, characterized in that: The determining, according to the visual event data, the fusion feature and the event noise feature, a second sub-loss value corresponding to the second encoder comprises: Inputting the fusion feature and the event noise feature into a second pre-trained decoder, and outputting decoded visual event data; determining a second reconstruction loss value corresponding to the second encoder according to a difference between the decoded visual event data and the visual event data; Determine a second similarity between the event noise feature and the fusion feature, and determine a second classification loss value corresponding to the second encoder according to the second similarity; The second reconstruction loss value and the second classification loss value are added to obtain a second sub-loss value corresponding to the second encoder.

5. The visual decision model training method according to claim 2, characterized in that: The determining, according to the image noise feature, the event noise feature and the fusion feature, a third sub-loss value corresponding to the third encoder comprises: Determine a first similarity between the event noise feature and the fusion feature, and determine a second similarity between the event noise feature and the fusion feature; Determine a third classification loss value corresponding to the third encoder according to the first similarity and the second similarity; A third sub-loss value corresponding to the third encoder is determined according to the third classification loss value.

6. The visual decision model training method according to claim 5, characterized in that: The determining, according to the third classification loss value, a third sub-loss value corresponding to the third encoder comprises: Inputting the fused features and the target action into a third pre-trained decoder, and outputting the decoded fused features at the next moment; Determine a first prediction loss value corresponding to the third encoder according to a difference between the decoded fusion feature and the fusion feature at the next moment; Inputting the fused feature and the target action into a fourth pre-trained decoder, and outputting a decoding reward value at the next moment; Determining a second prediction loss value corresponding to the third encoder according to a difference between the decoded reward value and the reward value corresponding to the target action; The first prediction loss value, the second prediction loss value and the third classification loss value are added to obtain a third sub-loss value corresponding to the third encoder.

7. The visual decision model training method according to claim 2, characterized in that: When the first loss does not satisfy a first preset loss condition, adjusting a model parameter corresponding to the encoding sub-model according to the first loss, and returning to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model, including: When any one of the first sub-loss value, the second sub-loss value, and the third sub-loss value does not satisfy a first preset loss condition, determining a sub-loss value that does not satisfy the first preset loss condition; Determining a target encoder corresponding to the sub-loss value among the first encoder, the second encoder, and the third encoder; The model parameters corresponding to the target encoder are adjusted according to the sub-loss value, and the image data and the visual event data are input into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model, wherein the trained encoding sub-model includes a trained third encoder.

8. The visual decision model training method according to claim 1, characterized in that: The decision sub-model includes a policy network and a value network, and the fusion feature is input into the decision sub-model to output a target action and an evaluation value corresponding to the target action, including: Input the fused features into the strategy network and output the target action; The target action and the fusion feature are input into the value network, and the evaluation value corresponding to the target action is output.

9. The visual decision model training method according to claim 8, characterized in that: The determining, according to the target action and the evaluation value, a second loss corresponding to the decision sub-model includes: Determining an action value corresponding to the fusion feature and the target action; Determine a target first sub-loss value corresponding to the strategy network according to the action value; A target second sub-loss value corresponding to the value network is determined according to the action value and the evaluation value, and the second loss includes the target first sub-loss value and the target second sub-loss value.

10. The visual decision model training method according to claim 8, characterized in that: When the second loss does not satisfy the second preset loss condition, adjusting the model parameters corresponding to the decision sub-model according to the second loss, returning to execute inputting the fusion feature into the decision sub-model to iteratively train the decision sub-model until the second loss satisfies the second preset loss condition, and obtaining the trained decision sub-model, including: When either the target first sub-loss value or the target second sub-loss value does not satisfy a second preset loss condition, determining a target sub-loss value that does not satisfy the second preset loss condition; Determining a target network corresponding to the target sub-loss value in the strategy network and the value network; The model parameters corresponding to the target network are adjusted according to the target sub-loss value, and the fusion feature is input into the decision sub-model to perform iterative training on the decision sub-model until the second loss satisfies the second preset loss condition, thereby obtaining the trained decision sub-model.

11. A visual decision model training device, characterized in that: The visual decision model includes a coding sub-model and a decision sub-model, the coding sub-model includes a first encoder, a second encoder and a third encoder, and the device includes: A first input module, used for acquiring image data and visual event data, inputting the image data into the first encoder, and outputting image noise features; Inputting the visual event data into a second encoder and outputting event noise features; Inputting the image data and the visual event data into a third encoder and outputting fusion features; A second input module is used to input the fusion feature into the decision sub-model, and output the target action and the evaluation value corresponding to the target action; A first determination module, configured to determine a first loss corresponding to the encoding sub-model according to the image data, the visual event data and the fusion feature; A second determination module, used to determine a second loss corresponding to the decision sub-model according to the target action and the evaluation value; a first iteration module, configured to adjust a model parameter corresponding to the encoding sub-model according to the first loss when the first loss does not satisfy a first preset loss condition, and return to execute inputting the image data and the visual event data into the encoding sub-model to iteratively train the encoding sub-model until the first loss satisfies the first preset loss condition, thereby obtaining a trained encoding sub-model; a second iteration module, configured to adjust the model parameters corresponding to the decision sub-model according to the second loss when the second loss does not satisfy the second preset loss condition, and return to execute inputting the fusion feature into the decision sub-model to iteratively train the decision sub-model until the second loss satisfies the second preset loss condition, thereby obtaining a trained decision sub-model; The trained visual decision model includes the trained encoding sub-model and the trained decision sub-model.

12. A method for processing visual data, characterized in that: include: Obtain target image data and target visual event data corresponding to the visual task; Inputting the target image data and the target visual event data into the trained encoding sub-model, and outputting target fusion features; Inputting the target fusion feature into the trained decision sub-model, and outputting the target execution action corresponding to the visual task; Among them, the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model, and the trained visual decision model is trained according to the visual decision model training method according to any one of claims 1-10.

13. A visual data processing device, characterized in that: include: An acquisition module is used to acquire target image data and target visual event data corresponding to a visual task; A fusion module, used for inputting the target image data and the target visual event data into the trained encoding sub-model, and outputting target fusion features; A decision module, used for inputting the target fusion feature into the trained decision sub-model, and outputting the target execution action corresponding to the visual task; Among them, the trained visual decision model includes the trained encoding sub-model and the trained decision sub-model, and the trained visual decision model is trained according to the visual decision model training method according to any one of claims 1-10.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, which are suitable for loading by a processor to execute the visual decision model training method described in any one of claims 1 to 10 or the visual data processing method described in claim 12.

15. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the visual decision model training method described in any one of claims 1 to 10 or the visual data processing method described in claim 12 is implemented.

Citation Information

Patent Citations

  • Visual model pre-training method and device, electronic equipment and storage medium

    CN117994617A

  • Routing decision model training method and device, computer equipment and storage medium

    CN118784550A