Training method, reasoning method and related device of stream matching generative model

Through adaptive loss functions and back-propagation algorithms, the flow matching generation model can adaptively select action states at different motion stages during the learning process, solving the problem of insufficient accuracy in action sequence prediction in existing technologies and achieving higher prediction accuracy and consistency.

CN120597950AActive Publication Date: 2025-09-05BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

Patent Information

Application Number
CN202510771202.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Existing generative strategies find it difficult to take into account the motion information of both low-dynamic motion stages and high-dynamic motion stages during the learning process, resulting in insufficient accuracy in motion sequence prediction.

Method used

Adaptive loss function and back-propagation algorithm are used to adaptively constrain the difference between the predicted velocity field vector and the action, adaptively learn the action state of different motion stages, and improve the accuracy of predicted action sequences.

Benefits of technology

The flow matching generation model's ability to learn motion states at different motion stages is improved, thereby improving the prediction accuracy and consistency of motion sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597950A_ABST
    Figure CN120597950A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a training method, a reasoning method and a related device for a stream matching generation model, which are used for improving the accuracy of an action sequence predicted by the trained stream matching generation model. The method provided by the embodiment of the invention comprises the following steps: acquiring noise, a moment t and an environment characteristic of a first action, wherein the environment characteristic of the first action at least comprises an observation value of the first action; inputting the noise, the moment t and the environment characteristics into an initialized flow matching generation model to obtain an output predicted velocity field vector of the conditional probability path at the moment t + 1; calculating the loss between the predicted velocity field vector and the real velocity field vector by using a preset loss function, wherein the preset loss function comprises at least one of a first loss function and a second loss function, and a third loss function; and training the initialized stream matching generation model by using a loss and back propagation algorithm until the stream matching generation model converges to obtain a trained stream matching generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data generation, and in particular to a training method, an inference method and related devices for a stream matching generation model. Background Art

[0002] In the field of embodied intelligence, existing generative policies (visuomotor policies or vision-language-action policies) all follow the imitation learning paradigm. Within this paradigm, policies primarily predict continuous actions based on observations. The core action generation module of these generative policies is primarily adapted from algorithms in the field of image generation.

[0003] Action sequences differ significantly from ordinary images in that high-frequency information is more important. Therefore, they are often used in conjunction with low-pass filters, which either retain only high-frequency information or retain all-frequency information. However, in action sequences, the frequency distribution within each action block changes dynamically as the task progresses.

[0004] The root of this change is that the robot operation sequence is usually composed of alternating static and non-static motion stages. The alternating motion stages generally include low-dynamic motion stages and high-dynamic motion stages. Among them, only some action dimensions in the low-dynamic motion stage show significant changes, and the remaining dimensions remain relatively smooth; while in the high-dynamic motion stage, the high-frequency changes are more significant and information-rich. However, during the learning process, the existing generative strategies either choose actions in the low-dynamic motion stage or the high-dynamic motion stage, and it is difficult to take into account all action information. Summary of the Invention

[0005] Embodiments of the present invention provide a training method, an inference method, and related devices for a flow matching generation model, which are used to adaptively select action states of different motion stages for learning during the training process of the flow matching generation model, thereby improving the accuracy of the action sequence predicted by the trained flow matching generation.

[0006] A first aspect of an embodiment of the present application provides a method for training a flow matching generation model, the method comprising:

[0007] Acquire noise, time t, and environmental characteristics of the first action, wherein the environmental characteristics of the first action at least include an observation value of the first action;

[0008] Inputting the noise, the time t, and the environmental characteristics into an initialized flow matching generation model to obtain a predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model;

[0009] Calculating the loss between the predicted velocity field vector and the true velocity field vector using a preset loss function, wherein the preset loss function includes at least one of a first loss function and a second loss function, and a third loss function, wherein the first loss function is used to adaptively constrain the difference between the predicted velocity field vectors at any two moments between the moment t and the moment t+1 to be minimized, the second loss function is used to adaptively constrain the predicted actions of actions starting at different moments between the moment t and the moment t+1 at the future moment u1 to converge to the same position, and the third loss function is used to constrain the loss between the predicted velocity field and the true velocity field to be minimized;

[0010] The initialized flow matching generation model is trained using the loss and back propagation algorithm until the flow matching generation model converges to obtain a trained flow matching generation model.

[0011] As an optional embodiment, the method further includes:

[0012] Determine the s1 moment and the r1 moment among the two arbitrary moments;

[0013] Obtaining the first predicted velocity field vector at the time s1;

[0014] Obtaining the second predicted velocity field vector at the time r1;

[0015] The first loss function is specifically used for:

[0016] According to the difference between the first predicted velocity field vector and the second predicted velocity field vector, an adaptive weighting strategy is adopted to calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector.

[0017] As an optional embodiment, the method further includes:

[0018] Determine any time s1 and r1 among the different time periods;

[0019] Obtaining a first predicted action when transferring from the time s1 to the time u1;

[0020] Obtaining a second predicted action when transitioning from the time r1 to the time u1;

[0021] The second loss function is specifically used for:

[0022] According to the difference between the first predicted action and the second predicted action, an adaptive weighting strategy is adopted to calculate the loss between the first predicted action and the second predicted action.

[0023] As an optional embodiment, after obtaining the predicted velocity field, the method further includes:

[0024] Projecting the predicted velocity field vector into a frequency domain space to obtain a frequency spectrum coefficient of the predicted velocity field vector;

[0025] The first loss function includes a fourth loss function, and the fourth loss function is used to:

[0026] Adaptively constraining the difference in spectral coefficients between the predicted velocity field vectors at any two moments between the moment t and the moment t+1 to be minimized according to the difference in frequency domain space between the predicted velocity field vectors at any two moments between the moment t and the moment t+1;

[0027] The second loss function includes a fifth loss function, and the fifth loss function is used to:

[0028] According to the difference in frequency domain space between the predicted actions at the future time u1 of the actions started at different times between the time t and the time t+1, the consistency of the spectral coefficients of the predicted actions at the future time u1 of the actions started at the different times is adaptively constrained.

[0029] As an optional embodiment, the method further includes:

[0030] Determine the s1 moment and the r1 moment among the two arbitrary moments;

[0031] Obtaining the first predicted velocity field vector at the time s1;

[0032] Obtaining the second predicted velocity field vector at the time r1;

[0033] Converting the first predicted velocity field vector from the time domain to the frequency domain to obtain a first frequency spectrum coefficient at time s1;

[0034] Converting the second predicted velocity field vector from the time domain to the frequency domain to obtain a second frequency spectrum coefficient at time r1;

[0035] The fourth loss function is specifically used for:

[0036] According to the difference between the first spectrum coefficient and the second spectrum coefficient, an adaptive weighting strategy is adopted to calculate the loss between the first spectrum coefficient and the second spectrum coefficient.

[0037] As an optional embodiment, the method further includes:

[0038] Determine any time s1 and r1 among the different time periods;

[0039] Obtaining a first predicted action from the time s1 to the time u1;

[0040] Obtaining a second predicted action when transferring from the time r to the time u1;

[0041] Converting the first prediction action from the time domain to the frequency domain to obtain a third spectrum coefficient of the first prediction action when transferring from the time s1 to the time u1;

[0042] Converting the second prediction action from the time domain to the frequency domain to obtain a fourth spectrum coefficient of the second prediction action when transferring from the time r1 to the time u1;

[0043] The fifth loss function is specifically used for:

[0044] According to the difference between the third spectrum coefficient and the fourth spectrum coefficient, an adaptive weighting strategy is adopted to calculate the loss between the third spectrum coefficient and the fourth spectrum coefficient.

[0045] As an optional embodiment, the initialized flow matching generation model includes an initialized flow matching one-step generation model.

[0046] A second aspect of the embodiments of the present application provides an inference method based on a flow matching generation model, the method comprising:

[0047] Acquire noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action at least include an observation value of the first action;

[0048] The noise, the time t and the environmental characteristics are input into the trained flow matching generation model as described in any one of claims 1 to 7 to obtain the second action predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.

[0049] As an optional embodiment, if the trained flow matching generation model includes a trained flow matching one-step generation model;

[0050] Inputting the noise, the time t, and the environmental characteristics of the first action into the trained flow matching generation model to obtain a second action obtained by a conditional probability path predicted by the trained flow matching generation model using a preset velocity field vector includes:

[0051] The noise, the time t, and the environmental characteristics of the first action are input into the trained flow matching one-step generation model to obtain the second action generated by the trained flow matching one-step generation model using the conditional probability path predicted by the preset velocity field vector in one step.

[0052] A third aspect of an embodiment of the present application provides a training device for a flow matching generation model, the device comprising:

[0053] an acquiring unit, configured to acquire noise, time t, and environmental characteristics of the first action, wherein the environmental characteristics of the first action at least include an observation value of the action;

[0054] An input unit, configured to input the noise, the time t, and the environmental characteristics into an initialized flow matching generation model to obtain a predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model;

[0055] a calculation unit, configured to calculate a loss between the predicted velocity field vector and the actual velocity field vector using a preset loss function, wherein the preset loss function includes at least one of a first loss function and a second loss function, and a third loss function, wherein the first loss function is configured to adaptively constrain a difference between the predicted velocity field vectors at any two moments between the moment t and the moment t+1 to be minimized, the second loss function is configured to adaptively constrain actions starting at different moments between the moment t and the moment t+1 to converge to a same position at the future moment u1 based on a difference between predicted actions at the future moment u1 of actions starting at different moments between the moment t and the moment t+1, and the third loss function is configured to constrain the loss between the predicted velocity field and the actual velocity field to be minimized;

[0056] A training unit is used to train the initialized flow matching generation model using the loss and back propagation algorithm until the flow matching generation model converges to obtain a trained flow matching generation model.

[0057] A fourth aspect of the embodiments of the present application provides an inference device based on a flow matching generation model, the device comprising:

[0058] an acquiring unit, configured to acquire noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action at least include an observation value of the first action;

[0059] A prediction unit is used to input the noise, the time t and the environmental characteristics into the trained flow matching generation model provided in the first aspect of the embodiment of the present application to obtain the second action predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.

[0060] A fifth aspect of an embodiment of the present application provides a computer device, comprising a processor, which, when executing a computer program stored in a memory, is used to implement the training method of the flow matching generation model provided in the first aspect of the embodiment of the present application, or the reasoning method based on the flow matching generation model provided in the second aspect of the embodiment of the present application.

[0061] The sixth aspect of the embodiments of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is used to implement the training method of the flow matching generation model provided in the first aspect of the embodiments of the present application, or the reasoning method based on the flow matching generation model provided in the second aspect of the embodiments of the present application.

[0062] The seventh aspect of the embodiments of the present application provides a computer program product having a computer program stored thereon. When the computer program is executed by a processor, it is used to implement the training method of the flow matching generation model provided in the first aspect of the embodiments of the present application, or the reasoning method based on the flow matching generation model provided in the second aspect of the embodiments of the present application.

[0063] It can be seen from the above technical solutions that the embodiments of the present invention have the following advantages:

[0064] In an embodiment of the present application, during the training process of the initialized flow matching generation model, not only the existing third loss function is used to calculate the loss between the predicted velocity field and the actual velocity field, but also the first loss function is used to adaptively constrain the difference between the predicted velocity field vectors at any two moments between time t and time t+1 to minimize the difference between the predicted velocity field vectors at any two moments, and / or, the second loss function is used to adaptively constrain the modal objects starting from different moments between time t and time t+1 to converge to the same position at the future time u1 based on the difference between the predicted actions at different moments between time t and time t+1. In this way, the initialized flow matching generation model can adaptively learn the action states of different motion stages during the backward denoising training process, thereby improving the accuracy of the action sequence predicted by the initialized flow matching generation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a schematic diagram of an embodiment of a method for training a flow matching generation model in an embodiment of the present application;

[0066] Figure 2 Schematic diagram of the motion state values ​​at different motion stages in the embodiment of the present application;

[0067] Figure 3 This is a schematic diagram of an embodiment of the process of calculating loss using the first loss function in an embodiment of the present application;

[0068] Figure 4 This is a schematic diagram of an embodiment of the process of calculating loss using the second loss function in an embodiment of the present application;

[0069] Figure 5 This is a schematic diagram of another embodiment of a method for training a flow matching generation model in an embodiment of the present application;

[0070] Figure 6 This is a schematic diagram of an embodiment of the process of calculating loss using the fourth loss function in an embodiment of the present application;

[0071] Figure 7 This is a schematic diagram of an embodiment of the process of calculating loss using the fifth loss function in an embodiment of the present application;

[0072] Figure 8 Schematic diagram of an embodiment of a reasoning method based on a flow matching generation model in an embodiment of the present application;

[0073] Figure 9 This is a schematic diagram of an embodiment of a training device for a flow matching generation model in an embodiment of the present application;

[0074] Figure 10 This is a schematic diagram of an embodiment of an inference device for a flow matching generation model in an embodiment of the present application. DETAILED DESCRIPTION

[0075] Embodiments of the present invention provide a training method, an inference method, and related devices for a flow matching generation model, which are used to adaptively select action states of different motion stages for learning during the training process of the flow matching generation model, thereby improving the accuracy of the action sequence predicted by the trained flow matching generation.

[0076] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0077] The terms "first," "second," "third," "fourth," and the like in the specification and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0078] For ease of understanding, the following describes the training method of the flow matching generation model in this application. Figure 1 , an embodiment of a method for training a flow matching generation model in an embodiment of the present application includes:

[0079] 101. Obtain noise, time t, and environmental characteristics of the first action, wherein the environmental characteristics of the first action at least include an observation value of the first action;

[0080] In existing generative learning strategies, especially the flow matching generative model, during the backward reasoning training process, either actions in the low-dynamic motion stage or the high-dynamic motion stage are selected for learning, making it difficult to take into account all action information.

[0081] To address this issue, the present application adaptively learns actions at different motion stages to improve the accuracy of the actions predicted in the prediction stage.

[0082] Specifically, the training process of the flow matching generative model generally includes a forward noise addition process and a backward reasoning process. The forward noise addition process is used to gradually (i.e., multi-step) add noise to the real modal object until the real modal object is evolved into a noise object that satisfies the Gaussian distribution. The backward reasoning process is to sample from a noise object that satisfies the Gaussian distribution, and then gradually denoise the sampled object to obtain the real modal object.

[0083] The forward denoising process in this application is similar to the forward denoising process in the prior art and will not be repeated here. However, what is different from the prior art is that the backward denoising process in this application adopts a loss function that is different from the prior art to improve the accuracy of the predicted objects generated by the flow matching generation model after training in this application.

[0084] Furthermore, in the backward reasoning training process of the initialized flow matching generation model, it is necessary to first obtain the noise, time t, and the environmental characteristics of the first action collected from the noise. The time t here can be regarded as the sampling moment, and the two adjacent moments are regarded as one-step denoising, such as from time t to time t+1, it is regarded as the first step denoising, and from time t+1 to time t+2, it is regarded as the second step denoising. The noise here refers to a noise object that obeys the Gaussian distribution, and the environmental characteristics of the first action generally include the action observation value at time t, and the environmental characteristics of the action observation value. The environmental characteristics here are generally related to the predicted action of the initialized flow matching generation model. For example, the environmental characteristics can be a process video of the robot performing the action, or a process picture of the robot performing the action, etc. As long as the environmental characteristics can help the robot improve the accuracy of the predicted action, they are within the scope of protection of this application.

[0085] 102. Input the noise, the time t, and the environmental characteristics into an initialized flow matching generation model to obtain a predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model;

[0086] After obtaining the environmental characteristics of the noise, time t, and the first action, the noise, time t, and the environmental characteristics of the first action are input into the initialized flow matching generation model to obtain the predicted velocity field of the conditional probability path at time t+1 output by the initialized flow matching generation model.

[0087] Among them, the predicted velocity field here refers to the predicted movement speed (including the predicted velocity magnitude and velocity direction) of the conditional probability path from time t to time t+1. Because in the training process of backward denoising of the flow matching generation model, it is often expected that the flow matching generation model can be restored from the noise object to the predicted action in a straight line motion. Therefore, in this application, it is hoped that the predicted velocity field is a constant, because only when the predicted velocity field is a constant can the conditional probability path be explained as a straight line.

[0088] 103. Calculate the loss between the predicted velocity field vector and the true velocity field vector using a preset loss function, wherein the preset loss function includes at least one of a first loss function and a second loss function, and a third loss function; the first loss function is used to adaptively constrain the difference between the predicted velocity field vectors at any two moments between the time t and the time t+1 to be minimized; the second loss function is used to adaptively constrain actions initiated at different moments between the time t and the time t+1 to converge to the same position at the future time u1 based on the difference between the predicted actions at the future time u1 of the actions initiated at different moments; and the third loss function is used to constrain the loss between the predicted velocity field and the true velocity field to be minimized;

[0089] During the backward denoising training of the flow matching generative model, after obtaining the predicted velocity field vector output by the flow matching generative model, the loss between the predicted velocity field vector and the true velocity field vector is calculated, where the true velocity field is the difference between the first action at time t and the predicted second action at time t+1.

[0090] What is different from the prior art is that in addition to using a third loss function similar to the prior art to constrain the loss between the predicted velocity field vector and the actual velocity field vector to be minimized, the present application also sets at least one of the first loss function and the second loss function to constrain the actual velocity field, wherein the first loss function is used to adaptively constrain the difference between the predicted velocity field vectors at any two moments between time t and the said time t+1 to be minimized, and the second loss function is used to adaptively constrain the predicted actions of actions starting from different moments between time t and time t+1 at the future time u1 to converge to the same position based on the difference between the predicted actions of actions starting from different moments between time t and time t+1 at the future time u1.

[0091] For ease of understanding, Figure 2 The action state values ​​of different motion stages are given, among which action state 6 belongs to the high dynamic motion stage, while action states 0-5 belong to the low dynamic motion stage. By setting at least one of the first loss function and the second loss function, the present application can enable the initialized flow matching generation model to not only adaptively learn the action states of different dynamic motion stages, but also improve the accuracy of the action sequence predicted by the trained flow matching generation model.

[0092] Because the first loss function in this application, when constraining the predicted velocity field vectors at any two moments, is based on the difference between the predicted velocity field vectors at any two moments, and adaptively constrains the difference between the predicted velocity field vectors at any two moments to be minimized, that is, the user does not need to manually add filters to select the action states at different moments (the different moments here generally include high-dynamic motion stages and low-dynamic motion stages) according to the motion stage, thereby improving the accuracy of the predicted action sequence, and the second loss function is based on the difference between the predicted actions of actions starting from different moments between time t and time t+1 at the future time u1, and adaptively constrains the predicted actions starting from the different moments at the future time u1 to converge to the same position, and does not require the user to add filters to select different action states according to the motion stage, further improving the accuracy of the predicted action sequence.

[0093] In the embodiments of the present application, the process of how to adaptively constrain will be described in the following embodiments and will not be repeated here.

[0094] 104. Utilize the loss and back-propagation algorithm to train the initialized flow matching generation model until the flow matching generation model converges, so as to obtain a trained flow matching generation model.

[0095] After obtaining the loss calculated by the first loss function, and / or the loss calculated by the second loss function, and the loss calculated by the third loss function in step 103, the initialized flow matching generation model is trained using the loss and the back propagation algorithm until the flow matching generation model converges to obtain the trained flow matching generation model.

[0096] In an embodiment of the present application, during the training process of the initialized flow matching generation model, not only the existing third loss function is used to calculate the loss between the predicted velocity field and the actual velocity field, but also the first loss function is used to adaptively constrain the difference between the predicted velocity field vectors at any two moments between time t and time t+1 to minimize the difference between the predicted velocity field vectors at any two moments, and / or, the second loss function is used to adaptively constrain the modal objects starting from different moments between time t and time t+1 to converge to the same position at the future time u1 based on the difference between the predicted actions at different moments between time t and time t+1. In this way, the initialized flow matching generation model can adaptively learn the action states of different motion stages during the backward denoising training process, thereby improving the accuracy of the action sequence predicted by the initialized flow matching generation model.

[0097] based on Figure 1The following describes the process of calculating the loss using the first loss function. Figure 3 :

[0098] 301. Determine time s1 and time r1 among the two arbitrary time moments;

[0099] Specifically, this application needs to determine any two moments between time t and time t+1, assuming that time s1 and time r1 are any two moments between time t and time t+1.

[0100] 302. Obtain the first predicted velocity field vector at the time s1;

[0101] After determining the time s1, the first predicted velocity field vector at the time s1 is obtained. When obtaining the first predicted velocity field vector at the time s, the action state generated at the time s, the time s and the environmental characteristics of the action state at the time s are input into the initialized flow matching generation model, and the first predicted velocity field vector at the time s output by the flow matching generation model can be obtained.

[0102] 303. Obtain the second predicted velocity field vector at the time r1;

[0103] Similarly, the present application can also obtain the second predicted velocity field vector at time r1.

[0104] 304. Calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector using an adaptive weighting strategy according to the difference between the first predicted velocity field vector and the second predicted velocity field vector.

[0105] After obtaining the first predicted velocity vector field vector and the second predicted velocity field vector, the present application uses an adaptive weighting strategy to calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector based on the difference between the first predicted velocity vector field vector and the second predicted velocity field vector. When calculating the loss between the first predicted velocity field vector and the second predicted velocity field vector, the L1 loss function or the L2 loss function can be used for calculation, and no specific limitation is imposed herein.

[0106] The following formulas 1 and 2 give examples of the first loss function, where the first loss function calculates the expected value of the L2 loss between the first predicted velocity field vector and the second predicted velocity field vector.

[0107]

[0108] In formula 1, r and s represent two arbitrary moments collected from the normalized moments, and a rrepresents the action collected from the noise D at time r, a s represents the action collected from the noise D at time s, and v θ (s,a s ) represents the first predicted velocity field vector at time s, v θ (r,a r ) represents the second predicted velocity field vector at time r, and W k The weight coefficient representing the loss between the first predicted velocity field vector of each dimension and the second predicted velocity field vector of each dimension.

[0109] In Formula 2, j belongs to the dimension of the predicted velocity field vector. Assuming that the predicted velocity field vector is a vector with one row and five columns, the value of j is 5.

[0110] In the embodiment of the present application, the calculation process of the first loss function is described in detail, and the first loss function designs the predicted velocity fields at any two moments s1 and r1 between time t and time t+1, so that according to the difference between the predicted velocity fields at any two moments, the loss between the predicted velocity field vectors at any two moments is adaptively calculated, so that the initialized flow matching generation model can learn the predicted velocity field vectors of the action states in different motion stages, thereby improving the accuracy of the action sequence predicted by the trained flow matching generation model.

[0111] The following describes the process of calculating the loss using the second loss function in the embodiment of the present application. Figure 4 :

[0112] 401. Determine any time s1 and r1 among different time periods;

[0113] In the embodiment of the present application, it is necessary to determine different moments between moment t and moment t+1. For the convenience of description, the embodiment of the present application also determines moment s1 and moment r1 among the different moments between moment t and moment t+1.

[0114] 402. Obtain the first predicted action when transferring from time s1 to time u1;

[0115] After determining the s1 moment, the u1 moment that lags behind the s1 moment is determined between the t moment and the t+1 moment. In order to obtain the first predicted action generated by the initialized flow matching generation model when the initialized flow matching generation model is transferred from the s1 moment to the u1 moment, the present application inputs the modal object generated at the s1 moment, the s1 moment, and the environmental features of the action at the s1 moment into the initialized flow matching generation model, and the first predicted action at the s1 moment output by the flow matching generation model can be obtained.

[0116] 403. Obtain the second predicted action when transferring from time r1 to time u1;

[0117] Similar to step 302 , the present application can also obtain the second predicted action when transferring from time r1 to time u1 .

[0118] 404. Calculate the loss between the first predicted action and the second predicted action using an adaptive weighting strategy according to the difference between the first predicted action and the second predicted action.

[0119] In the above steps, after obtaining the first predicted action and the second predicted action, the loss between the first predicted action and the second predicted action is further calculated using an adaptive weighted strategy based on the difference between the first predicted action and the second predicted action. When calculating the loss, the L1 loss function or the L2 loss function can also be used. There is no specific restriction on the type of the second loss function for calculating the loss.

[0120] For ease of understanding, examples of the second loss function are given below using Formula 3 and Formula 4, where the second loss function calculates the expected value of the L2 loss between the first predicted action and the second predicted action.

[0121]

[0122] In formula 3, r1, s1, and u1 represent three random moments collected from the normalized moments, and a r1 represents the action collected from the noise D at time r1, a s1 represents the action collected from the noise D at time s1, a u1 represents the action collected from the noise D at time u1, and v θ (s1,a s1 ) represents the first predicted velocity field vector at time s1, v θ (s1,a r1 ) represents the second predicted velocity field vector at time r1, and (a s1 +(u1-s1)v θ (s1,a s1 )) represents the first predicted action vector, and (a r1 +(u1-r1)v θ (s1,a r1 )) represents the second predicted motion vector.

[0123] In Formula 4, j belongs to the dimension of the predicted velocity field vector. Assuming that the predicted velocity field vector is a vector with one row and five columns, the value of j is 5.

[0124] In an embodiment of the present application, the calculation process of the second loss function is described in detail, and the second loss function uses an adaptive weighting strategy to calculate the loss between the first predicted action vector and the second predicted action vector based on the difference between the first predicted action vector and the second predicted action vector from any moment between time t and time t+1 to the future time u1, so that the initialized flow matching generation model can learn the action states of different motion stages, thereby improving the accuracy of the action sequence predicted by the trained flow matching generation model.

[0125] In the above embodiment, the characteristics of the action in the frequency domain are not taken into account when predicting the action. To address this problem, the embodiment of the present application can also perform the following steps to adaptively constrain the first predicted velocity field vector and the second predicted velocity field vector in the frequency domain. Figure 5 :

[0126] 501. Obtain noise, time t, and environmental characteristics of the first action, wherein the environmental characteristics of the first action at least include an observation value of the action;

[0127] 502. Input the noise, the time t, and the environmental characteristics into the initialized flow matching generation model to obtain a predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model;

[0128] Steps 501 to 502 in the embodiment of the present application are the same as Figure 1 The descriptions of steps 101 to 102 in the embodiment are similar and will not be repeated here.

[0129] 503. Project the predicted velocity field vector into the frequency domain to obtain a frequency spectrum coefficient of the predicted velocity field vector.

[0130] In order to improve the timing consistency of the predicted action, the embodiment of the present application further projects the predicted velocity field into the frequency domain space, that is, converts the predicted velocity field generated in step 502 from the time domain space to the frequency domain space. In the conversion process, the Fourier transform formula or the discrete cosine transform DCT can be used. There is no specific limitation on the process of converting from the time domain to the frequency domain.

[0131] For ease of understanding, this application uses discrete cosine transform (DCT) as an example to describe the process of converting the predicted velocity field from the time domain to the frequency domain using formula 5:

[0132]

[0133] In Formula 5, k represents the dimension of the predicted velocity field vector.

[0134] 504. Calculate the loss between the predicted velocity field vector and the true velocity field vector using a preset loss function, wherein the preset loss function includes at least one of a fourth loss function and a fifth loss function, and the third loss function; the fourth loss function is used to adaptively constrain the difference in spectral coefficients between the predicted velocity field vectors at any two moments between time t and time t+1 to be minimized based on the difference in frequency domain space between the predicted velocity field vectors at any two moments between time t and time t+1; the fifth loss function is used to adaptively constrain the consistency of spectral coefficients of predicted actions at the future time u1 for actions starting at different moments between time t and time t+1 to be consistent based on the difference in frequency domain space between predicted actions at the future time u1 for actions starting at different moments between time t and time t+1; and the third loss function is used to constrain the loss between the predicted velocity field and the true velocity field to be minimized;

[0135] After projecting the predicted velocity field vector into the frequency domain space, the role of the third loss function in this application remains unchanged, and is still used to constrain the loss between the predicted velocity field and the actual velocity field to be minimized, while the first loss function correspondingly evolves into the fourth loss function in this application, and the second loss function correspondingly evolves into the fifth loss function in this application, wherein the fourth loss function is used to adaptively constrain the difference in spectral coefficients between the predicted velocity field vectors at any two moments between time t and time t+1 to be minimized based on the difference in frequency domain space between the predicted velocity field vectors at any two moments between time t and time t+1, and the fifth loss function is used to adaptively constrain the consistency of spectral coefficients of the predicted action at the future time u1 of the action starting from different moments between time t and time t+1 based on the difference in frequency domain space between the predicted action at the future time u1 of the action starting from different moments.

[0136] As for the calculation process of the fourth loss function and the fifth loss function, it will be described in the following embodiments and will not be repeated here.

[0137] 505. The initialized flow matching generation model is trained using the loss and back propagation algorithm until the flow matching generation model converges to obtain a trained flow matching generation model.

[0138] In step 504, after the loss is calculated according to the fourth loss function, the fifth loss function and the third loss function, the initialized flow matching generation model is trained using the loss and the back propagation algorithm until the flow matching generation model converges to obtain the trained flow matching generation model.

[0139] When calculating the loss using the fourth loss function in the embodiment of the present application, the velocity field vectors of any two moments are converted from the time domain to the frequency domain to obtain the spectral coefficients corresponding to the velocity field vectors of any two moments. Then, based on the difference between the spectral coefficients at any two moments, the difference between the spectral coefficients between the predicted velocity field vectors at any two moments is adaptively constrained to be minimized, so that the initialized flow matching generation model can learn the spectral coefficients of the predicted velocity field of the action state at different motion stages. When calculating the loss, the fifth loss function in the present application converts the predicted action of the action starting from different moments at the future u1 moment from the time domain to the frequency domain, and based on the difference in the spectral coefficients converted from the time domain to the spectral coefficients of the predicted action of the action starting from different moments at the future u1 moment, the consistency of the spectral coefficients of the predicted action of the action starting from different moments at the future u1 moment is adaptively constrained, so that the initialized flow matching generation model can learn the spectral coefficients of the action state at different motion stages, thereby further improving the temporal consistency of the generated predicted action sequence on the basis of improving the accuracy of the action sequence predicted by the trained flow matching generation model.

[0140] based on Figure 5 The following describes the process of calculating the loss using the fourth loss function. Figure 6 :

[0141] 601. Determine time s1 and time r1 among the two arbitrary time moments;

[0142] 602. Obtain the first predicted velocity field vector at the time s1;

[0143] 603. Obtain the second predicted velocity field vector at the time r1;

[0144] It should be noted that steps 601 to 603 in the embodiment of the present application are the same as Figure 3 The descriptions of steps 301 to 303 in the embodiment are similar and will not be repeated here.

[0145] 604. Convert the first predicted velocity field vector from the time domain to the frequency domain to obtain a first frequency spectrum coefficient at time s1.

[0146] In order to further improve the temporal continuity of the predicted action and to improve the temporal consistency of the generated action sequence, the embodiment of the present application converts the first predicted velocity field vector from the time domain to the frequency domain to obtain the first spectrum coefficient at time s1. The conversion process of the first spectrum coefficient can refer to Formula 5. For the convenience of description, the embodiment of the present application converts the first spectrum coefficient at time s1 to F(v s1 ).

[0147] 605. Convert the second predicted velocity field vector from the time domain to the frequency domain to obtain a second frequency spectrum coefficient at time r1;

[0148] Similarly, the embodiment of the present application also converts the second predicted velocity field vector from the time domain to the frequency domain to obtain the second spectrum coefficient at time r. For the convenience of description, the second spectrum coefficient at time r1 is also recorded as F(v r1 ).

[0149] 606. Calculate the loss between the first spectrum coefficient and the second spectrum coefficient using an adaptive weighting strategy according to the difference between the first spectrum coefficient and the second spectrum coefficient.

[0150] After obtaining the first spectrum coefficient and the second spectrum coefficient, the embodiment of the present application further adopts an adaptive weighting strategy to calculate the loss between the first spectrum coefficient and the second spectrum coefficient according to the difference between the first spectrum coefficient and the second spectrum coefficient.

[0151] For ease of understanding, the following describes the process of calculating the loss between the first spectrum coefficient and the second spectrum coefficient using the adaptive weighting strategy, taking Formula 6 and Formula 7 as examples:

[0152]

[0153] Sim1(v r1 ,v s1 )=Sim1(v θ (s1,a s1 ),v θ (r1,a r1 ))=W k ·||F(v r1 )-F(v s1 )|| 2 (Formula 7)

[0154] In Formula 7, W k The calculation formula of can be found in Formula 4, which will not be repeated here.

[0155] In an embodiment of the present application, the calculation process of the fourth loss function is described in detail, and the fourth loss function calculates the spectral coefficients of the predicted velocity field at any two moments s and t between time t and time t+1, and adaptively constrains the loss of the spectral coefficients of the predicted velocity field at any two moments s and t to be minimized based on the difference between the spectral coefficients of the predicted velocity field at any two moments s and t, so that the initialized flow matching generation model can adaptively learn the frequency domain information of the velocity field vector of the predicted action in different motion stages, thereby further improving the temporal consistency of the generated predicted action sequence on the basis of improving the accuracy of the action sequence predicted by the trained flow matching generation model.

[0156] based on Figure 5 The following describes the process of calculating the loss using the fifth loss function. Figure 7 :

[0157] 701. Determine any time s1 and r1 among the different time periods;

[0158] 702. Obtain a first predicted action from the time s1 to the time u1;

[0159] 703. Obtain a second predicted action when transferring from the time r to the time u1;

[0160] It should be noted that steps 701 to 703 in the embodiment of the present application are the same as Figure 4 The descriptions of steps 401 to 403 in the embodiment are similar and will not be repeated here.

[0161] 704. Convert the first prediction action from the time domain to the frequency domain to obtain a third spectrum coefficient of the first prediction action when transferring from time s1 to time u1.

[0162] Here, the process of converting the first prediction action from the time domain to the frequency domain can refer to Formula 5. For the convenience of description, the third spectrum coefficient when transferring from time s1 to time u1 is recorded as F(v(a s1 +(u1-s1)v θ (s1,a s1 ))), where a s1 Indicates the action at time s1, a s1 +(u1-s1)v θ (s1,a s1 ) represents the predicted action at time u1 when the trigger is triggered from time s1 to time u1.

[0163] 705. Convert the second prediction action from the time domain to the frequency domain to obtain a fourth frequency spectrum coefficient of the second prediction action when transferring from time r1 to time u1.

[0164] Similarly, in the embodiment of the present application, the second prediction action can also be converted from the time domain to the frequency domain to obtain the fourth spectrum coefficient of the second prediction action when transferring from the time r1 to the time u1. For the convenience of description, the third spectrum coefficient when transferring from the time r1 to the time u1 is recorded as F(v(a r1 +(u1-r1)v θ (r1,a r1 ))), where a r1 represents the action at time r1, a r1 +(u1-r1)vθ (r1,a r1 ) When the trigger is triggered from time r1 to time u1, the predicted action at time u1.

[0165] 706. Calculate the loss between the third spectrum coefficient and the fourth spectrum coefficient using an adaptive weighting strategy according to the difference between the third spectrum coefficient and the fourth spectrum coefficient.

[0166] After obtaining the third spectrum coefficient and the fourth spectrum coefficient, the loss between the third spectrum coefficient and the fourth spectrum coefficient is further calculated using an adaptive weighting strategy based on the third spectrum coefficient and the fourth spectrum coefficient, wherein the loss calculation can be an L1 function or an L2 function, and there is no restriction on the loss calculation process here.

[0167] For ease of understanding, the following describes the process of calculating the loss between the third spectrum coefficient and the fourth spectrum coefficient using Formula 8 and Formula 9 as examples:

[0168]

[0169] Sim1(v(a s1 +(u1-s1)v θ (s1,a s1 ),v(a r1 +(u1-r1)v θ (r1,a r1 ))=

[0170] W k ·||F(v(a s1 +(u1-s1)v θ (s1,a s1 ))-F(v(a r1 +(u1-r1)v θ (r1,a r1 ))|| 2 (Formula 9)

[0171] W k The calculation process of can refer to Formula 4, which will not be repeated here.

[0172] In an embodiment of the present application, the calculation process of the fifth loss function is described in detail, and the fifth loss function adaptively constrains the consistency of the third spectral coefficient and the fourth spectral coefficient from any time between time t and time t+1 to the future time u1 based on the difference between the spectral coefficients of the first predicted action and the second predicted action from any time between time t and time t+1 to the future time u1, thereby ensuring that the initialized flow matching generation model can adaptively learn the frequency domain information of the action state in different motion stages, thereby further improving the temporal consistency of the generated predicted action sequence on the basis of improving the accuracy of the action sequence predicted by the trained flow matching generation model.

[0173] based on Figures 1 to 7 The described embodiment adaptively learns the action states of different motion stages (such as low-dynamic motion stage and high-dynamic motion stage) during the backward denoising training process, while the initialized flow matching generation model in the embodiment of the present application generally adopts a multi-step method for training during the training process. In order to further improve the training speed of the initialized flow matching generation model, the present application can also improve the initialized flow matching generation model to an initialized flow matching one-step generation model, so that in the backward denoising training process, the real action state can be restored from the noise in one step, thereby further improving the efficiency of obtaining the real action state.

[0174] The above describes the training process of the flow matching generation model in the embodiment of the present application. The following describes the reasoning process of the flow matching generation model in the embodiment of the present application. Figure 8 An embodiment of the inference method of the flow matching generation model in the embodiment of the present application includes:

[0175] 801. Obtain noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action at least include an observation value of the first action;

[0176] Specifically, step 801 here is similar to step 101 and will not be repeated here.

[0177] 802. Input the noise, the time t, and the environmental characteristics into the trained flow matching generation model to obtain a second action predicted by the trained flow matching generation model based on a conditional probability path predicted by a preset velocity field vector.

[0178] After obtaining the noise, time t, and environmental characteristics of the first action in step 801, the noise, time t, and environmental characteristics of the first action are input into the flow matching generation model trained according to the above method embodiment to obtain the second action predicted by the conditional probability path predicted by the trained flow matching generation model using the preset velocity field vector.

[0179] In an embodiment of the present application, after obtaining the trained flow matching generation model, the noise, time t, and environmental characteristics of the first action are input into the trained flow matching generation model, and then the second action predicted by the flow matching generation model according to the conditional probability path predicted by the preset predicted velocity field vector can be obtained. Furthermore, because the flow matching generation model initialized in the above training process adaptively learns the action states of different motion stages, the trained flow matching generation model in the embodiment of the present application can improve the accuracy of the predicted action sequence.

[0180] based on Figure 8 In the example described, when the flow matching generation model is a flow matching one-step generation model, after the noise, time t, and environmental characteristics of the first action are input into the trained flow matching one-step generation model, the second action predicted by the trained flow matching one-step generation model can be obtained.

[0181] An embodiment of the present application also provides a computer program product having a computer program stored thereon. When the computer program is executed by a processor, it is used to implement the training method of the flow matching generation model provided in the above method embodiment of the present application, or to implement the reasoning method based on the flow matching generation model provided in the above method embodiment of the present application.

[0182] The above describes in detail the training process of the flow matching generation model in the embodiment of the present application and the reasoning process based on the flow matching generation model. Next, the training device of the flow matching generation model in the embodiment of the present application is described. Figure 9 An embodiment of a training device for a flow matching generation model in an embodiment of the present application includes:

[0183] An acquiring unit 901 is configured to acquire noise, time t, and environmental characteristics of the first action, wherein the environmental characteristics of the first action at least include an observation value of the action;

[0184] An input unit 902 is configured to input the noise, the time t, and the environmental characteristics into an initialized flow matching generation model to obtain a predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model;

[0185] a calculation unit 903 configured to calculate the loss between the predicted velocity field vector and the actual velocity field vector using a preset loss function, wherein the preset loss function includes at least one of a first loss function and a second loss function, and a third loss function; the first loss function is configured to adaptively constrain the difference between the predicted velocity field vectors at any two moments between the time t and the time t+1 to be minimized based on the difference between the predicted velocity field vectors at the two moments; the second loss function is configured to adaptively constrain the predicted actions of actions starting at different moments between the time t and the time t+1 at the future time u1 to converge to the same position based on the difference between the predicted actions of actions starting at different moments between the time t and the time t+1 at the future time u1; and the third loss function is configured to constrain the loss between the predicted velocity field and the actual velocity field to be minimized;

[0186] The training unit 904 is configured to train the initialized flow matching generation model using the loss and back propagation algorithm until the flow matching generation model converges, so as to obtain a trained flow matching generation model.

[0187] As an optional embodiment, the acquiring unit 901 is further configured to:

[0188] Determine the s1 moment and the r1 moment among the two arbitrary moments;

[0189] Obtaining the first predicted velocity field vector at the time s1;

[0190] Obtaining the second predicted velocity field vector at the time r1;

[0191] The calculation unit 903 is specifically used for:

[0192] According to the difference between the first predicted velocity field vector and the second predicted velocity field vector, an adaptive weighting strategy is adopted to calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector.

[0193] As an optional embodiment, the acquiring unit 901 is further configured to:

[0194] Determine any time s1 and r1 among the different time periods;

[0195] Obtaining a first predicted action when transferring from the time s1 to the time u1;

[0196] Obtaining a second predicted action when transitioning from the time r1 to the time u1;

[0197] The calculation unit 903 is specifically used for:

[0198] According to the difference between the first predicted action and the second predicted action, an adaptive weighting strategy is adopted to calculate the loss between the first predicted action and the second predicted action.

[0199] As an optional embodiment, the calculating unit 903 is further configured to:

[0200] Projecting the predicted velocity field vector into a frequency domain space to obtain a frequency spectrum coefficient of the predicted velocity field vector;

[0201] The first loss function includes a fourth loss function, and the fourth loss function is used to:

[0202] Adaptively constraining the difference in spectral coefficients between the predicted velocity field vectors at any two moments between the moment t and the moment t+1 to be minimized according to the difference in frequency domain space between the predicted velocity field vectors at any two moments between the moment t and the moment t+1;

[0203] The second loss function includes a fifth loss function, and the fifth loss function is used to:

[0204] According to the difference in frequency domain space between the predicted actions at the future time u1 of the actions started at different times between the time t and the time t+1, the consistency of the spectral coefficients of the predicted actions at the future time u1 of the actions started at the different times is adaptively constrained.

[0205] As an optional embodiment, the acquiring unit 901 is further configured to:

[0206] Determine the s1 moment and the r1 moment among the two arbitrary moments;

[0207] Obtaining the first predicted velocity field vector at the time s1;

[0208] Obtaining the second predicted velocity field vector at the time r1;

[0209] The calculation unit 903 is further configured to:

[0210] Converting the first predicted velocity field vector from the time domain to the frequency domain to obtain a first frequency spectrum coefficient at time s1;

[0211] Converting the second predicted velocity field vector from the time domain to the frequency domain to obtain a second frequency spectrum coefficient at time r1;

[0212] According to the difference between the first spectrum coefficient and the second spectrum coefficient, an adaptive weighting strategy is adopted to calculate the loss between the first spectrum coefficient and the second spectrum coefficient.

[0213] As an optional embodiment, the acquiring unit 901 is further configured to:

[0214] Determine any time s1 and r1 among the different time periods;

[0215] Obtaining a first predicted action from the time s1 to the time u1;

[0216] Obtaining a second predicted action when transferring from the time r to the time u1;

[0217] The calculation unit 903 is further configured to:

[0218] Converting the first prediction action from the time domain to the frequency domain to obtain a third spectrum coefficient of the first prediction action when transferring from the time s1 to the time u1;

[0219] Converting the second prediction action from the time domain to the frequency domain to obtain a fourth spectrum coefficient of the second prediction action when transferring from the time r1 to the time u1;

[0220] According to the difference between the third spectrum coefficient and the fourth spectrum coefficient, an adaptive weighting strategy is adopted to calculate the loss between the third spectrum coefficient and the fourth spectrum coefficient.

[0221] As an optional embodiment, the initialized flow matching generation model includes an initialized flow matching one-step generation model.

[0222] It should be noted that the functions of the above-mentioned units are similar to those described in the corresponding method embodiments and will not be repeated here.

[0223] In an embodiment of the present application, during the training process of the initialized flow matching generation model, not only is the loss between the predicted velocity field and the actual velocity field calculated by the calculation unit 903 using the existing third loss function, but the first loss function is also used to adaptively constrain the difference between the predicted velocity field vectors at any two moments between time t and time t+1 to be minimized based on the difference between the predicted velocity field vectors at any two moments, and / or, the second loss function is used to adaptively constrain the modal objects starting from different moments between time t and time t+1 to converge to the same position at the future time u1 based on the difference between the predicted actions at different moments between time t and time t+1. In this way, the initialized flow matching generation model can adaptively learn the action states of different motion stages during the backward denoising training process, thereby improving the accuracy of the action sequence predicted by the initialized flow matching generation model.

[0224] Next, the inference device based on the flow matching generation model in the embodiment of the present application is described. Figure 10 In one embodiment of the present application, an inference device based on a flow matching generation model includes:

[0225] An acquisition unit 1001 is configured to acquire noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action include at least an observation value of the first action;

[0226] The prediction unit 1002 is used to input the noise, the time t and the environmental characteristics into the trained flow matching generation model to obtain the second action predicted by the trained flow matching generation model based on the conditional probability path predicted by the preset velocity field vector.

[0227] As an optional embodiment, if the trained flow matching generation model includes a trained flow matching one-step generation model;

[0228] The prediction unit 1002 is specifically configured to:

[0229] The noise, the time t, and the environmental characteristics of the first action are input into the trained flow matching one-step generation model to obtain the second action generated by the trained flow matching one-step generation model using the conditional probability path predicted by the preset velocity field vector in one step.

[0230] It should be noted that the functions of the above-mentioned units are similar to those described in the corresponding method embodiments and will not be repeated here.

[0231] In an embodiment of the present application, after obtaining the trained flow matching generation model, the noise, time t, and environmental characteristics of the first action are input into the trained flow matching generation model. Then, the second action predicted by the flow matching generation model according to the conditional probability path predicted by the preset predicted velocity field vector can be obtained through the prediction unit 1002. Furthermore, because the flow matching generation model initialized in the above training process adaptively learns the action states of different motion stages, the trained flow matching generation model in the embodiment of the present application can improve the accuracy of the predicted action sequence.

[0232] The above describes the embodiment of the present invention from the perspective of modular functional entities. The following describes the computer device in the embodiment of the present invention from the perspective of hardware processing:

[0233] The computer device is used to implement the functions of the training device of the flow matching generation model. In one embodiment of the present invention, the computer device includes:

[0234] processor and memory;

[0235] When the memory is used to store computer programs and the processor is used to execute the computer programs stored in the memory, the various steps in the embodiment of the training method for the above-mentioned flow matching generation model can be implemented.

[0236] The computer device is also used to implement the functions of the inference device based on the flow matching generation model. Another embodiment of the computer device in the embodiment of the present invention includes:

[0237] processor and memory;

[0238] The memory is used to store computer programs, and when the processor is used to execute the computer programs stored in the memory, each step in the above-mentioned embodiment of the reasoning method based on the flow matching generation model can be implemented.

[0239] It can be understood that when the processor in the computer device described above executes the computer program, it can also realize the functions of the various units in the corresponding device embodiments described above, which will not be repeated here. Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program in the training device of the flow matching generation model / the reasoning device based on the flow matching generation model. For example, the computer program can be divided into the various units in the training device of the flow matching generation model described above, and each unit can realize the specific functions described in the training device of the corresponding flow matching generation model described above.

[0240] The computer device may be a computing device such as a desktop computer, laptop, PDA, or cloud server. The computer device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that a processor and memory are merely examples of computer devices and do not constitute a limitation of the computer device. The computer device may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, and the like.

[0241] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, connecting various parts of the entire computer device using various interfaces and lines.

[0242] The memory can be used to store the computer programs and / or modules, and the processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0243] The present invention also provides a computer-readable storage medium, which is used to implement the functions of the training device side of the flow matching generation model. A computer program is stored thereon. When the computer program is executed by a processor, the processor can be used to implement each step in the above-mentioned flow matching generation model training method embodiment.

[0244] The present invention also provides another computer-readable storage medium, which is used to implement the functions of an inference device based on a flow matching generation model. A computer program is stored thereon. When the computer program is executed by a processor, the processor can be used to implement each step in an embodiment of an inference method based on a flow matching generation model.

[0245] It is understood that if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned corresponding embodiment methods, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0246] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0247] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0248] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0249] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for training a flow matching generation model, characterized in that: The method comprises: Acquire noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action at least include an observation value of the first action; Inputting the noise, the time t, and the environmental characteristics into an initialized flow matching generation model to obtain a predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model; Calculating the loss between the predicted velocity field vector and the true velocity field vector using a preset loss function, wherein the preset loss function includes at least one of a first loss function and a second loss function, and a third loss function, wherein the first loss function is used to adaptively constrain the difference between the predicted velocity field vectors at any two moments between the moment t and the moment t+1 to be minimized, the second loss function is used to adaptively constrain the predicted actions of actions starting at different moments between the moment t and the moment t+1 at the future moment u1 to converge to the same position, and the third loss function is used to constrain the loss between the predicted velocity field and the true velocity field to be minimized; The initialized flow matching generation model is trained using the loss and back propagation algorithm until the flow matching generation model converges to obtain a trained flow matching generation model.

2. The training method according to claim 1, characterized in that The method further comprises: Determine the s1 moment and the r1 moment among the two arbitrary moments; Obtaining the first predicted velocity field vector at the time s1; Obtaining the second predicted velocity field vector at the time r1; The first loss function is specifically used for: According to the difference between the first predicted velocity field vector and the second predicted velocity field vector, an adaptive weighting strategy is adopted to calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector.

3. The training method according to claim 1, characterized in that The method further comprises: Determine any time s1 and r1 among the different time periods; Obtaining a first predicted action when transferring from the time s1 to the time u1; Obtaining a second predicted action when transitioning from the time r1 to the time u1; The second loss function is specifically used for: According to the difference between the first predicted action and the second predicted action, an adaptive weighting strategy is adopted to calculate the loss between the first predicted action and the second predicted action.

4. The training method according to claim 1, characterized in that After obtaining the predicted velocity field, the method further includes: Projecting the predicted velocity field vector into a frequency domain space to obtain a frequency spectrum coefficient of the predicted velocity field vector; The first loss function includes a fourth loss function, and the fourth loss function is used to: Adaptively constraining the difference in spectral coefficients between the predicted velocity field vectors at any two moments between the moment t and the moment t+1 to be minimized according to the difference in frequency domain space between the predicted velocity field vectors at any two moments between the moment t and the moment t+1; The second loss function includes a fifth loss function, and the fifth loss function is used to: According to the difference in frequency domain space between the predicted actions at the future time u1 of the actions started at different times between the time t and the time t+1, the consistency of the spectral coefficients of the predicted actions at the future time u1 of the actions started at the different times is adaptively constrained.

5. The method according to claim 4, characterized in that The method further comprises: Determine the s1 moment and the r1 moment among the two arbitrary moments; Obtaining the first predicted velocity field vector at the time s1; Obtaining the second predicted velocity field vector at the time r1; Converting the first predicted velocity field vector from the time domain to the frequency domain to obtain a first frequency spectrum coefficient at time s1; Converting the second predicted velocity field vector from the time domain to the frequency domain to obtain a second frequency spectrum coefficient at time r1; The fourth loss function is specifically used for: According to the difference between the first spectrum coefficient and the second spectrum coefficient, an adaptive weighting strategy is adopted to calculate the loss between the first spectrum coefficient and the second spectrum coefficient.

6. The method according to claim 4, characterized in that The method further comprises: Determine any time s1 and r1 among the different time periods; Obtaining a first predicted action from the time s1 to the time u1; Obtaining a second predicted action when transferring from the time r to the time u1; Converting the first prediction action from the time domain to the frequency domain to obtain a third spectrum coefficient of the first prediction action when transferring from the time s1 to the time u1; Converting the second prediction action from the time domain to the frequency domain to obtain a fourth spectrum coefficient of the second prediction action when transferring from the time r1 to the time u1; The fifth loss function is specifically used for: According to the difference between the third spectrum coefficient and the fourth spectrum coefficient, an adaptive weighting strategy is adopted to calculate the loss between the third spectrum coefficient and the fourth spectrum coefficient.

7. The method according to claim 1, characterized in that The initialized flow matching generation model includes an initialized flow matching one-step generation model.

8. A reasoning method based on a flow matching generative model, characterized in that: The method comprises: Acquire noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action at least include an observation value of the first action; The noise, the time t and the environmental characteristics are input into the trained flow matching generation model as described in any one of claims 1 to 7 to obtain the second action predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.

9. The inference method according to claim 8, characterized in that: If the trained stream matching generation model includes a trained stream matching one-step generation model; Inputting the noise, the time t, and the environmental characteristics of the first action into the trained flow matching generation model to obtain a second action obtained by a conditional probability path predicted by the trained flow matching generation model using a preset velocity field vector includes: The noise, the time t, and the environmental characteristics of the first action are input into the trained flow matching one-step generation model to obtain the second action generated by the trained flow matching one-step generation model using the conditional probability path predicted by the preset velocity field vector in one step.

10. A training device for a flow matching generation model, characterized in that: The device comprises: an acquiring unit, configured to acquire noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action at least include an observation value of the first action; An input unit, configured to input the noise, the time t, and the environmental characteristics into an initialized flow matching generation model to obtain a predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model; a calculation unit, configured to calculate a loss between the predicted velocity field vector and the actual velocity field vector using a preset loss function, wherein the preset loss function includes at least one of a first loss function and a second loss function, and a third loss function, wherein the first loss function is configured to adaptively constrain a difference between the predicted velocity field vectors at any two moments between the moment t and the moment t+1 to be minimized, the second loss function is configured to adaptively constrain actions starting at different moments between the moment t and the moment t+1 to converge to a same position at the future moment u1 based on a difference between predicted actions at the future moment u1 of actions starting at different moments between the moment t and the moment t+1, and the third loss function is configured to constrain the loss between the predicted velocity field and the actual velocity field to be minimized; A training unit is used to train the initialized flow matching generation model using the loss and back propagation algorithm until the flow matching generation model converges to obtain a trained flow matching generation model.

11. An inference device based on a flow matching generation model, characterized in that: The device comprises: an acquiring unit, configured to acquire noise, time t, and environmental characteristics of a first action, wherein the environmental characteristics of the first action at least include an observation value of the first action; A prediction unit, configured to input the noise, the time t, and the environmental characteristics into the trained flow matching generation model according to any one of claims 1 to 7, so as to obtain a second action predicted by the conditional probability path predicted by the trained flow matching generation model using a preset velocity field vector.

12. A computer device comprising a processor, characterized in that: When executing the computer program stored in the memory, the processor is used to implement the training method of the flow matching generation model according to any one of claims 1 to 7, or the reasoning method based on the flow matching generation model according to any one of claims 8 to 9.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the training method of the flow matching generation model according to any one of claims 1 to 7, or the reasoning method based on the flow matching generation model according to any one of claims 8 to 9.

14. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the training method of the flow matching generation model according to any one of claims 1 to 7, or the reasoning method based on the flow matching generation model according to any one of claims 8 to 9.

Citation Information

Patent Citations

  • Speed field prediction method and device, model training method and electronic equipment

    CN117452526A

  • Generative stream model training method, action prediction method and device

    CN117540784A

  • Scene detection model training method and device, electronic equipment and storage medium

    CN118365992A

  • Molecular generation method and device based on flow matching model

    CN118800362A

  • System and Method for Estimating a Future Traffic Density in an Environment

    US20240304081A1

Cited By

  • Behavior risk perception detection method based on stream matching, electronic equipment and medium

    CN122067321A

  • Robot control method and robot

    CN122323191A