A training method and an inference method of a flow matching generation model and related devices
By introducing multiple loss function constraints in the training of the flow matching generation model, the inefficiency caused by path curves during training is solved, achieving a more efficient training process and ensuring that the model can quickly recover to the real modal object during the backward denoising process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
- Filing Date
- 2025-06-10
- Publication Date
- 2026-07-24
Smart Images

Figure CN120597951B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal data generation technology, and in particular to a training method, inference method, and related apparatus for a stream matching generation model. Background Technology
[0002] Currently, generative modeling-based visual motion strategy frameworks have extended the powerful capabilities of generative models in text-to-image synthesis to the field of robot operation imitation learning, achieving significant breakthroughs. The mainstream approach is based on diffusion models, which have the advantage of modeling complex multimodal distributions in high-dimensional action spaces. Recently, flow matching models have emerged as a generalization of diffusion models. By constructing probabilistic paths from simple prior distributions to complex data distributions to learn vector fields, they offer simpler optimization objectives and a more stable and efficient training process, gradually becoming the dominant generative model.
[0003] In the existing flow matching generation model, the conditional probability path from noise to real modal object recovery during the back-inference training process is often curved, resulting in a long training time and low training efficiency. Summary of the Invention
[0004] This invention provides a training method, an inference method, and related apparatus for a stream matching generation model, which improves the training efficiency of the stream matching generation model training process.
[0005] The first aspect of this application provides a method for training a stream matching generation model, comprising:
[0006] Acquire noise, time t, a first modal object collected from the noise, and environmental characteristics of the first modal object;
[0007] The noise, the time t, the first modal object, and the environmental features are input into the initialized flow matching generation model to obtain the predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model.
[0008] The loss between the predicted velocity field vector and the true velocity field vector is calculated using a preset loss function. The preset loss function includes at least one of a first loss function and a second loss function, and a third loss function. The first loss function is used to constrain the minimization of the difference between the predicted velocity field vectors at any two times between time t and time t+1. The second loss function is used to constrain the convergence of modal objects starting from different times between time t and time t+1 to the same position at time u in the future. The third loss function is used to constrain the minimization of the loss between the predicted velocity field and the true velocity field.
[0009] The initial flow matching generation model is trained using the loss and backpropagation algorithm until it converges, thus obtaining the trained flow matching generation model.
[0010] As an optional embodiment, the method further includes:
[0011] Obtain the first predicted velocity field vector at time s among any two time points;
[0012] Obtain the second predicted velocity field vector at time r from any two time points;
[0013] The first loss function is specifically used for:
[0014] Calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector.
[0015] As an optional embodiment, the method further includes:
[0016] Determine any two times s and r among the different times;
[0017] Obtain the first prediction mode object from time s to time u;
[0018] Obtain the second prediction mode object from time r to time u;
[0019] The second loss function is specifically used for:
[0020] Calculate the loss between the first prediction modality and the second prediction modality.
[0021] As an optional embodiment, after obtaining the predicted velocity field, the method further includes:
[0022] The predicted velocity field vector is projected into the frequency domain to obtain the spectral coefficients of the predicted velocity field vector.
[0023] The first loss function includes a fourth loss function, which is used for:
[0024] The constraint is to minimize the difference in spectral coefficients between the velocity field vectors at any two time points;
[0025] The second loss function includes a fifth loss function, which is used for:
[0026] Constrain the consistency of the spectral coefficients of modal objects starting from different times between time t and time t+1 in the future time u.
[0027] As an optional embodiment, the method further includes:
[0028] Obtain the first predicted velocity field vector at time s among any two time points;
[0029] Obtain the second predicted velocity field vector at time r from any two time points;
[0030] The first predicted velocity field vector is converted from the time domain to the frequency domain to obtain the first spectral coefficients at time s;
[0031] The second predicted velocity field vector is converted from the time domain to the frequency domain to obtain the second spectral coefficients at time r;
[0032] The fourth loss function is specifically used for:
[0033] Calculate the loss between the first spectral coefficient and the second spectral coefficient.
[0034] As an optional embodiment, the method further includes:
[0035] Determine any two times s and r among the different times;
[0036] Obtain the first prediction mode object that transitions from time s to time u;
[0037] Obtain the second prediction mode object when transitioning from time r to time u;
[0038] The first predicted mode object is converted from the time domain to the frequency domain to obtain the third spectral coefficients from time s to time u;
[0039] The second predicted mode object is converted from the time domain to the frequency domain to obtain the fourth spectral coefficients from time r to time u;
[0040] The fifth loss function is specifically used for:
[0041] Calculate the loss between the third spectral coefficient and the fourth spectral coefficient.
[0042] As an optional embodiment, the initialized stream matching generation model includes an initialized stream matching one-step generation model; the first modal object includes text, image, voice, or robot action state.
[0043] A second aspect of this application provides a reasoning method based on a stream matching generation model, the method comprising:
[0044] Acquire noise, time t, a first modal object collected from the noise, and environmental characteristics of the first modal object;
[0045] The noise, the time t, the first modal object, and the environmental features are input into the trained flow matching generation model provided in the first aspect of the present application to obtain the second modal object predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.
[0046] As an optional embodiment, the trained flow matching generation model may include a trained one-step flow matching generation model;
[0047] The step of inputting the noise, the time t, the first modal object, and the environmental features into the trained flow matching generation model to obtain the second modal object obtained by the trained flow matching generation model using the conditional probability path predicted by the preset velocity field vector includes:
[0048] The noise, the time t, the first modal object, and the environmental features are input into the trained flow matching one-step generation model to obtain the second modal object generated by the trained flow matching one-step generation model in one step by the conditional probability path predicted by the preset velocity field vector.
[0049] A third aspect of this application provides a training apparatus for a stream matching generation model, the apparatus comprising:
[0050] The acquisition unit is used to acquire noise, time t, a first modal object collected from the noise, and environmental features of the first modal object;
[0051] The input unit is used to input the noise, the time t, the first modal object and the environmental features into the initialized flow matching generation model to obtain the predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model.
[0052] The calculation unit is used to calculate the loss between the predicted velocity field vector and the true velocity field vector using a preset loss function. The preset loss function includes at least one of a first loss function and a second loss function, and a third loss function. The first loss function is used to constrain the minimization of the difference between the predicted velocity field vectors at any two times between time t and time t+1. The second loss function is used to constrain the convergence of modal objects starting from different times between time t and time t+1 to the same position at a future time u. The third loss function is used to constrain the minimization of the loss between the predicted velocity field and the true velocity field.
[0053] The training unit is used to train the initialized flow matching generation model using the loss and backpropagation algorithm until the flow matching generation model converges, so as to obtain the trained flow matching generation model.
[0054] A fourth aspect of this application provides an inference apparatus based on a stream matching generation model, the apparatus comprising:
[0055] The acquisition unit is used to acquire noise, time t, a first modal object collected from the noise, and environmental features of the first modal object;
[0056] The prediction unit is used to input the noise, the time t, the first modal object and the environmental features into the trained flow matching generation model to obtain the second modal object predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.
[0057] A fifth aspect of this application provides a computer device including a processor. When the processor executes a computer program stored in a memory, it is used to implement a training method for a stream matching generation model provided in the first aspect of this application, or an inference method based on a stream matching generation model provided in the second aspect of this application.
[0058] A sixth aspect of this application provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program is used to implement the training method of the stream matching generation model provided in the first aspect of this application, or the inference method based on the stream matching generation model provided in the second aspect of this application.
[0059] A seventh aspect of this application provides a computer program product having a computer program stored thereon. When the computer program is executed by a processor, it is used to implement the training method of the stream matching generation model provided in the first aspect of this application, or the inference method based on the stream matching generation model provided in the second aspect of this application.
[0060] As can be seen from the above technical solutions, the embodiments of the present invention have the following advantages:
[0061] In this embodiment, during the training of the initialized flow matching generation model, not only is the loss between the predicted velocity field and the real velocity field calculated using the existing third loss function, but also the difference between the predicted velocity field vectors at any two times between time t and t+1 is minimized using the first loss function, and / or the modal objects starting from different times between time t and t+1 converge to the same position at time u in the future. This enables the initialized flow matching generation model to recover the real modal objects with the shortest path during the backward denoising training process, thereby improving the training efficiency of the initialized flow matching generation model. Attached Figure Description
[0062] Figure 1 This is a schematic diagram of an embodiment of the training method for the flow matching generation model in this application.
[0063] Figure 2 This is a schematic diagram illustrating the first and second loss functions in the embodiments of this application;
[0064] Figure 3 This is a schematic diagram of an embodiment of calculating loss using a first loss function in this application.
[0065] Figure 4 This is a schematic diagram of an embodiment of calculating loss using a second loss function in this application.
[0066] Figure 5 This is a schematic diagram of another embodiment of the training method for the flow matching generation model in this application;
[0067] Figure 6 This is a schematic diagram of an embodiment of calculating loss using a fourth loss function in this application.
[0068] Figure 7 This is a schematic diagram of an embodiment of calculating loss using the fifth loss function in this application.
[0069] Figure 8 This is a schematic diagram of an embodiment of the inference method based on the stream matching generation model in this application.
[0070] Figure 9 This is a schematic diagram of the inference process of the one-step generation model based on stream matching in the embodiments of this application;
[0071] Figure 10 This is a schematic diagram of an embodiment of the training device for the flow matching generation model in this application.
[0072] Figure 11 This is a schematic diagram of one embodiment of the inference device based on the stream matching generation model in this application. Detailed Implementation
[0073] This invention provides a training method, an inference method, and related apparatus for a stream matching generation model, which improves the training efficiency of the stream matching generation model training process.
[0074] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0075] The terms "first," "second," "third," "fourth," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0076] For ease of understanding, the training method of the flow matching generation model in the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the training method for the flow matching generation model in this application includes:
[0077] 101. Acquire noise, time t, a first modal object collected from the noise, and environmental characteristics of the first modal object;
[0078] The training process of a stream matching generation model generally includes a forward noise addition process and a backward inference process. The forward noise addition process is used to gradually (i.e., in multiple steps) add noise to the real modal object until the real modal object is transformed into a noisy object that satisfies a Gaussian distribution. The backward inference process is to sample from a noisy object that satisfies a Gaussian distribution and then gradually denoise the sampled object to obtain the real modal object.
[0079] Specifically, the modal objects in this application can be text, images, voice, or actions, etc., and no specific restrictions are placed on the modal objects in this application.
[0080] Furthermore, the forward noise addition process in this application is similar to the forward noise addition process in the prior art, and will not be described in detail here. However, unlike the prior art, the backward noise reduction process in this application adopts a loss function that is different from that in the prior art, so as to improve the training efficiency of the modal objects generated by the flow matching generation model after training in this application.
[0081] In the training of the backward denoising process of the flow matching generation model in this embodiment, it is necessary to first acquire noise, time t, the first modal object collected from the noise, and the environmental features of the first modal object. Here, time t can be regarded as the sampling time, and two adjacent times are regarded as one step of denoising. For example, from time t to time t+1 is regarded as the first step of denoising, and from time t+1 to time t+2 is regarded as the second step of denoising. Here, noise refers to noise objects that follow a Gaussian distribution, and the first modal object collected from the noise is a noise point randomly collected from the noise objects. The environmental features of the first modality are the environmental features associated with the predicted second modal object. For example, when the first modal object is the first action, the second modal object is the second action associated with the first modal object in time. Here, the environmental features are the environmental features of the first action, such as the type of action currently being executed (grabbing action, throwing action, etc.) and the environment in which the action is executed (such as the object of the grabbing action or the object of the throwing action). That is, the environmental features here are associated with the predicted second action to improve the accuracy of the predicted second action.
[0082] 102. Input the noise, the time t, the first modal object and the environmental features into the initialized flow matching generation model to obtain the predicted velocity field of the conditional probability path at time t+1 output by the initialized flow matching generation model.
[0083] After obtaining the noise, time t, first modal object and environmental features, the noise, time t, first modal object and environmental features are input into the initialized flow matching generation model to obtain the predicted velocity field of the conditional probability path at time t+1 output by the initialized flow matching generation model.
[0084] Here, the predicted velocity field refers to the predicted motion velocity (including the predicted velocity magnitude and direction) of the conditional probability path from time t to time t+1. This is because during the training process of backward denoising of the flow matching generation model, it is often expected that the flow matching generation model can recover from the noise object to the real modal object in a straight line. Therefore, in this application, it is desirable that the predicted velocity field is a constant, because only when the predicted velocity field is a constant can it be said that the conditional probability path is a straight line.
[0085] 103. Calculate the loss between the predicted velocity field and the true velocity field using a preset loss function, wherein the preset loss function includes at least one of a first loss function and a second loss function, and a third loss function. The first loss function is used to constrain the minimization of the difference between the predicted velocity field vectors at any two times between time t and time t+1. The second loss function is used to constrain the convergence of modal objects starting from different times between time t and time t+1 to the same position at time u in the future. The third loss function is used to constrain the minimization of the loss between the predicted velocity field and the true velocity field.
[0086] During the backward denoising training of the flow matching generation model, after obtaining the predicted velocity field vector output by the flow matching generation model, the loss between the predicted velocity field and the true velocity field is calculated. Here, the true velocity field is the difference between the first modal object at time t and the predicted second modal object at time t+1.
[0087] Unlike existing technologies, this application, in addition to using a third loss function similar to that in existing technologies to constrain the minimization of the loss between the predicted velocity field and the actual velocity field, also sets at least one of a first loss function and a second loss function to constrain the actual velocity field. The first loss function is used to constrain the minimization of the difference between the predicted velocity field vectors at any two times between time t and t+1. That is, it is desired that the velocity field vectors at any two times between time t and t+1 are the same, because only by ensuring that the difference between the velocity field vectors at any two times is minimized (e.g., the difference is 0) can the prediction at time t be guaranteed. The predicted velocity field of the conditional probability path from time t to time t+1 is a constant, meaning the conditional probability path from time t to time t+1 is a straight line. In addition, the second loss function is used to constrain modal objects starting from different times between time t and time t+1 to converge to the same position at time u in the future. That is, it is hoped that the modal objects predicted at time u by modal objects starting from any time between time t and time t+1 are the same. This is because only by ensuring that the modal objects predicted at time u by modal objects starting from any time are the same can the loss between the predicted modal objects and the real modal objects be reduced.
[0088] For ease of understanding, Figure 2 A schematic diagram illustrating the first and second loss functions is provided. Assuming the modal object in this application is an action, the flow matching generation model predicts future actions from noise. Figure 2In this context, for any two times, r and s, the first loss function is used to constrain the velocity field at time r and the predicted velocity field at time s, minimizing the difference between them. The second loss function, on the other hand, constrains the predicted modal objects that transition from time r to the future time u to converge with the predicted modal objects that transition from time s to the future time u at the same location. This reduces the loss between the predicted action and the actual action.
[0089] 104. Using the loss and backpropagation algorithm, train the initialized flow matching generation model until the initialized flow matching generation model converges, so as to obtain the trained flow matching generation model.
[0090] In step 103, after calculating the loss between the predicted velocity field and the real velocity field using at least one of the first loss function and the second loss function, and the third loss function, the initial flow matching generation model is trained using the loss and the backpropagation algorithm until the initial flow matching generation model converges, so as to obtain the trained flow matching generation model.
[0091] In this embodiment, during the training of the initialized flow matching generation model, not only is the loss between the predicted velocity field and the real velocity field calculated using the existing third loss function, but also the difference between the predicted velocity field vectors at any two times between time t and t+1 is minimized using the first loss function, and / or the modal objects starting from different times between time t and t+1 converge to the same position at time u in the future. This enables the initialized flow matching generation model to recover the real modal objects with the shortest path during the backward denoising training process, thereby improving the training efficiency of the initialized flow matching generation model.
[0092] based on Figure 1 The following describes the process of calculating the loss using the first loss function in the aforementioned embodiment. Please refer to [link to previous text]. Figure 3 :
[0093] 301. Obtain the first predicted velocity field vector at time s among any two time points;
[0094] Specifically, this application assumes that any two times between time t and time t+1 are time s and time t. This application obtains the first predicted velocity field vector at time s. In obtaining the first predicted velocity field vector at time s, the modal object generated at time s, the environmental features of the modal object at time s, and the modal object at time s are input into the initialized flow matching generation model. Then, the first predicted velocity field vector at time s output by the flow matching generation model can be obtained.
[0095] 302. Obtain the second predicted velocity field vector at time r among any two time points;
[0096] Similar to step 301, this application also needs to obtain the second predicted velocity field vector at time r between any two time points. The generation process of the second predicted velocity field vector at time r is similar to that of the first predicted velocity field vector at time s, and will not be described again here.
[0097] 303. Calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector.
[0098] After obtaining the first predicted velocity field vector and the second predicted velocity field vector, the loss between the first predicted velocity field vector and the second predicted velocity field vector is calculated. Here, the loss between the first predicted velocity field vector and the second predicted velocity field vector can be calculated using the L1 loss function or the L2 loss function, etc., without any specific restrictions.
[0099] Specifically, when the modal object in this application is an action, since the predicted velocity field is a multidimensional function, the embodiments of this application can further calculate the expected value of the loss function of each dimension during the calculation of the predicted velocity field.
[0100] For ease of understanding, an example of the first loss function is given below using Formula 1, where the first loss function calculates the expected value of the L2 loss between the first predicted velocity field vector and the second predicted velocity field vector.
[0101]
[0102] Specifically, in Formula 1, r,s represent two arbitrary timestamps collected from the normalized time set, and a r Let a represent the modal object acquired from noise D at time r. s Let v represent the modal object acquired from the noise D at time s, and v θ (s,a s Then v represents the first predicted velocity field vector at time s. θ (r,a r Then ) represents the second predicted velocity field vector at time r.
[0103] In this embodiment of the application, the calculation process of the first loss function is described in detail. The first loss function is designed to minimize the loss of the predicted velocity field at any two times s and t between time t and time t+1, thereby ensuring that the conditional probability path from time t to time t+1 is a straight line.
[0104] The following describes the process of calculating the loss using the second loss function in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 4 :
[0105] 401. Determine any two times, s and r, among the different times mentioned above;
[0106] Specifically, this application requires determining multiple different times between time t and time t+1. For ease of explanation, let's assume that any two times between time t and time t+1 are time s and time r. Of course, time s and time r are just examples of different times and not restrictions. In actual calculation, not only can any number of times be determined from different times, such as three or four times, but the number of times between different times is not specifically limited, nor is the position of the times specifically limited.
[0107] 402. Obtain the first prediction mode object when transitioning from time s to time u;
[0108] Furthermore, in this embodiment of the application, the first predicted modal object is obtained when the time transitions from time s to the future time u. Specifically, when the modal object in this application is an action vector, the first predicted action vector when the time transitions from time s to the future time u is obtained.
[0109] In obtaining the first predicted modal object, the modal object generated at time s, the environmental features of the modal object at time s, and the modal object at time s are input into the initialized flow matching generation model, so as to obtain the first predicted action vector at time s output by the flow matching generation model.
[0110] 403. Obtain the second prediction mode object from time r to time u;
[0111] Furthermore, in this embodiment of the application, it is also necessary to obtain the second predicted modal object when transitioning from time r to the future time u. Specifically, when the modal object in this application is an action vector, the second predicted action vector when transitioning from time r to the future time u is obtained. The generation process of the second predicted action vector is similar to that of the first predicted action vector, and will not be repeated here.
[0112] 404. Calculate the loss between the first prediction modality and the second prediction modality.
[0113] After obtaining the first and second prediction modal objects in the above steps, the loss between the first and second prediction modal objects is further calculated. When calculating the loss, both the L1 loss function and the L2 loss function can be used. Here, there is no specific restriction on the type of the second loss function used to calculate the loss.
[0114] Furthermore, when the modal object is an action vector, the loss between the first predicted action vector and the second predicted action vector is calculated. Since the first predicted action vector and the second predicted action vector are multidimensional functions, the embodiments of this application can also calculate the expected value of the multidimensional function loss when calculating the loss.
[0115] For ease of understanding, an example of the second loss function is given below using Formula 2, where the second loss function calculates the expected value of the L2 loss between the first predicted action vector and the second predicted action vector.
[0116]
[0117] Specifically, in Formula 2, r, s, u represent three arbitrary timestamps collected from the normalized time set, and a r Let a represent the modal object acquired from noise D at time r. s Let a represent the modal object acquired from noise D at time s. u Let v represent the modal object acquired from the noise D at time u, and v θ (s,a s Then v represents the first predicted velocity field vector at time s. θ (r,a r ) represents the second predicted velocity field vector at time r, while (a s +(us)v θ (s,a s )) represents the first predicted action vector, while (a r +(ur)v θ (s,a r Then )) represents the second predicted action vector.
[0118] In this embodiment of the application, the calculation process of the second loss function is described in detail. The second loss function is designed to minimize the loss between the first and second predicted modal objects when starting from any time between time t and time t+1 and when transitioning to the future time u. This ensures the consistency of the predicted modal objects when transitioning from the aforementioned arbitrary time to the future time u, and also ensures the consistency of the predicted velocity field vector when transitioning from the aforementioned arbitrary time to the future time u.
[0119] The above embodiments, when predicting modal objects, fail to consider the temporal continuity of the predicted modal objects, resulting in insufficient temporal consistency of the generated modal objects. For example, when the modal object is an action vector, the temporal continuity of the action trajectory (such as the coherence of the action and the smoothness of the speed) is not considered, leading to insufficient temporal consistency of the generated action vectors. This can result in jitter or abrupt changes in the action over long periods. To address this issue, embodiments of this application can further perform the following steps to improve the temporal consistency of the generated modal objects. Please refer to [link to relevant documentation] for details. Figure 5 Another embodiment of the training method for the flow matching generation model in this application includes:
[0120] 501. Acquire noise, time t, a first modal object collected from the noise, and environmental characteristics of the first modal object;
[0121] 502. Input the noise, the time t, the first modal object and the environmental features into the initialized flow matching generation model to obtain the predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model.
[0122] Steps 501 to 502 in the embodiments of this application are... Figure 1 The descriptions of steps 101 to 102 in the embodiments are similar and will not be repeated here.
[0123] 503. Project the predicted velocity field vector into the frequency domain space to obtain the spectral coefficients of the predicted velocity field line;
[0124] To improve the temporal consistency of the predicted modal objects, this embodiment further projects the predicted velocity field to the frequency domain, that is, transforms the predicted velocity field generated in step 502 from the time domain to the frequency domain. In the process of transformation, Fourier transform formula or discrete cosine transform (DCT) can be used. Here, there are no specific restrictions on the process of transforming from the time domain to the frequency domain.
[0125] For ease of understanding, this application uses Discrete Cosine Transform (DCT) as an example to describe the process of converting the predicted velocity field from the time domain to the frequency domain using Equation 3:
[0126]
[0127] In Equation 3, k represents the dimension of the predicted velocity field vector.
[0128] 504. Calculate the loss between the predicted velocity field and the true velocity field using a preset loss function, wherein the preset loss function includes at least one of a fourth loss function and a fifth loss function, and a third loss function, the fourth loss function includes a fourth function used to constrain the minimization of the difference between the spectral coefficients of the predicted velocity field vector at any two times between time t and time t+1, the fifth loss function used to constrain the consistency of the spectral coefficients of the modal objects starting from different times between time t and time t+1 at future time u, and the third loss function used to constrain the minimization of the loss between the predicted velocity field and the true velocity field;
[0129] After projecting the predicted velocity field vector into the frequency domain, the role of the third loss function in this application remains unchanged, still constraining the minimization of the loss between the predicted velocity field and the true velocity field. The first loss function is then transformed into the fourth loss function in this application, and the second loss function is transformed into the fifth loss function in this application. The fourth loss function is used to constrain the minimization of the difference between the spectral coefficients of the velocity field vector at any two times between time t and t+1. The fifth loss function is used to constrain the consistency of the spectral coefficients of modal objects starting from different times between time t and t+1 at future time u. The calculation process of the fourth and fifth loss functions is described in the following embodiments and will not be repeated here.
[0130] 505. Using the loss and backpropagation algorithm, train the initialized flow matching generation model until the flow matching generation model converges, so as to obtain the trained flow matching generation model.
[0131] In this embodiment of the application, after obtaining the losses calculated by the third, fourth, and fifth loss functions, the initialized flow matching generation model is trained using these losses and the backpropagation algorithm until the flow matching generation model converges, so as to obtain the trained flow matching generation model.
[0132] In calculating the loss, this embodiment transforms the velocity field vectors at any two moments from the time domain to the frequency domain, and uses a fourth loss function to constrain the minimization of the difference between the spectral coefficients of the velocity field vectors at any two moments. Furthermore, it transforms the predicted modal objects from different moments to the frequency domain at the future time u, and uses a fifth loss function to constrain the consistency of the spectral coefficients of the modal objects from different moments at the future time u, thereby satisfying the temporal continuity of the predicted modal objects and improving the temporal consistency of the generated modal objects.
[0133] based on Figure 5The following is an example of calculating the loss using the fourth loss function. Please refer to [link to example]. Figure 6 :
[0134] 601. Obtain the first predicted velocity field vector at time s from any two time points;
[0135] 602. Obtain the second predicted velocity field vector at time r from any two time points;
[0136] It should be noted that steps 601 to 602 here are different from... Figure 3 The steps 301 to 302 in the embodiments are described similarly and will not be repeated here.
[0137] 603. Convert the first predicted velocity field vector from the time domain to the frequency domain to obtain the first spectral coefficients at time s;
[0138] To improve the temporal continuity of the predicted modal objects and thus enhance the temporal consistency of the generated modal objects, this embodiment converts the first predicted velocity field vector from the time domain to the frequency domain to obtain the first spectral coefficient at time s. The conversion process can be referred to in Formula 3. For ease of description, this embodiment denotes the first spectral coefficient at time s as F(v s ).
[0139] 604. Convert the second predicted velocity field vector from the time domain to the frequency domain to obtain the second spectral coefficients at time r;
[0140] Similarly, in this embodiment, the second predicted velocity field vector is also converted from the time domain to the frequency domain to obtain the second spectral coefficient at time r. For ease of description, the second spectral coefficient at time r is also denoted as F(v r ).
[0141] 605. Calculate the loss between the first spectral coefficient and the second spectral coefficient.
[0142] After obtaining the first and second spectral coefficients, the loss between the first and second spectral coefficients is further calculated. The loss can be calculated using either an L1 function or an L2 function, and there are no restrictions on the calculation process.
[0143] For ease of understanding, the following uses Equations 4 and 5 as examples to describe the process of calculating the loss between the first and second spectral coefficients:
[0144]
[0145] Sim(v r ,v s ) = Sim(v θ (s,a s),v θ (r,a r ))=||F(v r )-F(v s )|| 2 (Formula 5)
[0146] In this embodiment of the application, the calculation process of the fourth loss function is described in detail. The fourth loss function calculates the spectral coefficients of the predicted velocity field at any two times s and t between time t and time t+1, so that the loss of the spectral coefficients of the predicted velocity field at any two times is minimized, thereby ensuring the temporal continuity of the predicted modal object from time t to time t+1 and improving the temporal consistency of the predicted modal object.
[0147] Furthermore, when the predicted modal object is an action, it also improves the temporal continuity of the predicted action and avoids sudden changes or jitter in the action.
[0148] based on Figure 5 The following is an example of calculating the loss using the fifth loss function. Please refer to [link to example]. Figure 7 :
[0149] 701. Determine any two times s and r among the different times;
[0150] 702. Obtain the first prediction mode object that transitions from time s to time u;
[0151] 703. Obtain the second prediction mode object when transitioning from time r to time u;
[0152] It should be noted that the descriptions of steps 701 to 703 are consistent with... Figure 4 The descriptions in the embodiments are similar and will not be repeated here.
[0153] 704. Convert the first predicted mode object from the time domain to the frequency domain to obtain the third spectral coefficients when transitioning from time s to time u;
[0154] The process of transforming the first modal object from the time domain to the frequency domain can be referred to in Equation 3. For ease of description, the third spectral coefficient at the time of transition from time s to time u is denoted as F(v(a)). s +(us)v θ (s,a s ))), where a s a represents the modal object at time s. s +(us)v θ (s,a s ) represents the modal object at time u.
[0155] 705. Convert the second predicted mode object from the time domain to the frequency domain to obtain the fourth spectral coefficients from time r to time u;
[0156] Similar to step 704, this embodiment converts the second modal object from the time domain to the frequency domain to obtain the fourth spectral coefficients from time r to time u. For ease of description, the third spectral coefficients from time r to time u are denoted as F(v(a)). r +(us)v θ (r,a r ))), where a r a represents the modal object at time r. r +(us)v θ (r,a r ) represents the modal object at time u.
[0157] 706. Calculate the loss between the third spectral coefficient and the fourth spectral coefficient.
[0158] After obtaining the third and fourth spectral coefficients, the loss between the third and fourth spectral coefficients is further calculated. The loss can be calculated using either an L1 function or an L2 function, and there are no restrictions on the calculation process.
[0159] For ease of understanding, the following uses Equations 5 and 6 as examples to describe the process of calculating the loss between the third and fourth spectral coefficients:
[0160]
[0161] In this embodiment of the application, the calculation process of the fifth loss function is described in detail. The fifth loss function is designed to ensure the consistency of the spectral coefficients between the first and second predicted modal objects from any time between time t and time t+1 to the future time u. This ensures the temporal continuity of the predicted modal objects from the aforementioned arbitrary time to the future time u, thereby improving the temporal consistency of the predicted modal objects. On the other hand, it also ensures the consistency and continuity of the predicted velocity field vector from the aforementioned arbitrary time to the future time u.
[0162] based on Figures 1 to 7In the aforementioned embodiment, because a linear conditional probability path is used to train the initialized flow matching generation model during the backward denoising training process, the training efficiency of the initialized flow matching generation model is improved. When training the initialized flow matching generation model using a linear conditional probability path, a multi-step training method is generally used. In order to further improve the training speed of the initialized flow matching generation model, this application can also improve the initialized flow matching generation model into a one-step initialized flow matching generation model, so that the real modal object can be recovered from the noise in one step during the backward denoising training process, thereby further improving the efficiency of obtaining the real modal object.
[0163] The training process of the stream matching generation model in the embodiments of this application has been described above. The inference process of the stream matching generation model in the embodiments of this application will be described below. Please refer to [link / reference]. Figure 8 An embodiment of the inference method for the flow matching generation model in this application includes:
[0164] 801. Acquire noise, time t, a first modal object collected from the noise, and environmental characteristics of the first modal object;
[0165] Specifically, step 801 here is similar to step 101, and will not be repeated here.
[0166] 802. Input the noise, the time t, the first modal object and the environmental features into the trained flow matching generation model to obtain the second modal object predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.
[0167] After obtaining the noise, time t, first modal object and environmental features in step 801, the noise, time t, first modal object and environmental features are input into the flow matching generation model trained according to the above method embodiment to obtain the second modal object predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.
[0168] Specifically, the modal objects in the first and second modal objects can also be text, images, voice, or actions. There are no specific restrictions on the specific categories of the modal objects in the first and second modal objects.
[0169] In this embodiment, after obtaining the trained flow matching generation model, noise, time t, the first modal object, and environmental features are input into the trained flow matching generation model. Then, the second modal object predicted by the flow matching generation model according to the conditional probability path predicted by the preset prediction velocity field vector can be obtained. Furthermore, since the second modal object is predicted by a straight conditional probability path during the training process, the trained flow matching generation model in this application also predicts the second modal object by a straight conditional probability path, thereby improving the efficiency of second modal object generation.
[0170] based on Figure 8 In the aforementioned embodiment, when the flow matching generation model is a one-step flow matching generation model, after inputting noise, time t, the first modal object, and environmental features into the trained one-step flow matching generation model, the predicted modal object output by the trained one-step flow matching generation model can be obtained.
[0171] For ease of understanding, Figure 9 The diagram illustrates the process of predicting actions obtained by inputting noise, time t, robot action state, and robot environmental features into a trained flow matching one-step generation model (one-step action generator) when the modal object is robot action and the environmental features are a series of observed action feature vectors of the robot when performing action sequence.
[0172] This application also provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the training method of the flow matching generation model provided in the above-described method embodiments of this application, or to implement the inference method based on the flow matching generation model provided in the above-described method embodiments of this application.
[0173] The training process and inference process based on the stream matching generation model in the embodiments of this application have been described in detail above. The training apparatus for the stream matching generation model in the embodiments of this application will be described below. Please refer to [link to relevant documentation]. Figure 10 One embodiment of the training apparatus for the flow matching generation model in this application includes:
[0174] The acquisition unit 1001 is used to acquire noise, time t, a first modal object collected from the noise, and environmental features of the first modal object;
[0175] Input unit 1002 is used to input the noise, the time t, the first modal object and the environmental features into the initialized flow matching generation model to obtain the predicted velocity field of the conditional probability path at time t+1 output by the initialized flow matching generation model.
[0176] The calculation unit 1003 is used to calculate the loss between the predicted velocity field and the true velocity field using a preset loss function. The preset loss function includes at least one of a first loss function and a second loss function, and a third loss function. The first loss function is used to constrain the minimization of the difference between the predicted velocity field vectors at any two times between time t and time t+1. The second loss function is used to constrain the convergence of modal objects starting from different times between time t and time t+1 to the same position at time u in the future. The third loss function is used to constrain the minimization of the loss between the predicted velocity field and the true velocity field.
[0177] Training unit 1004 is used to train the initialized flow matching generation model using the loss and backpropagation algorithm until the flow matching generation model converges, so as to obtain the trained flow matching generation model.
[0178] As an optional embodiment, the acquisition unit 1001 is further configured to:
[0179] Obtain the first predicted velocity field vector at time s among any two time points;
[0180] Obtain the second predicted velocity field vector at time r from any two time points;
[0181] The computing unit 1003 is specifically used for:
[0182] Calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector.
[0183] As an optional embodiment, the acquisition unit 1001 is further configured to:
[0184] Determine any two times s and r among the different times;
[0185] Obtain the first prediction mode object from time s to time u;
[0186] Obtain the second prediction mode object from time r to time u;
[0187] The computing unit 1003 is specifically used for:
[0188] Calculate the loss between the first prediction modality and the second prediction modality.
[0189] As an optional embodiment, the computing unit 1003 is further configured to:
[0190] The predicted velocity field is projected into the frequency domain to obtain the spectral coefficients of the predicted velocity field.
[0191] The constraint is to minimize the difference in spectral coefficients between the velocity field vectors at any two time points;
[0192] Constrain the consistency of the spectral coefficients of modal objects starting from different times between time t and time t+1 in the future time u.
[0193] As an optional embodiment, the acquisition unit 1001 is further configured to:
[0194] Obtain the first predicted velocity field vector at time s among any two time points;
[0195] Obtain the second predicted velocity field vector at time r from any two time points;
[0196] Calculation unit 1003 is specifically used for:
[0197] The first predicted velocity field vector is converted from the time domain to the frequency domain to obtain the first spectral coefficients at time s;
[0198] The second predicted velocity field vector is converted from the time domain to the frequency domain to obtain the second spectral coefficients at time r;
[0199] Calculate the loss between the first spectral coefficient and the second spectral coefficient.
[0200] As an optional embodiment, the acquisition unit 1001 is further configured to:
[0201] Determine any two times s and r among the different times;
[0202] Obtain the first prediction mode object that transitions from time s to time u;
[0203] Obtain the second prediction mode object when transitioning from time r to time u;
[0204] Calculation unit 1003 is specifically used for:
[0205] The first predicted mode object is converted from the time domain to the frequency domain to obtain the third spectral coefficients from time s to time u;
[0206] The second predicted mode object is converted from the time domain to the frequency domain to obtain the fourth spectral coefficients from time r to time u;
[0207] Calculate the loss between the third spectral coefficient and the fourth spectral coefficient.
[0208] As an optional embodiment, the initialized stream matching generation model includes an initialized stream matching one-step generation model; the first modal object includes text, image, voice, or robot action state.
[0209] It should be noted that the functions of each of the above units are similar to those described in the above method embodiments, and will not be repeated here.
[0210] The following description details the inference apparatus based on the stream matching generation model in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 11 One embodiment of the inference device based on the stream matching generation model in this application includes:
[0211] Acquisition unit 1101 is used to acquire noise, time t, a first modal object collected from the noise, and environmental features of the first modal object;
[0212] The prediction unit 1102 is used to input the noise, the time t, the first modal object and the environmental features into the trained flow matching generation model to obtain the second modal object predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.
[0213] As an optional embodiment, if the trained stream matching generation model is a one-step stream matching generation model, then the prediction unit 1102 is specifically used for:
[0214] The noise, the time t, the first modal object, and the environmental features are input into the trained flow matching one-step generation model to obtain the second modal object generated by the trained flow matching one-step generation model in one step by the conditional probability path predicted by the preset velocity field vector.
[0215] It should be noted that the functions of each unit in this application embodiment are similar to those described in the above method embodiments, and will not be repeated here.
[0216] This application also provides a computer program product having a computer program stored thereon. When executed by a processor, the computer program is used to implement the various steps in the above-described training method embodiment of the stream matching generation model, and / or to implement the various steps in the above-described inference method embodiment based on the stream matching generation model.
[0217] The embodiments of the present invention have been described above from the perspective of modular functional entities. The computer device in the embodiments of the present invention will now be described from the perspective of hardware processing:
[0218] This computer device is used to implement the functions of the training device side of the stream matching generation model. One embodiment of the computer device in this invention includes:
[0219] Processor and memory;
[0220] The memory is used to store computer programs, and when the processor executes the computer programs stored in the memory, it can implement the various steps in the above-described training method embodiment for the flow matching generation model.
[0221] The computer device is also used to implement the functions of the inference device side based on the stream matching generation model. Another embodiment of the computer device in this invention includes:
[0222] Processor and memory;
[0223] The memory is used to store computer programs, and when the processor executes the computer programs stored in the memory, it can implement the various steps in the above-described inference method embodiment based on the stream matching generation model.
[0224] It is understood that when the processor in the computer device described above executes the computer program, it can also implement the functions of each unit in the corresponding device embodiments described above, which will not be repeated here. For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the training device / inference device based on the stream matching generation model. For example, the computer program can be divided into units in the training device of the stream matching generation model described above, and each unit can implement the specific functions described in the corresponding training device of the stream matching generation model.
[0225] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the processor and memory are merely examples of a computer device and do not constitute a limitation on the computer device. It may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.
[0226] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the computer device, connecting various parts of the computer device via various interfaces and lines.
[0227] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0228] The present invention also provides a computer-readable storage medium for implementing the functions of a training device side of a stream matching generation model. The computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the processor can be used to implement the various steps in the above-described embodiment of the training method for the stream matching generation model.
[0229] The present invention also provides another computer-readable storage medium for implementing the functions of an inference device based on a stream matching generation model, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the processor can be used to implement various steps in the embodiments of the inference method based on the stream matching generation model.
[0230] It is understood that if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0231] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0232] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0233] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0234] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a flow matching generation model, characterized in that, The method includes: Acquire noise, time t, a first modal object collected from the noise, and environmental features of the first modal object, wherein the first modal object includes robot action state, and the environmental features of the first modal object include at least one of robot action type and robot action execution environment; The noise, the time t, the first modal object, and the environmental features are input into the initialized flow matching generation model to obtain the predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model. The loss between the predicted velocity field vector and the true velocity field vector is calculated using a preset loss function. The preset loss function includes at least one of a first loss function and a second loss function, and a third loss function. The first loss function is used to constrain the minimization of the difference between the predicted velocity field vectors at any two times between time t and time t+1. The second loss function is used to constrain the convergence of modal objects starting from different times between time t and time t+1 to the same position at a future time u. The third loss function is used to constrain the minimization of the loss between the predicted velocity field vector and the true velocity field vector. The initial flow matching generation model is trained using the loss and backpropagation algorithm until the flow matching generation model converges, so as to obtain the trained flow matching generation model. The method further includes: After obtaining the predicted velocity field vector, the predicted velocity field vector is projected into the frequency domain space to obtain the spectral coefficients of the predicted velocity field vector. The first loss function includes a fourth loss function, which is used for: The constraint is to minimize the difference in spectral coefficients between the predicted velocity field vectors at any two moments. The second loss function includes a fifth loss function, which is used for: Constrain the consistency of the spectral coefficients of modal objects starting from different times between time t and time t+1 in the future time u.
2. The training method according to claim 1, characterized in that, The method further includes: Obtain the first predicted velocity field vector at time s among any two time points; Obtain the second predicted velocity field vector at time r from any two time points; The first loss function is specifically used for: Calculate the loss between the first predicted velocity field vector and the second predicted velocity field vector.
3. The training method according to claim 1, characterized in that, The method further includes: Determine any two times s and r among the different times; Obtain the first prediction mode object when transitioning from time s to time u; Obtain the second prediction mode object when transitioning from time r to time u; The second loss function is specifically used for: Calculate the loss between the first prediction modality and the second prediction modality.
4. The method according to claim 1, characterized in that, The method further includes: Obtain the first predicted velocity field vector at time s among any two time points; Obtain the second predicted velocity field vector at time r from any two time points; The first predicted velocity field vector is converted from the time domain to the frequency domain to obtain the first spectral coefficients at time s; The second predicted velocity field vector is converted from the time domain to the frequency domain to obtain the second spectral coefficients at time r; The fourth loss function is specifically used for: Calculate the loss between the first spectral coefficient and the second spectral coefficient.
5. The method according to claim 1, characterized in that, The method further includes: Determine any two times s and r among the different times; Obtain the first prediction mode object that transitions from time s to time u; Obtain the second prediction mode object when transitioning from time r to time u; The first predicted mode object is converted from the time domain to the frequency domain to obtain the third spectral coefficients from time s to time u; The second predicted mode object is converted from the time domain to the frequency domain to obtain the fourth spectral coefficients from time r to time u; The fifth loss function is specifically used for: Calculate the loss between the third spectral coefficient and the fourth spectral coefficient.
6. The method according to claim 1, characterized in that, The initialized flow matching generation model includes an initialized one-step flow matching generation model.
7. A reasoning method based on a flow matching generation model, characterized in that, The method includes: Acquire noise, time t, a first modal object collected from the noise, and environmental characteristics of the first modal object; The noise, the time t, the first modal object, and the environmental features are input into the trained flow matching generation model as described in any one of claims 1 to 6 to obtain the second modal object predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.
8. The reasoning method according to claim 7, characterized in that, If the trained flow matching generation model includes a trained one-step flow matching generation model; The step of inputting the noise, the time t, the first modal object, and the environmental features into the trained flow matching generation model to obtain the second modal object obtained by the trained flow matching generation model using the conditional probability path predicted by the preset velocity field vector includes: The noise, the time t, the first modal object, and the environmental features are input into the trained flow matching one-step generation model to obtain the second modal object generated by the trained flow matching one-step generation model in one step by the conditional probability path predicted by the preset velocity field vector.
9. A training apparatus for a stream matching generation model, characterized in that, The device includes: The acquisition unit is used to acquire noise, time t, a first modal object collected from the noise, and environmental features of the first modal object. The first modal object includes a robot action state, and the environmental features of the first modal object include at least one of the robot action type and the execution environment of the robot action. The input unit is used to input the noise, the time t, the first modal object and the environmental features into the initialized flow matching generation model to obtain the predicted velocity field vector of the conditional probability path at time t+1 output by the initialized flow matching generation model. The calculation unit is used to calculate the loss between the predicted velocity field vector and the true velocity field vector using a preset loss function. The preset loss function includes at least one of a first loss function and a second loss function, and a third loss function. The first loss function is used to constrain the minimization of the difference between the predicted velocity field vectors at any two times between time t and time t+1. The second loss function is used to constrain the convergence of modal objects starting from different times between time t and time t+1 to the same position at a future time u. The third loss function is used to constrain the minimization of the loss between the predicted velocity field vector and the true velocity field vector. The training unit is used to train the initialized flow matching generation model using the loss and backpropagation algorithm until the flow matching generation model converges, so as to obtain the trained flow matching generation model. The computing unit is also used for: The predicted velocity field vector is projected into the frequency domain to obtain the spectral coefficients of the predicted velocity field. The constraint is to minimize the difference in spectral coefficients between the predicted velocity field vectors at any two time points. Constrain the consistency of the spectral coefficients of modal objects starting from different times between time t and time t+1 in the future time u.
10. An inference device based on a stream matching generation model, characterized in that, The device includes: The acquisition unit is used to acquire noise, time t, a first modal object collected from the noise, and environmental features of the first modal object; The prediction unit is used to input the noise, the time t, the first modal object and the environmental features into the trained flow matching generation model as described in any one of claims 1 to 6, so as to obtain the second modal object predicted by the conditional probability path predicted by the trained flow matching generation model with a preset velocity field vector.
11. A computer device comprising a processor, characterized in that, When the processor executes a computer program stored in the memory, it is used to implement the training method of the stream matching generation model as described in any one of claims 1 to 6, or the inference method based on the stream matching generation model as described in any one of claims 7 to 8.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the training method of the stream matching generation model as described in any one of claims 1 to 6, or the inference method based on the stream matching generation model as described in any one of claims 7 to 8.
13. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the training method of the stream matching generation model as described in any one of claims 1 to 6, or the inference method based on the stream matching generation model as described in any one of claims 7 to 8.