Intuitive world model and visual reasoning method and device based on intuitive world model
Through the search and interaction module of the intuitive world model and combined with physical law learning, the problem of low accuracy of visual reasoning in multi-object motion interaction scenarios in the existing technology is solved, and the autonomous discovery and interpretation of physical events is achieved, and the accuracy of visual reasoning is improved.
Patent Information
- Application Number
- CN202510360048.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-15
AI Technical Summary
The existing world model designed based on artificial intelligence technology is difficult to dig out actions or phenomena that conform to the laws of the physical world when objects move in multi-object motion interaction scenarios, resulting in low accuracy of visual reasoning tasks.
The intuitive world model is adopted, and the static and dynamic features in the video data of the target scene are extracted through the search module, and the intuitive interaction module is used to decompose and calculate potential physical properties. Combined with the physical laws to learn the model, they independently discover and explain physical events.
It improves the visual inference accuracy of the world model in multi-object motion interaction scenarios, can independently discover and explain physical laws, and enhances the causal interpretation ability and modeling accuracy of interaction laws.
Smart Images

Figure CN120494085A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to an intuitive world model, and a visual reasoning method and device based on the intuitive world model. Background Art
[0002] Existing world models designed based on artificial intelligence technology are generally good at recognizing visual features and patterns, such as identifying the appearance of objects, predicting their positions or motion trajectories, etc. These tasks usually rely on large-scale labeled data and statistical associations. When faced with more complex multi-object motion interaction scenarios, existing world models often show significant limitations.
[0003] In related technologies, traditional world models usually perform pattern recognition and time series prediction tasks based on simple physical features (such as position or speed changes). However, world models usually operate as "black boxes" and are difficult to provide intuitive and explainable outputs. They are also unable to describe actions and phenomena that conform to the laws of the physical world from observation data, or it is difficult to infer physical laws from visual information. For example, in a multi-object interaction scene, the motion trajectories of multiple objects are predicted according to simple physical features through world models. However, due to the interaction between multiple objects due to their respective deep physical properties (for example, charge repulsion occurs between two objects), the actual speed, acceleration or running direction of the objects will be affected, resulting in low reliability of trajectory prediction or physical phenomenon analysis based solely on simple physical features, thereby resulting in low accuracy of visual reasoning tasks based on world models. Summary of the Invention
[0004] The present invention provides an intuitive world model, a visual reasoning method and device based on the intuitive world model, so as to solve the problem that the world model in the prior art is difficult to discover the actions or phenomena that conform to the laws of the physical world when the object moves when performing visual reasoning tasks based on simple physical features, resulting in low reliability of the reasoning and analysis process based on the world model, and thus low accuracy of visual reasoning tasks based on the world model, thereby improving the visual reasoning accuracy of the world model.
[0005] The present invention provides an intuitive world model, comprising: a search module configured to perform visual reasoning on video data of a target scene to obtain latent variables of a plurality of objects in the target scene, wherein the latent variables are used to represent at least one of a static feature, a dynamic feature, and a latent physical property of each object, wherein the latent physical property includes at least one of a mass and a charge of the plurality of objects; An intuitive interaction module is used to decompose and calculate the latent variables using a display modeling method to obtain interaction information between different objects, and update the motion status of the multiple objects based on the interaction information; derive the motion law parameters of each object based on the motion status of the multiple objects to perform a target reasoning task; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon interpretation and human-computer interaction.
[0006] According to an intuitive world model provided by the present invention, the static features include color, texture, shape, and position, the dynamic features include velocity direction and velocity magnitude, and the potential physical properties include mass and charge; The search module includes: A feature extraction module, configured to extract the static features and the dynamic features from the video data; A derivation module is used to deduce the physical properties of each object based on the static features and the dynamic features to obtain the potential physical properties, and adjust the parameters of the intuitive world model based on the potential physical properties.
[0007] According to an intuitive world model provided by the present invention, the latent variables include accelerations of the plurality of objects; The intuitive interaction module is specifically used to decompose the acceleration to obtain acceleration attention, acceleration direction and acceleration magnitude; calculate the target acceleration based on the acceleration attention, acceleration direction and acceleration magnitude, and confirm that the target acceleration is the real-time acceleration of different objects.
[0008] According to an intuitive world model provided by the present invention, the search module is further configured to perform visual reasoning on the video data through an iterative probabilistic reasoning mechanism; The probabilistic reasoning mechanism through iteration is expressed by the following formula: ; in, represents the latent variable, is the input scene video frame, is the generated video frame mean, is the corresponding object mask, Represents the gradient feedback of the object distribution.
[0009] According to an intuitive world model provided by the present invention, deducing the physical property law of each object based on the motion state of the multiple objects includes: Inputting the motion states of the multiple objects into a physical law learning model to obtain the motion law parameters; The physical law learning model is obtained by performing unsupervised training on a neural network using sample motion law parameters as training samples and an objective function as a loss function, wherein the objective function is determined based on a reconstruction error and a regularization term.
[0010] The present invention also provides a visual reasoning method based on an intuitive world model, comprising: Obtain visual data to be processed; Reasoning is performed on the visual data to be processed based on an intuitive world model to obtain a visual reasoning result, and a target reasoning task is performed according to the visual reasoning result; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon explanation, and human-computer interaction.
[0011] The present invention also provides a visual reasoning device based on an intuitive world model, comprising: A data acquisition module, used to acquire visual data to be processed; A visual reasoning module is used to reason about the visual data to be processed based on an intuitive world model to obtain a visual reasoning result, and to perform a target reasoning task based on the visual reasoning result; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon interpretation, and human-computer interaction.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the visual reasoning method based on the intuitive world model as described above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described visual reasoning methods based on the intuitive world model.
[0014] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described visual reasoning methods based on an intuitive world model.
[0015] The intuitive world model, the visual reasoning method and the device based on the intuitive world model provided by the present invention perform visual reasoning on the video data of the target scene through a search module to obtain the potential variables of multiple objects in the target scene, and can autonomously discover implicit physical properties; the intuitive interaction module is used to decompose and calculate the potential variables using a display modeling method to obtain interaction information between different objects, and update the motion status of multiple objects based on the interaction information. Finally, the motion law parameters of multiple objects are determined based on the motion status of multiple objects to perform target reasoning tasks. They can autonomously discover physical laws and display explanations of physical events, thereby improving the visual reasoning accuracy of the world model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is one of the structural diagrams of the intuitive world model provided by the present invention.
[0018] Figure 2 This is a schematic diagram of the physical phenomena displayed and physical law extracted when two objects collide with each other, provided by the present invention.
[0019] Figure 3 This is the second structural diagram of the intuitive world model provided by the present invention.
[0020] Figure 4 This is one of the diagrams for visually representing Coulomb's law learned by the IWM model provided by the present invention.
[0021] Figure 5 This is the second diagram of the visualization representation of Newton's laws learned by the IWM model provided by the present invention.
[0022] Figure 6 It is a flowchart of the visual reasoning method based on the intuitive world model provided by the present invention.
[0023] Figure 7 It is a structural diagram of the visual reasoning device based on the intuitive world model provided by the present invention.
[0024] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention.
[0025] Reference numerals: 110: Search module; 120: Intuitive interaction module. DETAILED DESCRIPTION
[0026] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0027] The following combination Figure 1-Figure 5 The intuitive world model, the visual reasoning method and the device based on the intuitive world model of the present invention are described.
[0028] Figure 1 This is a schematic diagram of the structure of the intuitive world model provided by the present invention. Figure 1 As shown, the intuitive world model includes: a search module 110 and an intuitive interaction module 120.
[0029] The search module 110 is used to perform visual reasoning on the video data of the target scene to obtain latent variables of multiple objects in the target scene. The latent variables are used to represent at least one of the static characteristics, dynamic characteristics and potential physical properties of each object. The potential physical properties include at least one of the mass and charge of the multiple objects.
[0030] In this embodiment, the target scenes include simple object motion scenes, such as scenes of object motion speed, position or trajectory prediction; the target scenes also include complex multi-body interaction scenes, such as collision events of multiple objects, Coulomb force interaction and three-body motion.
[0031] For example, target scenarios include the following scenarios or events: (1) Collision events: used to analyze momentum transfer and energy distribution between objects; (2) Coulomb force interaction scenario: used to describe the attraction or repulsion between charged objects; (3) Multi-object mixed interaction scenarios: including simultaneous collision and Coulomb force interactions; (4) Three-body motion scenario: the Coulomb force interaction motion of three objects.
[0032] In this embodiment, the latent variables include static features of the object, such as appearance features such as the object's shape, color, texture or material, and position features such as the object's two-dimensional or three-dimensional coordinates in the scene; they also include dynamic features, such as speed features such as the object's speed, acceleration, and angular velocity, and implicit physical properties of the object, such as the object's mass, charge or friction coefficient, which are not explicitly labeled but can be learned from the interaction process.
[0033] In this embodiment, high-quality video data containing multi-object interactions is sampled as input. These video scenes cover a wide range of typical physical events, including collisions, Coulomb force interactions, and mixed events combining the two. Each video contains static appearance features of the object (such as color, size, shape, and material) as well as dynamic motion information (such as velocity, acceleration, and displacement).
[0034] In this embodiment, the static appearance features obtained from the video data provide the basis for object identification for subsequent models, while the dynamic features provide important clues for deducing implicit physical properties and interaction laws. Since the video data itself does not display labeled physical properties (such as the mass and charge of the object), the subsequent models can automatically deduce these implicit characteristics through unsupervised learning methods and learn general physical laws based on diverse scene events.
[0035] In this embodiment, the search module 110 is used to derive the value of the variable based on the feedback of the prediction error, and to construct "new variables" when necessary, so as to model interactive behaviors that are more complex than linear motion by learning (discovering) new variables.
[0036] In this embodiment, the search module 110 can also use the optimized latent variables to accurately describe the individual characteristics of each object, that is, gradually converge by optimizing the consistency with the input data, thereby ensuring accurate representation of the physical characteristics of each object in the scene.
[0037] In this embodiment, the search module 110 is further configured to perform visual reasoning on the video data through an iterative probabilistic reasoning mechanism; the iterative probabilistic reasoning mechanism is expressed as follows: ; in, represents the latent variable, is the input scene video frame, is the generated video frame mean, is the corresponding object mask, Represents the gradient feedback of the object distribution.
[0038] The intuitive interaction module 120 is used to decompose and calculate latent variables using a display modeling method to obtain interaction information between different objects, and update the motion status of multiple objects based on the interaction information; derive the motion law parameters of each object based on the motion status of multiple objects to perform target reasoning tasks; wherein the target reasoning tasks include at least one of trajectory prediction, physical phenomenon explanation and human-computer interaction.
[0039] In this embodiment, the interaction module can model the forces between objects. For example, the intuitive interaction module 120 displays and decomposes the acceleration of the objects obtained according to the above steps to quantify the interactions between objects due to implicit physical properties.
[0040] Specifically, the interactive acceleration between modeled objects displayed by the intuitive interaction module 120 can capture the causal relationship and dynamic changes between objects by decomposing and calculating the three components of the acceleration: attention, direction, and magnitude.
[0041] In an embodiment, the above display decomposition method enhances the causal explanation capability of the intuitive world model and improves the modeling accuracy of multi-object interaction rules.
[0042] In this embodiment, the motion law parameters include mass, charge state, friction coefficient and other parameters that conform to the physical laws of the objective world.
[0043] In this embodiment, the intuitive world model includes the following multi-stage training mechanism: (1) Phase 1: The model is trained in a non-interactive scenario, and the static features of objects, including appearance and location distribution, are learned through the search module 110; (2) Second stage: In the interactive scene, the static feature prediction error feedback of the fixed object is obtained through the intuitive interaction module 120, and the implicit physical properties (such as mass and charge) are further learned based on the prediction results; (3) The third stage: The physical law generation module 130 combines different types of interaction events (such as collision and Coulomb force) to obtain the motion laws of multiple objects and perform the target reasoning task.
[0044] The intuitive world model provided by the embodiment of the present invention performs visual reasoning on the video data of the target scene through a search module to obtain the latent variables of multiple objects in the target scene, and can autonomously discover implicit physical properties; the intuitive interaction module is used to decompose and calculate the latent variables using a display modeling method to obtain interaction information between different objects, and update the motion status of multiple objects based on the interaction information. Finally, the motion law parameters of multiple objects are determined based on the motion status of multiple objects to perform target reasoning tasks. It can autonomously discover physical laws and display explanations of physical events, thereby improving the visual reasoning accuracy of the world model.
[0045] In some embodiments, static features include color, texture, shape, and position; dynamic features include velocity direction and velocity magnitude; and potential physical properties include mass and charge; the search module 110 includes: a feature extraction module and a derivation module.
[0046] The feature extraction module is used to extract static features and dynamic features from video data.
[0047] In this embodiment, in the initial stage, the feature extraction module of the search module 110 extracts static features (such as appearance and position) and dynamic features (such as speed) of objects from each frame of video data. These features constitute the initial input of the model and are used for subsequent physical property derivation and interaction modeling.
[0048] The derivation module is used to deduce the physical properties of each object based on static and dynamic features, obtain potential physical properties, and adjust the parameters of the intuitive world model based on the potential physical properties.
[0049] In this embodiment, the motion information of an object is used by a deduction module to gradually deduce its implicit physical properties (such as mass and charge). This deduction process is implemented through a deep learning framework, and the model parameters are adjusted through iterative optimization to ensure that the derived properties are consistent with the actual behavior of the object.
[0050] The intuitive world model provided by the embodiment of the present invention, by setting a search module including a feature extraction module, extracts static features and dynamic features from video data, sets a derivation module to deduce the physical properties of the object based on the static features and dynamic features to obtain potential physical properties, and adjusts the parameters of the intuitive world model based on the potential physical properties. It can fully explore the surface features and deep features of the object, thereby improving the efficiency of the intuitive world model in automatically acquiring implicit physical properties.
[0051] In some embodiments, the latent variables include accelerations of multiple objects; the intuitive interaction module 120 is specifically used to decompose the acceleration to obtain acceleration attention, acceleration direction, and acceleration magnitude (i.e., interaction information); calculate the target acceleration based on the acceleration attention, acceleration direction, and acceleration magnitude, and confirm that the target acceleration is the real-time acceleration of different objects.
[0052] In this embodiment, acceleration is decomposed into three components: acceleration attention (used to determine the intensity of interaction between objects and reflect the weight distribution of forces between different objects), acceleration direction (used to determine the direction of acceleration through the normalized position vector difference and used to represent the spatial distribution of interaction forces), and acceleration magnitude (used to quantify the amplitude of the force between objects and the ReLU function can be used to ensure that the calculation result is non-negative).
[0053] In this embodiment, the intuitive interaction module 120 uses a display modeling method to decompose the acceleration between objects, and the steps include: (1) Attention calculation: The interaction weight between object pairs is determined by the following formula: in, is the Sigmoid activation function, represents the attention calculation function, is the embedded feature of object i and object j; (2) Direction calculation: Calculate the direction of acceleration by normalizing the position vector difference between objects: in, Calculate function for direction; (3) Calculation: Calculate the magnitude of the acceleration using the following formula, ensuring it is non-negative: in, It is a size calculation function and ReLU is used to constrain the activation to be non-negative.
[0054] The resulting acceleration between the objects is defined by the following formula: The acceleration value is used to update the velocity characteristics of the object, thereby achieving an accurate description of the object's motion dynamics.
[0055] The intuitive world model provided by the embodiment of the present invention decomposes acceleration through an intuitive interaction module to obtain acceleration attention, acceleration direction and acceleration magnitude, calculates the target acceleration based on the acceleration attention, acceleration direction and acceleration magnitude, and confirms that the target acceleration is the real-time acceleration of different objects. The display decomposition method is used to enhance the causal interpretation ability of the intuitive world model and effectively improve the modeling accuracy of the interaction rules.
[0056] In some embodiments, deducing the physical property laws of each object based on the motion states of multiple objects includes: inputting the motion states of the multiple objects into a physical law learning model to obtain the motion law parameters; wherein the physical law learning model is obtained by unsupervised training of a neural network using sample motion law parameters as training samples and an objective function as a loss function, and the objective function is determined based on the reconstruction error and the regularization term.
[0057] In this embodiment, the physical law learning model can reversely deduce the approximate expressions of general physical laws (such as Newton's law, Coulomb's law, conservation of momentum, or the relationship between the direction and magnitude of force) from the motion states of multiple objects. For example, it can automatically learn the mathematical form of conservation of momentum from a vehicle collision video, and then explain the collision phenomenon based on the motion law parameters corresponding to the conservation of momentum, draw a force analysis diagram, and perform reasoning tasks in vehicle collision scenarios.
[0058] In this embodiment, the physical law learning model is trained by optimizing the loss function; the reconstruction error in the objective function is used to measure the difference between the model generation result and the input video observation data, ensuring that the object appearance and motion trajectory are consistent with the actual situation; the regularization term in the objective function is used to guide the latent variables to conform to the physical laws or the intrinsic structure of the data.
[0059] In this embodiment, during the multiple rounds of iterative training of the physical law learning model, the understanding and prediction capabilities of physical events are continuously improved by updating the properties and interaction laws of objects. That is, after each round of iteration, the model will recalculate the loss value based on the current prediction results, and further adjust the model parameters to gradually converge to the optimal state.
[0060] In this embodiment, after model training is completed, the performance and applicability of the model are evaluated through a verification process. The verification content includes the following aspects: (1) Prediction accuracy: On the validation data, the model first infers the implicit properties and interaction patterns of the object, and then uses these patterns to predict the object's future trajectory. By comparing the predicted results with the actual trajectory, the model's prediction error is evaluated.
[0061] (2) Consistency of physical laws: Verify whether the physical laws derived from the model are consistent with classical physical laws (such as Newton's laws, conservation of momentum, the relationship between the direction and magnitude of force, etc.). This assessment ensures that the model can not only predict motion behavior but also generate explanations that conform to physical logic.
[0062] In this embodiment, the objective function is optimized using video data, and unsupervised learning is performed on a neural network (such as a deep neural network, a convolutional neural network, or other neural networks) by adopting the above objective function. The obtained physical law learning model can extract physical laws that match the scene from the data. Finally, the optimized model is used to explain the physical phenomena, such as analyzing the source, magnitude, and direction of the force, and predicting the motion trajectory and interaction pattern of objects in future scenes.
[0063] In this embodiment, the objective function is expressed by the following formula: 𝐿 ; in, represents the generated distribution, represents the posterior distribution, β is the regularization parameter, is the regularization term; N is the number of video frames, t is the specific frame number, is an image of a specific frame, is the potential attributes of all objects corresponding to the t-frame scene, is the prior distribution of all physical properties.
[0064] The intuitive world model provided by the embodiment of the present invention performs unsupervised training on the neural network through objective function training determined by the generative distribution, the posterior distribution and the regularization parameter, thereby obtaining a physical law learning model, thereby realizing the law extraction and phenomenon explanation of the movement of objects that conform to objective physical laws in different scenarios, and improving the ability of the intuitive world model to autonomously discover physical laws and display and explain physical events.
[0065] Figure 2 This is a diagram showing the physical phenomena and physical law extraction when two objects collide. Figure 2 In the embodiment shown, the collision video of two marbles is observed and recorded, and the video is input into the intuitive world model for reasoning to obtain the reasoning results, which are used to explain physical events: including the forces on the two marbles (two forces F 1. F 2 directions and velocities of the two marbles v 1. v 2), and discovered physical principles that approximate Newton's laws (≈Newton's law).
[0066] Figure 3 This is the second structural diagram of the intuitive world model provided by the present invention. Figure 3 In the embodiment shown, the operation mechanism of the intuitive world model includes two stages ( Stage 1 and Stage 2, in, Stage 1 represents single-object dynamics learning, without adaptive properties, Stage 2 represents multi-object interaction dynamics learning, optionally using adaptive attributes), in Stage 1 During the process: Input: motion trajectory data of a single object (X).
[0067] Process: The observation data (X) is mapped into the latent space through the search module, and the representation of appearance (green) and speed (yellow) attributes is obtained.
[0068] Key: The velocity attribute is transformed through a linear layer to calculate the position offset (Δ) caused by the velocity.
[0069] Specifically, by adding the position offset to the appearance attribute, the attribute prediction at the next moment is obtained, and then the predicted attribute is input into the decoder module to reconstruct the motion trajectory (X') at the next moment.
[0070] The training objective is to minimize the difference between the reconstructed trajectory (X') and the true trajectory (X), so that the model can learn the motion law of a single object (for example, uniform linear motion). Adaptive properties are not involved in this stage.
[0071] exist Stage 2 During the process: Input: motion trajectory data of multiple objects (X).
[0072] Process: Similar to stage 1, the appearance and velocity attributes of each object are extracted through the search module.
[0073] These attributes are input into the Interactive Information Modeling (IIM) module (i.e., the direct interaction module mentioned above); the IIM module can optionally incorporate “adaptive attribute values” (blue solid border blocks, such as the mass and charge of an object).
[0074] Specifically, the IIM module outputs the updated appearance attribute representation, which is input to the decoder module to reconstruct the motion trajectory (X') at the next moment.
[0075] The training objective is to minimize the difference between the reconstructed trajectory (X') and the true trajectory (X), allowing the model to learn the interactions between multiple objects (e.g., collisions, Coulomb forces). Adaptive properties come into play at this stage, enhancing the model's ability to model these interactions.
[0076] exist Figure 3 In the illustrated embodiment, the training method of the IIM module (enlarged image on the right) includes: (1) Attribute extraction: Concatenate the appearance attributes (objecti, objectj) of each object i and j; combine them with the adaptive attribute values (blue dotted border, such as the distance rij between objects).
[0077] (2) Calculation of interactive information: Direction (Dir): Calculates the normalized direction vector (L2 norm) of the relative position between objects.
[0078] Strength (Mag): The ReLU function is used to process the strength of some interaction between objects (for example, the inverse of the distance, or other quantities related to physical laws).
[0079] Attention (Attn): The Sigmoid function is used to calculate the attention weight, which is used to adjust the intensity of the interaction between different objects.
[0080] (3) Attribute update: Multiply the direction vector, intensity scalar, and attention weight to obtain an interaction representation; add this interaction representation (acceleration) to the original velocity attribute to obtain an updated velocity attribute, and use the updated velocity attribute to predict the appearance attribute position offset at the next moment.
[0081] Figure 4 This is one of the diagrams for visualizing Coulomb's law learned by the IWM model provided by the present invention. Figure 4 In the illustrated embodiment, the acceleration characteristics of Objects 1 and 2 are plotted in a polar coordinate system (angle represents the direction of acceleration, and radius represents the magnitude of acceleration). Each data point represents the acceleration of Objects 1 and 2 under a specific "charge correlation" condition. "Charge correlation" ranges from negative to positive values. The acceleration changes of objects under different charge correlation conditions reflect the following physical laws: (1) Coulomb's law: As the charge correlation changes from negative (attraction) to positive (repulsion), the acceleration direction of two objects changes smoothly from moving toward each other (0° to 180°) to moving away from each other (180° to 0° / 360°), reflecting the effect of electrostatic force.
[0082] (2) Action and reaction: Regardless of the charge, the accelerations of the two objects are the same in magnitude but opposite in direction (the data points in the two figures are centrally symmetrical), which conforms to Newton's third law.
[0083] Figure 5 This is the second diagram of the visualization representation of Newton's law learned by the IWM model provided by the present invention. Figure 5 In the illustrated embodiment, a rectangular coordinate system is used for plotting. In the left figure, the horizontal axis represents the velocity of object 1, and the vertical axis represents the acceleration amplitude. In the right figure, the horizontal axis represents the mass of object 1, and the vertical axis represents the acceleration amplitude. Circles represent the acceleration of object 1, and squares represent the acceleration of object 2. This reflects the following physical laws: (1) Newton's Second Law (F=ma): Left: Object 1's velocity increases, indicating an increase in its total momentum before the collision. To conserve momentum, the magnitude of the acceleration of both objects increases after the collision. Right: Key correction: When the mass of object 1 increases, in order to maintain momentum conservation during a collision, the acceleration of object 1 decreases, while the acceleration of object 2 increases. This is because total momentum must remain constant, but the increase in mass of object 1 requires a decrease in its velocity change (acceleration), while the corresponding increase in the velocity change (acceleration) of object 2 must increase. This demonstrates that, given a constant interaction force, acceleration is inversely proportional to mass.
[0084] (2) Newton's Third Law (Action and Reaction): Although the accelerations of the two objects change in different directions (one increasing, the other decreasing), the interaction force between them is always equal in magnitude and opposite in direction. The trend in the right figure reflects this.
[0085] (3) Conservation of momentum: The left figure shows that before a collision, total momentum increases. After a collision, the acceleration of both objects must increase to ensure conservation of total momentum. The right figure shows that during a collision, the mass of one object increases, its acceleration decreases, and the acceleration of the other object increases accordingly, in order to ensure conservation of total momentum.
[0086] The visual reasoning method based on the intuitive world model provided by the present invention is described below. The visual reasoning method based on the intuitive world model described below and the intuitive world model described above can be referenced to each other.
[0087] Figure 6 This is a flow chart of the visual reasoning method based on the intuitive world model provided by the present invention. Figure 6 As shown, the visual reasoning method based on the intuitive world model includes the following steps: Step 610: Obtain visual data to be processed.
[0088] In this step, the visual data to be processed can be obtained from a video database or captured in real time.
[0089] In this embodiment, the visual data to be processed may include static appearance features of the object (such as color, size, shape, and material) and dynamic motion information (such as speed, acceleration, and displacement).
[0090] Step 620: Reasoning the visual data to be processed based on the intuitive world model to obtain a visual reasoning result, and performing a target reasoning task based on the visual reasoning result; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon interpretation, and human-computer interaction.
[0091] In this step, the intuitive world model includes: a search module, an intuitive interaction module and a physical law generation module.
[0092] In this embodiment, the search module is used to perform visual reasoning on the video data of the target scene to obtain latent variables of multiple objects in the target scene. The latent variables are used to represent at least one of the static characteristics, dynamic characteristics and potential physical properties of each object. The potential physical properties include at least one of the mass and charge of the multiple objects.
[0093] In this embodiment, the intuitive interaction module is used to decompose and calculate latent variables using a display modeling method, obtain interaction information between different objects, and update the motion states of multiple objects according to the interaction information.
[0094] In this embodiment, the intuitive interaction module is also used to derive the motion law parameters of each object based on the motion states of multiple objects to perform target reasoning tasks; wherein the target reasoning tasks include at least one of trajectory prediction, physical phenomenon explanation and human-computer interaction.
[0095] It should be noted that the embodiments corresponding to the search module and intuitive interaction module in the above intuitive world model are the same as those in the previous embodiment. Figure 1 The embodiments of each module in the intuitive world model correspond one to one, and will not be repeated in this embodiment.
[0096] In this embodiment, the target scenes include simple object motion scenes, such as scenes of object motion speed, position or trajectory prediction; the target scenes also include complex multi-body interaction scenes, such as collision events of multiple objects, Coulomb force interaction and three-body motion.
[0097] For example, target scenarios include the following scenarios or events: (1) Collision events: used to analyze momentum transfer and energy distribution between objects; (2) Coulomb force interaction scenario: used to describe the attraction or repulsion between charged objects; (3) Multi-object mixed interaction scenarios: including simultaneous collision and Coulomb force interactions; (4) Three-body motion scenario: the Coulomb force interaction motion of three objects.
[0098] In a feasible embodiment, the visual data to be processed is given video data, which contains the motion trajectories of multiple objects (such as collisions, motion under Coulomb force). After the above visual data is input into the intuitive world model, the intuitive world model extracts the appearance features and velocity features of the objects through the search module, and infers the implicit physical properties of the objects (mass difference or charge) through interactive modeling learning. Finally, the search module outputs the implicit physical properties (such as mass, charge) of each object. These properties are stored in the latent space of the model in an interpretable form. The output implicit physical properties can be used for corresponding target reasoning tasks, for example, calculating momentum conservation based on the inferred mass, or predicting the force between charges based on charge.
[0099] In this embodiment, the interaction between objects is modeled based on the implicit physical properties through the intuitive interaction module, the acceleration of the object is calculated, and the impact of changes in physical properties on acceleration is explored through counterfactual analysis. The learned laws are expressed in the form of curves through the physical law generation module, such as the inverse relationship between mass and acceleration (Newton's second law) or the inverse square relationship between charge and force (Coulomb's law). Finally, mathematical relationships or visual charts describing the physical laws are generated, and corresponding target reasoning tasks are performed. For example, the discovery of physical laws can be used to verify whether the model conforms to known physical laws or provide a prediction basis for unknown scenarios.
[0100] In a feasible embodiment, given video data (such as the repulsive motion between two charged objects) is input into the intuitive world model, and the acceleration of the objects is obtained through the search module; the three components of the acceleration between the objects (acceleration attention, direction, and magnitude) are calculated through the intuitive interaction module, and these components are combined into a final acceleration vector; an acceleration analysis diagram and force vector visualization are generated based on the final acceleration to display and explain the current physical event; and the following target reasoning tasks are performed: educational tool demonstration (intuitively displaying physical phenomena), robot interaction (explaining the physical reasons behind the action), or unmanned driving system (explaining why the motion trajectory of an object is predicted), etc.
[0101] The visual reasoning method based on the intuitive world model provided by the embodiment of the present invention uses the intuitive world model to reason on the visual data to be processed to obtain visual reasoning results, thereby improving the accuracy of visual reasoning for complex scenes.
[0102] The visual reasoning method based on the intuitive world model provided by the present invention is described below. The visual reasoning method based on the intuitive world model described below and the intuitive world model described above can be referenced to each other.
[0103] Figure 7 This is a schematic diagram of the structure of the visual reasoning device based on the intuitive world model provided by the present invention. Figure 7 As shown, the visual reasoning device based on the intuitive world model includes: a data acquisition module 710 and a reasoning module 720.
[0104] A data acquisition module 710 is used to acquire visual data to be processed; The visual reasoning module 720 is used to reason about visual data based on the intuitive world model, obtain visual reasoning results, and perform target reasoning tasks based on the visual reasoning results; wherein the target reasoning tasks include at least one of trajectory prediction, physical phenomenon interpretation, and human-computer interaction.
[0105] The visual reasoning device based on the intuitive world model provided by the embodiment of the present invention uses the intuitive world model to reason on the visual data to be processed to obtain visual reasoning results, thereby improving the accuracy of visual reasoning for complex scenes.
[0106] Figure 8 An example of a physical structure diagram of an electronic device is shown below. Figure 8 As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute a visual reasoning method based on an intuitive world model, the method comprising: obtaining visual data to be processed; reasoning on the visual data to be processed based on the intuitive world model to obtain a visual reasoning result; and performing a target reasoning task based on the visual reasoning result; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon interpretation, and human-computer interaction.
[0107] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0108] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the visual reasoning method based on the intuitive world model provided by the above methods, and the method includes: obtaining visual data to be processed; reasoning on the visual data to be processed based on the intuitive world model to obtain visual reasoning results, and performing target reasoning tasks based on the visual reasoning results; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon explanation and human-computer interaction.
[0109] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the visual reasoning method based on the intuitive world model provided by the above-mentioned methods, the method comprising: obtaining visual data to be processed; reasoning on the visual data to be processed based on the intuitive world model to obtain a visual reasoning result, and performing a target reasoning task based on the visual reasoning result; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon explanation and human-computer interaction.
[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0111] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An intuitive world model, characterized by: include: a search module configured to perform visual reasoning on video data of a target scene to obtain latent variables of a plurality of objects in the target scene, wherein the latent variables are used to represent at least one of a static feature, a dynamic feature, and a latent physical property of each object, wherein the latent physical property includes at least one of a mass and a charge of the plurality of objects; An intuitive interaction module is used to decompose and calculate the latent variables using a display modeling method to obtain interaction information between different objects, and update the motion status of the multiple objects based on the interaction information; derive the motion law parameters of each object based on the motion status of the multiple objects to perform a target reasoning task; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon interpretation and human-computer interaction.
2. The intuitive world model according to claim 1, characterized in that The static features include color, texture, shape, and position, the dynamic features include velocity direction and velocity magnitude, and the potential physical properties include mass and charge; The search module includes: A feature extraction module, configured to extract the static features and the dynamic features from the video data; A derivation module is used to deduce the physical properties of each object based on the static features and the dynamic features to obtain the potential physical properties, and adjust the parameters of the intuitive world model based on the potential physical properties.
3. The intuitive world model according to claim 1, characterized in that The latent variables include accelerations of the plurality of objects; The intuitive interaction module is specifically used to decompose the acceleration to obtain acceleration attention, acceleration direction and acceleration magnitude; calculate the target acceleration based on the acceleration attention, acceleration direction and acceleration magnitude, and confirm that the target acceleration is the real-time acceleration of different objects.
4. The intuitive world model according to claim 1, characterized in that The search module is further configured to perform visual reasoning on the video data through an iterative probabilistic reasoning mechanism; The probabilistic reasoning mechanism through iteration is expressed by the following formula: ; in, represents the latent variable, is the input scene video frame, is the generated video frame mean, is the corresponding object mask, Represents the gradient feedback of the object distribution.
5. The intuitive world model according to claim 1, characterized in that Deducing the physical property rules of each object based on the motion states of the multiple objects includes: Inputting the motion states of the multiple objects into a physical law learning model to obtain the motion law parameters; The physical law learning model is obtained by performing unsupervised training on a neural network using sample motion law parameters as training samples and an objective function as a loss function, wherein the objective function is determined based on a reconstruction error and a regularization term.
6. A visual reasoning method based on an intuitive world model, characterized in that: include: Obtain visual data to be processed; Reasoning is performed on the visual data to be processed based on the intuitive world model as described in any one of claims 1 to 5 to obtain a visual reasoning result, and a target reasoning task is performed according to the visual reasoning result; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon interpretation and human-computer interaction.
7. A visual reasoning device based on an intuitive world model, characterized in that: include: A data acquisition module, used to acquire visual data to be processed; A visual reasoning module, configured to reason on the visual data to be processed based on the intuitive world model as described in any one of claims 1 to 5, obtain a visual reasoning result, and perform a target reasoning task based on the visual reasoning result; wherein the target reasoning task includes at least one of trajectory prediction, physical phenomenon interpretation, and human-computer interaction.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the visual reasoning method based on the intuitive world model as claimed in claim 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the visual reasoning method based on the intuitive world model as claimed in claim 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the visual reasoning method based on the intuitive world model as claimed in claim 6 is implemented.