One-dimensional space-time atlas generation method based on comparative predictive coding and subsequent representation

By comparing the combination of predictive coding and subsequent characterization, a one-dimensional spatial time map is directly generated from mouse visual input, solving the problem of limited input data and incomplete computing framework in the existing SR method, and achieving a more accurate spatial navigation simulation and an efficient computing framework.

CN120277349APending Publication Date: 2025-07-08SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510215418.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing SR method mainly uses one-dimensional discrete position sequences as input, making it difficult to simulate the rich sensory information of mice in the natural environment, and lacks an end-to-end computing framework, resulting in the lack of spatial time characterization and the computational model that cannot reflect the navigation mechanism of mice.

Method used

The contrast predictive coding (CPC) model was used to extract temporal dynamic features from the visual input of mice, and combined with successive characterization (SR) to generate a one-dimensional spatial temporal map to build an end-to-end computing framework to directly generate a spatial temporal map from the perceived data.

Benefits of technology

It improves the biological rationality and applicability of the model, can more accurately simulate the spatial navigation behavior of mice, reduces manual intervention, adapts to different experimental conditions, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277349A_ABST
    Figure CN120277349A_ABST
Patent Text Reader

Abstract

The invention provides a one-dimensional space-time atlas generation method based on comparative predictive coding and subsequent representation. The method comprises the following steps: S1, collecting and preprocessing visual input data, actual position data and neural activity data of an experimental mouse; s2, performing time dynamic feature extraction on the preprocessed visual input data through a comparison predictive coding model to generate predictive representation; s3, aligning the generated predictive characterization with the preprocessed actual position data, taking the aligned data as input, performing spatial state estimation through a subsequent characterization model, and generating a subsequent characterization matrix of different positions of the experimental mouse; s4, mapping the subsequent representation matrix to a corresponding neuron activity mode, and generating a one-dimensional space time map of the experimental mouse hippocampal neurons; and S5, comparing the one-dimensional space time atlas with the preprocessed neural activity data to verify the biological rationality of the predictive coding model and the subsequent representation model. According to the method, the space-time atlas is efficiently generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a one-dimensional spatio-temporal map generation method based on contrastive predictive coding and successor representation. Background Art

[0002] Spatial navigation is an important function in the mammalian cognitive system. Especially in the hippocampus of rodents such as mice, there are a large number of neurons related to spatial representation, such as place cells, grid cells, and head direction cells. The traditional cognitive map theory holds that the hippocampus constructs an internal spatial representation through the activity patterns of these neurons, enabling animals to navigate and search for goals in complex environments. However, how to effectively simulate this process in a computational model, especially how to combine biological plausibility and computational efficiency, has always been a core issue in the cross-field of neuroscience and artificial intelligence.

[0003] In recent years, the successor representation (SR) model, as a new state representation method under the reinforcement learning (RL) framework, has provided new ideas for understanding spatial navigation and cognitive maps. The core idea of SR is to construct the representation of a certain state not only limited to the current position, but by predicting the future states that may be reached in this state, so as to provide richer spatial information. The proposal of this method has provided a new perspective for the fields of computational neuroscience and artificial intelligence, enabling researchers to more accurately simulate the navigation strategies of animals and their learning processes in complex environments. However, existing SR methods usually only take a one-dimensional discrete position sequence as input, and it is difficult to simulate the rich sensory information received by mice in natural environments, such as visual input. This limitation of existing methods leads to the lack of spatio-temporal representation, making the computational model unable to truly reflect the navigation mechanism of mice and the formation of cognitive maps.

[0004] The existing technologies have the following disadvantages:

[0005] (1) The input data of existing SR methods is limited, and it is difficult to truly simulate the spatial navigation mechanism of mice.

[0006] In recent years, the successor representation methods mainly use one-dimensional discrete position sequences as inputs. However, during the actual navigation of mice, the information received by their perception systems is much more complex than a single position sequence: Mice mainly rely on multi-modal information such as vision, smell, and touch during movement. Existing SR methods only use discrete position inputs and ignore the role of these perception signals in spatial navigation. Visual input is one of the main means for mice to perceive the environment, but existing methods have not fully utilized visual input for spatial representation, resulting in a large gap between simulation results and real physiological data. Since the successor representation methods are mainly based on discrete state spaces, they cannot effectively capture the navigation patterns of mice in continuous spaces and lack a deep understanding of environmental geometric information.

[0007] (2) Existing methods lack an end-to-end computational framework and it is difficult to generate spatio-temporal maps from perceptual data.

[0008] Current spatial navigation modeling methods usually adopt a step-by-step processing strategy. First, features are extracted from perceptual inputs, and then spatial representation learning is carried out separately. This approach has the following problems: The disconnection between the perceptual layer and the representation layer leads to information loss, making the spatial representation unable to fully utilize perceptual information. Existing computational methods are mainly based on the reinforcement learning framework, and reinforcement learning methods usually require predefined states and rewards, making it difficult to directly learn spatio-temporal maps from perceptual data. Since existing methods do not optimize the entire system end-to-end, the generalization ability of computational models to different experimental conditions is weak.

[0009] (3) The efficiency of extracting spatial information from visual input is low.

[0010] The visual input data of mice usually has the characteristics of high dimensionality and strong temporal correlation. High dimensionality means that a single-frame visual data contains a large amount of pixel information, and directly inputting it into the SR model incurs too high a computational cost. Strong temporal correlation means that the visual input of mice has strong correlation in a short period of time. Therefore, how to extract key temporal information and perform efficient compression is a challenge. Summary of the Invention

[0011] In view of this, the embodiments of the present invention provide a one-dimensional spatio-temporal map generation method based on contrastive predictive coding and successor representation to solve at least the above problems.

[0012] According to the first aspect of the embodiments of the present invention, a one-dimensional spatio-temporal atlas generation method based on contrastive predictive coding and successor representation is provided, including: Step S1: Collect and preprocess the visual input data, actual position data, and neural activity data of experimental mice; Step S2: Extract the temporal dynamic features of the preprocessed visual input data through a contrastive predictive coding model to generate predictive representations; Step S3: Use the generated predictive representations aligned with the preprocessed actual position data as inputs, and perform spatial state estimation through a successor representation model to generate successor representation matrices at different positions of the experimental mice; Step S4: Map the successor representation matrices to corresponding neuron activity patterns to generate a one-dimensional spatio-temporal atlas of the hippocampal neurons of the experimental mice; Step S5: Compare the one-dimensional spatio-temporal atlas with the preprocessed neural activity data to verify the biological rationality of the contrastive predictive coding model and the successor representation model.

[0013] Optionally, the visual input data, actual position data, and neural activity data of the experimental mice are collected in the following manner: When the experimental mice move in the experimental environment, a micro head-mounted camera or a fixed environmental camera is used to record the sequence of visual frames seen by the experimental mice; The change of the position of the experimental mice in the experimental environment over time is recorded through an optical tracking system or a motion trajectory detection method; The firing conditions of the hippocampal place cells of the experimental mice at different times are recorded through optical imaging technology.

[0014] Optionally, the visual input data, actual position data, and neural activity data of the experimental mice are preprocessed in the following manner: Image normalization, color space adjustment, and frame sequence division are performed on the visual input data; Trajectory smoothing and time alignment are performed on the actual position data; Signal denoising and peak detection are performed on the neural activity data.

[0015] Optionally, when training the contrastive predictive coding model, a four-layer convolutional neural network is used to extract the spatial features of each frame of visual input data, a gated recurrent unit is used to learn the long-term dependencies in the time series, and the contrastive loss function is used to optimize the contrastive predictive coding model.

[0016] Optionally, during the training process of the contrastive predictive coding model, the visual input data at each time step is used to predict the representations at multiple future time steps.

[0017] Optionally, after the contrastive predictive coding model is trained, the visual input data at each time step will be converted into a predictive representation, and these predictive representations include the current perceptual information of the experimental mice in the experimental environment and the prediction information of the future state.

[0018] Optionally, the successor representation matrix represents the probability distribution of the states that the experimental mice may visit at different future time steps when at the current position.

[0019] According to the second aspect of the embodiments of the present invention, there is provided a one-dimensional spatio-temporal map generation system based on contrastive predictive coding and successor representation, including: a neural data acquisition module for acquiring neural activity data of experimental mice; a behavior data recording module for acquiring actual position data of experimental mice; a visual data recording module for acquiring visual input data of experimental mice; a contrastive predictive coding module for learning predictive representations of the visual input of experimental mice; a successor representation learning module for generating successor representation matrices at different positions of experimental mice and constructing a one-dimensional spatio-temporal map; a contrastive analysis module for comparing the one-dimensional spatio-temporal map with the preprocessed neural activity data; and a computing platform for storing data of each module, training models, and performing large-scale data processing and computing tasks.

[0020] According to the third aspect of the embodiments of the present invention, there is provided an electronic device including a processor and a memory storing a program. Among them, the program includes instructions that, when executed by the processor, cause the processor to perform the steps performed by the method in the first aspect as described above.

[0021] According to the fourth aspect of the embodiments of the present invention, there is provided a computer storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements the method in the first aspect as described above.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0023] (1) Break through the limitation of the SR method on discrete position input and improve biological rationality.

[0024] Current SR methods usually assume that the state of a mouse is a discrete spatial position, that is, researchers need to manually define the state space first and learn on this state set. The main problem with this method is that it ignores the real perceptual input of the mouse and cannot directly learn spatial representations from the visual information of the mouse. The present invention uses CPC to learn the visual input representation of the mouse and uses it as the state input of the SR model, thus no longer relying on the artificially divided discrete state space, making the modeling more in line with the biological mechanism of the mouse. In the method of the present invention, the visual input of the mouse in the experimental environment is directly used for state representation, avoiding the step of artificially specifying the state space and improving the generalization ability of the method. Since the CPC representation can adapt to different environments and visual inputs, the computational framework of the present invention can be applied to different experimental conditions without redefining the state set for each experiment. This improvement enables the model to more naturally learn the spatial cognitive process of the mouse and enhances its applicability in neuroscience experiments.

[0025] (2) Combine CPC and SR to achieve effective fusion of short-term and long-term spatial information.

[0026] Although existing SR methods can model the long-term state access probability, they lack the ability to model short-term temporal dynamics. For example, traditional SR methods usually assume that state transitions are based on fixed time steps. However, in the actual environment, the movement pattern of mice is highly non-uniform. They may move quickly during some time periods and remain stationary during other time periods, which makes the performance of SR methods poor in short-term prediction. The present invention combines the short-term prediction ability of CPC with the long-term prediction ability of SR to construct a computational framework that can extract short-term sequence information and predict long-term states. CPC performs excellently in time series modeling. It can learn short-term temporal dependencies through GRU and adopt a time step prediction with k = 8, enabling the model to have higher prediction accuracy on short time scales. While the SR method is responsible for learning the long-term state access probability of mice at different positions to ensure that the model can correctly model long-term spatial dynamics. In this way, the present invention makes up for the deficiency of SR methods in short-term temporal dynamics modeling, enabling the model to more accurately simulate the behavior of mice.

[0027] (3) End-to-end computational framework, improving the degree of automation and reducing manual intervention.

[0028] The most advanced existing SR methods usually rely on the reinforcement learning framework for training, which means that researchers need to manually define the state space and design the reward signal to enable the model to perform effective training. A major problem with this method is the over-reliance on the setting of external parameters rather than directly learning spatio-temporal representations from the data. In addition, since reinforcement learning methods usually require a long training time and are sensitive to hyperparameters, their adaptability in different experimental environments is poor. In contrast, the present invention constructs a complete end-to-end computational framework. From the visual input of mice to the generation of the final spatio-temporal map, the whole process is completely data-driven, without the need to manually specify the state space or design the reinforcement learning reward function. CPC can directly learn representations from the visual input of mice, while SR makes long-term state predictions based on these representations, making the whole computational process more natural and efficient. This improvement significantly reduces the need for manual intervention, improves the generality of the method, and enables the computational framework of the present invention to be more easily applicable to different experimental environments. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0030] Figure 1It is a flowchart of the steps of the one-dimensional spatio-temporal map generation method based on contrastive prediction coding and successor representation of the present invention;

[0031] Figure 2 For Figure 1 The corresponding overall flowchart;

[0032] Figure 3 It is a schematic structural diagram of the one-dimensional spatio-temporal map generation system based on contrastive prediction coding and successor representation of the present invention;

[0033] Figure 4 It is an experimental virtual reality environment;

[0034] Figure 5 It is the visualization result of the representation of contrastive prediction coding;

[0035] Figure 6 They are the environmental front-end simulation map and the real map;

[0036] Figure 7 They are the environmental front-end simulation map and the real map;

[0037] Figure 8 They are the environmental teleportation position simulation map and the real map;

[0038] Figure 9 They are the environmental middle-end simulation map and the real map;

[0039] Figure 10 They are the environmental middle-back-end simulation map and the real map;

[0040] Figure 11 They are the environmental middle-back-end simulation map and the real map;

[0041] Figure 12 It is the network structure diagram of the contrastive prediction coding model. Detailed implementation manners

[0042] In order to have a clearer understanding of the technical features, objectives, and effects of the embodiments of the present invention, the specific implementation manners of the embodiments of the present invention are now described with reference to the accompanying drawings.

[0043] In this article, "exemplarily" means "serving as an instance, example, or illustration", and any illustration or implementation manner described as "exemplarily" in this article should not be construed as a more preferred or more advantageous technical solution.

[0044] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present invention.

[0045] For ease of understanding, before describing the specific embodiments of the present invention in detail, an exemplary description of the prior art of the method of the present invention is first given.

[0046] Prior art solution:

[0047] Successor representation is a method based on reinforcement learning. Its core idea is to extend the representation of a state to the states that may be accessed in the future. This means that the representation of a state depends not only on the current position, but is determined by the set of states it may reach in the future. The proposal of this method enables the reinforcement learning system to quickly update the state value when the environment changes without having to retrain completely.

[0048] The basic principle of successor representation is to calculate the probability of each state being accessed in the future and use this as the representation of the state. This method can be regarded as a compromise between model-free learning and model-based learning. It has both the computational efficiency of model-free learning and retains a certain environmental prediction ability, so it performs well in complex navigation tasks.

[0049] In spatial navigation tasks, successor representation is widely used to simulate the neural activities of the hippocampus. It has been found that the firing pattern of place cells not only encodes the current position, but also reflects the possible future trajectories of the animal, which is consistent with the prediction characteristics of the successor representation method. In addition, successor representation can be adaptively adjusted under different environmental conditions, making the spatial representation more flexible.

[0050] The core idea of SR is to define the representation of a state as the weighted sum of its future states. That is, at a certain state s, SRM(s, s′) predicts the probability of accessing state s′ in future time steps. Its mathematical expression is as follows:

[0051]

[0052] where s is the current state, s′ is the state that may be accessed in the future, γ is the discount factor, used to control the influence degree of the future state on the current state, and I(st = s′) is the indicator function, which takes the value of 1 when the state st = s′, and 0 otherwise.

[0053] SR can well model the spatial representation of mice because it can encode long-term spatial relationships, not limited to the current position; it can be quickly updated in a dynamic environment without the need to completely retrain the model; its structure is similar to the firing pattern of hippocampal place cells.

[0054] From the foregoing description, it can be seen that on the one hand, the current SR methods mainly rely on one-dimensional or two-dimensional discrete position inputs, that is, it is assumed that the position of the mouse in the environment is represented by a series of discrete states, and the goal of the model is to learn the successor relationship between different states. However, in a real biological system, a mouse does not navigate solely based on discrete position information, but relies on continuous visual, olfactory, and tactile inputs to form its spatial cognition. Therefore, the existing methods have the following problems in terms of input data: Ignoring perceptual input: Existing SR methods usually assume that the state of the mouse is known and directly use discrete position coordinates as input. This method ignores how the mouse learns spatial representation from perceptual data (such as vision), resulting in simulation results that are difficult to match biological experimental data. Lack of connection between perception and representation: The hippocampal representation of a real mouse is not directly based on discrete positions, but high-dimensional features obtained after being processed by the perceptual system. However, existing SR methods do not consider this process, making it difficult for the model to be applicable to real environments. Difficulty in adapting to different environments: Since the input is pre-defined discrete states, SR methods often need to re-adjust the state space partition in different environments and cannot adapt to different visual scenarios as flexibly as biological systems.

[0055] On the other hand, the current SR methods lack an end-to-end learning framework and are difficult to achieve a complete modeling from perception to representation. Traditional SR methods usually require manually constructing a state space and then learning the successor representation in a reinforcement learning framework. However, a real mouse does not pre-define a state space, but constructs a spatial representation by continuously perceiving environmental information. Therefore, the limitations of the existing methods are mainly reflected in: Manually constructing the state space with low generalization ability: Existing methods usually require researchers to pre-define the state space in advance and learn based on it, making the SR method less adaptable in different experimental environments and difficult to generalize to complex environments. Lack of an end-to-end training mechanism: Since traditional SR methods rely on discrete state modeling, it usually takes multiple steps to extract the spatial representation from perceptual data, while the present invention hopes to achieve automatic modeling from perception to spatial representation through an end-to-end learning framework. Dependence on reinforcement learning rewards and difficulty in simulating exploration behavior: SR methods are usually applied in a reinforcement learning framework and need to be optimized based on reward signals. However, a mouse's spatial learning often does not solely rely on external rewards but constructs spatial cognition through self-exploration. Therefore, the existing methods have poor adaptability in reward-free tasks.

[0056] Therefore, aiming at the deficiencies of existing successor representation methods in terms of input data, temporal dynamics modeling, computational framework integrity, and computational efficiency, the method proposed by the present invention aims to learn spatio-temporal representations from the real visual input of mice and generate neural activity patterns at the cellular level to more accurately simulate the information processing mechanism of the hippocampus during spatial navigation. Based on existing research, the present invention introduces a computational method that is more in line with biological reality, enabling the model to not only perform feature learning at the perceptual level but also perform efficient reasoning at the spatial cognition level while optimizing the computational cost to meet the requirements of large-scale data processing and real experimental environments.

[0057] a. Construct an end-to-end computational framework from perception to representation

[0058] The primary objective of the present invention is to break through the dependence of existing successor representation methods on discrete position inputs and achieve end-to-end learning from the visual input of mice to spatial representations. Compared with the traditional SR method of manually constructing the state space, the present invention directly extracts spatio-temporal features from the continuous visual perception data of mice, avoiding the artificial division of the state space and making the model more widely applicable.

[0059] Introduce CPC for temporal feature extraction of visual input: Since the visual perception of mice is temporally correlated, the present invention uses Contrastive Predictive Coding (CPC) to process visual input, enabling the model to learn the environmental perception features of mice at different time steps. This method can not only reduce the redundant information in visual data but also improve the model's adaptability to environmental dynamic changes.

[0060] Combine SR for long-term spatial prediction: CPC can extract short-term temporal features, while SR is responsible for learning long-term spatial representations. The present invention combines CPC and SR, enabling the model to not only learn the current environmental information but also predict future spatial states, thus forming a spatio-temporal map that is more in line with biological mechanisms.

[0061] Avoid manually defining the state space: Existing SR methods usually require researchers to manually construct the state space and perform learning on this state set. However, in real biological systems, animals do not pre-divide the state space but obtain spatial information through continuous learning of the perceptual system. The end-to-end method of the present invention enables the model to directly construct spatial representations from perceptual data without human intervention, improving the generalization ability.

[0062] b. Improve the temporal dynamics modeling ability and adapt to non-uniform temporal changes

[0063] During the spatial navigation of mice, the state transition pattern between different time steps may be non-uniform. For example, a mouse may move faster in a familiar environment, while in an unfamiliar environment, it may slow down or even stop to better collect environmental information. Traditional SR methods assume that state transitions are based on fixed time steps, which does not hold true in real biological systems, making it difficult to accurately simulate the navigation behavior of mice.

[0064] Another core objective of the present invention is to improve the time dynamics modeling so that the model can adapt to the spatial exploration behavior of mice at different time scales. Specifically, the present invention proposes the following improvement solutions:

[0065] Introduce an adaptive time modeling mechanism: The present invention adopts a dynamic discount factor adjustment method to dynamically adjust the time weight according to the behavior pattern of mice, enabling the model to flexibly switch between short-term prediction and long-term prediction. For example, when a mouse accelerates exploration, the model can adopt a larger time step to improve prediction efficiency; while when a mouse stops to observe, the model can reduce the time step to improve the prediction accuracy of short-term behavior.

[0066] Combine CPC to strengthen time-dependent modeling: CPC can learn time-related features, enabling the model to make more flexible predictions in the time dimension. The present invention uses CPC to extract the perceptual features of mice at different time steps and combines the long-term prediction ability of SR to achieve an organic combination of short-term behavior and long-term goals.

[0067] c. Optimize the computational efficiency and improve the adaptability of the model to large-scale data

[0068] Although SR performs well in spatial navigation tasks, its computational complexity is relatively high. Especially in large-scale environments, the growth of the number of states will lead to a sharp increase in the storage and computational costs of the SR matrix. The computational costs of traditional SR methods mainly come from the high-dimensional problems of matrix calculations, such as matrix inversion and the storage of state transition matrices. These problems will significantly reduce the computational efficiency in complex environments.

[0069] Another important objective of the present invention is to optimize the computational cost so that the SR method can operate efficiently in larger-scale environments. To this end, the present invention proposes the following optimization solutions:

[0070] The present invention adopts a matrix decomposition method to decompose the SR matrix into a low-rank approximation matrix to reduce the computational overhead. For example, in the case of a large visual state space, CPC can be used to reduce the dimension of the state representation, thereby reducing the storage requirements of the SR matrix.

[0071] Through this improvement, the present invention can more realistically simulate the spatial cognitive process of mice, making the computational model more in line with biological laws and reducing the dependence on data preprocessing, thereby improving the application value of the model.

[0072] The following further illustrates the specific implementation of the embodiments of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention.

[0073] See Figure 1 and Figure 2 The one-dimensional spatio-temporal map generation method based on contrastive predictive coding and successor representation provided by the present invention mainly includes the following steps:

[0074] Step S1: Collect and preprocess the visual input data, actual position data, and neural activity data of the experimental mice;

[0075] Step S2: Extract the temporal dynamic features of the preprocessed visual input data through a contrastive predictive coding model to generate a predictive representation;

[0076] Step S3: Align the generated predictive representation with the preprocessed actual position data and use it as input, and perform spatial state estimation through a successor representation model to generate a successor representation matrix at different positions of the experimental mice;

[0077] Step S4: Map the successor representation matrix to the corresponding neuron activity pattern to generate a one-dimensional spatio-temporal map of the hippocampal neurons of the experimental mice;

[0078] Step S5: Compare the one-dimensional spatio-temporal map with the preprocessed neural activity data to verify the biological rationality of the contrastive predictive coding model and the successor representation model.

[0079] It should be understood that the one-dimensional spatio-temporal map predicts the neuron activity patterns of the experimental mice at different positions.

[0080] Optionally, the visual input data, actual position data, and neural activity data of the experimental mice are collected in the following ways: when the experimental mice move in the experimental environment, record the visual frame sequence seen by the experimental mice through a miniature head-mounted camera or a fixed environmental camera; record the change of the position of the experimental mice in the experimental environment over time through an optical tracking system or a motion trajectory detection method; record the firing conditions of the hippocampal place cells of the experimental mice at different times through an optical imaging technique.

[0081] Optionally, the visual input data, actual position data, and neural activity data of the experimental mice are preprocessed in the following ways: perform image normalization, color space adjustment, and frame sequence division on the visual input data; perform trajectory smoothing and time alignment on the actual position data; perform signal denoising and peak detection on the neural activity data.

[0082] Optionally, when training the contrastive prediction coding model, a four-layer convolutional neural network is used to extract the spatial features of each frame of visual input data, a gated recurrent unit is used to learn the long-term dependencies in the time series, and the contrastive prediction coding model is optimized through a contrastive loss function.

[0083] Optionally, during the training process of the contrastive prediction coding model, the visual input data at each time step is used to predict the representations at multiple future time steps.

[0084] Optionally, after the contrastive prediction coding model is trained, the visual input data at each time step is converted into a predictive representation, which includes the current perceptual information of the experimental mouse in the experimental environment and the predictive information about the future state.

[0085] Optionally, the successor representation matrix represents the probability distribution of the states that the experimental mouse may visit at different future time steps when it is at the current position.

[0086] Specifically, the solution of the present invention is further described according to the following examples:

[0087] (1) Data collection and preprocessing

[0088] In real neuroscience experiments, mice are usually placed in a specific experimental environment (such as a runway, maze, or free exploration scenario), and researchers record the perceptual input and neural activity information of the mice through a multimodal data acquisition system. The data of the present invention mainly includes the following sources:

[0089] a. Visual input, i.e., visual input data: When the experimental mouse moves in the environment, a sequence of visual frames seen by the mouse is recorded through a miniature head-mounted camera (or a fixed environmental camera). These visual frames reflect the visual stimuli received by the mouse at different positions.

[0090] b. Actual position of the mouse, i.e., actual position data: The change in the position of the mouse in the experimental environment over time is recorded through an optical tracking system (such as DeepLabCut) or other motion trajectory detection methods.

[0091] c. Calcium signal data, i.e., neural activity data: Using optical imaging techniques (such as two-photon imaging or a miniature calcium imaging probe), the firing of hippocampal place cells at different times is recorded. These signals represent the activity patterns of neurons and are the experimental data finally used for comparison.

[0092] Since the original experimental data has a high dimension and complexity, in order to ensure the stability and efficiency of the computational model, it is necessary to preprocess the data. The specific steps are as follows:

[0093] a. Visual input data preprocessing

[0094] Image normalization: Unify the size of visual frames to ensure that data at different time steps has the same input format.

[0095] Color space adjustment: Since mice mainly rely on light and dark contrast, grayscale images can be used in the experiment to reduce computational complexity.

[0096] Frame sequence division: Convert video frames into a time series at a fixed time step to adapt to the input format of CPC.

[0097] b. Actual position data processing

[0098] Trajectory smoothing: Use Kalman filtering or low-pass filtering to smooth the movement trajectory of mice to reduce the influence of noise.

[0099] Time alignment: Ensure that visual frames, position trajectories, and neural data are completely aligned on the time axis.

[0100] c. Calcium signal data preprocessing

[0101] Signal denoising: Use DF / F calculation to standardize the calcium signal and remove baseline drift.

[0102] Peak detection: Extract neuron firing events for comparison with predicted neural activities.

[0103] (2) Comparison prediction coding model training

[0104] See Figure 12 , after the data preprocessing is completed, the present invention uses a contrastive predictive coding (CPC) model to learn the time-related representations of visual inputs. CPC is a self-supervised learning method that can extract predictive information from time series data without manual annotation. Through the learning process of CPC, the model can capture the temporal dynamic features in the spatial environment from the visual inputs of mice, enabling the subsequent SR model to perform spatio-temporal predictions based on these representations.

[0105] In the CPC model, the feature extraction of visual inputs consists of a four-layer convolutional neural network (CNN), which is used to extract local spatial features and convert the visual input of each frame into a low-dimensional representation. After being processed by the CNN, the feature vector of each frame is input into a gated recurrent unit (GRU) to learn the long-term dependencies of the time series. GRU can integrate historical information over multiple time steps, enabling the model to predict future states based on past visual inputs.

[0106] During the training process of CPC, the visual input at each time step is used to predict the representations for the next 8 time steps (k = 8). To optimize the model, CPC adopts the contrastive loss (InfoNCE Loss). By maximizing the similarity of the correctly predicted representations and minimizing the similarity of the wrongly predicted representations, the model can learn to extract useful temporal dynamic features from past information. After the training of the CPC model is completed, the visual input at each time step is converted into a temporally related representation. These representations not only contain the current visual features but also the prediction information about future states.

[0107] It should be understood that setting k to k = 8 in the experimental verification of the present invention's solution is only for illustrative purposes and not a limitation to the present invention's solution. In practical applications, there is no restriction on k, that is, the number of time steps is multiple.

[0108] The present invention does not impose a rigid limit on the dimension of the CPC representation, but it is recommended that its dimension be proportional to the number of time steps of the neural activity recorded in the experiment, so that in the final generation process of the spatio-temporal map, it can be directly compared with the experimental data. This representation method ensures the biological rationality of the computational model and enables it to better match the firing patterns of real neurons.

[0109] (3) Conversion of visual input to predictive representation

[0110] After the CPC training is completed, the visual input at each time step is converted into a high-dimensional representation vector, that is, the predictive representation. These representations not only encode the current perceptual information of the mouse in the environment but also contain the prediction of future states. Next, these representations learned by CPC will be used as the state input of the SR model to model the long-term spatial cognition of the mouse.

[0111] In this process, the present invention aligns the CPC representation with the actual position information of the mouse, enabling the SR model to make predictions based on these temporally related representations. At the same time, the temporal continuity of the CPC representation enables SR to learn more stable state transition relationships, thereby improving the quality of the spatio-temporal map generation. Through this conversion process, the perceptual features extracted by CPC are successfully mapped to the spatial exploration state of the mouse, providing input data for the training of the SR model.

[0112] (4) The successor representation model learns the one-dimensional spatio-temporal map of cells

[0113] In the training stage of the SR model, the goal of the present invention is to establish a successor representation matrix based on the temporally related representations learned by CPC, for predicting the future states that the mouse may visit at different positions. The core idea of the SR model is to use the representation of the current state to predict the possible future positions and construct the long-term access relationships between states.

[0114] During the training process, the SR model learns the state transition matrix of the CPC representation, such that the state at each time step contains not only the current spatial information but also encodes possible future trajectories. As training progresses, SR is able to generate a prediction matrix for each position, which represents the probability distribution of the states that the mouse may visit at different future time steps when at the current position.

[0115] The key innovation of the present invention lies in mapping each state of the SR prediction matrix to the corresponding neuronal activity, thereby generating a one-dimensional spatio-temporal map of the mouse hippocampal neurons. Since the number of time steps of the CPC representation is proportional to the number of time steps of the neural activity recorded in the experiment, the finally generated spatio-temporal map can be compared with the real calcium signal data to verify its accuracy.

[0116] During the verification process of the simulation results, the present invention compares the neuronal activity patterns generated by SR with the calcium imaging data recorded in the experiment, and analyzes whether their firing patterns are consistent at different positions. Through this comparison process, the present invention can quantitatively calculate the biological plausibility of the computational model and optimize the learning strategy of SR to make it more in line with the spatial information processing mechanism of the hippocampus.

[0117] See Figure 3 , the embodiment of the present invention also provides a one-dimensional spatio-temporal map generation system based on contrastive predictive coding and successor representation. The system includes multiple data acquisition, representation learning, spatial prediction, and contrastive analysis modules, constructing a complete computational framework to realize the whole process from experimental data acquisition to spatio-temporal map generation. The system includes the following core modules:

[0118] A neural data acquisition module for acquiring neural activity data of experimental mice;

[0119] A behavioral data recording module for acquiring actual position data of experimental mice;

[0120] A visual data recording module for acquiring visual input data of experimental mice;

[0121] A contrastive predictive coding module for learning the predictive representation of the visual input of experimental mice;

[0122] A successor representation learning module for generating a successor representation matrix at different positions of experimental mice and constructing a one-dimensional spatio-temporal map;

[0123] A contrastive analysis module for comparing the one-dimensional spatio-temporal map with the preprocessed neural activity data;

[0124] A computing platform (offline planning) for storing data of each module, training the contrastive predictive coding model and the successor representation model, and performing large-scale data processing and computing tasks.

[0125] It should be understood that the one-dimensional spatio-temporal map generation system based on contrastive predictive coding and successor representation in the embodiments of the present invention is used to implement the corresponding methods in the foregoing multiple method embodiments and has the beneficial effects of the corresponding method embodiments.

[0126] This system has end-to-end computing capabilities and can directly generate spatio-temporal maps from the visual input of mice without the need for manually defined state spaces, improving the biological rationality and adaptability of the computational model.

[0127] The present invention is verified through a real virtual reality environment, demonstrating its feasibility, and the experimental results show that this method has significant effectiveness and practicality. The following are the specific experimental and simulation results:

[0128] See Figure 4 , in this experiment, a virtual reality environment with a total length of 100 cm was used, and mice were allowed to run 88 laps in this environment. In the first 60 laps at the beginning of the experiment, the mice ran along a normal trajectory, and their position changes in this environment fully conform to the continuity hypothesis. After the 60th lap, a teleportation point was set in the experiment, that is, when the mouse reached the 10 cm position, the system would force it to be instantly teleported to the 35 cm position without passing through the intermediate spatial area. The core purpose of this design is to study whether there will be significant changes in the firing patterns of the place cells and spatio-temporal representations of the mice after experiencing spatial teleportation.

[0129] During the experiment, using the method proposed in the present invention, the position representations of mice in the normal movement mode (without teleportation) and the teleportation mode (with teleportation) were respectively learned, and the neural coding features of the two modes were compared through a neural network model. The experimental results show that in the teleportation mode, there is an obvious gap in the spatio-temporal representation of the mice, which is consistent with the teleportation setting of the experiment, as Figure 5 shown. Specifically, in the normal mode, the representations of the place cells show a continuous distribution and can evenly cover the entire trajectory area; but in the teleportation mode, the representation information between 10 cm and 35 cm is completely missing, indicating that there is an obvious break in the spatial cognition of the mice in this area, and this phenomenon is consistent with the lack of environmental perception caused by the teleportation mechanism.

[0130] Further analyze the one-dimensional spatio-temporal map of the mice learned using the Successor Representation (SR) model and compare it with the real calcium signals recorded in the experiment, as Figures 6 to 11As shown. The experimental results show that the place cells encoding this position in the front side (10 cm) of the teleportation point also showed spiking activities near the teleportation termination point (35 cm). This indicates that after the mouse experienced teleportation, its brain still retained the neural activities of the position before teleportation and wrongly generated spiking at the new position (near 35 cm). This expansibility of neural activities is a cognitive compensation mechanism of the mouse when dealing with discontinuous spatial jumps. In addition to the spiking phenomenon at the teleportation termination point, the spiking patterns of cells at other positions also changed. Specifically, first, the discharge centers of some place cells shifted: for the cells encoding certain positions, their spiking activities moved backward by a certain distance, indicating a dislocation in the mouse's overall cognition of the environment. Second, new spiking regions emerged: some neurons that did not spike in the normal mode showed new spiking activities in the teleportation mode, indicating that teleportation not only affects local spatial encoding but may also induce a wider range of neural adaptive adjustments.

[0131] These results suggest that after the mouse experienced teleportation, the place cell network in its hippocampus was reorganized, resulting in a shift in its cognition of the spatial environment. This phenomenon is consistent with existing neuroscience research, that is, when the hippocampus processes discontinuous spatial jumps, it will rely on existing spatial memories for compensation and readjust the spatial representation at the new position.

[0132] Through comparative analysis, it is found that the method proposed in the present invention can accurately learn the changes in the spatio-temporal maps of hippocampal place cells before and after the mouse experienced teleportation, and these changes are highly consistent with the real neural data. This indicates that this method has high reliability and biological rationality in studying the spatial cognitive mechanism of mice, providing a powerful tool for further studying the information processing mechanism of the hippocampus in discontinuous spatial environments.

[0133] In addition, it should be noted that the following are the alternative technical means for some parts of the present invention:

[0134] a. Using different representation learning methods to replace contrastive predictive coding (CPC)

[0135] Method based on variational autoencoder (VAE): If more precise control of the representation distribution is required, VAE can be used to replace CPC to perform low-dimensional encoding of the mouse's visual input, while ensuring good distribution continuity of the representation to enhance the generalization ability of the model to different environments.

[0136] Method based on autoregressive Transformer: In prediction tasks on a longer time scale, Transformer can be used to replace GRU to model the time-dependent structure of the mouse with a stronger attention mechanism and improve the ability to capture predictive representations.

[0137] Method based on Generative Adversarial Network (GAN): If further optimization of the separability of the representation is required, GAN can be used to learn predictive representations, making the feature distributions of different states more distinct and enhancing the interpretability of the model.

[0138] b. Replace the successor representation (SR) with different spatial representation methods

[0139] Method based on Markov Decision Process (MDP): If it is necessary to model the path selection of mice in a reinforcement learning environment, MDP can be used for modeling, and the state representation can be optimized through value iteration or policy gradient methods to make it more suitable for dynamic environment prediction tasks.

[0140] Method based on Graph Neural Network (GNN): In a more complex spatial structure, GNN can be used to learn the transition relationships between states to enhance the model's adaptability to irregular environments and be applicable to a wider range of neural navigation tasks.

[0141] c. Optimize the model structure using different computational frameworks

[0142] State aggregation optimization strategy: In the case of limited computing resources, K-means clustering can be used to automatically compress the state space of SR to reduce computational overhead and improve computational efficiency on large-scale experimental data.

[0143] It should be understood that the core technology of the present invention is not only applicable to mouse spatial cognition modeling, but can also be extended to multiple research and application fields, such as neuroscience experimental design, reinforcement learning, robot navigation, cognitive computing, medical informatics, etc. Specifically:

[0144] a. Neuroscience and brain function modeling

[0145] The present invention can be used to study the spatial cognitive function of the hippocampus and extended to the representation learning of other brain regions, such as the information encoding of the prefrontal cortex during the decision-making process. In addition, this method can be applied to Brain-Computer Interface (BCI) for analyzing neural activity patterns and optimizing brain signal decoding techniques.

[0146] b. Reinforcement learning and agent path optimization

[0147] The method of the present invention can be directly applied to the state representation in reinforcement learning, optimizing the state modeling of the reinforcement learning algorithm through the CPC-SR framework, enabling the agent to more efficiently learn the long-term dynamics of the environment. In robot path planning, agent policy learning, and autonomous driving decision-making, this method can be used to enhance the state prediction ability and improve the navigation efficiency of the agent.

[0148] c. Robot navigation and autonomous driving

[0149] In the fields of drone navigation, autonomous vehicle path planning, and intelligent warehousing systems, the method of the present invention can be used to learn the spatio-temporal map of the environment, improving the real-time performance and robustness of path planning. Combined with multi-agent cooperation algorithms, this method can be used for multi-robot scheduling, enhancing the efficiency of tasks such as logistics distribution and factory automation.

[0150] d. Cognitive Computing and Intelligent Prediction

[0151] This method can be used for cognitive computing tasks such as user behavior prediction, market trend analysis, and financial modeling. By learning the temporal representations of long-term dependencies, it optimizes the prediction performance. For example, in an intelligent recommendation system, this method can be used to optimize user interest prediction, improving the personalization of content recommendations.

[0152] e. Medical Informatics and Brain Disease Research

[0153] The present invention can be used for modeling neurodegenerative diseases such as Alzheimer's disease (AD) or Parkinson's disease (PD), analyzing the functional changes in the patient's hippocampus, and constructing a computational model of cognitive decline. In addition, this method can be applied to brain-computer interfaces (BCIs) to improve the accuracy of neural signal decoding, providing technical support for brain disease rehabilitation and neuromodulation.

[0154] f. Logistics and Supply Chain Management

[0155] In tasks such as drone delivery, intelligent warehousing, and supply chain optimization, the present invention can be used to learn the environmental topology and optimize path planning based on SR, improving the operating efficiency of the logistics system. For example, in city-level delivery tasks, this method can be used to optimize the paths of delivery drones, enhancing the efficiency and reliability of express delivery.

[0156] The ultimate goal of the present invention is to compare the one-dimensional spatio-temporal map generated by SR with the real calcium signal data recorded in experiments to evaluate the fitness of the computational model for mouse neural activity. This comparison process can not only verify whether the computational method of the present invention can effectively simulate the spatial information processing mechanism of the hippocampus but also provide computational prediction support for future experimental designs, enabling neuroscience research to more precisely understand the neural representation process of mice when exploring the environment.

[0157] In summary, the present invention has significant advantages over traditional methods. First, it breaks through the limitation that the SR method only uses discrete position inputs and directly uses the visual input of mice for learning, making the biological rationality of the model stronger. Second, by combining CPC and SR, the model can learn short-term and long-term spatio-temporal dependencies simultaneously, ensuring that the generated neural activity patterns are more in line with the navigation behavior of mice in real environments. In addition, the present invention constructs a complete end-to-end computing framework, avoiding the need for manual construction of the state space, and by optimizing the computing efficiency, the model can be applied to larger-scale datasets and prediction tasks with longer time scales.

[0158] As another example, an embodiment of the present invention also provides an electronic device, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0159] The electronic device may include: a processor, a communications interface, a memory, and a communication bus.

[0160] The processor, the communications interface, and the memory communicate with each other through the communication bus. The communications interface is used to communicate with other electronic devices or servers.

[0161] The processor is used to execute a program, and specifically can execute the relevant steps in the above method embodiments.

[0162] Specifically, the program may include program code, and the program code includes computer operation instructions.

[0163] The processor may be a processor CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0164] A memory for storing programs. The memory may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.

[0165] When executed by a processor, the program is used to cause an electronic device to execute the one-dimensional spatio-temporal map generation method based on contrastive prediction coding and subsequent representations of the present invention.

[0166] In addition, for the specific implementation of each step in the program, reference may be made to the corresponding steps and descriptions in the corresponding units in the above method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0167] An exemplary embodiment of the present invention also provides a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, the methods of the embodiments of the present invention are implemented. Reference may be made to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0168] The methods according to the embodiments of the present invention described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and stored in a local recording medium, so that the methods described herein can be processed by such software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0169] So far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.

[0170] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0171] Finally, it should be noted that the above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant art can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.

Claims

1. A one-dimensional spatio-temporal map generation method based on contrastive predictive coding and successor representation, characterized in that Including: Step S1: Collect and preprocess the visual input data, actual position data, and neural activity data of experimental mice; Step S2: Extract temporal dynamic features from the preprocessed visual input data through a contrastive predictive coding model to generate predictive representations; Step S3: Align the generated predictive representations with the preprocessed actual position data and use them as inputs to perform spatial state estimation through a successor representation model to generate successor representation matrices at different positions of the experimental mice; Step S4: Map the successor representation matrices to corresponding neuron activity patterns to generate a one-dimensional spatio-temporal map of the hippocampal neurons of the experimental mice; Step S5: Compare the one-dimensional spatio-temporal map with the preprocessed neural activity data to verify the biological rationality of the contrastive predictive coding model and the successor representation model.

2. The method according to claim 1, wherein Collect the visual input data, actual position data, and neural activity data of experimental mice in the following ways: When the experimental mice move in the experimental environment, record the visual frame sequence seen by the experimental mice through a miniature head-mounted camera or a fixed environmental camera; Record the change of the position of the experimental mice in the experimental environment over time through an optical tracking system or a motion trajectory detection method; Record the firing conditions of the hippocampal place cells of the experimental mice at different times through optical imaging technology.

3. The method according to claim 1, characterized in that, Preprocess the visual input data, actual position data, and neural activity data of experimental mice in the following ways: Perform image normalization, color space adjustment, and frame sequence division on the visual input data; Perform trajectory smoothing and time alignment on the actual position data; Perform signal denoising and peak detection on the neural activity data.

4. The method according to claim 1, wherein When training the contrastive predictive coding model, a four-layer convolutional neural network is used to extract the spatial features of each frame of visual input data, a gated recurrent unit is used to learn the long-term dependencies in the time series, and the contrastive predictive coding model is optimized through a contrastive loss function.

5. The method according to claim 4, characterized in that, During the training process of the contrastive predictive coding model, the visual input data at each time step is used to predict the representations at multiple future time steps.

6. The method according to claim 5, wherein After the contrastive predictive coding model is trained, the visual input data at each time step will be converted into a predictive representation, and these predictive representations include the current perception information of the experimental mice in the experimental environment and the prediction information of the future state.

7. The method according to claim 6, wherein The successor representation matrix represents the probability distribution of the states that the experimental mice may visit at different future time steps when at the current position.

8. A one-dimensional spatio-temporal atlas generation system based on contrastive predictive coding and successor representation, characterized in that Including: A neural data acquisition module for collecting the neural activity data of experimental mice; A behavior data recording module for collecting the actual position data of experimental mice; A visual data recording module for collecting the visual input data of experimental mice; A contrastive predictive coding module for learning the predictive representation of the visual input of experimental mice; A successor representation learning module for generating successor representation matrices at different positions of experimental mice and constructing a one-dimensional spatio-temporal map; A contrastive analysis module for comparing the one-dimensional spatio-temporal map with the preprocessed neural activity data; A computing platform for storing the data of each module, training models, and performing large-scale data processing and computing tasks.

9. An electronic device, characterized in that, Including: A processor; A memory for storing programs; Among them, the program includes instructions that, when executed by the processor, cause the processor to perform the steps performed by the method according to any one of claims 1-7.

10. A computer storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1-7 is implemented.