Device for detecting behavior of experimental animal after taking medicine based on multi-modal fusion
Through a multimodal fusion experimental animal behavior detection device, multi-spectral cameras and thin-film pressure sensors combined with machine learning and large language models, the shortcomings of traditional monitoring methods are solved, automated and continuous behavior monitoring and report generation are achieved, and data reliability and efficiency of drug research are improved.
Patent Information
- Application Number
- CN202510476482.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional experimental animal behavior monitoring methods rely on manual observation, which have problems such as strong subjectivity, low efficiency, difficulty in achieving 24-hour uninterrupted monitoring and poor environmental adaptability, resulting in low data reliability and efficiency.
Using a multimodal fusion behavior detection device, data is collected simultaneously through a multispectral camera and a thin film pressure sensor, combined with machine learning and large language models, automation, continuous monitoring and experimental reports are achieved.
It realizes high-precision and automated detection of experimental animal behaviors, generates reliable experimental reports, and significantly improves the data reliability and efficiency of drug research.
Smart Images

Figure CN120336985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of animal experiment behavior detection equipment, and particularly to a behavior detection device for experimental animals after taking medicine based on multi-modal fusion, which is used to detect the behavior changes of experimental animals after taking medicine. Background Art
[0002] In the fields of drug R & D process and toxicology research, mice, rats, etc. are extremely commonly used research models. Accurately analyzing the behavior changes of experimental animals after taking medicine plays a decisive role in accurately evaluating the efficacy and safety of drugs. However, most traditional experimental animal behavior monitoring methods rely on direct manual observation or simple recording devices, and these methods have significant drawbacks. For example, they are highly subjective, and manual interpretation is prone to introducing individual experience differences, resulting in systematic data deviations; there is insufficient continuity, and it is difficult to achieve 24-hour uninterrupted monitoring, which may miss key behavior patterns; the environmental adaptability is poor, and it is difficult to clearly capture behavior details under night or low-light conditions. In view of this, developing an automated, multi-modal fusion behavior monitoring system that can continuously and objectively record the movement characteristics (such as standing, running, lying on the side) of experimental animals after taking medicine, environmental interaction behaviors, and generate standardized experimental reports based on large language models will significantly improve the data reliability and experimental efficiency of drug research, and provide a quantifiable and traceable scientific basis for pharmacological mechanism analysis. Summary of the Invention
[0003] The purpose of the present invention is to provide a behavior detection device for experimental animals after taking medicine based on multi-modal fusion. By integrating the data collected by thin-film pressure sensors and multi-spectral cameras, and using the powerful data analysis ability of language models, high-precision, automated detection and in-depth analysis of the behavior of experimental animals after taking medicine are realized, providing a reliable basis for drug research.
[0004] To achieve the above purpose, the present invention provides the following technical solution: A behavior detection device for experimental animals after taking medicine based on multi-modal fusion.
[0005] Preferably, it includes the following modules: (1) Data acquisition module, configured to synchronously acquire the initial images of the behavior of animals after taking medicine and the thin-film pressure sensor data, including: (1a) The initial images are captured by a multi-spectral camera deployed at the top of the experiment, including visible light and infrared bands, and can adapt to the shooting of animal behaviors under different lighting conditions; (1b) The thin-film pressure sensor data is generated by a thin-film pressure sensor array embedded at the bottom of the experimental cage. The thin-film pressure sensor array covers no less than 90% of the bottom area of the experimental cage, the sampling frequency is not less than 100Hz, and the sensor resolution is 0.1g / cm²; (1c) The data acquisition module aligns the acquisition times of images and sensor data through the IEEE 1588 Precision Time Protocol (PTP), with a time error of no more than ±5 ms, achieving multi-modal data synchronization at the hardware level; (2) A data processing module, configured to preprocess and extract features from the initial image and thin-film pressure sensor data, including: (2a) Perform background subtraction and moving target segmentation on the initial image to extract the contour, skeletal key points, and movement trajectory of the experimental animal. Among them, the skeletal key points are identified by the OpenPose model with transfer learning, and the contour is obtained through morphological opening and closing operations, with an identification accuracy of ≥90%; (2b) Perform filtering and normalization processing on the thin-film pressure sensor data to extract the contact area, pressure peak, and spatial distribution entropy value. The entropy value is calculated through the pressure distribution probability; (3) A behavior determination module, configured to classify the preprocessed multi-modal data, including: (3a) Divide the contour, movement trajectory, contact area, and spatial distribution entropy value into a training set and a test set; (3b) Train a classification model based on the training set samples. The classification model dynamically fuses multi-modal feature vectors through an attention mechanism, and the output is a class label; (3c) Use the classification model to determine the behavior of the test set samples and record the behavior data aligned with the time stamp; (4) A report generation module, configured to convert the behavior data into structured text and generate an experimental report through a large language model, including: (4a) Encode the time series data into a JSON structured input containing drug dosage, behavior category, and pressure distribution; (4b) Invoke a pre-trained LLM based on the Transformer architecture to parse the input. The LLM injects parameters in the field of experimental animal behavior analysis through the Low-Rank Adaptation (LoRA) technique to generate a natural language report containing statistical charts, anomaly warnings, and speculation on pharmacological mechanisms.
[0006] Preferably, the thin-film pressure sensor array adopts a cross-electrode design, with a unit pitch of 2 mm × 2 mm and a pressure response linearity error of ≤1.5%.
[0007] Preferably, the attention mechanism adopts a Transformer encoder structure, and the dynamic weight allocation ratio of visual features to pressure features is adjustable from 7:3 to 3:7.
[0008] Preferably, the JSON structured input includes: (1) Experimental information fields: drug dosage (unit: mg / kg), timestamp (format: YYYY-MM-DDTHH:mm:ss); (2) Behavioral data fields: behavior category, pressure distribution array.
[0009] Preferably, the training process of the machine learning classifier includes: (1) Using the radial basis function (RBF) as the kernel function; (2) Optimizing the hyperparameter combination through grid search, where the hyperparameters include the penalty coefficient C and the kernel function parameter γ.
[0010] Preferably, the large language model is a pre-trained model based on the Transformer architecture, and through domain adaptation fine-tuning, including: (1) Using the experimental animal ethology experiment report text and multimodal data to construct a fine-tuning dataset; (2) Injecting adaptation parameters using the low-rank adaptation (LoRA) technique.
[0011] In the above technical solution, a behavior detection device for experimental animals after taking medicine based on multimodal fusion provided by the present invention has the following beneficial effects: It can automatically, continuously and objectively monitor the behavioral changes of mice after taking medicine, and can generate experiment reports in a timely manner. This not only effectively avoids the defects of traditional methods, but also greatly improves the experiment efficiency and data accuracy, provides more reliable and efficient technical support for the research on the behavior of mice after taking medicine, and helps to promote the in-depth development of related scientific research work. Description of the Drawings
[0012] Figure 1 It is the initial action image acquisition device for experimental animals of the present invention;
[0013] Figure 2 It is the flowchart of experimental animal behavior image acquisition;
[0014] Figure 3 It is the flowchart of experimental animal pressure data acquisition;
[0015] Figure 4 It is the flowchart of extracting the body posture characteristics of experimental animals;
[0016] Figure 5 It is the flowchart of behavior recognition for the test set. Detailed Embodiments
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Unless otherwise specified, the technical means used in the implementation examples are conventional means well known to those skilled in the art.
[0019] There are many deficiencies in the behavior detection methods of experimental animals after taking medicine. For example, the monitoring means are single, making it difficult to comprehensively obtain the behavior information of experimental animals; the data processing lacks depth and accuracy, and it is impossible to effectively analyze the impact of drugs on the behavior of experimental animals; and most rely on manual operations, with low efficiency and subjective errors. In view of these problems, the present invention discloses a behavior detection device for experimental animals after taking medicine based on multimodal fusion, which uses multimodal data acquisition and advanced machine learning algorithms to accurately classify and analyze the behavior of experimental animals.
[0020] The present invention provides a behavior detection device for experimental animals after taking medicine based on multimodal fusion. Its working process is as follows: the data acquisition module synchronously acquires the initial images of the behavior of mice after taking medicine and the data of the thin-film pressure sensor; the data processing module preprocesses and extracts features from the initial images and the data of the thin-film pressure sensor; the behavior determination module classifies the behavior of the preprocessed multimodal data; the report generation module converts the behavior data into structured text and generates an experimental report through a large language model.
[0021] Specifically, it includes the following steps:
[0022] Step 1, as shown in Figure 1 for reference, build an experimental cage for experimental animals, using a transparent material. Taking mice as an example, the experimental cage is 30 cm long, 30 cm wide, and 15 cm high. This size can be adjusted conventionally according to actual experimental needs. Deploy a multispectral camera on the top of the experimental cage. This camera is [FLIR Blackfly S], which has the function of shooting in visible light and infrared bands, with a frame rate of not less than 30 frames per second and a resolution of 1920×1080 pixels, and can adapt to the shooting of mouse behavior under different lighting conditions. Embed a thin-film pressure sensor array at the bottom of the experimental cage. The resolution of the sensor array is 0.1 g / cm², the sampling frequency is not less than 100 Hz, and the area covered by the sensor array at the bottom of the experimental cage is not less than 90%, which can accurately measure the foot contact area and pressure distribution of mice. The data acquisition module performs hardware-level timestamp synchronization through the IEEE 1588 Precision Time Protocol (PTP) to ensure that the acquisition times of the initial images and the data of the thin-film pressure sensor are aligned, and the synchronization error does not exceed ±5 ms.
[0023] Step 2, as shown in Figure 2 After collecting the initial image with a multispectral camera, use Python combined with the OpenCV library to process the initial image. First, perform Gaussian filtering on the initial image to remove image noise and enhance image quality. Convert the RGB image to the HSV color space, extract the saturation channel for normalization processing, and then use the Otsu algorithm to dynamically determine the binary threshold to binarize the normalized image, obtaining a preliminary experimental animal contour image. Then, perform morphological opening operations on the preliminary mouse contour image to remove small noise points and interference regions in the image, and then perform closing operations to fill the holes inside the mouse contour, thereby obtaining a clear and accurate mouse contour image.
[0024] Step 3, as shown in Figure 3 After collecting the initial image with a multispectral camera, use Python combined with the OpenCV library to process the initial image. First, perform Gaussian filtering on the initial image to remove image noise and enhance image quality. Convert the RGB image to the HSV color space, extract the saturation channel for normalization processing, and then use the Otsu algorithm to dynamically determine the binary threshold to binarize the normalized image, obtaining a preliminary experimental animal contour image. Then, perform morphological opening operations on the preliminary mouse contour image to remove small noise points and interference regions in the image, and then perform closing operations to fill the holes inside the mouse contour, thereby obtaining a clear and accurate mouse contour image.
[0025] Step 4, as shown in Figure 4 Extract the body pose features of the mouse from the clear and accurate mouse contour image. Use deep learning algorithms, such as the OpenPose algorithm, to extract the skeletal key points of the mouse, and then calculate features such as body length, width, angle, and limb extension degree. At the same time, obtain the centroid position of the body by calculating the centroid of the contour, obtain the body area by calculating the number of pixels contained in the contour, obtain the body contour perimeter by calculating the perimeter of the contour, and determine the body center line through principal component analysis.
[0026] Step 5, as shown in Figure 5 Analyze the normalized pressure sensor data and extract features such as contact area, pressure peak value, and spatial distribution entropy value. The contact area is calculated by counting the number of sensor units with pressure values greater than a certain threshold; the pressure peak value is directly found as the maximum value from the data; the spatial distribution entropy value is used to measure the uniformity of the pressure distribution by calculating the entropy of the pressure distribution.
[0027] Step 6, integrate the extracted image features and pressure sensor features to obtain a multi-modal raw feature set. Perform normalization processing on the multi-modal raw feature set using the Min - Max normalization method to map all feature values to the [0,1] interval to ensure that different features have the same scale for subsequent model training.
[0028] Step 7: Divide the normalized multi-modal raw feature set into a training set and a test set, with the ratio of the number of training set samples to the number of test set samples being 7:3. The division process uses a random division method, while ensuring that both the training set and the test set contain various behavior categories such as mice standing, walking, running, lying on their sides, etc., to ensure the generalization ability of the model.
[0029] Step 8: Train a classification model based on the training set samples. Perform kernel function mapping on the multi-modal feature vectors, and use the radial basis function (RBF) as the kernel function. Optimize the hyperparameter combination through grid search, including the penalty coefficient C and the kernel function parameter γ, with the search ranges being [0.1, 1, 10] and [0.01, 0.1, 1] respectively. Use the Scikit-learn library in Python to implement the training of the machine learning classifier, and obtain the trained SVM model.
[0030] Step 9: Use the trained SVM model to determine the behavior of the test set samples. Input the multi-modal feature vectors of the test set samples into the SVM model, and the model outputs the behavior category labels, while recording the behavior data aligned with the timestamps. Calculate evaluation metrics such as the accuracy and recall rate of the behavior determination to evaluate the performance of the model.
[0031] Step 10: Encode the time series behavior data and experimental parameters (drug dosage, time node) into JSON format for input. Use the JSON library in Python to convert the data into a standardized JSON string to ensure the unity of the data format, which is convenient for large language models to process.
[0032] Step 11: Invoke a pre-trained large language model based on the Transformer architecture, such as GPT-Neo. This model is fine-tuned through domain adaptation. Use the mouse behavior experiment report text and the corresponding multi-modal data to construct a fine-tuning data set, and adopt the low-rank adaptation LoRA (r = 8, α = 16) technology to inject the adaptation parameters in the field of mouse behavior analysis. Pass the JSON-formatted input data to the large language model, and the model parses the input data to generate a natural language report containing behavior statistics, abnormal behavior warnings, and mechanism speculation.
[0033] Step 11: Invoke a pre-trained large language model based on the Transformer architecture, inject the parameters in the field of experimental animal behavior analysis through low-rank adaptation (LoRA) technology, and fine-tune 1000 groups of experimental animal behavior report texts and the corresponding multi-modal data, annotating behavior labels such as standing and lying on the side and abnormal types; the LoRA parameters are rank r = 8 and scaling factor α = 16, and only the parameters of the adaptation layer are trained, keeping the weights of the base model frozen.
[0034] Step 12, support through the natural language query interface, the user inputs a natural language description of a specific time window, such as "View the behavioral data of mice 1 - 2 hours after taking medicine", the system will query the corresponding behavioral data, and use the Matplotlib library in Python to generate visualization charts such as line charts, bar charts, and heat maps, and return them to the user.
[0035] Embodiment
[0036] Use the detection device of the present invention to conduct a post - medication behavior detection experiment on 20 mice. The experiment lasts for 24 hours, and the drug dose is set to [fluoxetine 10mg / kg]. Use different feature combinations to train the machine learning classifier respectively, and run it independently 30 times to obtain the average accuracy and average consumption time.
[0037] The experimental simulation platform is Python 3.8, the operating system: Windows 10, the processor is Intel (R) Core(TM) I7 - 10700 CPU, and the memory: 16GB.
[0038] Table 1 Accuracy and consumption time of different feature extraction methods
[0039] Feature extraction method Average accuracy rate (%) Average consumption time (s) Only body features 78.5 255 Only pressure features 79.2 268 Combined image and pressure sensor 86.2 215.5
[0040] It can be seen from Table 1 that the behavior recognition rate obtained by only image features or only pressure sensor features is lower than that obtained by combining image and pressure sensor features. The time consumed for behavior recognition with only image features is 40.5s more than that for behavior recognition by combining image and pressure sensor features, and the time consumed for behavior recognition with only pressure sensor features is 53.5s more than that for behavior recognition by combining image and pressure sensor features. Therefore, the behavior recognition effect of combining image and pressure sensor features in the present invention is better, with less time consumption and high accuracy.
[0041] The above - described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An experimental animal behavior detection device based on multimodal fusion, characterized in that, Comprising: (1) A data acquisition module configured to synchronously acquire the initial images of the animal's behavior after taking the medicine and the thin-film pressure sensor data, including: (1a) The initial images are captured by a multispectral camera deployed at the top of the experiment, including visible light and infrared bands, and can adapt to the shooting of animal behavior under different lighting conditions; (1b) The thin-film pressure sensor data is generated by an array of thin-film pressure sensors embedded at the bottom of the experimental cage. The thin-film pressure sensor array covers no less than 90% of the bottom area of the experimental cage, the sampling frequency is not less than 100 Hz, and the sensor resolution is 0.1 g / cm²; (1c) The data acquisition module aligns the acquisition times of the images and the sensor data through the IEEE 1588 Precision Time Protocol (PTP), and the time error does not exceed ±5 ms, realizing multi-modal data synchronization at the hardware level; (2) A data processing module configured to preprocess and extract features from the initial images and the thin-film pressure sensor data, including: (2a) Perform background subtraction and moving target segmentation on the initial images, and extract the contours, skeletal key points, and movement trajectories of the experimental animals. Among them, the skeletal key points are identified by the OpenPose model based on transfer learning, and the contours are obtained through morphological opening and closing operations, and the recognition accuracy is ≥90%; (2b) Perform filtering and normalization processing on the thin-film pressure sensor data, and extract the contact area, pressure peak value, and spatial distribution entropy value. The entropy value is calculated through the pressure distribution probability; (3) A behavior determination module configured to classify the preprocessed multi-modal data, including: (3a) Divide the contours, movement trajectories, contact area, and spatial distribution entropy value into a training set and a test set; (3b) Train a classification model based on the training set samples. The classification model dynamically fuses multi-modal feature vectors through an attention mechanism, and the output is a class label; (3c) Use the classification model to determine the behavior of the test set samples and record the behavior data aligned with the time stamp; (4) A report generation module configured to convert the behavior data into structured text and generate an experimental report through a large language model, including: (4a) Encode the time series data into a JSON structured input including drug dosage, behavior category, and pressure distribution; (4b) Invoke a pre-trained LLM based on the Transformer architecture to parse the input. The LLM injects parameters in the field of experimental animal behavior analysis through the Low-Rank Adaptation (LoRA) technique, and generates a natural language report including statistical charts, anomaly warnings, and speculation on pharmacological mechanisms.
2. The behavior detection device for experimental animals after taking medicine based on multi-modal fusion according to claim 1, wherein The thin-film pressure sensor array adopts a cross-electrode design, the unit pitch is 2 mm × 2 mm, and the linearity error of the pressure response is ≤1.5%.
3. The behavior detection device for experimental animals after taking medicine based on multimodal fusion according to claim 1, characterized in that, The attention mechanism adopts a Transformer encoder structure, and the dynamic weight distribution ratio of visual features and pressure features is adjustable from 7:3 to 3:
7.
4. The behavior detection device for experimental animals after taking medicine based on multi-modal fusion according to claim 1, characterized in that, The training process of the classification model includes: (1) Using the Radial Basis Function (RBF) as the kernel function; (2) Optimizing the hyperparameter combination through grid search. The hyperparameters include the penalty coefficient C and the kernel function parameter γ.
5. The behavior detection device for experimental animals after taking medicine based on multimodal fusion according to claim 1, characterized in that, The JSON structured input includes: (1) Experimental information fields: drug dose (unit: mg / kg), timestamp (format: YYYY-MM-DDTHH:mm:ss); (2) Behavioral data fields: behavior category, pressure distribution array.
6. The behavior detection device for experimental animals after taking medicine based on multi-modal fusion according to claim 1, wherein, The large language model is a pre-trained model based on the Transformer architecture and is fine-tuned through domain adaptation, including: (1) Using the text of mouse behavioral experiment reports and multi-modal data to construct a fine-tuning dataset; (2) Adopting the Low-Rank Adaptation (LoRA) technique to inject adaptation parameters.