Multi-view synchronous behavior capture method and system, and electronic device

By adopting the multi-eye synchronous behavioral capture method in the multi-eye camera acquisition system, synchronous acquisition and real-time compression processing of image data are solved, and the frame drop problem that is prone to occur when multi-eye cameras collect biological behavior in the prior art is improved, and the integrity and reliability of image data are improved.

WO2025129591A1PCT designated stage expired Publication Date: 2025-06-26SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/140775
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing multi-eye cameras are prone to frame loss when collecting biological behaviors, resulting in information loss and reconstruction errors.

Method used

The multi-objective synchronous behavioral capture method is adopted to create the acquisition process and compression process corresponding to each camera, and realize the synchronous acquisition and real-time compression processing of image data to ensure the integrity and reliability of the target image data.

Benefits of technology

The synchronous acquisition of the target object image data is realized, frame dropping is avoided, the integrity and reliability of image data is improved, and the coverage and resolution of three-dimensional spatial reconstruction are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140775_26062025_PF_FP_ABST
    Figure CN2023140775_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of ethology, and in particular relates to a multi-view synchronous behavior capture method and system, and an electronic device. The method comprises: creating a control process, and creating an acquisition process and a compression process for each camera; each camera capturing a frame of image data; sending the image data to the corresponding acquisition process, which then sends the image data to the corresponding compression process, incrementing a semaphore in the control process by a value equal to the number of acquisition processes that have received images, and blocking the corresponding acquisition processes; determining whether the semaphore acquired in the control process is greater than or equal to the total number of cameras, and if so, proceeding to the next step, and if not, repeating the previous step; determining whether the number of frames acquired in each acquisition process is equal to a preset number, and if so, completing capture, and if not, awakening the acquisition processes and repeating the previous three steps. By means of the described configuration, the present invention remedies the defect of existing multi-view cameras being prone to frame loss during biological behavior acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Multi-eye synchronous behavioral capture method, system and electronic equipment Technical Field

[0001] The present invention relates to the field of animal behavior technology, and in particular to a multi-eye synchronous behavior capture method, system and electronic equipment. Background Art

[0002] Behavioral science plays an important role in scientific research. Generally speaking, experimental targets produce different behavioral manifestations compared with control groups in different paradigms, which can correspond to the inhibition or activation of different neural circuits.

[0003] Currently, researchers primarily use monocular and multi-camera cameras to capture target behaviors, observing and analyzing the collected data to explore the underlying mechanisms of biological behavior. However, given the richness of motion in three-dimensional space, monocular camera data capture can lose important information in one dimension, making detailed behavioral analysis difficult. Furthermore, monocular camera acquisition can be subject to self-occlusion and mutual occlusion between multiple subjects, leading to information loss and reconstruction errors. Multi-camera acquisition can alleviate these issues to a certain extent, but requires simultaneous recording, which can easily lead to frame dropouts due to data bandwidth and computing resource constraints.

[0004] Summary of the Invention

[0005] In order to solve the defect that existing multi-cameras are prone to frame loss when capturing biological behaviors, the present invention proposes a multi-camera synchronous behavioral capture method, system and electronic equipment.

[0006] The technical solution adopted by the present invention is a multi-eye synchronous behavioral capture method, which includes: creating a control process and creating an acquisition process and a compression process for each camera; each camera captures a frame of image data of the target object; sending the frame of image data captured by each camera to its corresponding acquisition process, and the acquisition process sends the image data to its corresponding compression process for compression processing, adding a signal amount equal to the number of acquisition processes that receive the image to the control process, and blocking the corresponding acquisition process; judging whether the signal amount collected by the control process is greater than / equal to the total number of the cameras, if so, proceeding to the next step; if not, repeating the previous step; judging whether the number of frames collected by each acquisition process is equal to the set number of frames, if so, completing the capture; if not, waking up all acquisition processes and repeating the first three steps.

[0007] Preferably, each camera captures a frame of image of the target object, specifically including: 8 cameras respectively capture a frame of image of the target object from different perspectives.

[0008] Preferably, the acquisition process sends the image data to its corresponding compression process for compression processing, which specifically includes: the acquisition process sends the image data to its corresponding compression process for compression processing through queue communication.

[0009] Preferably, before each camera captures a frame of image data of the target object, the method further includes: estimating the intrinsic parameter matrix and relative position of each camera using Zhang's calibration method.

[0010] Preferably, all resources are released after completion and the acquisition is completed, which specifically includes: releasing all resources after completion, using machine learning means to track the key points of the body of the target object, and using a triangulation method based on computer vision to perform three-dimensional reconstruction of the key points of the body to complete the capture.

[0011] The present invention also proposes a multi-camera synchronous behavioral capture system, which is used to implement the above method, including: an acquisition module, which includes multiple cameras, and the multiple cameras are used to acquire original image data; a fixing module, which is used to fix the multiple cameras at intervals and make the imaging directions of the multiple cameras face the target object; and a processing module, which is used to receive the original image data acquired by the multiple cameras and compress and store the multiple original image data in real time.

[0012] Preferably, the acquisition module includes 8 cameras, and the 8 cameras are all depth cameras with a resolution of 1280×720 and an acquisition frame rate of 30 frames per second.

[0013] Preferably, the fixing module comprises an aluminum fixing frame.

[0014] Preferably, data signals are transmitted between the processing module and the acquisition module via a universal serial bus.

[0015] The present invention also proposes an electronic device comprising: at least one processor, at least one memory, and at least one communication bus, wherein the memory stores a computer program, and the processor reads the computer program in the memory through the communication bus; when the computer program is executed by the processor, the above-mentioned multi-eye synchronous behavioral capture method is implemented.

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] 1. The acquisition process sends the image data to its corresponding compression process through queue communication for compression processing. The compression module can compress and store the original image data collected by multiple cameras in real time. The control process controls the blocking and waking of the eight acquisition processes, which ensures the synchronization of the target object image data acquisition and avoids frame loss.

[0018] 2. Eight cameras are arranged at different positions of the target object to capture image data from different perspectives of the target object, increasing the number of observation windows, expanding the coverage of three-dimensional space reconstruction, and improving the acquisition resolution and data bandwidth. By fusing information from eight perspectives, data in missing areas can be completed, which helps to reduce the impact of self-occlusion or multi-body occlusion, avoid information loss and reconstruction errors caused by occlusion, and improve the integrity and reliability of image data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention is described in detail below with reference to the embodiments and accompanying drawings, in which:

[0020] FIG1 is a schematic diagram of a multi-camera synchronous behavioral capture system;

[0021] FIG2 is a flow chart of a multi-eye synchronous behavioral capture method;

[0022] FIG3 is a structural block diagram of an electronic device. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions, and advantages of the present invention more apparent, embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar components or components having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0024] In one embodiment, a multi-camera synchronized behavioral capture system includes: an acquisition module, a fixing module, and a processing module. The acquisition module is used to acquire raw image data; the fixing module is used to fix the acquisition module and align its acquisition direction with the target object; the processing module is used to receive the raw image data acquired by the acquisition module and to compress and store the raw image data in real time. Specifically, the acquisition module includes multiple cameras 100, which are used to acquire raw image data; the fixing module is used to fix the multiple cameras 100 at intervals around the target object, aligning the imaging direction of the cameras 100 with the target object; and the processing module is used to receive the raw image data acquired by the multiple cameras 100 and to compress and store the multiple raw image data in real time. Specifically, the processing module can be an acquisition host 120 equipped with a display screen. The specific structure of the acquisition host can be seen in the electronic device 3000 in the subsequent embodiments.

[0025] The compression module can compress and store the original image data collected by multiple cameras 100 in real time. That is, during the collection process of the multiple cameras 100 of the collection module, the compression module can complete real-time compression, which ensures the synchronization of the collection of the target object image data and avoids frame loss.

[0026] In one embodiment, as shown in Figure 1, the acquisition module includes eight cameras 100, each of which uses a depth camera 100 with a resolution of 1280×720 and an acquisition frame rate of 30 frames per second. Compared to a capture system consisting of four or five cameras 100, eight cameras 100 offer a wider range of viewing angles for the target object, higher information coverage, and better reduction of self-occlusion and inter-occlusion, resulting in more reliable image data. The 1280×720 resolution approaches standard definition, and the images captured by the cameras 100 offer high quality, facilitating feature extraction and information extraction by subsequent computer vision algorithms. The 30 frames per second acquisition rate effectively meets the requirements of real-time acquisition and processing, ensuring continuous acquisition of dynamic scenes or rapid motion. The depth cameras 100 not only capture images but also directly obtain distance information, demonstrating three-dimensional perception. Compared to higher-spec cameras 100, they achieve a good balance between performance and price.

[0027] In other embodiments, other numbers of cameras 100 such as 10 or 12 may be used, and the resolution, frame rate, and type of the cameras 100 may also be determined according to the actual acquisition requirements of the target object.

[0028] In one embodiment, the fixed module includes an aluminum fixed frame 110. Aluminum has high rigidity and light weight, is easy to install and move, and has good positioning stability. In addition, the aluminum material processing technology is simple and convenient, which is conducive to the finalized production of the fixed frame 110. The fixed frame 110 includes a rectangular frame and four pillars fixed directly below the four corners of the rectangular frame. Four of the eight cameras 100 are respectively located at the four corners of the rectangular frame, and the other four cameras 100 are respectively located in the middle of the four sides of the rectangular frame. The four pillars of the fixed frame 110 are erected on the circumference of the target object, and the lenses of the eight cameras 100 are facing the target object in the middle of the lower part of the rectangular frame. In other embodiments, the fixed frame 110 can also be made of stainless steel, resin plastic, etc.

[0029] In one embodiment, data signals are transmitted between the processing module and the acquisition module via a universal serial bus 130 (USB). Specifically, data is transmitted between the processing module and the acquisition module via a USB 3.0 data cable, which provides faster data transmission speeds. No external timer or trigger is required, nor is a high-speed solid-state drive required. Simply connecting a data cable is sufficient, reducing the cost of additional hardware.

[0030] In one embodiment, as shown in FIG2 , a multi-camera synchronous behavioral capture method using the multi-camera synchronous behavioral capture system of the above embodiment is provided. The method includes the following steps:

[0031] Step 310: Create a control process, and create a capture process and a compression process for each camera.

[0032] In one possible implementation, resources such as communication queues and semaphores are initialized, process communication rules are set, and cameras, acquisition processes, and compression processes of each camera are initialized, with cameras, acquisition processes, and compression processes corresponding to each other.

[0033] In step 320 , each camera captures a frame of image data of the target object.

[0034] In one possible implementation, eight cameras each capture a frame of image from a different perspective of the target object. The eight cameras are arranged at different positions on the target object to capture image data from different perspectives of the target object. This increases the number of observation windows, expands the coverage of three-dimensional spatial reconstruction, and improves both the acquisition resolution and data bandwidth. By fusing information from eight perspectives, data in missing areas is completed, which helps mitigate the impact of self-occlusion or multi-body occlusion, avoids information loss and reconstruction errors caused by occlusion, and improves the integrity and reliability of image data.

[0035] In step 330, a frame of image data captured by each camera is sent to its corresponding acquisition process, which sends the image data to its corresponding compression process for compression processing, adds a signal quantity to the control process equal to the number of acquisition processes that receive the image, and blocks the corresponding acquisition process.

[0036] In one possible implementation, the acquisition process sends image data to its corresponding compression process via queue communication for compression processing. This enables real-time compression of image data collected by the acquisition process, ensuring synchronization of acquisition across eight cameras and preventing frame drops. The acquisition and compression processes can operate asynchronously, eliminating the need for synchronous waiting and improving efficiency. The queue can buffer large amounts of image data, reducing the pressure between acquisition and processing. The number and priority of processes can be flexibly adjusted based on queue load, fully utilizing CPU resources.

[0037] Step 340: Determine whether the signal quantity collected by the control process is greater than or equal to the total number of cameras. If so, proceed to step 350 to achieve the purpose of synchronous acquisition. Cooperating with the compression process, complete synchronous recording among the eight cameras is achieved. If not, repeat step 330.

[0038] Step 350 determines whether the number of frames collected by each acquisition process is equal to the set frame number. If so, capture is completed. If not, all acquisition processes are awakened and steps 320, 330, and 340 are repeated. The set frame number refers to the number of frames required by the experimenter to complete the behavioral study.

[0039] In one possible implementation, all resources are released upon completion, and machine learning is used to track the target subject's body key points. Computer vision-based triangulation is then used to perform a 3D reconstruction of these key points, completing the capture. Using machine learning to track the target subject's body key points effectively identifies and tracks key points such as target body parts. The algorithm's performance improves with increasing data volume, adapting to diverse shooting scenarios. Key point identification is achieved without manual labeling, resulting in highly automated results. Computer vision-based triangulation is used to perform a 3D reconstruction of key points, leveraging multi-viewpoint information to restore the true 3D spatial coordinates of these key points with high accuracy.

[0040] If so, the capture is completed, specifically including: stopping the acquisition and waiting for the compression process to complete the compression, and releasing all resources after completion.

[0041] In one embodiment, before each camera captures a frame of image data of the target object in step 320, the process further includes: estimating the intrinsic parameter matrix and relative position of each camera using Zhang's calibration method, determining the intrinsic parameter and posture relationship of each camera, providing basic data support for key point tracking and three-dimensional reconstruction, and avoiding errors introduced by incorrect camera parameters when fusing multi-view information.

[0042] In one embodiment, as shown in FIG3 , an electronic device 3000 includes at least one processor 3001 , at least one communication bus 3002 , and at least one memory 3003 .

[0043] The processor 3001 and the memory 3003 are connected, for example, via a communication bus 3002. Optionally, the electronic device 3000 may further include a transceiver 3004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 3004 is not limited to one, and the structure of the electronic device 3000 does not constitute a limitation on the embodiments of the present application.

[0044] Processor 3001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 3001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0045] Communication bus 3002 may include a path for transmitting information between the aforementioned components. Communication bus 3002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Communication bus 3002 may be divided into an address bus, a data bus, a control bus, and so on. For ease of illustration, FIG3 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0046] The memory 3003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.

[0047] The memory 3003 stores a computer program, and the processor 3001 reads the computer program stored in the memory 3003 through the communication bus 3002 .

[0048] When the computer program is executed by the processor 3001, the multi-eye synchronous behavioral capture method in the above embodiments is implemented.

[0049] In one embodiment, the multi-camera synchronous behavioral capture system is applied in behavioral data collection and analysis of abnormal behavior detection.

[0050] In this specification, the use of terms such as "Embodiment 1," "this embodiment," and "in one embodiment" indicates that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in the invention or at least one embodiment or example of the invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example; furthermore, the specific features, structures, materials, or characteristics described may be appropriately combined in any one or more embodiments or examples.

[0051] In this specification, the terms "connect," "install," "fix," "dispose," and "have" are to be understood broadly. For example, "connect" can mean a fixed connection, a detachable connection, or an integral connection; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; and it can mean internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in this application based on the specific circumstances.

[0052] In the description of this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0053] The above description of the embodiments is to facilitate ordinary technicians in this technical field to understand and apply the technology of this case. People familiar with the technology in this field can obviously make various modifications to these examples easily and apply the general principles described here to other embodiments without having to go through creative work. Therefore, this case is not limited to the above embodiments. Modifications to the following situations should all be within the scope of protection of this case: ① A new technical solution implemented based on the technical solution of the present invention and combined with existing common knowledge, the technical effect produced by the new technical solution does not exceed the technical effect of the present invention; ② The equivalent replacement of some features of the technical solution of the present invention with the known technology, the technical effect produced is the same as the technical effect of the present invention; ③ The technical solution of the present invention is expandable, and the substantive content of the expanded technical solution does not exceed the technical solution of the present invention; ④ The equivalent transformation made by the content of the description and drawings of the present invention is directly or indirectly applied to other related technical fields.

Claims

1. A multi-camera synchronous ethology capture method, characterized in that, The method includes: Create a control process, and create an acquisition process and a compression process for each camera; Each of the cameras captures a frame of image data of the target object; Send a frame of image data captured by each of the cameras to its corresponding acquisition process. The acquisition process sends the image data to its corresponding compression process for compression processing, increments the semaphore in the control process by the same number as the number of acquisition processes that received the image, and blocks the corresponding acquisition processes; Determine whether the semaphore collected by the control process is greater than or equal to the total number of cameras. If so, proceed to the next step; if not, repeat the previous step; Determine whether the number of frames acquired by each acquisition process is equal to the set number of frames. If so, the capture is completed; if not, wake up all the acquisition processes and repeat the previous three steps.

2. The multi-view synchronous ethology capture method according to claim 1, wherein Each of the cameras captures a frame of image of the target object, specifically including: 8 cameras respectively capture a frame of image from different perspectives of the target object.

3. The multi-view synchronous ethology capture method according to claim 1, characterized in that The acquisition process sends the image data to its corresponding compression process for compression processing, specifically including: the acquisition process sends the image data to its corresponding compression process for compression processing through queue communication.

4. The multi-view synchronous ethology capture method according to claim 1, characterized in that Before each of the cameras captures a frame of image data of the target object, it also specifically includes: estimating the internal parameter matrix and relative position of each of the cameras using the Zhang calibration method.

5. The multi-view synchronous ethology capture method according to claim 1, wherein After completion, release all resources, and after the acquisition is completed, specifically including: release all resources after completion, use machine learning means to track the body key points of the target object, use computer vision-based triangulation method to perform three-dimensional reconstruction on the body key points, and complete the capture.

6. A multi-view synchronous ethology capture system, which is used to implement the method described in any one of claims 1-5, and is characterized in that, Includes: An acquisition module, which includes multiple cameras, and the multiple cameras are used to acquire raw image data; A fixing module, which is used to fix the multiple cameras at intervals and make the imaging directions of the multiple cameras face the target object; A processing module, which is used to receive the raw image data acquired by the multiple cameras and process the multiple raw image data for real-time compression and storage.

7. The multi-camera synchronous ethology capture system according to claim 6, characterized in that, The acquisition module includes 8 cameras, and all 8 cameras are depth cameras with a resolution of 1280×720 and an acquisition frame rate of 30 frames per second.

8. The multi-eye synchronous ethology capture system according to claim 6, wherein The fixing module includes an aluminum fixing frame.

9. The multi-camera synchronous ethology capture system according to claim 6, characterized in that, The processing module and the acquisition module transfer data signals through a universal serial bus.

10. An electronic device, characterized in that, Includes: At least one processor, at least one memory, and at least one communication bus, where A computer program is stored on the memory, and the processor reads the computer program in the memory through the communication bus; When the computer program is executed by the processor, it implements the multi-view synchronous ethology capture method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-ocular camera system, device and synchronization method

    CN104539931A

  • Synchronous acquisition system and method of multi-channel cameras

    CN107241553A

  • Multi-camera synchronous working method and device, medium and electronic equipment

    CN111182226A

  • Method and device for synchronizing view fields among different devices and automatic driving vehicle

    CN115967775A

  • Image Capture Control Apparatus, Image Capture Control Method, and Image Capture Control Program

    US20180376050A1