Multi-camera visual motion capture method, system and equipment suitable for acquisition time sequence asynchronization and medium
By obtaining the relative position relationship matrix and deep learning fusion model between multiple cameras, the problem of timing asynchrony in the multi-camera visual motion capture system is solved, the accurate fusion and time synchronization of multi-perspective information are achieved, and the application scope of visual motion capture technology is expanded.
Patent Information
- Application Number
- CN202510937271.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-28
AI Technical Summary
In a multi-camera visual motion capture system, it is difficult to accurately align multi-view information due to the asynchronous acquisition timing of each camera. In particular, the errors caused by timing differences in fast-motion scenes affect the practicality of the system.
By obtaining the relative position relationship matrix between multiple cameras, three-dimensional reconstruction is performed on images captured by multiple cameras from different perspectives, structural feature information is extracted, and a deep learning fusion model is used to achieve time-synchronized motion capture results. The time sliding window and random perturbation technology are combined to enhance timing robustness.
Accurately fusing multi-view information even with misaligned timestamps solves the timing discrepancy problem in traditional methods, expands the application scope of visual motion capture technology, and makes it suitable for more practical application scenarios.
Smart Images

Figure CN120853256A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method, system, device, and medium for multi-camera visual motion capture with asynchronous acquisition time, and belongs to the field of computer vision technology. Background Technology
[0002] Human motion capture technology has wide applications in medical rehabilitation, sports training, and animation production, primarily encompassing infrared camera and image vision technologies. While infrared camera-based motion capture systems offer high accuracy, they require reflective markers on the human body surface, and the equipment is not portable, limiting their application scenarios. Image-based motion capture systems, due to their portability and lack of marker requirements, are gradually becoming a research hotspot.
[0003] Existing visual motion capture methods are mainly divided into two categories: human pose estimation methods and human mesh reconstruction methods. Compared with traditional human pose estimation methods that only extract joint positions, human mesh reconstruction methods can provide more comprehensive 3D human information. However, current monocular camera-based methods struggle to obtain accurate 3D spatial information. While multi-camera systems can provide richer spatial information, the dispersed distribution of portable cameras often leads to asynchronous acquisition times, posing a challenge to the fusion of multi-view information. In multi-camera visual motion capture systems, the sparse distribution of human keypoints in the image, coupled with differences in camera acquisition times, makes accurate alignment of multi-view information difficult. This is especially pronounced in fast-moving scenes, where the errors caused by temporal differences are further amplified, impacting the system's practicality.
[0004] Currently, although deep learning technology has made significant progress in the field of computer vision, research on the problem of temporal asynchrony in multi-camera motion capture systems is still insufficient. Summary of the Invention
[0005] The present invention aims to at least solve one of the technical problems existing in the prior art. Therefore, in response to the above-mentioned problems, the object of the present invention is to provide a method, system, device, and medium for multi-camera visual motion capture with asynchronous acquisition times, capable of solving the problem of multi-view information fusion under conditions of asynchronous acquisition times of multiple cameras.
[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:
[0007] In a first aspect, the present invention provides a multi-camera visual motion capture method suitable for acquisition at asynchronous times, comprising:
[0008] Obtain the relative positional relationship matrix among multiple cameras;
[0009] A three-dimensional mesh model of the human body surface is obtained by reconstructing all images acquired from different perspectives of multiple cameras. Structured feature information is extracted based on the three-dimensional mesh model of the human body surface, and the set of words projected from different camera perspectives to the main camera coordinate system is obtained through the relative position relationship matrix.
[0010] By inputting a set of lexical units into a deep learning fusion model, the output of the time-synchronized motion capture results is achieved, enabling multi-view data fusion under conditions of asynchronous time.
[0011] Some possible implementations include obtaining the relative positional relationship matrix between multiple cameras, including:
[0012] Place the checkerboard calibration board within the common field of view of multiple cameras;
[0013] Using the corner points of the chessboard as feature points, the relative positional relationship matrix between multiple cameras is calculated based on Zhang Zhengyou's calibration method.
[0014] In some possible implementations, a 3D mesh model of the human body surface is obtained by reconstructing all images acquired from different perspectives by multiple cameras. The specific process is as follows:
[0015] All images from different perspectives of multiple cameras are reconstructed using a deep learning-based 3D mesh model. The input to the 3D mesh model is the image and the intrinsic parameters of each camera, and the output is a set of 3D mesh models of the human body surface. At the same time, the timestamp of the image is assigned to the corresponding 3D mesh model of the human body surface.
[0016] Some possible implementations also include a step of grouping the acquired images according to a set time sliding window, and segmenting each group into words, where each word represents the observation result of a camera on a human body at a certain point in time, wherein the word includes camera source information, time information, key point features and virtual marker point features.
[0017] In some possible implementations, camera source information is defined by encoding the camera number to identify which camera the term originates from; time information is recorded by recording the time difference between the observation frame and the start time of the current time window; key point features are defined by uniformly transforming the estimated key point position information in the frame to the main camera coordinate system using the relative position relationship matrix between cameras; and virtual marker point features are defined by extracting several representative virtual marker point positions from the three-dimensional mesh model of the human body surface and similarly transforming them to the main camera coordinate system using the relative position relationship matrix between cameras.
[0018] In some possible implementations, the training process of a deep learning fusion model includes:
[0019] Obtain the vocabulary from all cameras as training data;
[0020] The training data with misaligned time is input into the deep learning fusion model. By jointly modeling the content between all camera lexies, the Transformer network automatically learns the temporal relationship and potential alignment between lexies, extracts motion information from the lexy information, completes the automatic alignment of timestamps, and outputs a unified and time-synchronized target marker trajectory sequence. Here, the target marker refers to the virtual marker after time synchronization.
[0021] In some possible implementations, obtaining the lexical data from all cameras as training data also includes a step of temporal robustness enhancement of the temporal information in the lexical data, specifically:
[0022] The camera ID, the captured motion image, and the timestamp are saved as a triplet to form a camera image sequence.
[0023] Based on the camera ID, apply a random perturbation to the timestamp of the image;
[0024] The captured images are grouped according to the timestamp after perturbation, based on a set time sliding window.
[0025] Secondly, the present invention also provides a multi-camera visual motion capture system suitable for acquisition with asynchronous acquisition times, comprising:
[0026] The camera view difference calculation module is configured to obtain the relative position relationship matrix between multiple cameras;
[0027] The feature extraction module is configured to perform 3D reconstruction on all images acquired from different perspectives of multiple cameras to obtain a 3D mesh model of the human body surface, extract structured feature information based on the 3D mesh model of the human body surface, and obtain the set of words projected from different camera perspectives onto the main camera coordinate system through the relative position relationship matrix.
[0028] The multi-view fusion module is configured to input a set of lexical units into a deep learning fusion model and output time-synchronized motion capture results, thereby achieving multi-view data fusion under time-asynchronous conditions.
[0029] Thirdly, the present invention provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to enable the processor to perform the method described thereon.
[0030] Fourthly, the present invention also provides a computer-readable storage medium for storing one or more programs, said one or more programs including computer instructions for causing a computer to perform the method.
[0031] Because the present invention adopts the above technical solution, it has the following characteristics:
[0032] 1. This invention utilizes a human body mesh reconstruction method to extract key point information and motion features from various perspectives. This method makes full use of the human motion information contained in a single camera sequence, and can accurately fuse multi-view information even when timestamps are not aligned, effectively solving the problem of temporal synchronization in traditional methods.
[0033] 2. This invention sets a temporal robustness enhancement strategy during the training data generation process in multi-view fusion training. By enhancing the temporal data, the system's robustness to temporal differences is improved, thus solving the dependence of traditional multi-view motion capture systems on the synchronization of acquisition timing.
[0034] 3. Based on the temporal continuity characteristics of human motion in a single camera image sequence, and combined with timestamp random perturbation technology, this invention can maintain stable performance in practical application environments, making it applicable to more practical application scenarios and expanding the application scope of visual motion capture technology.
[0035] In summary, this invention can be widely applied to multi-camera visual motion capture. Attached Figure Description
[0036] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. In the drawings:
[0037] Figure 1 This is a schematic diagram of a multi-camera visual motion capture system according to an embodiment of the present invention.
[0038] Figure 2 This is a structural diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0039] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0040] Although terms such as first, second, third, etc., may be used in this document to describe multiple elements, components, regions, layers, and / or segments, these elements, components, regions, layers, and / or segments should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, or segment from another. Unless the context clearly indicates otherwise, terms such as "first," "second," and other numerical terms used herein do not imply order or sequence. Therefore, the first element, component, region, layer, or segment discussed below may be referred to as the second element, component, region, layer, or segment without departing from the teachings of the exemplary embodiments.
[0041] For ease of description, spatial relative terms can be used in the text to describe the relationship of one element or feature relative to another element or feature, as shown in the figure. These relative terms include, for example, "inside," "outside," "middle," "outer," "below," "above," etc. Such spatial relative terms are intended to include different orientations of the system in use or operation, in addition to those depicted in the figure.
[0042] This invention addresses the shortcomings in research on temporal synchronization in multi-camera motion capture systems. It provides a method, system, and medium for multi-camera visual motion capture with asynchronous acquisition times. The method includes: acquiring a relative positional relationship matrix between multiple cameras; reconstructing a 3D human body model from all images acquired from different viewpoints of the multiple cameras using a 3D human body mesh; extracting structured feature information from the 3D human body model; and obtaining a set of words projected from different camera viewpoints onto the main camera coordinate system using the relative positional relationship matrix. The word set is then input into a deep learning fusion model to achieve multi-view data fusion under asynchronous temporal conditions, and outputting a time-synchronized motion capture result. Therefore, this invention overcomes the dependence of traditional multi-view motion capture systems on synchronous acquisition times, expanding the application scope of visual motion capture technology.
[0043] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the invention and to fully convey the scope of the invention to those skilled in the art.
[0044] Example 1: This example provides a multi-camera visual motion capture method suitable for acquisition with asynchronous timing, including:
[0045] S1. Obtain the relative position relationship matrix between multiple cameras.
[0046] In this embodiment, the relative positional relationship matrix between multiple cameras is obtained, and the specific implementation process is as follows:
[0047] Place the checkerboard calibration board within the common field of view of multiple cameras;
[0048] Using the corner points of the chessboard as feature points, the relative positional relationship matrix between multiple cameras is calculated based on Zhang Zhengyou's calibration method.
[0049] S2. Reconstruct all images from different perspectives to obtain a 3D mesh model of the human body surface. Extract structured feature information based on the 3D mesh model of the human body surface and obtain the projection of different camera perspectives onto the main camera coordinate system based on the relative position relationship matrix. The main camera coordinate system is the camera coordinate system of a pre-selected camera.
[0050] In this embodiment, all images from different perspectives are reconstructed using a deep learning-based 3D mesh model. The input to the 3D mesh model is the image and the intrinsic parameters of each camera, and the output is a set of 3D mesh models of the human body surface. The timestamp of the image is assigned to the corresponding 3D mesh model of the human body surface.
[0051] In this embodiment, the acquired images are grouped according to a set time sliding window, and each group is segmented. Segmentation refers to converting the original observation data within each time sliding window into a set of discrete units with a fixed structure (called "tokens") for subsequent unified modeling and analysis.
[0052] Specifically, for each time sliding window of data, one frame of observation is extracted from each unsynchronized camera and converted into a word.
[0053] Furthermore, each morpheme represents a camera's observation of the human body at a specific point in time, specifically including:
[0054] Camera origin information: By encoding the camera number, it can be determined which camera the term originated from;
[0055] Time information: Record the time difference between the observation frame and the start time of the current time window, and perform normalization processing to achieve time alignment between different cameras;
[0056] Key point features: The estimated key point position information in this frame is uniformly transformed to the main camera coordinate system using the relative position relationship matrix between cameras;
[0057] Virtual marker features: Several representative virtual marker positions are extracted from the three-dimensional mesh model of the human body surface, and the relative position relationship matrix between the cameras is used to transform them to the main camera coordinate system. The representative virtual markers are points with similar positions that are manually selected by experts on the three-dimensional mesh model of the human body surface based on the positions of physical reflective markers commonly used in optical motion capture.
[0058] S3. Input the lexical set into the deep learning fusion model to output the motion capture results, thereby achieving multi-view data fusion under asynchronous time conditions.
[0059] In this embodiment, the training process of the deep learning fusion model is achieved by integrating motion information from the image sequence, specifically as follows:
[0060] The vocabulary of all cameras is obtained as training data. The vocabulary is sorted according to the standardized temporal information to obtain a structurally consistent observation sequence.
[0061] These time-misaligned training data are input into a deep learning fusion network. By jointly modeling the content between all camera lexical units, the Transformer network automatically learns the temporal relationships and potential alignment methods between lexical units, achieving automatic timestamp alignment and outputting a unified and synchronized sequence of target marker trajectory points. Target marker points refer to the real spatial marker points corresponding to virtual marker points after time synchronization. In summary, this invention can extract motion information from lexical information using a deep learning fusion model, where motion information is a representation of the time sequence of key point features in lexical information.
[0062] Furthermore, artificially introducing random time offsets can enhance the model's robustness to time shifts, simulate the asynchronous problem between cameras in a real environment, and improve the system's adaptability to time asynchrony. Therefore, the specific steps for enhancing the temporal robustness of the temporal information in the lexical units when generating training data are as follows:
[0063] The camera image sequence is constructed by saving the camera number, the captured motion images, and the timestamp as a triple.
[0064] Based on the camera number, the timestamp of the image is randomly perturbed. Assume there are two cameras, numbered 1 and 2. Camera 1 captures image 1 at time t1, and camera 2 captures image 2 at time t2. Two time variables, Δ1 and Δ2 (which can be positive or negative), are randomly generated, and the timestamps of image 1 and image 2 are modified to t1+Δ1 and t2+Δ2, respectively.
[0065] The collected images are grouped according to the set time sliding window: Assuming the set time sliding window length is T = 2s, the captured photos are grouped according to the perturbed timestamp into all photos taken by cameras within 0s to 2s, all photos taken by cameras within 2s to 4s, and so on.
[0066] Example 2: Following the method described in Example 1 for multi-camera visual motion capture with asynchronous acquisition timing, this example provides a system for multi-camera visual motion capture with asynchronous acquisition timing. The system provided in this example can implement the method from Example 1 for multi-camera visual motion capture with asynchronous acquisition timing. This system can be implemented through software, hardware, or a combination of both. For ease of description, this example is described by dividing the system into functional units. Of course, in practice, the functions of each unit can be implemented in one or more software and / or hardware components. For example, the system may include integrated or separate functional modules or units to execute the corresponding steps in the methods of Example 1. Since the system in this example is fundamentally similar to the method example, the description process is relatively simple. Relevant details can be found in the description of Example 1. The example of the multi-camera visual motion capture system with asynchronous acquisition timing provided by this invention is merely illustrative.
[0067] Specifically, the present invention provides a multi-camera visual motion capture system suitable for acquisition with asynchronous acquisition time, including a camera viewpoint difference calculation module, which is configured to acquire the relative positional relationship matrix between multiple cameras;
[0068] The feature extraction module is configured to perform 3D reconstruction on all images acquired from different perspectives of multiple cameras to obtain a 3D mesh model of the human body surface, extract structured feature information based on the 3D mesh model of the human body surface, and obtain the set of words projected from different camera perspectives onto the main camera coordinate system through the relative position relationship matrix.
[0069] The multi-view fusion module is configured to input a set of lexical units into a deep learning fusion model and output time-synchronized motion capture results, thereby achieving multi-view data fusion under time-asynchronous conditions.
[0070] Furthermore, a time-series robustness enhancement module is also included, which is used to perturb and group time during the training of deep learning fusion models, thereby improving the system's ability to adapt to time-series asynchrony. This module plays a key role in the training phase, enabling the system to learn to handle different time-series differences, which will not be elaborated on in detail.
[0071] The following detailed embodiments illustrate the specific applications of the present invention in a multi-camera visual motion capture system with asynchronous acquisition timing.
[0072] like Figure 1 As shown, this embodiment provides a visual motion capture system based on four cameras. In this embodiment, cameras 1 and 2 are located on the left and right sides of the front of the area, and cameras 3 and 4 are located on the left and right sides of the back of the area. The fields of view of the four cameras overlap to ensure the capture effect.
[0073] The camera viewpoint difference calculation module calculates the camera viewpoint difference. A 12x9 checkerboard calibration board with alternating black and white colors is placed within the common field of view of the four cameras. Using the corner points of the checkerboard as feature points, the relative positional relationship matrix between the four cameras is calculated. The accuracy of the camera viewpoint difference calculation is ensured through the acquisition and processing of multiple sets of calibration data. Specifically, when acquiring motion data, all four cameras acquire data at a frame rate of 30 frames per second. Due to the lack of hardware synchronization, there is a time synchronization issue among the four cameras, with a maximum difference of up to 200 milliseconds. Furthermore, due to transmission limitations, there is also significant frame drop, further increasing the time synchronization problem.
[0074] The grouped images within each time sliding window are processed as follows: First, a 3D human body mesh reconstruction based on deep learning is performed on all images from different viewpoints to obtain a 3D mesh model of the human body surface. Then, structured feature information is extracted based on the reconstruction result, and a set of words projected onto the main camera coordinate system from different camera viewpoints is obtained based on the relative position relationship matrix.
[0075] During the training phase, the temporal robustness enhancement module saves the images and timestamps acquired by each camera as triples to form a camera image sequence, and randomly perturbs the timestamps within a range of ±300 milliseconds to generate training data. A 2-second time sliding window is set to group the perturbed image sequence, with each group including multiple frames from four cameras. For example, the first group includes a moving image 1 acquired by camera 1 and the perturbation timestamp 1. The training phase aims to obtain a deep learning fusion model capable of adapting to temporal differences under multi-view observation. This process is implemented through a temporal post-processing network based on the Transformer architecture. The training input consists of tokens extracted from different camera perspectives and projected onto the main camera coordinate system, and the target output is the aligned true position of the 3D marker points at the corresponding time points (real data obtained from a high-precision infrared optical motion capture system). During training, the deep learning fusion model learns the temporal relationship and spatial fusion strategy between different viewpoints by minimizing the mean squared error (MSE) loss between predicted and true marker points. During the deployment phase, the obtained lexical set is input into the deep learning fusion model trained to output time-synchronized motion capture results, thereby achieving multi-view data fusion under time-asynchronous conditions.
[0076] In summary, the test results show that under the above asynchronous conditions, the multi-camera visual motion capture system proposed in this invention can still maintain an average error of about 3cm, and the main error is generated by bias, which has little impact on downstream tasks.
[0077] Example 3: This example provides an electronic device corresponding to the multi-camera visual motion capture method with asynchronous acquisition timing provided in Example 1. The electronic device can be an electronic device for the client, such as a mobile phone, laptop, tablet computer, desktop computer, etc., to execute the method of Example 1.
[0078] like Figure 2 As shown, the electronic device includes a processor, a memory, a communication interface, and a bus. The processor, memory, and communication interface are connected via the bus to enable communication between them. The memory stores a computer program that can run on the processor. When the processor runs the computer program, it executes the method of Embodiment 1. The implementation principle and technical effects are similar to those of Embodiment 1, and will not be repeated here. Those skilled in the art will understand that... Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computing device on which the present application is applied. The specific computing device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0079] In a preferred embodiment, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), and optical discs.
[0080] In a preferred embodiment, the processor can be any type of general-purpose processor such as a central processing unit (CPU) or a digital signal processor (DSP), and is not limited thereto.
[0081] Example 4: This example provides a computer-readable storage medium for storing one or more programs, the one or more programs including computer instructions, which, when executed by a computer, cause the computer to perform the method provided in Example 1 above.
[0082] In a preferred embodiment, the computer-readable storage medium may be a tangible device for holding and storing instructions executable, such as, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof. The computer-readable storage medium stores computer program instructions that cause a computer to perform the method provided in Embodiment 1 above.
[0083] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.
[0084] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0086] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In the description of this specification, the terms "a preferred embodiment," "furthermore," "specifically," "in this embodiment," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the embodiments in this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for multi-camera visual motion capture applicable to acquisition with asynchronous acquisition times, characterized in that, include: Obtain the relative positional relationship matrix among multiple cameras; A three-dimensional mesh model of the human body surface is obtained by reconstructing all images acquired from different perspectives of multiple cameras. Structured feature information is extracted based on the three-dimensional mesh model of the human body surface, and the set of words projected from different camera perspectives to the main camera coordinate system is obtained through the relative position relationship matrix. By inputting a set of lexical units into a deep learning fusion model, the output of the time-synchronized motion capture results is achieved, enabling multi-view data fusion under conditions of asynchronous time.
2. The multi-camera visual motion capture method according to claim 1, applicable to acquisition times with asynchronous timing, is characterized in that, Obtain the relative positional relationship matrix between multiple cameras, including: Place the checkerboard calibration board within the common field of view of multiple cameras; Using the corner points of the chessboard as feature points, the relative positional relationship matrix between multiple cameras is calculated based on Zhang Zhengyou's calibration method.
3. The multi-camera visual motion capture method according to claim 1, applicable to acquisition times that are asynchronous, is characterized in that, A 3D mesh model of the human body surface is obtained by reconstructing all images acquired from different perspectives by multiple cameras. The specific process is as follows: All images from different perspectives of multiple cameras are reconstructed using a deep learning-based 3D mesh model. The input to the 3D mesh model is the image and the intrinsic parameters of each camera, and the output is a set of 3D mesh models of the human body surface. At the same time, the timestamp of the image is assigned to the corresponding 3D mesh model of the human body surface.
4. The multi-camera visual motion capture method according to claim 1, applicable to acquisition times that are asynchronous, is characterized in that, It also includes the step of grouping the acquired images according to a set time sliding window, and segmenting each group into words. Each word represents the observation result of a camera on the human body at a certain point in time. The word includes camera source information, time information, key point features and virtual marker point features.
5. The multi-camera visual motion capture method according to claim 4, characterized in that, Camera source information: By encoding the camera number, it is clear which camera the term comes from; Time information: Record the time difference between the observation frame and the start time of the current time window; Key point features: The estimated key point position information in this frame is uniformly transformed to the main camera coordinate system using the relative position relationship matrix between cameras; virtual Marker point features: Several representative virtual marker point positions are extracted from the three-dimensional mesh model of the human body surface, and then transformed to the main camera coordinate system using the relative position relationship matrix between cameras.
6. The multi-camera visual motion capture method according to claim 5, characterized in that, The training process of a deep learning fusion model includes: Obtain the vocabulary from all cameras as training data; The training data with misaligned time is input into the deep learning fusion model. By jointly modeling the content between all camera lexies, the Transformer network automatically learns the temporal relationship and potential alignment between lexies, extracts motion information from the lexy information, completes the automatic alignment of timestamps, and outputs a unified and time-synchronized target marker trajectory sequence. Here, the target marker refers to the virtual marker after time synchronization.
7. The multi-camera visual motion capture system according to claim 6, characterized in that, Obtaining all camera lexies as training data also includes a step of enhancing the temporal robustness of the temporal information in the lexies, specifically: The camera ID, the captured motion image, and the timestamp are saved as a triplet to form a camera image sequence. Based on the camera ID, apply a random perturbation to the timestamp of the image; The captured images are grouped according to the timestamp after perturbation, based on a set time sliding window.
8. A multi-camera visual motion capture system suitable for acquisition with asynchronous acquisition time, characterized in that, include: The camera view difference calculation module is configured to obtain the relative position relationship matrix between multiple cameras; The feature extraction module is configured to perform 3D reconstruction on all images acquired from different perspectives of multiple cameras to obtain a 3D mesh model of the human body surface, extract structured feature information based on the 3D mesh model of the human body surface, and obtain the set of words projected from different camera perspectives onto the main camera coordinate system through the relative position relationship matrix. The multi-view fusion module is configured to input a set of lexical units into a deep learning fusion model and output time-synchronized motion capture results, thereby achieving multi-view data fusion under time-asynchronous conditions.
9. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, which are executed by the processor to enable the processor to perform the method according to any one of claims 1-7.
10. A computer-readable storage medium for storing one or more programs, characterized in that, One or more programs include computer instructions for causing a computer to perform the method according to any one of claims 1-7.