FPGA-accelerated multi-source image synchronous acquisition system for panoramic conference camera
The panoramic conference camera multi-source image synchronous acquisition system accelerated by FPGA hardware constructs a three-layer collaborative architecture of FPGA hardware core processing + protocol conversion + host computer lightweight distortion correction, which solves the problems of high latency, poor real-time performance and poor compatibility of traditional panoramic conference equipment, and realizes low-latency stitching and audio and video hardware-level linkage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN TECH UNIV
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-05
AI Technical Summary
Traditional panoramic conferencing equipment suffers from problems such as high latency in multi-channel image stitching, poor real-time performance, poor compatibility, insufficient audio and video coordination, and high IO resource consumption.
The panoramic conference camera multi-source image synchronous acquisition system, which adopts FPGA hardware acceleration, acquires multi-source image data in real time through the data acquisition and processing module and processes it based on FPGA hardware. It constructs a three-layer collaborative architecture of FPGA hardware core processing + protocol conversion + host computer lightweight distortion correction. It uses FPGA hardware pipeline to replace CPU software to process multiple images, realizing low-latency stitching and audio and video hardware-level linkage.
It reduces the latency of multi-channel image stitching to the millisecond level, achieves wide device compatibility, avoids excessive computing resources being consumed by the host computer, and is suitable for panoramic conferencing application scenarios on various general PCs, solving the problems of high latency, poor real-time performance and poor compatibility of traditional solutions.
Smart Images

Figure CN122160465A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system. Background Technology
[0002] With the normalization of remote work and multi-location collaborative work, the demand for video conferencing equipment is shifting from simply "being able to see" to "panoramic coverage, intelligent tracking, and extremely simple deployment." In medium to large conference room scenarios, participants are scattered, and traditional single-lens cameras have limited field of view (usually <90°), making it difficult to cover the entire room.
[0003] In related technologies, common panoramic conferencing equipment typically employs a "multi-camera + software stitching" solution. Because this solution relies on a central processing unit (CPU) to serially process high-resolution, multi-channel images, it suffers from high stitching latency and poor real-time performance. Summary of the Invention
[0004] The main purpose of this application is to propose an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system, which aims to overcome the shortcomings of traditional solutions such as high latency and poor real-time performance in multi-channel image stitching.
[0005] To achieve the above objectives, the first aspect of this application proposes an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system, the system comprising: The data acquisition and processing module is used to acquire multi-source image data in real time and process the multi-source image data based on field-programmable gate array (FPGA) hardware to obtain high-definition multimedia interface image data. The data transmission and protocol conversion module is used to convert the high-definition multimedia interface image data into standard universal serial bus image data for transmission. The host computer auxiliary processing module is used to perform lightweight distortion correction processing on the standard universal serial bus image data to obtain panoramic conference image data.
[0006] In some embodiments, the data acquisition and processing module includes a multi-view image acquisition unit, through which the data acquisition and processing module acquires multi-source image data in real time; The data acquisition and processing module polls the registers of the multi-view image acquisition unit through a set of serial computer buses, so that the multi-view image acquisition unit can synchronously transmit the multi-source image data through an independent interface.
[0007] In some embodiments, the data acquisition and processing module includes a core processing unit based on FPGA hardware. The data acquisition and processing module processes the multi-source image data through the core processing unit according to a preset full hardware processing pipeline to obtain high-definition multimedia interface image data that has been stitched together and aligned with the conference speaker. The fully hardware processing pipeline includes: feature extraction, feature matching, mismatch removal, image fusion, and electronic rotation.
[0008] In some embodiments, the core processing unit is configured to perform weighted average fusion of overlapping regions of the multi-source image data with a preset fixed width based on the homography matrix of the multi-source image data.
[0009] In some embodiments, the core processing unit is further configured to calculate the starting address for reading panoramic wide-angle image data based on the microphone array sound field matrix and perform audio-visual coordination control; the panoramic wide-angle image data is obtained by weighted average fusion of the overlapping regions by the core processing unit.
[0010] In some embodiments, the data acquisition and processing module includes an audio acquisition and positioning unit. The data acquisition and processing module acquires multiple audio data through the audio acquisition and positioning unit, and performs phase difference calculation and orientation calculation on the multiple audio data through a generalized cross-correlation algorithm to obtain the sound source angle signal of the multi-source image data.
[0011] To achieve the above objectives, the second aspect of this application proposes an FPGA-accelerated method for synchronous acquisition of multi-source images from a panoramic conference camera, which is applied to an FPGA-accelerated system for synchronous acquisition of multi-source images from a panoramic conference camera. The system includes a data acquisition and processing module, a data transmission and protocol conversion module, and a host computer auxiliary processing module. The method includes: The data acquisition and processing module acquires multi-source image data in real time and processes the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data. The high-definition multimedia interface image data is converted into standard universal serial bus image data for transmission through the data transmission and protocol conversion module. The host computer auxiliary processing module performs lightweight distortion correction on the standard universal serial bus image data to obtain panoramic conference image data.
[0012] To achieve the above objectives, a third aspect of this application proposes an electronic device comprising a memory, a processor, and the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system described in the first aspect above. When the processor executes the computer program, it implements the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method described in the second aspect above.
[0013] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras described in the second aspect.
[0014] To achieve the above objectives, the fifth aspect of this application proposes a computer program product comprising a computer program that, when executed by a processor, implements the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras described in the second aspect.
[0015] This application discloses an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system, method, electronic device, computer-readable storage medium, and computer program product. The system includes: a data acquisition and processing module for real-time acquisition of multi-source image data and processing of the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data; a data transmission and protocol conversion module for converting the high-definition multimedia interface image data into standard universal serial bus image data for transmission; and a host computer auxiliary processing module for performing lightweight distortion correction processing on the standard universal serial bus image data to obtain panoramic conference image data.
[0016] Thus, in this embodiment, multi-source image data is acquired in real time by the data acquisition and processing module and processed by the FPGA hardware to obtain high-definition multimedia interface image data. Then, the high-definition multimedia interface image data is converted into standard universal serial bus image data by the data transmission and protocol conversion module for transmission. Finally, the standard universal serial bus image data is subjected to lightweight distortion correction by the host computer auxiliary processing module to obtain panoramic conference image data. A three-layer collaborative architecture of "FPGA hardware core processing + protocol conversion + host computer lightweight distortion correction" is constructed. The FPGA hardware pipeline replaces the CPU software to process multiple images, reducing the delay of multi-image stitching and rotation to the millisecond level, thereby solving the defects of high latency and poor real-time performance of traditional multi-image stitching solutions.
[0017] Furthermore, compared to traditional solutions, this application's embodiment deploys complex, real-time-critical processing tasks such as stitching and rotation onto FPGA hardware, with the protocol conversion module responsible for compatibility adaptation and the host computer performing only lightweight distortion correction. This ensures low-latency processing while achieving broad device compatibility and avoids excessive computing resources being consumed by the host computer, thus making it suitable for panoramic conferencing application scenarios on various general-purpose computers (PCs). Attached Figure Description
[0018] Figure 1 Architecture diagram of an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system provided in this application embodiment; Figure 2 A flowchart illustrating the steps of the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras provided in some embodiments of this application; Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0020] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0022] First, the technical terms used in the embodiments of this application will be explained.
[0023] Field-Programmable Gate Array (FPGA): A semiconductor device that includes programmable logic components and interconnects, allowing parallel image stitching and audio signal processing to be implemented in the embodiments of this application.
[0024] Direction of Arrival (DOA) estimation: an algorithm that calculates the location of a sound source by measuring the phase difference of signals received by a microphone array.
[0025] Mobile Industry Processor Interface (MIPIS): An interface standard used to connect camera sensors to the main processor.
[0026] Inter-Integrated Circuit (IIC): In this embodiment, it is used to configure multiple camera sensors.
[0027] USB Video Class (UVC): A standard USB device class that allows camera devices to be recognized by the operating system without the need to install a dedicated driver.
[0028] Random Sample Consensus (RANSAC): Used to estimate mathematical model parameters in a sample set containing outlier data. In this embodiment, it is used to remove erroneous points in image feature matching.
[0029] Coaxial / Ring Microphone Array: refers to multiple microphones arranged in a specific geometric shape (a ring in this embodiment) for collecting spatial audio.
[0030] Next, the overall concept of this application will be explained.
[0031] With the normalization of remote work and multi-location collaborative work, the demand for video conferencing equipment is shifting from simply "being able to see" to "panoramic coverage, intelligent tracking, and minimalist deployment." In medium to large conference room scenarios, participants are scattered, and traditional single-lens cameras have limited field of view (usually <90°), making it difficult to cover the entire room. Existing multi-lens stitching devices often rely on the main CPU of a high-performance PC for software stitching, resulting in high video latency (often >100ms) and severely consuming the computing resources of the conference terminal, causing the conference software to lag. In addition, the linkage of audio and video usually requires the intervention of the operating system for scheduling, resulting in slow response and failing to achieve "sound and picture in sync."
[0032] In related technologies, common panoramic conferencing equipment typically employs a "multiple USB cameras + PC software stitching" solution. For example, some commercially available panoramic cameras only perform simple image synchronization internally, transmitting multiple raw images to the computer via USB. The computer's driver or dedicated software handles image distortion correction, stitching, and rendering. The implementation logic is: Camera A / B / C -> USB Hub -> PC CPU / GPU -> Software Decompression -> Feature Point Matching -> Software Stitching -> Virtual Camera Driver -> Conferencing Software. This solution is highly dependent on PC performance; when running high-load software (such as Tencent Meeting or Zoom), the video stitching frame rate drops significantly. Furthermore, because it requires the installation of dedicated drivers or software, it cannot be plug-and-play on all operating systems (such as Linux).
[0033] In summary, the relevant technologies have the following significant drawbacks: High splicing latency and poor real-time performance: Relying on CPU serial processing of high-resolution multi-channel images, the splicing latency is usually ≥50ms, and it is easy to cause motion blur and stuttering in dynamic images.
[0034] Poor compatibility and high deployment costs: Most devices require the installation of dedicated drivers or customized clients to output panoramic images, cannot be directly connected to general meeting software (such as Lark and Tencent Meeting) via standard USB protocol, and are incompatible with some operating systems.
[0035] Insufficient audio-visual coordination: There is a lack of an underlying mechanism for linking sound sources and viewing angles. Sound localization is typically handled by software, which then instructs the camera to rotate; this process is too long and makes it difficult to quickly focus on the speaker.
[0036] High input / output (IO) resource consumption: Multiple cameras independently occupy the control bus, resulting in a shortage of hardware pin resources, which increases the routing difficulty of the printed circuit board (PCB) and the system power consumption.
[0037] To address the aforementioned shortcomings, this application proposes an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system. This is a panoramic conference camera system based on FPGA hardware acceleration, aiming to achieve the following objectives: Achieving low-latency panoramic stitching: Utilizing FPGA hardware pipelines to replace CPU software processing reduces the latency of multi-channel image stitching and rotation to the millisecond level.
[0038] Achieving driverless and wide compatibility: The innovative "High Definition Multimedia Interface (HDMI) to USB protocol bridge" architecture is introduced to convert the complex FPGA processing results into standard UVC signals, enabling plug-and-play functionality.
[0039] Achieve hardware-level audio-video linkage: Directly control the video reading logic using audio data within the FPGA to achieve ultra-fast sound source tracking.
[0040] Achieve IO resource optimization: Use bus multiplexing technology to connect multiple cameras, reducing hardware costs and design complexity.
[0041] In this embodiment, multi-source image data is acquired in real time by a data acquisition and processing module and processed by FPGA hardware to obtain high-definition multimedia interface image data. Then, the high-definition multimedia interface image data is converted into standard universal serial bus image data by a data transmission and protocol conversion module for transmission. Finally, the standard universal serial bus image data is subjected to lightweight distortion correction by a host computer auxiliary processing module to obtain panoramic conference image data. A three-layer collaborative architecture of "FPGA hardware core processing + protocol conversion + host computer lightweight distortion correction" is constructed. The FPGA hardware pipeline replaces the CPU software to process multiple images, reducing the latency of multi-image stitching and rotation to the millisecond level, thereby solving the defects of high latency and poor real-time performance of traditional multi-image stitching solutions.
[0042] Furthermore, compared to traditional solutions, this application's embodiment deploys complex, real-time-critical processing tasks such as stitching and rotation onto FPGA hardware, with the protocol conversion module responsible for compatibility adaptation and the host computer performing only lightweight distortion correction. This ensures low-latency processing while achieving broad device compatibility and avoids excessive computing resources being consumed by the host computer, thus making it suitable for panoramic conferencing application scenarios on various general-purpose computers (PCs).
[0043] Based on the overall concept of this application, specific embodiments of the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system, method, electronic device, computer-readable storage medium, and computer program product provided in this application are proposed.
[0044] It should be noted that the embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0045] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0046] Furthermore, in various specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Moreover, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. Additionally, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after explicitly obtaining the user's separate permission or consent is the necessary user-related data for the proper functioning of these embodiments acquired.
[0047] Furthermore, the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras provided in this application can be applied to terminals, servers, or software running on either the terminal or the server. In some embodiments, the terminal can be an all-in-one smart conferencing machine, a panoramic video conferencing terminal, a remote medical teaching device, etc., or it can be a terminal configured / associated with these individuals, such as smartphones, tablets, laptops, desktop computers, and other electronic devices. These terminals can communicate and interact with corresponding individuals via a network. The server can be the backend server terminal device of the terminal, which can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The software can be an application, a computer program, and a storage medium carrying the computer program that implements the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras. It should be understood that, based on different design needs of practical applications, the terminal, server, and software of the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method provided in this application may also be other forms not listed here, depending on the different feasible embodiments. The FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method provided in this application does not specifically limit these.
[0048] Furthermore, this application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: distributed security monitoring systems, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, personal computers (PCs), minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0049] Next, we will first describe in detail various specific embodiments of an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system of this application.
[0050] For ease of understanding and explanation, the following text will use a panoramic conference camera multi-source image synchronous acquisition system that replaces FPGA acceleration as an example for detailed explanation. The implementation of any of the above-mentioned topics can be referred to the operation process of the assisted driving system described later.
[0051] In some embodiments, the FPGA-accelerated panoramic conferencing camera multi-source image synchronous acquisition system provided in this application may include: The data acquisition and processing module is used to acquire multi-source image data in real time and process the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data. The data transmission and protocol conversion module is used to convert the high-definition multimedia interface image data into standard universal serial bus image data for transmission. The host computer auxiliary processing module is used to perform lightweight distortion correction processing on the standard universal serial bus image data to obtain panoramic conference image data.
[0052] The system can acquire multi-source image data in real time through the data acquisition and processing module during panoramic conference image acquisition, and process the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data. Then, the high-definition multimedia interface image data is converted into standard universal serial bus image data for transmission through the data transmission and protocol conversion module. Finally, the standard universal serial bus image data is subjected to light distortion correction processing through the host computer auxiliary processing module to obtain panoramic conference image data.
[0053] It should be noted that the input to the data transmission and protocol conversion module is the high-definition multimedia interface image data output from the data acquisition and processing module (FPGA side), i.e., the HDMI video signal. Additionally, the input to the host computer auxiliary processing module is the USB video stream, i.e., the standard universal serial bus image data transmitted by the data transmission and protocol conversion module.
[0054] In some embodiments, the hardware configuration of the data transmission and protocol conversion module can be an HDMI to USB capture card chip (or integrated module). The data transmission and protocol conversion module can act as a simple protocol conversion bridge, encoding the HDMI™DS signal into USB Video Class (UVC) protocol data packets in real time, thereby outputting a standard USB video stream (standard Universal Serial Bus image data). Therefore, the data transmission and protocol conversion module can be recognized as a universal camera by Windows / Linux / macOS. Furthermore, this process does not involve complex image processing, only format encapsulation.
[0055] In some embodiments, the software architecture of the host computer-aided processing module can be a Python program running on a general-purpose PC (e.g., edited based on OpenCV). Thus, when performing distortion correction, the host computer-aided processing module can read USB video frames, call preset intrinsic parameter matrices (e.g., Camera Matrix) and distortion coefficients (e.g., DistortionCoefficients), and execute the cv2.undistort operation to correct fisheye / panoramic distortion into a planar image suitable for human viewing. Furthermore, the host computer-aided processing module can also call a Haar cascade classifier every N frames to detect faces and overlay green focus boxes or name tags on the image to achieve intelligent assistance in the meeting. Finally, the host computer-aided processing module pushes the processed image to the virtual camera interface for use by specific meeting software (e.g., Tencent Meeting, Lark), thereby outputting the final meeting screen displayed to the user, i.e., panoramic meeting image data.
[0056] For example, such as Figure 1 As shown, the system can be a three-layer architecture system consisting of "FPGA hardware core processing + protocol conversion bridging + host computer software assistance". The data acquisition and processing module, specifically the FPGA hardware acquisition and core processing module, is the core of the system, responsible for high real-time data acquisition and processing. The data transmission and protocol conversion module acts as a simple protocol conversion bridge, encoding the HDMI TMDS signal output by the FPGA into UVC protocol data packets in real time. This process does not involve complex image processing, only format encapsulation. The host computer auxiliary processing module (software side) performs distortion correction on the USB video stream and performs virtual streaming to obtain the conference screen displayed to the user.
[0057] In this embodiment, a three-layer collaborative architecture of "FPGA hardware core processing + UVC protocol conversion + host computer lightweight distortion correction" is constructed. The FPGA hardware pipeline replaces the CPU software to process multiple images, reducing the latency of multi-channel image stitching and rotation to the millisecond level. This solves the defects of high latency and poor real-time performance of traditional multi-channel image stitching solutions.
[0058] Furthermore, compared to traditional solutions, this application's embodiment deploys complex, real-time-critical processing tasks such as stitching and rotation onto FPGA hardware, with the protocol conversion module handling compatibility adaptation and the host computer performing only lightweight distortion correction. This ensures low-latency processing while achieving broad device compatibility and avoids excessive computing resources being consumed by the host computer, making it suitable for various general-purpose PC-based panoramic conferencing application scenarios.
[0059] Furthermore, in this embodiment, the system adopts an HDMI to UVC protocol conversion bridging architecture. The HDMI to USB capture card chip (or integrated module) encodes the HDMI TMDS signal output by the FPGA into standard UVC protocol data packets in real time. This process only performs format encapsulation and does not involve complex image processing. While realizing video signal protocol conversion, it enables the device to be recognized by operating systems such as Windows / Linux / macOS without drivers. Therefore, it can access general conferencing software without dedicated drivers or customized clients. This solves the problems of poor compatibility and high deployment costs of related devices used in traditional panoramic conferencing. It is suitable for plug-and-play scenarios in multi-system environments.
[0060] In some embodiments, the data acquisition and processing module includes a multi-view image acquisition unit, through which the data acquisition and processing module acquires multi-source image data in real time; The data acquisition and processing module polls the registers of the multi-view image acquisition unit through a set of serial computer buses, so that the multi-view image acquisition unit can synchronously transmit the multi-source image data through an independent interface.
[0061] It should be noted that the input to the multi-view image acquisition unit is the ambient light signal, and its hardware configuration can consist of multiple image sensors. For example, the multi-view image acquisition unit can consist of three OV5640 image sensors (220° ultra-wide angle), connected via a custom tri-view adapter board.
[0062] During system operation, the data acquisition and processing module can acquire multi-source image data in real time through this multi-view image acquisition unit. Furthermore, to address the limited FPGA pin count, the system can employ a time-division multiplexing / address remapping mechanism. That is, the data acquisition and processing module polls the registers of the multi-view image acquisition unit via a set of serial computer buses, allowing the multi-view image acquisition unit to synchronously transmit the multi-source image data through an independent interface.
[0063] For example, three cameras share a set of IIC buses (SCL / SDA) on the FPGA. During system power-on initialization, the FPGA controls the multiplexer on the adapter board via GPIO to poll and configure the registers of the three cameras. After configuration, the three cameras synchronously send image data through independent MIPIS interfaces.
[0064] In this embodiment, the FPGA controls the multiplexer on the adapter board through GPIO to realize a time-division multiplexing / address remapping mechanism for multiple image sensors to share a set of IIC buses: when the system powers on, it polls and configures the registers of each camera. After the configuration is completed, the image data is transmitted synchronously through an independent MIPIS interface. This can solve the problem of tight FPGA pin resources, while taking into account the synchronization of multi-view image acquisition and the simplicity of hardware design, reducing the difficulty of PCB routing and system power consumption, thus making it suitable for hardware resource optimization scenarios of multi-view panoramic camera equipment.
[0065] In some embodiments, the data acquisition and processing module includes an audio acquisition and positioning unit. The data acquisition and processing module acquires multiple audio data through the audio acquisition and positioning unit, and performs phase difference calculation and orientation calculation on the multiple audio data through a generalized cross-correlation algorithm to obtain the sound source angle signal of the multi-source image data.
[0066] It should be noted that the input to the audio acquisition and positioning unit is a spatial sound wave signal, and its hardware configuration can be a multi-channel microphone array (such as a 6-channel ring-shaped MSM261S4030HOR microphone array). The sound source angle signal can also be called the sound source azimuth angle signal.
[0067] During system operation, the data acquisition and processing module can acquire multiple audio data streams simultaneously with the multi-view image acquisition unit, thereby obtaining multiple audio data streams corresponding to the multi-source image data. Furthermore, the data acquisition and processing module can perform phase difference calculation and azimuth determination on these multiple audio data streams using a generalized cross-correlation algorithm, thus obtaining the sound source azimuth angle signal of the multi-source image data.
[0068] For example, a 6-channel ring-shaped MSM261S4030HOR microphone array can read 6 channels of audio data at a fixed sampling rate through the I²S interface. Phase analysis is performed by the FPGA internal logic to calculate the phase difference (TDOA) of the signals received by different microphones. The azimuth calculation is specifically based on a hardware IP core of the generalized cross-correlation algorithm (GCC-PHAT), which outputs the angle value of the sound source relative to the center of the device in real time (0-360°, resolution of 12 directions), thereby obtaining the sound source azimuth angle signal (Angle_Trigger).
[0069] In this embodiment, by integrating a hardware IP core based on the Generalized Cross-Correlation Algorithm (GCC-PHAT) within the FPGA, phase difference (TDOA) calculation and orientation determination are performed on the audio signals collected by the 6-channel ring microphone array. This results in the real-time output of sound source angle signals with a resolution of 0-360° and 12 directions, realizing real-time sound source localization and corresponding logic in hardware-based DOA algorithm. This can replace traditional software localization solutions while achieving millisecond-level low-latency sound source localization, thus providing reliable support for audio and video hardware-level linkage. It solves the problem of slow sound source localization response in traditional devices and is suitable for scenarios such as panoramic conferences that require rapid focusing of conference speakers.
[0070] In some embodiments, the data acquisition and processing module includes a core processing unit based on FPGA hardware. The data acquisition and processing module processes the multi-source image data through the core processing unit according to a preset full hardware processing pipeline to obtain high-definition multimedia interface image data that has been stitched together and aligned with the conference speaker. The fully hardware processing pipeline includes: feature extraction, feature matching, mismatch removal, image fusion, and electronic rotation.
[0071] During system operation, the data acquisition and processing module can use the FPGA-based core processing unit to take multi-source image data and its sound source angle signals as input. Following a full hardware processing pipeline, it first performs feature extraction (parallel FAST corner detection is performed on multi-source image data to extract feature points), feature matching (BRIEF descriptors are generated and feature points of adjacent images are matched), mismatch removal (using hardware-implemented RANSAC logic to remove erroneous matching points), image fusion (weighted average fusion of overlapping areas to generate a panoramic wide-angle image), and electronic rotation (combining the sound source azimuth angle signal to calculate the starting address for reading the panoramic image to capture the image from a specific perspective in the panoramic wide-angle image). Finally, it outputs a panoramic / close-up video stream (HDMI format, i.e., high-definition multimedia interface image data) that has been stitched together and aligned with the conference speaker.
[0072] In this embodiment, a fully hardware-pipeline multi-view image stitching and electronic pan-tilt control architecture is constructed within the FPGA, which includes "feature extraction-matching-mismatch removal-fusion-electronic rotation". This architecture executes FAST corner detection, BRIEF descriptor generation, RANSAC mismatch removal, and weighted average fusion of overlapping regions in parallel. Furthermore, "electronic rotation" is achieved by calculating the panoramic image reading start address using the sound source azimuth angle signal. This reduces the latency of multi-channel image stitching to milliseconds while enabling hardware-level audio-visual linkage, thus solving the problems of high latency, motion blur, and insufficient audio-visual coordination caused by traditional CPU serial software stitching solutions. This approach is suitable for real-time stitching and speaker tracking scenarios in panoramic conferences.
[0073] In some embodiments, the core processing unit is configured to perform weighted average fusion of overlapping regions of the multi-source image data with a preset fixed width based on the homography matrix of the multi-source image data.
[0074] It should be noted that the preset fixed width can be 18 pixels.
[0075] During system operation, the data acquisition and processing module can use the core processing unit based on FPGA hardware to calculate the homography matrix of the multi-source image data when performing "mismatch removal". Then, when performing "image fusion" processing, it can perform weighted average fusion of the overlapping areas (preset 18-pixel width) according to the homography matrix to generate a panoramic wide-format image.
[0076] In this embodiment, the system adopts a weighted average fusion and panoramic generation method for overlapping regions of multi-view images. By pre-setting a fixed-width overlapping region, the overlapping parts of the multi-view ultra-wide-angle images are weighted average fused based on the homography matrix. The fusion logic is implemented by FPGA hardware. In this way, while ensuring stitching efficiency, it can avoid the occurrence of discontinuities or ghosting at the stitching point, and ensure that the panoramic image is free of motion blur in dynamic scenes. It is suitable for image stitching optimization scenarios of multi-view panoramic camera equipment.
[0077] In some embodiments, the core processing unit is further configured to calculate the starting address for reading panoramic wide-angle image data based on the microphone array sound field matrix and perform audio-visual coordination control; the panoramic wide-angle image data is obtained by weighted average fusion of the overlapping regions by the core processing unit.
[0078] During system operation, the data acquisition and processing module can use the FPGA-based core processing unit to perform image fusion processing on multi-source image data to generate a panoramic wide-angle image. When further performing "electronic rotation" processing on the panoramic wide-angle image, the module locates the sound source angle signal based on the microphone array sound field matrix and calculates the read start address (Read_Address_Offset) of the panoramic wide-angle image based on the angle value of the sound source angle signal. Based on this read start address, the module can extract a specific viewpoint of the panoramic wide-angle image data from the output buffer, thereby realizing hardware-linked audio and video collaborative control.
[0079] In this embodiment, the system is based on a hardware-linked audio-video collaborative control method. By directly utilizing the angle signal output by the sound source localization module inside the FPGA, it controls the viewpoint capture logic after image stitching, forming a hardware-level closed-loop linkage of "audio acquisition - orientation calculation - video rotation". This can eliminate the long link delay of "software positioning - command issuance - camera mechanical rotation" while avoiding the intervention and scheduling of the operating system, thereby achieving extremely fast focusing of the speaker's viewpoint and solving the problem of insufficient audio-video collaboration in traditional devices. It is suitable for scenarios such as panoramic conferences that require rapid audio-video linkage.
[0080] Next, a complete embodiment of the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system provided in this application is presented.
[0081] In this embodiment, the data acquisition and processing module in the system is a hardware acquisition and core processing module on the FPGA side. It is the core of the system and is responsible for high real-time data acquisition and processing. Specifically, it may include: Multi-view image acquisition unit: Input: Ambient light signal; Hardware configuration: 3 OV5640 image sensors (220° ultra-wide angle), connected via a custom trinocular adapter board; Key logic (IIC bus multiplexing): Based on the time-division multiplexing / address remapping mechanism, the three cameras share a set of IIC buses (SCL / SDA) on the FPGA. When the system is powered on and initialized, the FPGA controls the multiplexer on the adapter board through GPIO to poll and configure the registers of the three cameras. After configuration, the three cameras send image data synchronously through independent MIPIS interfaces.
[0082] Output: Three synchronized raw RGB / YUV image data streams.
[0083] Audio acquisition and positioning unit: Input: Spatial acoustic wave signal; Hardware configuration: 6-channel MSM261S4030HOR microphone array with a circular layout; Logic (hardware implementation of DOA algorithm), including: Data acquisition: Six audio data streams are read at a fixed sampling rate via the I²S interface; Phase analysis: The FPGA internal logic calculates the phase difference (TDOA) of signals received by different microphones; Orientation calculation: Based on the hardware IP core of the generalized cross-correlation algorithm (GCC-PHAT), the sound source is output in real time as an angle value (0-360°, resolution of 12 directions) relative to the center of the device. Output: Sound source azimuth angle signal (Angle_Trigger).
[0084] FPGA core processing unit (hardware splicing and rotation): Input: Three channels of raw image data, sound source azimuth angle signal; Logic (full hardware pipeline), including: Feature extraction: Perform FAST corner detection on the three images in parallel to extract feature points; Feature matching: Generate BRIEF descriptors and perform feature point matching between adjacent images; False match removal: Using hardware-implemented RANSAC logic, false match points are removed and the homography matrix is calculated; Image fusion: Based on the homography matrix, the overlapping areas (preset 18-pixel width) are weighted and averaged to generate a panoramic wide-format image; Electronic pan-tilt control: Receives the "sound source azimuth angle signal", calculates the reading start address (Read_Address_Offset) of the panoramic image based on the angle value, and then captures the image from a specific viewpoint in the output buffer to achieve "electronic rotation"; Output: A stitched and aligned panoramic / close-up video stream (HDMI format) of the speaker.
[0085] In addition, the data transmission and protocol conversion module in the system takes the HDMI video signal output by the FPGA as input. The hardware configuration adopts an HDMI to USB capture card chip (or integrated module). Its operating logic is as follows: as a simple protocol conversion bridge, it encodes the HDMI TMDS signal into USB Video Class (UVC) protocol data packets in real time (this process does not involve complex image processing, only format encapsulation), and outputs a standard USB video stream (which can be recognized as a universal camera by Windows / Linux / macOS).
[0086] Furthermore, the host computer auxiliary processing module (software side) in the system takes USB video stream as input. Its software structure uses a Python program (based on OpenCV) that runs on a general-purpose PC, and its operating logic includes: Distortion removal: Read USB video frames, call the preset intrinsic parameter matrix (Camera Matrix) and distortion coefficients (Distortion Coefficients), and perform the cv2.undistort operation to correct fisheye / panoramic distortion into a planar image suitable for human viewing; Intelligent assistance: (optional) Call the Haar cascade classifier every N frames to detect faces and overlay a green focus box or name tag on the image; Virtual streaming: Pushes the processed image to the virtual camera interface for use by software such as Tencent Meeting and Lark.
[0087] Finally, the host computer auxiliary processing module outputs the final meeting screen displayed to the user.
[0088] Compared to the traditional "multiple USB cameras + PC software stitching" solution, this embodiment achieves breakthrough improvements in four core dimensions: real-time performance, compatibility, audio-visual collaboration, and hardware resource utilization. In terms of real-time performance, this embodiment replaces traditional CPU serial processing with an FPGA hardware pipeline architecture, executing core operations such as feature extraction, matching, fusion, and electronic rotation of multi-view images in parallel. Combined with the hardware implementation of the RANSAC algorithm, the panoramic stitching latency is reduced from ≥50ms in existing technologies to the millisecond level, completely solving the problems of frame rate drop, dynamic image ghosting, and stuttering in high-load conference scenarios, perfectly adapting to the dynamic needs of participants moving in medium to large conference rooms. Furthermore, regarding compatibility and ease of deployment, this embodiment innovatively adopts an "HDMI to UVC protocol bridging" design, directly encapsulating the video stream processed by the FPGA into a standard USB video signal. Without the need for dedicated drivers or customized clients, it can be recognized by operating systems such as Windows, Linux, and macOS, seamlessly integrating with common conferencing software, significantly reducing cross-platform deployment costs and solving the pain point of limited compatibility of related technologies.
[0089] Meanwhile, this embodiment utilizes a hardware-level audio-video linkage mechanism to directly control the panoramic image's viewing angle capture within the FPGA using the DOA positioning results of a 6-channel ring microphone array. This forms a closed-loop response of "audio acquisition - orientation calculation - video rotation," eliminating the need for operating system intervention and achieving extremely rapid speaker focusing from sound to image. Compared to the long chain of "software positioning - command issuance - mechanical rotation" in related technologies, the response speed is improved several times. Furthermore, this embodiment uses IIC bus multiplexing technology to achieve synchronous configuration of multiple cameras, effectively reducing hardware pin usage, PCB routing complexity, and system power consumption. It balances the requirements of device miniaturization and low power consumption, making it more suitable for deployment requirements in embedded scenarios such as smart conference all-in-one machines and remote medical teaching devices.
[0090] Next, specific embodiments of the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras proposed in this application will be presented.
[0091] The FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras proposed in this application is applied to the aforementioned FPGA-accelerated multi-source image synchronous acquisition system for panoramic conferencing cameras. This system includes a data acquisition and processing module, a data transmission and protocol conversion module, and a host computer auxiliary processing module.
[0092] It should be noted that the specific structure of the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system and the specific functional implementation process of each functional module therein have been described in the above specific embodiments. In the specific embodiments of the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method proposed in this application, the same content will not be repeated.
[0093] Please refer to Figure 2 , Figure 2 The flowchart illustrates the steps of the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras provided in this application in some embodiments. It should be understood that, although... Figure 2 The flowcharts illustrating subsequent steps show the execution order of some method steps. However, based on different design needs in practical applications, the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras provided in this application can, of course, employ a different execution order of method steps than those shown in the figures. That is, Figure 2 The order of the steps shown does not constitute a limitation on the execution logic order of the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method provided in this application. Any other method based on... Figure 2 Reasonable changes to the sequence of steps shown should be included within the protection scope of the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method provided in this application.
[0094] like Figure 2As shown, in some embodiments, the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method provided in this application may include steps S201 to S203 as shown below.
[0095] Step S201: The data acquisition and processing module acquires multi-source image data in real time and processes the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data.
[0096] During the panoramic conference image acquisition process, the system can acquire multi-source image data in real time through the data acquisition and processing module, and process the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data.
[0097] Step S202: The high-definition multimedia interface image data is converted into standard universal serial bus image data for transmission through the data transmission and protocol conversion module.
[0098] During system operation, the high-definition multimedia interface image data output by the data acquisition and processing module can be further converted into standard universal serial bus image data for transmission through the data transmission and protocol conversion module.
[0099] Step S203: Perform lightweight distortion correction processing on the standard universal serial bus image data through the host computer auxiliary processing module to obtain panoramic conference image data.
[0100] After the system converts the standard Universal Serial Bus image data through the data transmission and protocol conversion module, it further performs lightweight distortion correction on the standard Universal Serial Bus image data through the host computer auxiliary processing module to obtain panoramic conference image data.
[0101] For example, the system can construct a three-layer architecture system of "FPGA hardware core processing + protocol conversion bridging + host computer software assistance" through a data acquisition and processing module, a data transmission and protocol conversion module, and a host computer auxiliary processing module. In this system, the data acquisition and processing module, specifically the FPGA hardware acquisition and core processing module, is the core of the system, responsible for high real-time data acquisition and processing. The data transmission and protocol conversion module acts as a simple protocol conversion bridge, encoding the HDMI TMDS signal output by the FPGA into UVC protocol data packets in real time. This process does not involve complex image processing, only format encapsulation. The host computer auxiliary processing module (software side) performs distortion correction on the USB video stream and performs virtual streaming to obtain the conference screen displayed to the user.
[0102] It should be noted that the system adopts a three-layer architecture of "FPGA hardware core processing + protocol conversion bridging + host computer software assistance". The design goals and performance considerations of the system are as follows: to solve the pain points of incomplete panoramic coverage, high stitching latency, audio and video asynchrony, and inconvenient deployment in medium and large-scale conference scenarios, support low power consumption and low latency operation, and ensure that the system can run efficiently and stably in different hardware platforms and multiple operating system environments. By optimizing the hardware architecture and algorithm logic, hardware resource consumption and power consumption are reduced, while improving the robustness and adaptability of the system. This ensures that high-quality panoramic video is output in conference scenarios such as multi-person movement and complex sound fields, meets the integration requirements of embedded devices, and is easy to deploy and expand.
[0103] In some embodiments, based on hardware implementation preferences, the system can primarily implement the core functions of panoramic processing on FPGA hardware. It adopts key designs such as IIC bus multiplexing, hardware pipeline, and protocol bridging, which not only reduces hardware pin occupation and PCB routing difficulty, but also ensures the real-time performance and stability of processing, and is easier to integrate and deploy on resource-constrained embedded devices such as smart conference all-in-one machines and remote medical teaching equipment.
[0104] In some embodiments, combining the advantages of pure hardware core processing, the system can achieve panoramic stitching and rapid speaker tracking through FPGA hardware acceleration and deep audio-visual linkage. Furthermore, the system adopts an end-to-end hardware processing architecture, with core algorithms (FAST corner detection, RANSAC mismatch removal, and DOA localization) implemented through FPGA hardware logic, eliminating reliance on traditional CPU software serial processing and improving processing efficiency. Utilizing the parallel processing capabilities of FPGA ensures millisecond-level low latency and low power consumption while maintaining good panoramic image quality and audio-visual synchronization.
[0105] In this embodiment, the system employs a three-layer architecture of "FPGA hardware core + protocol conversion bridging + host computer software assistance." This architecture deploys real-time-critical functions such as stitching, sound source localization, and audio-visual linkage on the FPGA hardware. Multi-channel image and audio data are processed in parallel through a fully hardware pipeline. Cross-system compatibility is achieved through HDMI to UVC protocol bridging. The host computer performs only lightweight distortion correction and intelligent assistance functions, reducing system latency while avoiding excessive consumption of terminal computing resources, significantly improving the smoothness and deployment convenience of panoramic conferencing. Furthermore, this embodiment optimizes hardware resources through IIC bus multiplexing, ensures real-time performance through a fully hardware pipeline, and enhances compatibility through UVC protocol bridging, further improving the system's robustness. Moreover, this embodiment achieves high-performance panoramic conferencing functionality on resource-constrained embedded devices, is compatible with operating systems such as Windows and Linux, and various mainstream conferencing software, providing a robust and reliable solution for intelligent conferencing imaging systems.
[0106] Compared to traditional solutions (fisheye cameras + software processing), this embodiment offers several advantages in optical design and distortion control. It utilizes three 220° ultra-wide-angle non-fisheye cameras, employing hardware stitching and minor software distortion correction to achieve a panoramic image with minimal distortion and intact edge details. Traditional fisheye cameras rely on optical distortion for a wide field of view, resulting in severe edge stretching and requiring complex software algorithms for distortion correction, which can lead to detail loss and image blurring. Furthermore, at the core functionality logic level, this embodiment can build a full hardware pipeline within the FPGA, integrating image stitching and DOA positioning algorithms to form a hardware-level linkage of "audio acquisition - orientation calculation - video rotation" with millisecond-level response. Traditional fisheye solutions, on the other hand, execute distortion correction, stitching, and positioning step-by-step in software, resulting in a lengthy link, a total latency ≥60ms, and low linkage accuracy. Moreover, regarding latency and power consumption, the FPGA hardware processing architecture in this embodiment features low latency (stitching latency ≤8ms, linkage response latency ≤20ms) and low power consumption, meeting the requirements of embedded conferencing equipment (traditional fisheye solutions rely on PCs). Software processing suffers from high latency and consumes significant computing resources, with a total power consumption of ≥50W (PC + fisheye camera), limiting its practicality. Finally, in terms of scalability and modularity, this embodiment can add functions such as gesture recognition and people counting through the programmable characteristics of FPGA, exhibiting good scalability. Furthermore, the number of cameras can also be expanded. In contrast, the functions of traditional fisheye solutions rely on software algorithm upgrades, and the single-lens design cannot expand the field of view coverage, resulting in poor flexibility.
[0107] This application also provides an electronic device, which includes a memory, a processor, and the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system described in any of the above embodiments. The memory stores a computer program, and the processor executes the computer program to implement the above-described FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method.
[0108] In some embodiments, the electronics can be intelligent conference all-in-one machines, panoramic video conferencing terminals, remote medical teaching devices, etc., or they can be individual configured / associated terminals, such as smartphones, tablets, laptops, desktop computers, and other electronic devices.
[0109] Please see Figure 3 , Figure 3 This illustration shows the hardware structure of an electronic device according to one embodiment. The electronic device includes: The processor 301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 302 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 302 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 302 and is called and executed by the processor 301 to execute a method for simulating the visual state of a patient with visual impairment using AR collaborative simulation according to an embodiment of this application. Input / output interface 303 is used to implement information input and output; The communication interface 304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 305 transmits information between various components of the device (e.g., processor 301, memory 302, input / output interface 303, and communication interface 304); The processor 301, memory 302, input / output interface 303, and communication interface 304 are connected to each other within the device via bus 305.
[0110] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described FPGA-accelerated method for synchronous acquisition of multi-source images from a panoramic conferencing camera.
[0111] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0112] This application also provides a computer program product, including a computer program, the steps of which, when executed by a processor, are basically the same as the specific embodiments of the above-described FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method, and will not be repeated here.
[0113] The embodiments described above are for the purpose of more clearly illustrating the technical solutions of this application and do not constitute a limitation on the technical solutions provided in the embodiments of this application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0114] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0115] The embodiments of this application described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0116] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0117] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0118] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0119] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0120] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0121] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0122] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0123] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An FPGA-accelerated panoramic conferencing camera multi-source image synchronous acquisition system, characterized in that, The system includes: The data acquisition and processing module is used to acquire multi-source image data in real time and process the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data. The data transmission and protocol conversion module is used to convert the high-definition multimedia interface image data into standard universal serial bus image data for transmission. The host computer auxiliary processing module is used to perform lightweight distortion correction processing on the standard universal serial bus image data to obtain panoramic conference image data.
2. The system according to claim 1, characterized in that, The data acquisition and processing module includes a multi-view image acquisition unit, through which the data acquisition and processing module acquires multi-source image data in real time; The data acquisition and processing module polls the registers of the multi-view image acquisition unit through a set of serial computer buses, so that the multi-view image acquisition unit can synchronously transmit the multi-source image data through an independent interface.
3. The system according to claim 1, characterized in that, The data acquisition and processing module includes a core processing unit based on FPGA hardware. The data acquisition and processing module processes the multi-source image data through the core processing unit according to a preset full hardware processing pipeline to obtain high-definition multimedia interface image data that has been stitched together and aligned with the conference speaker. The fully hardware processing pipeline includes: feature extraction, feature matching, mismatch removal, image fusion, and electronic rotation.
4. The system according to claim 3, characterized in that, The core processing unit is used to perform weighted average fusion of the overlapping regions of the multi-source image data with a preset fixed width based on the homography matrix of the multi-source image data.
5. The system according to claim 4, characterized in that, The core processing unit is also used to calculate the starting address for reading panoramic wide-angle image data based on the microphone array sound field matrix and perform audio-visual coordination control; the panoramic wide-angle image data is obtained by weighted average fusion of the overlapping areas by the core processing unit.
6. The system according to any one of claims 1 to 5, characterized in that, The data acquisition and processing module includes an audio acquisition and positioning unit. The data acquisition and processing module acquires multiple audio data through the audio acquisition and positioning unit, and performs phase difference calculation and orientation calculation on the multiple audio data through a generalized cross-correlation algorithm to obtain the sound source angle signal of the multi-source image data.
7. A method for synchronous acquisition of multi-source images from a panoramic conferencing camera using FPGA acceleration, characterized in that, A panoramic conference camera multi-source image synchronous acquisition system for FPGA acceleration, the system includes a data acquisition and processing module, a data transmission and protocol conversion module, and a host computer auxiliary processing module; The method includes: The data acquisition and processing module acquires multi-source image data in real time and processes the multi-source image data based on FPGA hardware to obtain high-definition multimedia interface image data. The high-definition multimedia interface image data is converted into standard universal serial bus image data for transmission through the data transmission and protocol conversion module. The host computer auxiliary processing module performs lightweight distortion correction on the standard universal serial bus image data to obtain panoramic conference image data.
8. An electronic device, characterized in that, The electronic device includes a memory, a processor, and an FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition system as described in any one of claims 1 to 6. The memory stores a computer program, and when the processor executes the computer program, it implements the FPGA-accelerated panoramic conference camera multi-source image synchronous acquisition method as described in claim 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras as described in claim 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the FPGA-accelerated multi-source image synchronous acquisition method for panoramic conferencing cameras as described in claim 7.