Multi-mode graphic image processing equipment

Through the multi-modal graphic image processing equipment integrating gesture recognition and speech recognition modules, the problem of single interaction mode of image processing equipment and inefficient format conversion is solved, and efficient human-machine collaboration and automated image processing are achieved.

CN120355560AInactive Publication Date: 2025-07-22SHAOYANG SENIOR TECHNICAL SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363826.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In prior art In multimedia editing, industrial design and medical image processing, image processing equipment has problems such as single interaction mode, multimodal splitting, and inefficient format conversion, making it difficult to achieve efficient input and automatic format optimization of complex instructions.

Method used

It adopts multi-modal graphic image processing equipment, integrates gesture recognition module, voice recognition module, multi-modal fusion control unit and output interface, supports dynamic gesture tracking, offline/cloud voice command analysis, real-time format conversion and adaptive optimization, and improves processing efficiency through hardware acceleration units.

Benefits of technology

It realizes efficient human-computer collaboration for offices in various occasions, improves the efficiency and automation of graphics processing, and solves the problems of single interaction modes and inefficient format conversion of image processing equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355560A_ABST
    Figure CN120355560A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode graphic image processing device comprising a gesture recognition module used for collecting a user gesture signal through a depth camera or an optical sensor; the voice recognition module is used for collecting a voice instruction through a microphone array and converting the voice instruction into a control signal; the image processing module is used for receiving an input image and executing format conversion, resolution adjustment and compression operations; the multi-mode fusion control unit is used for integrating the gesture signal, the voice instruction and the image processing logic to generate an output instruction; the output interface supports automatic generation of JPEG, PNG, PDF and SVG formatted files and outputs the formatted files through a wired or wireless transmission protocol; according to the invention, the gesture recognition module and the multi-mode fusion control unit are adopted, and efficient man-machine cooperation of office work in various occasions can be realized by integrating gestures, voice instructions and automatic image processing, so that the purpose of improving the graphic processing efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of computer vision and human - computer interaction technology, and particularly relates to a multi - modal graphic image processing device. This multi - modal graphic image processing device is an intelligent image processing device integrating gesture recognition, voice control, and multi - format output, and is applicable to multimedia editing, industrial design, and medical image processing scenarios. Background Art

[0002] In multimedia editing, industrial design, and medical imaging, where a large amount of image processing is required, the existing technology currently has the following several defects in image processing: First, the image interaction method is single: traditional devices rely on keyboard / touch operations and it is difficult to achieve efficient input of complex instructions (such as "simultaneously adjusting image rotation and transparency"); Second, multi - modal fragmentation: gesture and voice control modules operate independently and lack a priority coordination mechanism (for example, there is no processing logic when a gesture zoom conflicts with the voice command "save"); Third, format conversion is inefficient: it is necessary to manually select the output format and it cannot be automatically optimized according to the target scenario (such as PNG for web adaptation, PDF for printing adaptation). Therefore, further technological innovation and improvement are required for image processing devices. Summary of the Invention

[0003] The present invention aims at the above - mentioned problems and proposes a multi - modal graphic image processing device to solve these problems.

[0004] To achieve the above - mentioned technical purpose, the present invention adopts a multi - modal graphic image processing device, including: a gesture recognition module for collecting user gesture signals through a depth camera or an optical sensor;

[0005] a voice recognition module for collecting voice commands through a microphone array and converting them into control signals; an image processing module for receiving input images and performing format conversion, resolution adjustment, and compression operations;

[0006] a multi - modal fusion control unit for integrating gesture signals, voice commands, and image processing logic to generate output commands;

[0007] an output interface that supports automatically generating JPEG, PNG, PDF, and SVG format files and outputting them through wired or wireless transmission protocols.

[0008] As a preference of the present invention, the gesture recognition module supports dynamic gesture tracking, including:

[0009] a gesture classification model based on a convolutional neural network;

[0010] realizing three - dimensional space gesture positioning through skeleton key - point detection.

[0011] As a further preference of the present invention, the speech recognition module supports an offline instruction set and cloud semantic analysis, including: predefined image processing instructions such as "save as PDF" and "adjust resolution"; a natural language processing engine for parsing complex instructions.

[0012] As a further preference of the present invention, the image processing module supports real-time format conversion, including:

[0013] Automatically selecting the target format according to the output interface type;

[0014] An adaptive optimization algorithm based on the user's historical operation data.

[0015] Furthermore, the present invention further includes:

[0016] A visual interaction interface for displaying the gesture operation trajectory and real-time feedback of voice instructions;

[0017] A hardware acceleration unit such as a GPU or FPGA for improving the image processing speed.

[0018] The present invention adopts a gesture recognition module and a multi-modal fusion control unit, which can realize efficient human-machine collaboration in various office scenarios by integrating gestures, voice instructions and automated image processing, so as to achieve the purpose of improving the graphics processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The figure shows the device system architecture diagram of the present invention (specifically showing the linkage of sensors, processing units and output modules);

[0020] Figure 2 The figure shows the multi-modal control flow chart of the present invention (including the parallel processing logic of gestures and voice instructions);

[0021] Figure 3 The figure shows the schematic diagram of the format conversion algorithm of the present invention (a decision tree model based on format priority);

[0022] Figure 4 The figure shows the multi-modal instruction conflict resolution flow chart of the present invention (a timing alignment and weight calculation module);

[0023] Figure 5 The figure shows the format conversion adaptive decision tree of the present invention (including three-layer branches of device type, network status and user history);

[0024] Figure 6 The figure shows the hardware acceleration architecture diagram of the present invention (including GPU / FPGA task allocation and data pipeline); DETAILED DESCRIPTION OF THE INVENTION

[0025] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0026] A multimodal graphic image processing device specifically includes:

[0027] A gesture recognition module for collecting user gesture signals through a depth camera or an optical sensor;

[0028] A voice recognition module for collecting voice commands through a microphone array and converting them into control signals; An image processing module for receiving input images and performing format conversion, resolution adjustment, and compression operations;

[0029] A multimodal fusion control unit for integrating gesture signals, voice commands, and image processing logic to generate output commands;

[0030] An output interface that supports automatically generating files in JPEG, PNG, PDF, and SVG formats and outputting them through wired or wireless transmission protocols.

[0031] As a preference of the present invention, the gesture recognition module supports dynamic gesture tracking, including:

[0032] A gesture classification model based on a convolutional neural network (CNN);

[0033] Three-dimensional space gesture positioning is achieved through skeleton key point detection.

[0034] In the present invention, the voice recognition module supports an offline instruction set and cloud semantic analysis, including:

[0035] Predefined image processing instructions (such as "save as PDF", "adjust resolution");

[0036] A natural language processing (NLP) engine parses complex instructions.

[0037] In the present invention, the image processing module supports real-time format conversion, including:

[0038] Automatically select the target format according to the output interface type;

[0039] An adaptive optimization algorithm based on the user's historical operation data.

[0040] The present invention further includes:

[0041] A visual interaction interface for displaying the real-time feedback of gesture operation trajectories and voice commands;

[0042] A hardware acceleration unit (such as a GPU or an FPGA) for improving the image processing speed.

[0043] The core innovation of the present invention lies in the adoption of a multimodal collaborative control architecture, specifically including:

[0044] Dynamic priority allocation: Adjust the weights of gestures and voice commands based on the context scenario (such as design mode / medical mode).

[0045] Example: In the medical imaging mode, the voice command "Mark the lesion area" takes precedence over gesture operations. Conflict resolution algorithm: Solve the timing conflicts of multi-modal commands through timestamp comparison and semantic analysis.

[0046] The present invention also includes an adaptive format conversion engine, specifically including:

[0047] Scene-aware decision tree: Automatically select the format according to the output device type (such as monitor / printer) and network bandwidth;

[0048] Cross-modal feature fusion: Jointly map the keywords in the voice command (such as "high resolution") and the gesture trajectory data to the image processing parameters.

[0049] The present invention also includes a low-latency hardware acceleration solution, specifically including:

[0050] Heterogeneous computing architecture: GPU parallel processes image format conversion, and FPGA realizes real-time tracking of gesture skeleton points;

[0051] Edge-cloud collaboration: Complex semantic commands (such as "Generate a 3D preview image") call cloud computing power through a 5G link.

[0052] The hardware configuration of the present invention is as follows:

[0053] Gesture recognition: Use an Intel RealSense D455 depth camera, support 30FPS dynamic gesture tracking, with an error <1mm;

[0054] Voice recognition: Use a 6-microphone circular array, beamforming noise reduction, support offline instruction set (more than 100 commands);

[0055] Image processing: Use an NVIDIA Jetson AGX Xavier embedded GPU. 32 TOPS computing power, support TensorRT acceleration.

[0056] The software process of the present invention includes:

[0057] Gesture signal processing: Use the OpenPose model to extract 21 key points of the hand and generate a three-dimensional space trajectory;

[0058] Identify continuous gesture semantics through an LSTM network (such as "drawing a circle" corresponding to image cropping).

[0059] Voice command parsing: The local ASR engine (based on Wav2Vec 2.0) achieves low-latency conversion with a response time < 200ms;

[0060] The cloud NLP service parses complex commands (such as "sort all pictures by date and export as PDF").

[0061] Image format conversion: Call the FFmpeg library to achieve real-time transcoding, supporting emerging formats such as HEIC / WebP;

[0062] Learn user preferences based on historical data (e.g., designers prefer SVG, doctors prefer DICOM).

[0063] Specific application scenario embodiments of the present invention include but are not limited to the following:

[0064] Medical imaging processing scenario: Doctors gesture to draw the lesion boundary to the voice command "mark the malignant probability" to the device to automatically generate an annotated DICOM file; The system converts the image into a structured report that conforms to the DICOM SR standard according to the PACS (Picture Archiving and Communication System) interface requirements.

[0065] Figure 1 Shown is the device system architecture diagram of the present invention, which specifically shows the linkage of sensors, processing units, and output modules.

[0066] Figure 2 Shown is the multi-modal control flow chart of the present invention, including the parallel processing logic of gestures and voice commands.

[0067] Figure 3 Shown is the schematic diagram of the format conversion algorithm of the present invention, mainly a decision tree model based on format priority.

[0068] Figure 4 Shown is the multi-modal instruction conflict resolution flow chart of the present invention, which involves a timing alignment and weight calculation module. The flow chart is explained as follows:

[0069] Multi-modal input: Simultaneously receive instructions from the gesture recognition module, voice recognition module, and image processing module.

[0070] Multi-modal fusion control unit: This is the core of the entire process, responsible for integrating and processing multi-modal instructions.

[0071] Timing alignment module: Check the timestamps of each instruction to ensure that the instructions are synchronized. If the instructions are not synchronized, wait for the next instruction or re-check the timestamps until all instructions are synchronized.

[0072] Weight calculation module: Calculate the weight of each instruction according to preset rules (such as user preferences, system configuration, security, etc.). The weight can be dynamically adjusted to reflect priorities in different situations.

[0073] Select the instruction with the highest weight: After the weight calculation is completed, select the instruction with the highest weight as the output instruction.

[0074] Generate the output instruction: Generate the final output instruction according to the selected instruction, and this instruction will be sent to the output interface.

[0075] Output interface: Support automatic generation of files in JPEG, PNG, PDF, and SVG formats, and output through wired or wireless transmission protocols.

[0076] End: The process ends.

[0077] This flowchart shows the basic process of multimodal instruction conflict resolution, including timing alignment, weight calculation, and instruction selection. In actual applications, these steps may be more complex and may require considering more factors and conditions.

[0078] Figure 5 What is shown is the format conversion adaptive decision tree of the present invention, which mainly includes three layers of branches: device type, network status, and user history.

[0079] Regarding Figure 5 The logical description of the three-layer branches is as follows:

[0080] 1. Device type branch part;

[0081] (1) Decision objective: Optimize format selection according to device hardware capabilities;

[0082] (2) Mobile devices: Prioritize high compression rate formats (such as JPEG) to reduce processing load and storage occupancy;

[0083] (3) Desktop devices: Enable three modes according to user selection:

[0084] (4) High compression rate priority: JPEG (for fast transmission);

[0085] (5) Balanced mode: PNG (balanced quality and compression);

[0086] (6) Original image quality priority: SVG / PDF (vector or lossless format);

[0087] 2. Network status branch;

[0088] (1) Decision objective: Dynamically adjust the format according to the real-time bandwidth;

[0089] (2) Bandwidth > threshold (high bandwidth): Allows large file transfers. Prefer SVG / PDF (high-fidelity) or PNG (lossless compression).

[0090] (3) Bandwidth ≤ threshold (low bandwidth): Forces the use of JPEG (high compression rate) to reduce transmission latency.

[0091] 3. User history branch;

[0092] (1) Decision-making goal: Personalized recommendation based on user behavior data;

[0093] (2) User's preferred format: Directly output the historical preference format (e.g., the user uses PDF in 80% of the scenarios);

[0094] (3) No historical record: According to the default priority (JPEG > PNG > PDF > SVG);

[0095] Now, an example of the adaptive decision-making process is presented.

[0096] 1. Scenario: The user is operating on a mobile device, the current network bandwidth is low, and the historical record shows that the commonly used format is SVG;

[0097] 2. Device type branch: Mobile device triggers the high compression rate priority mode (default JPEG);

[0098] 3. Network status branch: Low bandwidth forces the use of JPEG;

[0099] 4. User history branch: The user commonly uses SVG, which conflicts with the network branch;

[0100] 5. Conflict resolution: The network status weight is greater than the user history weight, resulting in the final output of the JPEG dynamic weight rule (extensible);

[0101] 6. Output interface logic: Generate a target format file according to the decision tree result. If the target format is the same as the input format, skip the conversion step. Automatically append metadata (such as device type, network status marker). Transmit through a wireless protocol (Wi-Fi / 5G) or a wired interface (USB / HDMI). This decision tree achieves multi-dimensional adaptability through hierarchical logic, ensuring both real-time performance (network status first) and taking into account personalization (user history) and device adaptability. Branch rules can be extended or weight allocation can be adjusted according to actual needs.

[0102] Figure 6 The following shows the hardware acceleration architecture diagram of the present invention, which mainly includes GPU / FPGA task allocation and data pipeline. The task allocation strategy of this hardware acceleration architecture diagram is as follows:

[0103] 1. FPGA core tasks;

[0104] (1) High real-time requirements:

[0105] (1 Gesture signal optical flow calculation (hardware parallelized pixel processing);

[0106] (2 Voice endpoint detection (FIR filter hardware implementation);

[0107] (3 Wireless protocol encapsulation (MAC layer data packet generation);

[0108] (2) Low-power scenarios:

[0109] (1 Image preprocessing (YUV→RGB conversion);

[0110] (2 Sensor data verification (CRC hardware acceleration);

[0111] 2. GPU core tasks;

[0112] (1) Computation-intensive tasks:

[0113] (1 Gesture semantic recognition (CNN model inference based on TensorRT);

[0114] (2 Speech recognition (RNN-T model inference);

[0115] (3 Image compression (parallelization of DCT transform for JPEG encoding);

[0116] (2) Vector graphics processing:

[0117] (1 SVG path generation (CUDA-accelerated Bezier curve calculation)

[0118] (2 PDF document synthesis (multi-layer Alpha blending)

[0119] 3. Heterogeneous cooperation mechanism;

[0120] (1) Data pipeline optimization;

[0121] (1 The gesture coordinate data after FPGA preprocessing is directly written into the GPU video memory through DMA;

[0122] (2 The compressed image data of the GPU is transmitted to the FPGA output queue through zero-copy memory mapping;

[0123] (2) Task-level pipeline:

[0124] (1 Key hardware design details;

[0125] (2 FPGA logic unit allocation;

[0126] (3Regarding 20% of the resources for the sensor interface (MIPI CSI-2 / RGB protocol parsing);

[0127] (4Regarding 40% of the resources for the real-time processing pipeline (gesture + voice parallel pipeline);

[0128] (5Regarding 30% of the resources for the output protocol control (USB 3.0 PHY / IP core);

[0129] (6Regarding 10% of the resources for exception handling (watchdog timer / ECC check);

[0130] (7Batch processing strategy: Voice commands are segmented and batch-inferred every 50 ms to reduce the kernel startup overhead;

[0131] (8Memory multiplexing: Intermediate results of image processing are retained in the video memory to avoid CPU-GPU data transfer;

[0132] (9Mixed precision: The gesture recognition model uses FP16 + INT8 mixed quantization (TensorCore acceleration) for energy efficiency ratio control;

[0133] (10Regarding FPGA dynamic frequency adjustment: Switch between 200 MHz / 400 MHz modes according to the task load;

[0134] (11Regarding the GPU power wall: Limit the TDP to 75W (adjusted through nvidia-smi);

[0135] (12Hardware sleep during idle: Turn off unused GPU SM units and FPGA logic block performance;

[0136] An example of the metrics is shown in Table 1 of the following table:

[0137] Table 1: Example of performance metrics

[0138] Task FPGA Latency GPU Throughput Energy Efficiency Ratio (TOPS / W) Gesture Recognition 2.1ms - 8.2 (FPGA) Speech Recognition (1-second Audio) - 3200 Samples per Second 4.5 (GPU) JPEG Compression (4K Image) - 120 Frames per Second 6.8 (GPU) End-to-End Output Latency ≤25ms - -

[0139] This architecture maximizes the utilization of computing resources while ensuring real-time performance through the heterogeneous cooperation of FPGA and GPU. You can adjust the allocation ratio of logic units and CUDA cores according to the specific hardware model (such as Xilinx UltraScale+ FPGA or NVIDIA A2 GPU).

[0140] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can still be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A multimodal graphic image processing device, characterized in that, Including: A gesture recognition module for collecting user gesture signals through a depth camera or an optical sensor; A speech recognition module for collecting voice commands through a microphone array and converting them into control signals; An image processing module for receiving an input image and performing format conversion, resolution adjustment, and compression operations; A multi-modal fusion control unit for integrating gesture signals, voice commands, and image processing logic to generate output commands; An output interface that supports automatically generating files in JPEG, PNG, PDF, and SVG formats and outputting them through wired or wireless transmission protocols.

2. The multimodal graphic image processing device according to claim 1, wherein The gesture recognition module supports dynamic gesture tracking, including: A gesture classification model based on a convolutional neural network; Three-dimensional space gesture positioning is achieved through skeleton key point detection.

3. A multimodal graphic image processing device according to claim 1, wherein, The speech recognition module supports an offline instruction set and cloud semantic analysis, including: Pre-defined image processing instructions; A natural language processing engine for parsing complex instructions.

4. A multimodal graphic image processing device according to claim 1, characterized in that, The image processing module supports real-time format conversion, including: Automatically selecting a target format according to the output interface type; An adaptive optimization algorithm based on the user's historical operation data.

5. A multimodal graphic image processing device according to any one of claims 1-4, characterized in that It also includes: A visual interaction interface for displaying the real-time feedback of gesture operation trajectories and voice commands; A hardware acceleration unit for improving the image processing speed.