Edge deployment system and method for distributed visual language model

By separately deploying the visual coding module and the language inference module on edge devices and inference devices, and using dynamic security protocols for feature transmission, the problems of high inference delay and insufficient privacy protection in the prior art are solved, and efficient edge deployment and privacy protection are achieved.

CN120088624AInactive Publication Date: 2025-06-03ZHEJIANG SCI-TECH UNIV

Patent Information

Application Number
CN202510571622.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing distributed visual language model based on ViT and Transformer has high inference latency on edge devices, severe computing resource constraints, and lacks privacy protection mechanisms, which poses a risk of data leakage.

Method used

An edge deployment system of distributed visual language model is proposed. By separating and deploying the visual coding module and language inference module on edge devices and inference devices, using dynamic security protocols for feature transmission, realizing computational load separation and privacy protection.

Benefits of technology

It reduces the computing load of edge devices, ensures privacy and security, and realizes efficient feature extraction at the edge and joint inference with backend, meeting real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088624A_ABST
    Figure CN120088624A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and edge computing crossing technologies, in particular to an edge deployment system and method of a distributed visual language model, a storage medium and a processor. The invention provides an edge deployment system and method for a distributed visual language model, a storage medium and a processor. The edge deployment system comprises a distributed topological structure with a front end and a rear end working cooperatively. The front end is an edge device end, and the rear end is a reasoning device end; the visual coding module is deployed at an edge device end, and the language reasoning module is deployed at a reasoning device end; and the visual coding module and the language reasoning module carry out feature transmission through a dynamic security protocol. The visual coding module and the multi-modal reasoning module are separately deployed in different hardware devices, edge end efficient feature extraction and back end joint reasoning are achieved through modular design, and the calculation load of the edge devices is reduced while privacy security is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - technical field of artificial intelligence (AI) and edge computing, and particularly to a distributed vision - language model (VLM) architecture optimized for edge computing environments and its deployment method. Background Art

[0002] The architectures based on Vision Transformer (ViT) and Transformer have problems with computational resource constraints. This high computational complexity results in the inference latency of existing models on edge devices generally exceeding 500ms, which cannot meet the stringent real - time requirements of scenarios such as intelligent security and medical monitoring. The multi - modal fusion process further exacerbates the computational pressure, and the single - inference FLOPs of typical VLM architectures exceed 200G, making it almost infeasible to deploy on edge devices with less than 128MB of memory.

[0003] The lack of privacy protection mechanisms is another key defect. Traditional cloud - edge collaborative architectures adopt raw image or plain - text feature transmission schemes, and such methods have risks of privacy leakage. Attackers can recover the original file content through feature inversion techniques, which may cause the leakage of patients' biometric data in sensitive scenarios such as medical image analysis.

[0004] Existing model compression technologies have limitations in multi - modal scenarios. These methods generally require the support of dedicated accelerators and cannot adapt to heterogeneous edge computing environments, resulting in an increase in actual deployment costs.

[0005] Therefore, it is necessary to propose an edge deployment system, method, storage medium, and processor for a distributed vision - language model. Summary of the Invention

[0006] The purpose of the present invention is to overcome the above - mentioned deficiencies of the prior art and provide an edge deployment system, method, storage medium, and processor for a distributed vision - language model, aiming to solve any of the problems raised in the background art.

[0007] To achieve the above - mentioned purpose, in the first aspect, the present invention proposes an edge deployment system for a distributed vision - language model, including a distributed topology structure in which the front - end and the back - end work collaboratively; the front - end is the edge device side, and the back - end is the inference device side; the visual encoding module is deployed on the edge device side, and the language inference module is deployed on the inference device side; the visual encoding module and the language inference module transmit features through a dynamic security protocol.

[0008] Preferably, it further includes an image sensor module deployed at the edge device end to capture and preprocess the original RGB image data; a visual encoder module to construct an initial convolutional layer from the input preprocessed image data; a feature compression module deployed at the edge device end to serialize the features after the initial convolutional layer is stacked through multiple bottleneck modules and global feature aggregation; and an encoding and encryption module deployed at the edge device end to encrypt the encoding output from the feature compression module.

[0009] Preferably, a decoding and decryption module is deployed at the inference device end to deserialize and align the feature data packets from the edge device end; a multi-modal fusion module is deployed at the inference device end to input visual features and text features into a Transformer-based multi-modal fusion module; a language inference module is deployed at the inference device end to complete the inference and decision-making generation tasks; and an application interface module is deployed at the inference device end to encapsulate the text results generated at the inference device end.

[0010] In a second aspect, the present invention proposes an edge deployment method for a distributed vision-language model, which is applied to an edge deployment system of a distributed vision-language model, and includes the following steps: S1: Image sensor initialization and data acquisition, where the edge device end captures and preprocesses the original RGB image data; S2. Lightweight visual encoder forward inference, where the preprocessed image is input into the lightweight visual encoder module deployed at the edge device end S3. Feature serialization is performed after stacking through multiple bottleneck modules and global feature aggregation; S4. A long connection is established between the edge device end and the inference device end through the MQTT protocol, and the feature data packets are transmitted through a TLS encrypted channel; S5. Feature deserialization and alignment are completed through the decoding and decryption module; S6. Multi-modal joint encoding, where visual features and text features are input into a Transformer-based multi-modal fusion module; S7. Inference and decision-making generation tasks are completed through language inference; S8. The application interface encapsulates the text results generated at the inference device end and can perform downstream transmission and edge-end response.

[0011] Preferably, an image sensor module is built into the edge device end, and hardware initialization configuration is completed at the startup stage, including setting the resolution, frame rate, exposure parameters, and white balance. The original RGB image data is captured in real time through the image sensor interface and temporarily stored in the device memory buffer. For dynamic scenes, a multi-frame buffer mechanism is supported to eliminate motion blur and ensure that the input image quality meets the requirements of subsequent processing; The preprocessing process of the picture is as follows: Color space conversion: Convert the original Bayer format data to the standard RGB color space, and apply gamma correction and demosaicing algorithms to eliminate sensor noise; Size normalization: Scale the image to the preset input size through bilinear interpolation algorithm to adapt to the input requirements of the vision encoder; Pixel normalization: Perform mean subtraction and standard deviation normalization on each pixel channel , where is the normalized output value, is the original input, is the channel mean of the training dataset, is the standard deviation, so as to eliminate the influence of illumination changes.

[0012] Preferably, the processing flows in steps S2 and S3 are as follows: a. Initial convolution layer: Use a suitable convolution kernel and stride to expand the channels of the input image to obtain a channel feature map, and perform sampling in the spatial dimension; b. Stacking of multi-level bottleneck modules: Sequentially perform multiple groups of inverted residual bottleneck operations. Each bottleneck module includes an expansion convolution, a depthwise separable convolution, a squeeze-and-excitation module, and a projection layer. Among them, the squeeze-and-excitation module obtains channel attention weights through global average pooling and dynamically adjusts the importance of feature channels; c. Global feature aggregation: The final feature map is compressed in dimension by global average pooling of average pooling, and then mapped to a compact feature vector by a fully connected layer; d. Feature serialization: Divide the feature vector into multiple sub-vectors according to the preset length, and form a visual token sequence through dimension compression , where L is the sequence length and d is the feature dimension.

[0013] Preferably, the specific implementation method in S4 is as follows: Certificate mutual authentication: The device side and the server side exchange X.509 certificates to verify the legitimacy of the identities of both communication parties; Session key negotiation: Use the ECDHE key exchange algorithm to generate a temporary session key to ensure forward security; Packet fragmentation transmission: Large feature packets are fragmented and transmitted according to the MTU size, and the receiver verifies the integrity after recombination; The specific decryption and decoding processing method in S5 is as follows: Verification and parsing: Verify the CRC check code, parse the packet header metadata, and reconstruct the feature sequence Z; Modal alignment: Project the visual token sequence Z and the text embedding vector E into the same semantic space, and adjust the feature scale and offset through linear transformation.

[0014] Preferably, the multi-modal fusion process in S6 is as follows: Position encoding injection: Add learnable position encodings for visual tokens and text tokens to capture sequence order information; Cross-modal attention calculation: Through the multi-head attention mechanism, text tokens are used as Queries to retrieve relevant visual features, and visual tokens are used as Keys and Values to participate in the weight calculation to achieve information interaction between modalities; Hierarchical feature aggregation: Alternately stack self-attention layers and cross-attention layers to gradually fuse visual and text semantic information to generate a joint representation , where M is the length of the text sequence; In S7, large model inference is performed and the final result is output. The specific method is as follows: Autoregressive decoding: Based on the greedy search or beam search algorithm, generate text responses word by word, and calculate the probability distribution of the next word at each time step , and select the highest probability token and append it to the output sequence; Logical constraint injection: Apply a preset rule template to post-process the generated results to ensure domain compliance; Confidence calibration: Perform temperature scaling and Platt scaling on the output probability distribution to improve the accuracy of confidence estimation and filter out low-confidence predictions.

[0015] In a third aspect, the present invention proposes a storage medium, which includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the edge deployment method of the distributed vision-language model.

[0016] In a fourth aspect, the present invention proposes a processor, which is used to run a program. When the program runs, it executes the edge deployment method of the distributed vision-language model.

[0017] Compared with the prior art, the beneficial effects of an edge deployment system, method, storage medium, and processor for a distributed vision-language model provided by the present invention are as follows: Separate the visual encoding and multi-modal inference modules and deploy them on different hardware devices. Through modular design, efficient feature extraction at the edge and joint inference at the backend are achieved, reducing the computing load of edge devices while ensuring privacy and security. The specific implementation process includes four main stages: visual information acquisition and encoding at the edge, feature data transmission, multi-modal fusion and inference at the backend, and result feedback. Each stage works together to form a complete closed loop.

[0018] The features and advantages of the present invention will be described in detail through embodiments in combination with the accompanying drawings. Description of the Drawings

[0019] Figure 1 It is a technical architecture diagram of the model system.

[0020] Figure 2 It is the architecture diagram of the transmission protocol stack.

[0021] Figure 3 It is the architecture diagram of the visual feature encoding module.

[0022] Figure 4 It is the architecture diagram of decoding and inference. Detailed implementation manners

[0023] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0024] Refer to Figure 1 , an edge deployment system for a distributed vision-language model is provided in an embodiment of the present invention. It is characterized in that it includes a distributed topology structure in which the front end and the back end cooperate; the front end is the edge device end, and the back end is the inference device end; the visual encoding module is deployed at the edge device end, and the language inference module is deployed at the inference device end; the visual encoding module and the language inference module perform feature transmission through a dynamic security protocol.

[0025] Among them, the lightweight visual encoding module deployed at the edge device end is mainly used to extract compressed visual features. The multimodal inference module deployed at the back end (cloud or local server) receives the features through an encrypted channel and performs semantic inference.

[0026] Furthermore, it further includes an image sensor module deployed at the edge device end to capture and preprocess the original RGB image data; A visual encoder module to construct an initial convolutional layer for the input preprocessed image data; A feature compression module deployed at the edge device end to serialize the features after the initial convolutional layer is stacked through a multi-level bottleneck module and global feature aggregation; An encoding encryption module deployed at the edge device end to encrypt the encoding output from the feature compression module.

[0027] Furthermore, a decoding and decryption module is deployed at the inference device end to complete feature deserialization and alignment for the feature data packet at the edge device end; A multimodal fusion module is deployed at the inference device end to input the visual features and text features into a multimodal fusion module based on Transformer; A language inference module is deployed at the inference device end to complete the inference and decision-making generation tasks; The application interface module is deployed on the inference device side to encapsulate the text results generated on the inference device side.

[0028] Among them, the modules in the edge device side and the modules in the inference device side transmit features through a dynamic security protocol to achieve calculation load separation and privacy protection. Its technical architecture is as Figure 2 shown.

[0029] Among them, the following key mechanisms are included in the transmission process: Secure channel establishment: In the initialization phase, an encrypted channel is established through the TLS 1.3 protocol. The edge device sends the device fingerprint containing the HMAC-SHA256 signature, and the cloud returns the dynamically generated session key K_session = ECDH-256(Priv_key, Pub_key). The key validity period is bound to the device working cycle; Data transmission specification: Each frame of feature stream is encapsulated into a triple structure of [ciphertext || MAC || timestamp]. The packet header includes: frame sequence number (32-bit counter to prevent replay attacks), feature dimension identifier (4-bit encoding to support dynamic dimension switching), encryption mode flag (2-bit indicating optional algorithms such as AES-GCM / TWINE), and the transport layer uses the UDP protocol to accelerate, and the transmission reliability under ≤ 3 retransmissions is achieved through forward error correction coding (FEC); Quality of service guarantee, offline resume transmission mechanism: The edge device locally caches the feature data of the last 30 seconds and supports resume from breakpoint.

[0030] As an implementation of the above Figure 1-2 system embodiment shown, the embodiment of the present application provides a method for edge deployment of a distributed vision-language model, including the following steps: S1: Image sensor initialization and data acquisition, the edge device side captures the original RGB image data and performs preprocessing; S2. Lightweight vision encoder forward inference, the preprocessed image is input into the lightweight vision encoder module deployed on the edge device side. This encoder is constructed based on depthwise separable convolution and includes an initial convolution layer, multiple inverted residual bottleneck modules (Bottleneck), global average pooling, and fully connected mapping layers; S3. Feature serialization is performed after stacking multiple bottleneck modules and global feature aggregation; S4. The edge device side and the inference device side establish a long connection through the MQTT protocol, and the feature data packet is transmitted through the TLS encrypted channel; S5. Feature deserialization and alignment are completed through the decoding and decryption module; S6. Multimodal joint encoding, input the visual feature and text feature into the multimodal fusion module based on Transformer; S7. Complete the inference and decision-making generation tasks through language reasoning; S8. The application interface encapsulates the text results generated by the inference device side and can perform downlink transmission and edge-side response.

[0031] Furthermore, an image sensor module (such as CMOS or CCD) is built into the edge device side, and the hardware initialization configuration is completed during the startup phase, including setting the resolution (typical value is 224×224 pixels), frame rate (adjustable range 1 - 30fps), exposure parameters, and white balance. The original RGB image data is captured in real time through an image sensor interface (such as MIPI-CSI or I2C) and temporarily stored in the device memory buffer. For dynamic scenes, a multi-frame buffer mechanism is supported to eliminate motion blur and ensure that the input image quality meets the requirements of subsequent processing; The preprocessing process of the picture is as follows: Color space conversion: Convert the original Bayer format data to the standard RGB color space, and apply gamma correction and demosaicing algorithms to eliminate sensor noise; Size normalization: Scale the image to a preset input size (such as 224×224) through the bilinear interpolation algorithm to adapt to the input requirements of the vision encoder; Pixel normalization: Perform mean subtraction and standard deviation normalization on each pixel channel, , where is the normalized output value, is the original input, is the channel mean of the training dataset, is the standard deviation, thereby eliminating the influence of illumination changes.

[0032] The processing flows in steps S2 and S3 are as follows: a. Initial convolution layer, using a suitable convolution kernel (such as 3×3) and stride (such as 2), expand the channels of the input image to obtain a channel feature map (expand the input image from 3 channels to 16 channels to obtain a channel feature map), and perform spatial downsampling (downsample to 112×112); b. Stacking of multi-level bottleneck modules: Sequentially perform multiple groups (such as 11 groups) of inverted residual bottleneck operations. Each bottleneck module includes an expansion convolution (1×1 convolution to increase the number of channels), a depthwise separable convolution (3×3 or 5×5 kernel to perform spatial feature extraction), a squeeze-and-excitation module (SE Block), and a projection layer (1×1 convolution to reduce the dimension). Among them, the SE module obtains channel attention weights through global average pooling and dynamically adjusts the importance of feature channels; c. Global feature aggregation: The final feature map is compressed in dimension (such as compressed to 1×1×576 dimension) through global average pooling (such as 7×7) of average pooling, and then mapped to a compact feature vector (such as a 1024-dimensional compact feature vector) through a fully connected layer; d. Feature serialization: The feature vector is divided into multiple sub-vectors according to a preset length (such as 256), and a visual token sequence is formed through dimension compression. , where L is the sequence length and d is the feature dimension.

[0033] To improve the computing efficiency of edge devices, the visual encoder also adopts the following optimization measures: Operator fusion: Convolution, batch normalization, and activation functions (such as Hard-Swish) are combined into a single computing unit to reduce memory access overhead.

[0034] Weight quantization: 8-bit fixed-point quantization technology is adopted to map the model weights and activation values to low-bit representations, reducing the computational complexity and memory occupancy.

[0035] Hardware instruction optimization: The convolution kernel implementation is customized for edge processors (such as the AI acceleration instruction set of ESP32-S3), and SIMD instructions are used to process multiple data channels in parallel.

[0036] After the feature sequence output by the visual encoder is serialized, it is encapsulated into data packets according to a preset protocol in step 307. The encapsulation format includes: Packet header identifier: A 4-byte magic number is used to identify the packet type; Feature dimension metadata: 2 bytes store the sequence length and feature dimension; Feature data body: L×d floating-point numbers are stored in row-major order, using the IEEE 754 single-precision format; Checksum: A 4-byte CRC32 checksum is used to ensure the integrity of data transmission.

[0037] See Figure 2 , and the specific implementation method in S4 is as follows: Two-way certificate authentication: The device side and the server side exchange X.509 certificates to verify the legitimacy of the identities of both communication parties; Session key negotiation: The Elliptic Curve Diffie-Hellman Ephemeral (ECDHE) key exchange algorithm is used to generate a temporary session key to ensure forward secrecy; Fragmented transmission of data packets: Large feature data packets are fragmented and transmitted according to the MTU size, and the integrity is verified after recombination at the receiving end.

[0038] In the encrypted feature stream, to adapt to network fluctuations, the system implements the following optimization strategies: Adaptive bitrate adjustment: Dynamically adjust the feature quantization accuracy according to the real-time network RTT and packet loss rate, and enable lossy compression (such as PCA dimensionality reduction) during network congestion; Resume mechanism for interrupted transmission: Each data packet is attached with a sequence number, and the receiving end triggers a retransmission request after detecting packet loss; Priority queue scheduling: The feature data of high-priority tasks (such as security alarms) is preferentially transmitted to ensure real-time performance.

[0039] Furthermore, an edge device provides a visual feature encoding method, and its technical architecture is as Figure 3 shown, including: Construct a lightweight convolutional neural network (CNN), which includes a bottleneck module with dynamic channel expansion, and the expansion coefficient ; Among them, is the input feature map, H is the height, W is the width, C is the number of channels, GAP(•) is the global average pooling operation, and X is compressed into a vector, is the Sigmoid activation function.

[0040] The visual encoding module adopts a hybrid attention mechanism, including channel attention and spatial attention. The fused channel attention and spatial attention is .

[0041] Among them, , are the parameters of the channel attention fully connected layer, r is the compression ratio, is the ReLU activation function,

[0042] are the parameters of the 3×3 depthwise separable convolution kernel, C is the number of channels, c is the channel index, k represents the offset of the convolution kernel in the height direction (row direction), and l represents the offset of the convolution kernel in the width direction (column direction); i, j are the spatial position indices, is the element-wise multiplication, and max(m) is the maximum value of the spatial attention map, which is used for normalization.

[0043] The generated compressed feature vector is , and the dimension .

[0044] Among them, L is the sequence length, d is the feature dimension, H, W are the resolutions (height × width) of the input image, such as 640×480, is the logarithm of the total number of pixels of the image, which represents the spatial complexity, α is the dynamic channel expansion coefficient, which is adaptively adjusted according to the input resolution, is the floor operation to ensure that the dimension is an integer.

[0045] After receiving the feature data packet, the encrypted feature vector backend service in S5 performs dynamic decryption and decoding for the following processing: Verification and Parsing: Verify the CRC checksum, parse the header metadata, and reconstruct the feature sequence Z; Modal Alignment: Project the visual token sequence Z and the text embedding vector E (generated by the text encoder) into the same semantic space, and adjust the feature scale and offset through linear transformation.

[0046] The multi-modal fusion process in S6 is as follows: Position Encoding Injection: Add learnable position encoding to visual and text tokens to capture sequence order information; Cross-modal Attention Calculation: Through the multi-head attention mechanism, text tokens are used as Queries to retrieve relevant visual features, and visual tokens are used as Keys and Values to participate in weight calculation to achieve information interaction between modalities; Hierarchical Feature Aggregation: Alternately stack self-attention layers and cross-attention layers to gradually fuse visual and text semantic information and generate a joint representation , where M is the length of the text sequence; In S7, large model inference is performed and the final result is output. The specific method is as follows: Autoregressive Decoding: Based on the greedy search or beam search algorithm, generate text responses word by word, and calculate the probability distribution of the next word at each time step , select the word token with the highest probability and append it to the output sequence; Logical Constraint Injection: Apply preset rule templates (such as medical term libraries, legal provisions) to post-process the generated results to ensure domain compliance; Confidence Calibration: Perform temperature scaling and Platt Scaling on the output probability distribution to improve the accuracy of confidence estimation and filter out low-confidence predictions. Among them, Platt Scaling is a general method for making the output probability values of binary classification models.

[0047] In the back-end inference module, the back-end server uses the following technologies to improve inference efficiency: Model Parallelization: Distribute the decoder layers to multiple GPU devices and accelerate the calculation through pipeline parallelism and tensor parallelism; Video Memory Optimization: Enable dynamic video memory allocation and activation value checkpoint technology to reduce the video memory consumption of a single card; Just-in-Time Compilation (JIT): Use TVM or TorchScript to compile the model into hardware-specific instructions to improve the execution efficiency of operators.

[0048] Result Feedback and System Closed-loop.

[0049] Result Serialization and Encapsulation.

[0050] . Further, the text results generated by the back end in S8 (such as classification labels, Q&A answers) are encapsulated in JSON format and contain the following fields: Result body: a text string encoded in UTF-8; Timestamp: a timestamp generated with millisecond precision; Confidence score: a floating-point numerical value representing the prediction confidence; Error code: a reserved field to identify processing exceptions.

[0051] The application interface part also includes: downlink transmission and edge-side response. The result data packet is returned to the edge device through an encrypted channel, triggering the following actions: Local cache update: caching the high-frequency query results in the device flash memory to reduce repeated calculations; Actuator control: driving the actuator (such as a relay, motor) to complete physical operations according to the result content; User interface feedback: visually presenting the results to the end user through an LCD screen or a voice module.

[0052] Through the application interface module, the following can also be achieved: system self-monitoring and maintenance. For example: Health status reporting: the edge device periodically sends heartbeat packets to report the computing load, memory usage, and sensor status; Model hot update: the back end pushes an incremental model update package, and the edge side loads the new model after verifying the signature through a secure boot mechanism; Exception recovery mechanism: when a hardware failure or communication interruption is detected, the edge side switches to the degraded mode and only performs local basic vision processing.

[0053] Further, the inference device side also provides an efficient decoding and inference method, and its technical architecture is as Figure 4 shown, including: Decoding and decryption module, receiving the encrypted feature stream and performing decryption and decoding; Feature alignment mechanism, defining the visual feature matrix and text embedding , and realizing modality fusion through dual-path cross-attention; Dynamic weight allocation, introducing a learnable modality gating coefficient: ; where, is the modality classifier, is the gating parameter; Using the window attention mechanism to divide the input sequence into dynamic windows for local calculation, and using the mixed-precision calculation strategy to reduce the deployment and operation costs.

[0054] In an alternative embodiment, the present application provides a storage medium, the storage medium including a stored program, wherein, when the program runs, it controls the device where the storage medium is located to execute Figure 1-4 the edge deployment method of the distributed vision language model in

[0055] In an alternative embodiment, the present application provides a processor, the processor being used to run a program, wherein, when the program runs, it executes Figure 1-4 the edge deployment method of the distributed vision language model in

[0056] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0057] It can be understood that the relevant features in the above methods and systems can be referred to each other. Additionally, the "first", "second", etc. in the above embodiments are used to distinguish the respective embodiments, and do not represent the superiority or inferiority of the respective embodiments.

[0058] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein again.

[0059] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. Based on the above description, the structure required to construct such systems is obvious. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions made above for specific languages are for disclosing the best implementation manner of the present application.

[0060] In addition, the memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip.

[0061] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0062] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0063] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0064] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0065] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0066] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0067] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0068] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0069] Those skilled in the art should understand that the computing model provided in the embodiments of the present application is not limited to MiniCPM, Whisper, and DeepSeek, and that the effects of the present invention may be achieved by replacing it with other computing models in the same field. The above three models are merely a preferred embodiment of the present invention.

[0070] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0071] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent substitution or improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An edge deployment system for a distributed visual language model, characterized in that: The invention comprises a distributed topology structure in which the front-end and the back-end work in collaboration; the front-end is an edge device end, and the back-end is an inference device end; the visual coding module is deployed on the edge device end, and the language inference module is deployed on the inference device end; the visual coding module and the language inference module transmit features through a dynamic security protocol.

2. The edge deployment system of a distributed visual language model as claimed in claim 1, characterized in that: Also includes The image sensor module is deployed on the edge device to capture raw RGB image data and perform preprocessing; The visual encoder module constructs the initial convolutional layer from the input preprocessed image data; The feature compression module is deployed on the edge device side, and the initial convolutional layer is stacked through multi-level bottleneck modules and global feature aggregation to perform feature serialization; The encoding encryption module is deployed on the edge device to encrypt the code coming out of the feature compression module.

3. The edge deployment system of a distributed visual language model as claimed in claim 1, characterized in that: The decoding and decryption module is deployed on the inference device to deserialize and align the feature data packets on the edge device. The multimodal fusion module is deployed on the inference device side, and the visual features and text features are input into the Transformer-based multimodal fusion module; The language reasoning module is deployed on the reasoning device to complete the reasoning and decision-making tasks; The application interface module is deployed on the inference device side and encapsulates the text results generated by the inference device side.

4. A distributed visual language model edge deployment method, applied to the distributed visual language model edge deployment system according to any one of claims 1 to 3, characterized in that: The following steps are involved: S1: Image sensor initialization and data acquisition, the edge device captures the original RGB image data and performs preprocessing; S2, lightweight visual encoder forward reasoning, the pre-processed image is input to the lightweight visual encoder module deployed on the edge device; S3, feature serialization after multi-level bottleneck module stacking and global feature aggregation; S4. The edge device and the inference device establish a long connection through the MQTT protocol, and the feature data packet is transmitted through the TLS encrypted channel; S5. Complete feature deserialization and alignment through the decoding and decryption module; S6, multimodal joint encoding, inputting visual features and text features into the Transformer-based multimodal fusion module; S7, complete reasoning and decision-making tasks through language reasoning; S8. The application interface encapsulates the text results generated by the inference device, which can be transmitted downlink and responded to by the edge.

5. The edge deployment method of a distributed visual language model as claimed in claim 4, characterized in that: The edge device has a built-in image sensor module, which completes hardware initialization configuration during the startup phase, including setting resolution, frame rate, exposure parameters, and white balance. It captures raw RGB image data in real time through the image sensor interface and temporarily stores it in the device memory buffer. For dynamic scenes, it supports a multi-frame buffer mechanism to eliminate motion blur and ensure that the input image quality meets the requirements of subsequent processing. The image preprocessing process is as follows: Color space conversion: Convert the original Bayer format data to the standard RGB color space, apply gamma correction and demosaicing algorithm to remove sensor noise; Size normalization: The image is scaled to the preset input size through a bilinear interpolation algorithm to adapt to the visual encoder input requirements; Pixel normalization: perform mean subtraction and standard deviation normalization on each pixel channel. ,in is the normalized output value, is the original input, is the channel mean of the training data set, is the standard deviation to eliminate the influence of illumination changes.

6. The edge deployment method of a distributed visual language model as claimed in claim 4, characterized in that: The processing flow in steps S2 and S3 is as follows: a. Initial convolution layer, using appropriate convolution kernel and step size, expands the channel of the input image to obtain the channel feature map and samples it at the spatial size; b. Multi-level bottleneck module stacking: Multiple sets of inverted residual bottleneck operations are performed in sequence. Each bottleneck module contains dilated convolution, depthwise separable convolution, compressed excitation module and projection layer. The compressed excitation module obtains channel attention weights through global average pooling and dynamically adjusts the importance of feature channels. c. Global feature aggregation: The final feature map is compressed by global average pooling through average pooling, and then mapped to a compact feature vector by the fully connected layer; d. Feature serialization: Divide the feature vector into multiple sub-vectors according to the preset length, and form a visual token sequence through dimension compression , where L is the sequence length and d is the feature dimension.

7. The edge deployment method of a distributed visual language model as claimed in claim 4, characterized in that: The specific implementation method in S4 is as follows: Two-way certificate authentication: The device and the server exchange X.509 certificates to verify the legitimacy of the identities of both parties in communication; Session key negotiation: The ECDHE key exchange algorithm is used to generate temporary session keys to ensure forward security; Data packet fragment transmission: large characteristic data packets are transmitted in fragments according to the MTU size, and the receiving end reassembles them to verify the integrity; The specific decryption and decoding processing method in S5 is as follows: Verification and parsing: verify the CRC checksum, parse the header metadata, and reconstruct the characteristic sequence Z; Modality alignment: Project the visual token sequence Z and the text embedding vector E into the same semantic space, and adjust the feature scale and offset through linear transformation.

8. The edge deployment method of a distributed visual language model as claimed in claim 4, characterized in that: The multimodal fusion process in S6 is as follows: Positional encoding injection: Add learnable positional encodings to visual tokens and textual tokens to capture sequence order information; Cross-modal attention calculation: Through the multi-head attention mechanism, text tokens are used as queries to retrieve relevant visual features, and visual tokens are used as keys and values ​​to participate in weight calculation to achieve information interaction between modalities; Hierarchical feature aggregation: Alternately stack self-attention layers and cross-attention layers to gradually fuse visual and textual semantic information to generate a joint representation , where M is the length of the text sequence; S7 performs large model reasoning and generates final result output. The specific method is as follows: Autoregressive decoding: Generate text responses word by word based on greedy search or beam search algorithms, and calculate the probability distribution of the next word at each time step , select the highest probability token and append it to the output sequence; Logical constraint injection: Apply preset rule templates to post-process the generated results to ensure domain compliance; Confidence calibration: Temperature scaling and Pratt scaling are performed on the output probability distribution to improve the accuracy of confidence estimates and filter out low-confidence predictions.

9. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the edge deployment method of the distributed visual language model as described in any one of claims 4 to 8.

10. A processor, characterized in that: The processor is used to run a program, wherein the program, when running, executes the edge deployment method of a distributed visual language model as described in any one of claims 4 to 8.

Citation Information

Patent Citations

  • Multi-modal model visual perception ability enhancement method and device, and medium

    CN119809925A

Cited By

  • Language reasoning server, method for language reasoning, system for visual language big model reasoning, medium and product

    CN120725153A

  • Intelligent system based on cloud edge collaborative reasoning

    CN121525884A

  • An intelligent system based on cloud-edge collaborative inference

    CN121525884B

  • PLC operation behavior prediction method and device based on manifold learning, equipment and medium

    CN121857516A