Multi-modal large model cloud edge collaborative reasoning method, device and equipment based on semantic coding and medium

By deploying a multimodal semantic coding model on the side devices, multimodal data is converted into semantic information, solving the problems of time-consuming and privacy protection of video transmission in whole-house intelligent systems, and achieving efficient and secure data transmission and processing.

CN120430409APending Publication Date: 2025-08-05EAST CHINA NORMAL UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510523984.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

When the existing whole-house intelligent system processes video data, the video transmission code rate is high, resulting in a long data transmission time, increasing computing overhead, and unable to effectively protect user privacy.

Method used

By deploying a pre-trained multimodal semantic coding model on the side devices, multimodal data is converted into multimodal semantic information and processed in the cloud to generate control instructions, reducing data transmission and computing overhead while protecting privacy.

Benefits of technology

It realizes efficient and secure data transmission, reduces transmission code rate, reduces transmission time, avoids user privacy leakage, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430409A_ABST
    Figure CN120430409A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal large model cloud edge collaborative reasoning method and device based on semantic coding, equipment and a medium, and relates to the field of artificial intelligence. Comprising the steps that a side device receives multi-modal data sent by a terminal device, semantic coding is conducted on the multi-modal data through a pre-trained multi-modal semantic coding model, and multi-modal semantic information is obtained and sent to a cloud server; the cloud server processes the multi-modal semantics based on a pre-trained multi-modal large model to obtain a control instruction and sends the control instruction to the side device; and the side device sends the control instruction to the terminal device, so that the terminal device executes a corresponding instruction operation based on the control instruction, thereby realizing efficient and secure transmission of data between the terminal device and the cloud large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a multimodal large-model cloud-edge collaborative reasoning method, device, equipment and medium based on semantic coding. Background Art

[0002] In recent years, whole-home intelligence has been a key goal of smart home development. This model, which previously relied primarily on single-product functionality and simple linkage to provide user services, now offers centralized control of smart devices in the home, systematically fulfilling user intent, and intelligently enabling inter-device linkage. Multimodal large models, with their powerful capabilities for complex multimodal information processing and decision-making, are typically deployed on cloud servers to gain high computing power. They work with smart home devices to perceive the home environment and adaptively generate device control commands. For example, Amazon Alexa's LMM technology controls home devices through voice interaction. However, it lacks access to the most important surveillance video data in the home, preventing true whole-home intelligence.

[0003] Currently, existing whole-home smart big models primarily process voice and text commands, not video data. This is due to the following difficulties faced by multimodal big models in practical applications: the transmission bit rate of video is much higher than that of voice and text data. Transmitting video data collected by home terminal devices to the cloud-based multimodal big model for processing not only prolongs transmission time, affecting the user's real-time experience, but also significantly increases data transmission and big model computational overhead, making it unacceptable to users. Therefore, how to achieve efficient and secure data transmission between terminal devices and the cloud-based big model is a technical problem that needs to be solved in this paper. Summary of the Invention

[0004] Based on the above technical problems, the present invention provides a multimodal large model cloud-edge collaborative reasoning method, device, equipment and medium based on semantic coding, aiming to overcome the above problems or at least partially solve the above problems.

[0005] A first aspect of the present invention provides a multimodal large model cloud-edge collaborative reasoning method based on semantic coding, the method comprising: The edge device receives the multimodal data sent by the terminal device, semantically encodes the multimodal data through a pre-trained multimodal semantic encoding model, obtains multimodal semantic information and sends it to the cloud server; The cloud server processes the multimodal semantics based on a pre-trained multimodal large model, obtains a control instruction, and sends it to the edge device; The edge device sends the control instruction to the terminal device, so that the terminal device performs a corresponding instruction operation based on the control instruction.

[0006] A second aspect of the present invention provides a multimodal large model cloud-edge collaborative reasoning device based on semantic coding, the device comprising: The semantic encoding module is deployed on the edge device and is used to receive multimodal data sent by the terminal device, semantically encode the multimodal data through the pre-trained multimodal semantic encoding model, obtain multimodal semantic information, and send it to the cloud server; An inference module, deployed on the cloud server, is used to process the multimodal semantics based on a pre-trained multimodal large model, obtain control instructions, and send them to the edge device; The instruction control module is deployed on the edge device and is used to send the control instruction to the terminal device so that the terminal device performs the corresponding instruction operation based on the control instruction.

[0007] The third aspect of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the computer program is executed by the processor, the multimodal large model cloud-edge collaborative reasoning method based on semantic coding as described in the first aspect of the present invention is implemented.

[0008] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multimodal large model cloud-edge collaborative reasoning method based on semantic coding according to the first aspect of the present invention.

[0009] In the multimodal large model cloud-edge collaborative reasoning method based on semantic coding provided by the present invention, a new mode of encoding multimodal semantic information with an edge device and communicating with a multimodal large model on the cloud to generate control instructions is adopted. A pre-trained multimodal semantic coding model is deployed in the edge device, and no additional computing power is required for the terminal device. At the same time, semantic communication is utilized to convert multimodal data into multimodal semantic information through a pre-trained multimodal semantic coding model and transmit it to the cloud server to generate control instructions through the pre-trained multimodal large model. Compared with the classical communication processing bit transmission error, the semantic communication of the present invention pays more attention to the transmission error of semantic information. By converting the transmitted data symbols into a semantic space, the data transmission bit rate can be greatly reduced and the original information can be protected, which greatly reduces the overhead of data transmission and large model calculation. It solves the problem that video transmission in traditional methods requires a large amount of resources, reduces data transmission time, brings real-time experience to users, and avoids user privacy leakage, solving the problem of efficient and secure data transmission between cloud large models and edge devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0011] Figure 1 This is a flowchart of a multimodal large model cloud-edge collaborative reasoning method based on semantic encoding according to an embodiment of the present invention; Figure 2 is a schematic diagram of a video semantic coding model pre-training method shown in one embodiment of the present invention; Figure 3 1 is a schematic structural diagram of a multimodal semantic coding model according to an embodiment of the present invention; Figure 4 is a schematic diagram of a multimodal semantic encoding model pre-training method according to an embodiment of the present invention; Figure 5 is a schematic diagram of an audio semantic encoder pre-training method according to an embodiment of the present invention; Figure 6 1 is a schematic diagram of a multimodal large model pre-training method according to an embodiment of the present invention; Figure 7 is a schematic diagram of a lightweight fine-tuning method for a multimodal semantic coding model according to an embodiment of the present invention; Figure 8 1 is a schematic diagram illustrating a method for implementing semantic communication of a whole-house intelligent central control robot according to an embodiment of the present invention; Figure 9 This is a structural block diagram of a multimodal large model cloud-edge collaborative reasoning device based on semantic coding provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0013] Please refer to Figure 1 , Figure 1 This is a flowchart of a multimodal large model cloud-edge collaborative reasoning method based on semantic coding according to an embodiment of the present invention. Figure 1 As shown, the multimodal large model cloud-edge collaborative reasoning method based on semantic encoding provided by this embodiment includes at least the following steps: Step S11: The edge device receives the multimodal data sent by the terminal device, performs semantic encoding on the multimodal data through a pre-trained multimodal semantic encoding model, obtains multimodal semantic information and sends it to the cloud server.

[0014] In this embodiment, the environmental state of a real-world home scene is often composed of multimodal data (e.g., images, videos, audio, text, and other modalities) captured by different terminal devices. Whole-home intelligence aims to build a human-centric, dynamic device perception and robust environmental awareness system. The terminal devices in this embodiment can be home devices, including but not limited to speakers, cameras, robots, air conditioners, lights, electric curtains, washing machines, door locks, and so on. Multiple terminal devices can be interconnected via a local area network (LAN), such as Wi-Fi, to provide more comfortable and intelligent services.

[0015] In this embodiment, one or more terminal devices can collect multimodal data and send the collected multimodal data to the edge device, such as sending the multimodal data to the edge device through a local area network. Edge devices are usually located at the edge of the network, that is, close to the user and the data source, corresponding to the "cloud", and the edge devices are in the middle layer between the cloud server and the terminal device. The edge devices of this embodiment may include edge boxes, central control robots and other devices with higher computing power. The edge device is deployed with a pre-trained multimodal semantic coding model for semantically encoding multimodal data to obtain multimodal semantic information.

[0016] The edge device can receive multimodal data sent by the terminal device and semantically encode the received multimodal data (such as video data, image data, audio data, and text data) using a pre-trained multimodal semantic encoding model to obtain multimodal semantic information. The multimodal semantic information is then sent to the cloud server. Furthermore, in one embodiment, the edge device can also directly collect multimodal data, such as voice commands issued by the user.

[0017] In an optional embodiment, for a terminal device with AI processing capabilities, a semantic extraction model of the corresponding modality can be deployed in the terminal device with AI processing capabilities to perform semantic extraction on the original data of the corresponding modality collected by itself, obtain semantic information after dimensionality reduction processing of the original data, and transmit it to the side device.

[0018] Step S12: The cloud server processes the multimodal semantics based on the pre-trained multimodal large model, obtains a control instruction and sends it to the edge device.

[0019] In this embodiment, a pre-trained multimodal large model is deployed in the cloud server to understand the environmental and user states and make command decisions. After receiving multimodal semantic information from edge devices, the cloud server processes the multimodal semantics using the pre-trained multimodal large model, obtains control instructions (e.g., corresponding air conditioning temperature adjustment instructions, audio volume adjustment instructions, etc.), and sends the control instructions to the edge devices.

[0020] Step S13: The edge device sends the control instruction to the terminal device, so that the terminal device performs a corresponding instruction operation based on the control instruction.

[0021] In this embodiment, after receiving a control instruction from a cloud server, the edge device can manage the corresponding terminal device to execute the corresponding instruction operation based on the control instruction. Specifically, the edge device can send the control instruction to the corresponding terminal device, so that the terminal device executes the corresponding instruction operation based on the received control instruction to implement the corresponding function, thereby realizing a "user-free" intelligent home. For example, an air conditioner adjusts the air conditioner temperature based on the received air conditioner temperature adjustment instruction, such as raising the air conditioner temperature to 30°C.

[0022] In this embodiment, a pre-trained multimodal semantic coding model deployed on the edge device is used to compress the multimodal data collected by the terminal device during the whole-house intelligent implementation process, and the multimodal data is converted into multimodal semantic information and transmitted to the cloud server to generate control instructions through the pre-trained multimodal large model. This embodiment pays more attention to the transmission error of semantic information through semantic communication. By converting the sent symbols into the semantic space, the data transmission bit rate with the cloud server is reduced and the original information is protected, the overhead of data transmission and large model calculation is reduced, and the problem of video transmission requiring a large amount of resources in traditional methods is solved. The data transmission time is reduced, a real-time experience is brought to users, and user privacy leakage is avoided, solving the problem of efficient and secure data transmission between the cloud large model and the edge device.

[0023] In combination with the above embodiments, in one embodiment, the present invention further provides a multimodal large model cloud-edge collaborative reasoning method based on semantic coding. In this method, the pre-trained multimodal semantic coding model includes at least: a pre-trained video semantic coding model, and the pre-trained video semantic coding model includes at least: a pre-trained video semantic coding module and a pre-trained mask adaptation module; the above step S11 of "semantically encoding the multimodal data through the pre-trained multimodal semantic coding model to obtain multimodal semantic information and sending it to the cloud server" can specifically include steps S21 to S25: Step S21: When the multimodal data includes video data, mask processing is performed on a current video frame in the video data to obtain a current image mask feature.

[0024] In this embodiment, when the multimodal data includes video data, the video data can be considered to be generated by the evolution of multiple video frames (static images) over time. However, there is a large amount of repeated redundant information in the preceding and following video frames. Therefore, to save data and computing resources and eliminate temporal and spatial redundant information in the video data, this embodiment can perform masking on the current video frame in the video data. This masking removes redundant information in consecutive video frames to obtain the current image mask feature. The current video frame in this embodiment can be every video frame in the video data.

[0025] Step S22: semantically encode the current image mask features through the video semantic encoding module to obtain current video semantic information.

[0026] In this embodiment, after obtaining the current image mask feature, the current image mask feature can be input into a pre-trained video semantic coding module, and the video semantic coding module performs semantic coding on the current image mask feature to obtain current video semantic information.

[0027] Step S23: sorting the importance of the current video semantic information by the mask adaptive module, and dividing the current video frame into a first semantic layer, a second semantic layer and a third semantic layer according to the importance of the current video semantic information carried.

[0028] In order to further reduce the transmission bit rate of video semantic information and protect privacy data, this embodiment designs a mask adaptive module. After the video semantic encoding module outputs the current video semantic information, the current video semantic information can be input into the pre-trained mask adaptive module. The mask adaptive module is used to sort the current video semantic information by importance, and the current video frame is divided into the first semantic layer, the second semantic layer and the third semantic layer according to the importance of the current video semantic information it carries.

[0029] Among them, the first semantic layer is the image features corresponding to the important semantic information in the current video frame, which is mainly composed of foreground objects and background information with rich features such as some people, animals, and cars; the second semantic layer is the image features corresponding to the detailed semantic information in the current video frame, which is mainly composed of foreground objects and background information that can supplement the first semantic layer; the third semantic layer is the image features corresponding to the privacy information in the current video frame, among which privacy information is sensitive information that users are unwilling to share, such as personal privacy information such as faces and license plates.

[0030] For example, the mask adaptation module can use the transformer attention mechanism to sort the importance of the current video semantic information tokens in the current video frame, and divide the current video frame into the first semantic layer, the second semantic layer, and the third semantic layer according to the semantic information carried.

[0031] Step S24: semantically encode the first semantic layer, the second semantic layer, and the third semantic layer by the video semantic encoding module to obtain first video semantic information, second video semantic information, and third video semantic information, respectively.

[0032] In this embodiment, after obtaining the first semantic layer, the second semantic layer, and the third semantic layer, the video semantic encoding module can perform semantic encoding on the first semantic layer, the second semantic layer, and the third semantic layer to obtain first video semantic information corresponding to the first semantic layer, second video semantic information corresponding to the second semantic layer, and third video semantic information corresponding to the third semantic layer, respectively. The first video semantic information is more important than the second video semantic information, and the third video semantic information represents the private information in the current video frame.

[0033] Step S25: Through the mask adaptive module, when the current network status is not higher than the first status threshold, the first video semantic information, the second video semantic information and the third video semantic information are sent to the cloud server; when the current network status is higher than the first status threshold, the first video semantic information is sent to the cloud server.

[0034] In this embodiment, after obtaining the first video semantic information, the second video semantic information, and the third video semantic information, the mask adaptive module can determine the current network state and compare whether the current network state is above a preset first state threshold. In one example, the current network state can be represented by packet loss rate, throughput, bandwidth utilization, etc., and accordingly, the first state threshold can be represented by a packet loss rate threshold, a throughput threshold, a bandwidth utilization threshold, etc., which is not limited in this embodiment.

[0035] If the mask adaptive module determines that the current network state is not higher than the first state threshold, the current network state is determined to be good. In this case, the mask adaptive module controls the video semantic coding module to send the first video semantic information, the second video semantic information, and the third video semantic information to the cloud server, thereby supplementing the detailed semantic information of the current video frame with the second video semantic information. If the mask adaptive module determines that the current network state is higher than the first state threshold, the current network state is determined to be poor. In this case, the mask adaptive module controls the video semantic coding module to discard the second video semantic information and the third video semantic information to save bit rate, and only send the first video semantic information to the cloud server.

[0036] In an optional embodiment, when the video semantic encoding module semantically encodes the third semantic layer, it also encrypts it to obtain the third video semantic information. The third video semantic information is the encrypted semantic information. When the third video semantic information is transmitted to the cloud server, the corresponding secret key is required to restore the third video semantic information.

[0037] In this embodiment, by designing a mask adaptive module, the current video semantic information is sorted by importance, and the current video frame is divided into the first semantic layer, the second semantic layer and the third semantic layer. When the network signal jitters, the first semantic layer is retained first, and the second semantic layer is retained according to the attention importance. The semantic information of the third semantic layer can only be restored when the cloud server has the key, thereby achieving the purpose of adaptively adjusting the transmission bit rate and protecting privacy data.

[0038] In combination with the above embodiments, in one implementation, the present invention further provides a multimodal large model cloud-edge collaborative reasoning method based on semantic coding. In this method, the "masking the current video frame in the video data to obtain the current image mask feature" in the above step S21 can specifically include steps S31 to S33: Step S31: The mask adaptation module adjusts the initial mask area according to the current network state to obtain a first mask area.

[0039] In this embodiment, the mask adaptation module can adjust the initial mask area according to the current network status to obtain a first mask area. In this embodiment, the initial mask area is a pre-set mask area, which is the area covered by the masking operation. In this embodiment, the initial mask area can be adjusted according to the network status to obtain the first mask area.

[0040] Specifically, the mask adaptation module can determine the current network state and compare whether the current network state is above a preset second state threshold. In one example, the current network state can be represented by packet loss rate, throughput, bandwidth utilization, etc., and accordingly, the second state threshold can be represented by a packet loss rate threshold, a throughput threshold, a bandwidth utilization threshold, etc., which is not limited in this embodiment. The second state threshold can be the same as or different from the first state threshold, which is not limited.

[0041] If the mask adaptive module determines that the current network state is higher than the second state threshold, indicating that the current network condition is poor, the mask adaptive module may increase the initial mask area to obtain the first mask area. If the mask adaptive module determines that the current network state is not higher than the second state threshold, indicating that the current network condition is good, the mask adaptive module may reduce the initial mask area to obtain the first mask area.

[0042] Step S32: performing mask processing on the redundant pixel blocks in the first mask area of the current video frame to obtain a first image feature.

[0043] In this embodiment, mask processing is performed on the redundant pixel blocks in the first mask area of the current video frame, and the redundant pixel blocks in the first mask area of the current video frame are covered by masking to obtain the first image feature.

[0044] Step S33: splicing the first image feature with the preset pixel block feature to obtain the current image mask feature.

[0045] In this embodiment, the first image feature is concatenated with the preset pixel block feature to obtain a current image mask feature having the same size as the original input image (current video frame). The preset pixel block may be a preset learnable pixel block.

[0046] In combination with the above embodiments, in one embodiment, the present invention further provides a multimodal large model cloud-edge collaborative reasoning method based on semantic coding. In this embodiment, the pre-trained video semantic coding model is trained based on the initial video semantic coding model, and the initial video semantic coding model at least includes: an initial video semantic coding module and an initial mask adaptation module; the training step of the initial video semantic coding model includes the following steps: Step S41: adjusting the sample initial mask region according to the sample network state through the initial mask adaptive module to obtain the sample first mask region.

[0047] This embodiment takes into account that the initial video semantic coding model is difficult to converge when trained using a common image semantic coding model pre-training method. Based on this, this embodiment adopts a self-supervised paradigm and a masking-and-reconstruction-based encoding and decoding method to pre-train the video semantic extraction model.

[0048] First, the initial mask adaptation module can be used to adjust the sample initial mask region based on the sample network state to obtain the sample first mask region. The sample network state is the current network state during model training, and the sample initial mask region is a mask region pre-set during model training. This mask region is the region masked by the masking operation. This embodiment can adjust the sample initial mask region based on the sample network state to obtain the sample first mask region.

[0049] For example, if the initial mask adaptive module determines that the sample network state is higher than the sample state threshold (the second state threshold during model training), indicating that the current network condition is poor, the initial mask adaptive module can increase the sample initial mask area to obtain the sample first mask area. If the initial mask adaptive module determines that the sample network state is not higher than the sample state threshold, indicating that the current network condition is good, the initial mask adaptive module can reduce the sample initial mask area to obtain the sample first mask area.

[0050] Step S42: masking the redundant pixel blocks in the sample first mask area of the sample current video frame in the sample video to obtain the sample first image feature and splicing it with the sample preset pixel block feature to obtain the sample current image mask feature.

[0051] In this embodiment, a masking process can be performed on redundant pixel blocks in a sample first mask region of a sample current video frame in a sample video. By masking out the redundant pixel blocks in the sample first mask region of the sample current video frame, a sample first image feature can be obtained. The sample first image feature is then concatenated with the sample preset pixel block feature to obtain the sample current image mask feature. The sample video is the video data used for model training, the sample current video frame is each video frame in the sample video, and the sample preset pixel blocks are pixel blocks that are preset and learnable during the model training process.

[0052] Step S43: inputting the sample current image mask feature into the initial video semantic encoding module to obtain sample video semantic information.

[0053] In this embodiment, the obtained sample current image mask features may be input into an initial video semantic coding module, and the initial video semantic coding module performs semantic coding on the sample current image mask features to obtain sample video semantic information.

[0054] Step S44: sorting the sample video semantic information by importance through the initial mask adaptive module, and dividing the sample current video frame into a first sample semantic layer, a second sample semantic layer and a third sample semantic layer according to the importance of the sample video semantic information carried.

[0055] In this embodiment, in order to further reduce the transmission bit rate of video semantic information and protect privacy data, after the initial video semantic encoding module outputs the sample video semantic information, this embodiment can input the sample video semantic information into the initial mask adaptation module, and use the initial mask adaptation module to sort the sample video semantic information by importance, and divide the current video frame of the sample into the first sample semantic layer, the second sample semantic layer and the third sample semantic layer according to the importance of the sample video semantic information carried.

[0056] The first, second, and third sample semantic layers are the first, second, and third semantic layers, respectively, in the model training process. Specifically, the first sample semantic layer represents the image features corresponding to the important semantic information in the current video frame of the sample, primarily consisting of foreground objects and background information with rich features, such as some people, animals, and cars. The second sample semantic layer represents the image features corresponding to the detailed semantic information in the current video frame of the sample, primarily consisting of foreground objects and background information that can supplement the first semantic layer. The third sample semantic layer represents the image features corresponding to the private information in the current video frame of the sample, where private information is sensitive information that the user is unwilling to share, such as personal privacy information such as faces and license plates.

[0057] Step S45: semantically encode the first sample semantic layer, the second sample semantic layer, and the third sample semantic layer through the initial video semantic encoding module to obtain first sample video semantic information, second sample video semantic information, and third sample video semantic information, respectively.

[0058] In this embodiment, the initial video semantic encoding module may perform semantic encoding on the first sample semantic layer, the second sample semantic layer, and the third sample semantic layer to obtain first sample video semantic information, second sample video semantic information, and third sample video semantic information corresponding to the first sample semantic layer, the second sample semantic layer, and the third sample semantic layer, respectively. The first sample video semantic information is more important than the second sample video semantic information, and the third sample video semantic information represents the private information in the current video frame of the sample.

[0059] Step S46: inputting the first sample video semantic information, the second sample video semantic information and the third sample video semantic information into a decoder to obtain a current reconstructed image.

[0060] In this embodiment, during the pre-training process of the video semantic coding model, a decoder is designed to help the initial video semantic coding model better learn the deep semantic information in the sample video. Thus, the semantic information of the first sample video, the second sample video, and the third sample video can be input into the decoder to obtain the current reconstructed image. The decoder also possesses the key for decrypting and restoring the semantic information of the third sample video. In an optional example, the decoder can be a frame-reconstruction-based decoder, which can restore the entire frame based on partial semantic information. For information that is missing in previous and subsequent frames, the restored image will be blurry, and thus can serve as a basis for encrypting private data.

[0061] Step S47: Based on the current reconstructed image and the sample current video frame, the initial video semantic coding module and the initial mask adaptation module are trained to obtain the video semantic coding module and the mask adaptation module.

[0062] In this embodiment, loss calculation can be performed based on the current reconstructed image and the current video frame of the sample, and the network parameters of the initial video semantic coding module and the initial mask adaptation module are updated according to the calculated loss value until the loss converges, thereby obtaining a pre-trained video semantic coding module and a pre-trained mask adaptation module.

[0063] In one embodiment, if Figure 2 As shown, Figure 2 FIG. 1 is a schematic diagram of a video semantic coding model pre-training method according to an embodiment of the present invention. Figure 2 In the method, the mask adaptation module (i.e., the initial mask adaptation module) can adjust the sample initial mask area based on the network state to obtain the sample first mask area, and perform an image mask operation on the sample first mask area in each sample video frame in the video clip to obtain an image mask feature; the image mask feature is then input into the encoder (i.e., the initial video semantic encoding module) to obtain the sample video semantic information, and the initial mask adaptation module uses the transformer attention mechanism to sort the importance of the sample video semantic information tokens in the sample video screen, and divide the sample video frame into a semantic base layer (first sample semantic layer), a semantic enhancement layer (second sample semantic layer), and a semantic encryption layer (third sample semantic layer) according to the semantic information carried; the encoder then performs semantic encoding on the semantic base layer, the semantic enhancement layer, and the semantic encryption layer to obtain the first sample video semantic information, the second sample video semantic information, and the third sample video semantic information, which are input into the decoder for video reconstruction to obtain the current reconstructed image; finally, the encoder and the mask adaptation module are trained based on the current reconstructed image and the sample video frame to obtain the trained video semantic encoding module and the mask adaptation module, thereby realizing the pre-training of the video semantic coding model.

[0064] In combination with the above embodiments, in one embodiment, the present invention also provides a multimodal large model cloud-edge collaborative reasoning method based on semantic coding. In this method, the pre-trained multimodal semantic coding model includes at least: a text semantic encoder, an image semantic encoder and an audio semantic encoder, which are respectively used to semantically encode text data, image data and audio data in multimodal data; the pre-trained multimodal large model is obtained by training based on the initial multimodal large model, and the text semantic encoder, the image semantic encoder and the audio semantic encoder are respectively trained based on the initial text semantic encoder, the initial image semantic encoder and the initial audio semantic encoder, and the training steps of the initial text semantic encoder, the initial image semantic encoder and the initial audio semantic encoder include the following steps: Step S51: inputting the first sample text data, the first sample image data and the first sample audio data into the initial text semantic encoder, the initial image semantic encoder and the initial audio semantic encoder respectively to obtain first sample text semantic information, first sample image semantic information and first sample audio semantic information.

[0065] In this embodiment, the initial text semantic encoder, the initial image semantic encoder, and the initial audio semantic encoder respectively include: respective corresponding embedding layers and multiple transformer coding layers, and the original data can be compressed by using multiple transformer coding layers. In one embodiment, Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a multimodal semantic coding model shown in one embodiment of the present invention. Figure 3 In the figure, the left side shows the structure of the initial audio semantic encoder, the middle side shows the structure of the initial audio semantic encoder, and the right side shows the structure of the initial image semantic encoder.

[0066] In an optional example, the structure of the initial video semantic encoding module is the same as that of the initial image semantic encoder. Figure 3 The structure of the initial image semantic encoder (Image) on the right is the same.

[0067] In this embodiment, the first sample text data may be input into the initial text semantic encoder to obtain first sample text semantic information; the first sample image data may be input into the initial image semantic encoder to obtain first sample image semantic information; and the first sample audio data may be input into the initial audio semantic encoder to obtain first sample audio semantic information. The first sample text data, the first sample image data, and the first sample audio data are the text data, image data, and audio data used in the training process of the initial text semantic encoder, the initial image semantic encoder, and the initial audio semantic encoder, respectively.

[0068] Step S52: Input the sample device description coding information corresponding to the sample terminal device, the first sample text semantic information, the first sample image semantic information and the first sample audio semantic information into the initial multimodal large model, align the sample device description coding information, the first sample image semantic information and the first sample audio semantic information with the first sample text semantic information, and obtain the first sample instruction, the second sample instruction and the third sample instruction respectively.

[0069] In this embodiment, the text semantic encoder, image semantic encoder, and audio semantic encoder are obtained by jointly training the initial text semantic encoder, initial image semantic encoder, initial audio semantic encoder, and the initial multimodal large model, with the accuracy of generated instructions as the training goal. Therefore, in order for the initial multimodal large model to generate accurate instructions, in addition to inputting the first sample text semantic information, the first sample image semantic information, and the first sample audio semantic information into the initial multimodal large model, it is also necessary to input the sample device description encoding information corresponding to the sample terminal device into the initial multimodal large model.

[0070] Among them, the sample terminal device is the terminal device in the model training process, and the sample device description encoding information corresponding to the sample terminal device is the encoding information obtained by inputting the device description corresponding to the sample terminal device into the pre-trained device encoder, wherein the device description includes the basic information of the terminal device, usage guide information, etc.

[0071] This embodiment inputs the sample device description coding information, first sample text semantic information, first sample image semantic information and first sample audio semantic information corresponding to the sample terminal device into the initial multimodal large model, and aligns the sample device description coding information, first sample image semantic information and first sample audio semantic information with the first sample text semantic information, such as translating the audio into text, aligning the image to the text through CLIP, etc., so that in the semantic space of the model, multiple modalities can be semantically represented through an intermediate unified modality, and finally obtain the first sample instruction corresponding to the first sample text data, the second sample instruction corresponding to the first sample image data and the third sample instruction corresponding to the first sample audio data output by the initial multimodal large model.

[0072] In one optional embodiment, after training the multimodal semantic encoding model and the multimodal large model and applying the model, in addition to inputting the multimodal semantic information into the pre-trained multimodal large model, the device description encoding information corresponding to the terminal device needs to be input into the pre-trained multimodal large model so that the pre-trained multimodal large model outputs the corresponding control instructions. The device description encoding information corresponding to the terminal device can be obtained by inputting the device description corresponding to the terminal device into a pre-trained device encoder, where the device description includes basic information about the terminal device, user guide information, etc.

[0073] Step S53: Based on the first sample instruction and the label instruction corresponding to the first sample text data, the initial text semantic encoder is trained to obtain the text semantic encoder.

[0074] In this embodiment, the label instruction corresponding to the first sample text data is the real instruction corresponding to the first sample text data. Based on the first sample instruction and the label instruction corresponding to the first sample text data, the initial text semantic encoder is trained with the accuracy of the generated instruction as the training goal to obtain a trained text semantic encoder.

[0075] Step S54: Based on the second sample instruction and the label instruction corresponding to the first sample image data, the initial image semantic encoder is trained to obtain the image semantic encoder.

[0076] In this embodiment, the label instruction corresponding to the first sample image data is the real instruction corresponding to the first sample image data. Based on the second sample instruction and the label instruction corresponding to the first sample image data, the initial image semantic encoder is trained with the accuracy of the generated instruction as the training goal to obtain a trained image semantic encoder.

[0077] Step S55: Based on the third sample instruction and the label instruction corresponding to the first sample audio data, the initial audio semantic encoder is trained to obtain the audio semantic encoder.

[0078] In this embodiment, the label instruction corresponding to the first sample audio data is the real instruction corresponding to the first sample audio data. Based on the third sample instruction and the label instruction corresponding to the first sample audio data, the initial audio semantic encoder is trained with the accuracy of the generated instruction as the training goal to obtain a trained audio semantic encoder.

[0079] In this embodiment, the initial text semantic encoder, the initial image semantic encoder, the initial audio semantic encoder and the initial multimodal large model are jointly trained with the accuracy of generated instructions as the training goal. The large language model's powerful understanding and reasoning capabilities of text data are utilized to align the device group representation encoding information (i.e., device description encoding information), audio and image semantic information with the text representation, so that the large language model can simultaneously receive input forms of multiple different modal combinations and output control instructions that can be executed by the terminal device.

[0080] In one embodiment, if Figure 4 As shown, Figure 4 Schematic diagram of a multimodal semantic coding model pre-training method according to an embodiment of the present invention. Figure 4 In the example, the initial multimodal semantic encoding model includes at least: an initial text semantic encoder (text encoder), an initial image semantic encoder (image encoder) and an initial audio semantic encoder (audio encoder), and the initial multimodal large model includes: a pre-trained large language model and multiple linear layers to be trained (such as Figure 4 First, multiple device descriptions (such as device descriptions 1, 2, 3, etc.) can be input into a pre-trained smart home device encoder to obtain corresponding device codes (such as device codes 1, 2, 3, etc.). Furthermore, sample audio data, sample image data, and sample text data are respectively input into an audio encoder, an image encoder, and a text encoder to obtain audio codes (audio semantic information), image codes (image semantic information), and image codes (text semantic information), respectively. The device codes, audio codes, image codes, and text codes are then respectively input into corresponding linear layers to be trained and then into a pre-trained large language model to generate complex instructions (such as instructions 0, 1, 2, 3, etc.). Finally, based on the complex instructions and the label instructions corresponding to the sample audio data, sample image data, and sample text data, the network parameters of the audio encoder, image encoder, text encoder, and multiple linear layers to be trained are updated until the trained audio encoder, image encoder, and text encoder are obtained.

[0081] In addition, in conjunction with the above embodiments, in one embodiment, the present invention also provides a semantic coding-based multimodal large-scale cloud-edge collaborative inference method. In this method, the audio semantic encoder is no longer jointly trained with the text semantic encoder and image semantic encoder in conjunction with the initial multimodal large-scale model. Instead, it is trained separately, similar to the pre-training process of the video semantic encoding model.

[0082] The pre-training method of audio semantic encoder is similar to that of video semantic encoding model, such as Figure 5 As shown, Figure 5FIG. 1 is a schematic diagram of an audio semantic encoder pre-training method according to an embodiment of the present invention. Figure 5 In the audio data, you can first convert the audio clips into a spectrogram (also known as a spectrogram) for encoding and decoding. A spectrogram is a visual representation of the spectrum of a signal over time. It can simultaneously display the time, frequency, and amplitude of a signal, providing more information than other signal representations. In most sound-related research, the spectrogram can be considered one of the most important digital representations of sound.

[0083] In this embodiment, a mask adaptation module (the mask adaptation module corresponding to the audio semantic encoder) first masks the spectrogram based on the network status to obtain a spectrogram mask. The spectrogram mask is then input into the initial encoder (i.e., the initial audio semantic encoder) to obtain audio semantic information. The mask adaptation module then divides the spectrogram into three layers based on the importance of the audio semantic information: a semantic base layer, a semantic enhancement layer, and a semantic encryption layer. The initial encoder then semantically encodes the semantic base layer, semantic enhancement layer, and semantic encryption layer, and then inputs them into the decoder, producing a reconstructed spectrogram output by the decoder to restore the audio. The initial encoder and mask adaptation module are then trained based on the audio clip and the restored audio, resulting in a trained audio semantic encoder and mask adaptation module.

[0084] Among them, the semantic base layer is the minimum semantic information required for the model to reconstruct the spectrogram, which is mainly composed of useful sounds such as human speech and alarm sounds; the semantic enhancement layer is the useful sounds and background noise that can supplement the semantic base layer. When the network condition is poor, it can be discarded to save bit rate. When the network condition is good, it can be used to assist in reconstructing the detailed information of the image; the semantic encryption layer is personal conversations that users do not want to share. This information can be omitted, encrypted during encoding, and restored with a secret key during decoding.

[0085] In one embodiment, the audio encoding and decoding system in this embodiment consists of four parts: an audio sampling module, a masking module, an encoding module and a decoding module. The steps are as follows: 1) the audio acquisition module converts the sound signal of a period of time into a spectrogram; 2) the masking module selects an appropriate masking strategy based on the network conditions, referring to the video processing method; 3) the encoding module uses the vision Transformer model to extract the features (Tokens) of the spectrogram after masking and sends them to the decoding end; 4) the decoding module uses semantic information to reconstruct the spectrogram image, or perform downstream tasks (such as retrieval, crying monitoring, etc.). The spectrogram can be converted into a sound signal for playback.

[0086] In conjunction with the above embodiments, in one embodiment, the present invention further provides a semantically encoded multimodal large model cloud-edge collaborative reasoning method. In this method, the pre-trained multimodal large model is obtained based on the training of an initial multimodal large model, and the initial multimodal large model at least includes: a pre-trained large language model, a device description linear layer to be trained, and linear layers to be trained corresponding to each modality; the training steps of the initial multimodal large model at least include the following steps: Step S61: inputting the first sample multimodal data into the pre-trained multimodal semantic encoding model to obtain first sample multimodal semantic information.

[0087] In this embodiment, after obtaining a pre-trained multimodal semantic encoding model, the first sample multimodal data can be input into the pre-trained multimodal semantic encoding model to obtain the first sample multimodal semantic information output by the pre-trained multimodal semantic encoding model. The first sample multimodal data is the multimodal data used when training the initial multimodal large model.

[0088] Step S62: input the first sample multimodal semantic information into the corresponding linear layers to be trained, and obtain multiple sample modal representations respectively.

[0089] In this embodiment, the first sample multimodal semantic information can be input into the corresponding linear layers to be trained in the initial multimodal large model to obtain multiple sample modal representations. For example, if the first sample multimodal semantic information includes audio semantic information and video semantic information, the audio semantic information and video semantic information can be input into the first linear layer to be trained corresponding to audio and the second linear layer to be trained corresponding to video in the initial multimodal large model, respectively, to obtain the audio representation output by the first linear layer and the video representation output by the second linear layer, respectively.

[0090] Step S63: inputting the sample device description encoding information corresponding to the sample terminal device into the device description linear layer to be trained to obtain a sample device cluster representation.

[0091] In this embodiment, the sample terminal devices are terminal devices used in the model training process. The sample device description encoding information corresponding to the sample terminal devices is the encoded information obtained by inputting the device description corresponding to the sample terminal devices into a pre-trained device encoder. The device description includes basic information about the terminal devices, user guide information, etc. The sample device description encoding information corresponding to the sample terminal devices can be input into the device description linear layer to be trained in the initial multimodal large model to obtain a sample device cluster representation output by the device description linear layer to be trained.

[0092] Step S64: input the multiple sample modal representations and the sample device cluster representation into the pre-trained large language model to obtain multiple first sample instructions corresponding to the first sample multimodal data.

[0093] In this embodiment, multiple sample modal representations output by the linear layer to be trained corresponding to each modality and the sample device cluster representation output by the device description linear layer to be trained can be input into a pre-trained large language model to obtain multiple first sample instructions corresponding to the first sample multimodal data output by the pre-trained large language model. The first sample instructions are control instructions obtained when training the initial multimodal large model.

[0094] Step S65: Based on the multiple first sample instructions and the multiple label instructions corresponding to the first sample multimodal data, the device description linear layer to be trained and the linear layers to be trained corresponding to each modality are trained to obtain the trained device description linear layer and the trained linear layers corresponding to each modality.

[0095] In this embodiment, the multiple label instructions corresponding to the first sample multimodal data are multiple real instructions corresponding to the first sample multimodal data. Based on the multiple first sample instructions output by the pre-trained large language model and the multiple label instructions corresponding to the first sample multimodal data, the network parameters of the device description linear layer to be trained and the linear layers to be trained corresponding to each modality can be updated until the trained device description linear layer and the trained linear layers corresponding to each modality are obtained.

[0096] Step S66: Based on the pre-trained large language model, the trained device description linear layer and the trained linear layers corresponding to each modality, the pre-trained multimodal large model is obtained.

[0097] In this embodiment, a pre-trained multimodal large model can be obtained based on the pre-trained large language model, the trained device description linear layer and the trained linear layers corresponding to each modality, thereby achieving pre-training of the initial multimodal large model.

[0098] In one embodiment, if Figure 6 As shown, Figure 6 Schematic diagram of a multimodal large model pre-training method according to an embodiment of the present invention. Figure 6In this paper, the initial multi-module large model includes a pre-trained large language model and multiple linear layers to be trained. The edge device is a central control robot, which is equipped with a pre-trained multimodal semantic encoding model (e.g., including an audio encoder, a video encoder, etc.) and a pre-trained smart home device encoder. Multiple device descriptions (e.g., device descriptions 1, 2, 3, etc.) can be input into the pre-trained smart home device encoder to obtain corresponding device codes (e.g., device codes 1, 2, 3, etc.). Sample audio data and sample video data are input into the audio encoder and video encoder, respectively, to obtain audio semantic information and video semantic information. The device codes, audio semantic information, and video semantic information are then input into the corresponding linear layers to be trained (e.g., linear layers 1, 2, 3), and then into the pre-trained large language model to generate complex instructions (e.g., instructions 0, 1, 2, 3, etc.). Finally, based on the generated complex instructions and the label instructions corresponding to the sample audio data, sample image data, and sample video data, the network parameters of the multiple linear layers to be trained are updated until the multiple linear layers are trained, thus obtaining the pre-trained multimodal large model.

[0099] In conjunction with the above embodiments, in one embodiment, the present invention further provides a multimodal large model cloud-edge collaborative reasoning method based on semantic encoding. In this method, in addition to the above steps, steps S71 to S77 may also be included, and the above step S11 may specifically include step S78: Step S71: Determine the number M of updated neurons based on the resource information of the edge device.

[0100] In this embodiment, to adapt a pre-trained multimodal semantic coding model (such as a video semantic coding model) to cloud-based large model instruction generation and to enable the deployment and operation of the pre-trained multimodal semantic coding model on edge devices with varying computing power, after obtaining the pre-trained multimodal semantic coding model and before deploying it on the edge device, the pre-trained multimodal semantic coding model to be deployed on the edge device must be fine-tuned. Specifically, the number of neuron updates, M, can be determined based on the resource information of the edge device. This resource information includes, but is not limited to, computing power, available memory capacity, available computing resources, and so on.

[0101] Step S72: semantically encode the second sample multimodal data through the pre-trained multimodal semantic encoding model to obtain the second sample multimodal semantic information and send it to the pre-trained multimodal large model.

[0102] In this embodiment, the second sample multimodal data can be semantically encoded using a pre-trained multimodal semantic encoding model to obtain second sample multimodal semantic information output by the pre-trained multimodal semantic encoding model, and the second sample multimodal semantic information is sent to the pre-trained large multimodal model. The second sample multimodal data is the multimodal data used to fine-tune the multimodal semantic encoding model.

[0103] Step S73: Process the second sample multimodal semantics through the pre-trained multimodal large model to obtain multiple second sample instructions corresponding to the second sample multimodal data.

[0104] In this embodiment, the second sample multimodal semantics can be processed by the pre-trained multimodal large model to obtain multiple second sample instructions corresponding to the second sample multimodal data output by the pre-trained multimodal large model. The second sample instructions are control instructions obtained when fine-tuning the multimodal semantic encoding model.

[0105] Step S74: performing loss calculation based on the multiple second sample instructions and the multiple label instructions corresponding to the second sample multimodal data to obtain a first loss value.

[0106] In this embodiment, a loss calculation can be performed based on multiple second sample instructions and multiple label instructions (i.e., real instructions) corresponding to the second sample multimodal data to obtain a first loss value. The first loss value is the loss value obtained by calculating the loss for all neurons in the last layer of the pre-trained multimodal large model.

[0107] Step S75: After setting the parameters of a neuron in the last layer of the pre-trained multimodal semantic encoding model to 0, the loss is calculated again based on the multiple second sample instructions and the multiple label instructions corresponding to the second sample multimodal data to obtain a second loss value.

[0108] In this embodiment, during the fine-tuning training of the multimodal semantic encoding model in conjunction with the cloud-based large model, it is necessary to calculate the importance of all neurons in the last layer of the pre-trained multimodal large model to determine which neurons are more important for the execution of the task (such as semantic encoding of the corresponding modality). The loss function measures the difference between the model's predicted results and the actual results. The loss of the model on a certain task reflects its importance to the task performance. In this embodiment, the effect of "removing" neurons is achieved by setting the parameters of the neurons to 0, thereby taking the impact of removing neurons on the loss value as the importance of the neuron.

[0109] Specifically, this embodiment may set the parameters of a neuron in the last layer of the pre-trained multimodal semantic encoding model to 0, and then perform loss calculation based on the multiple second sample instructions and the multiple label instructions corresponding to the second sample multimodal data to obtain a second loss value. The second loss value is the loss value obtained by removing the neuron from the last layer of the pre-trained multimodal large model from participating in the loss calculation.

[0110] Step S76: Based on the difference between the first loss value and the second loss value, determine the importance of neurons whose parameters are set to 0, until the importance of all neurons in the last layer of the pre-trained multimodal semantic coding model is obtained.

[0111] In this embodiment, the importance of the neuron whose parameters are set to 0 can be determined based on the difference between the first loss value and the second loss value, that is, the importance of the "removed" neuron can be determined, until the importance of all neurons in the last layer of the pre-trained multimodal semantic coding model is obtained.

[0112] For example, there are 10 neurons in the last layer. The parameters of each of the 10 neurons can be set to 0 in turn, and then the second loss value can be calculated, so as to determine the importance of each of the 10 neurons based on the difference between the 10 second loss values and the first loss value.

[0113] In an optional embodiment, the effect of "removing" neurons is achieved by setting the parameter w of the neurons in the last layer to 0, and the impact of removing neurons on the loss is used as the importance of the neurons, expressed as: ; Among them, T is the transpose, is the first loss value of the model containing neuron parameters, To set the neuron parameter to 0, that is, the second loss value of the model after removing the neuron. is the gradient of the loss function L with respect to the parameter w. The gradient is a vector whose direction indicates the direction in which the function rises fastest at that point, and its modulus indicates the rate of increase. To remove the influence of neurons on the loss, that is, to remove the importance of neurons. The larger the value, the more important the neuron is.

[0114] Step S77: Based on the first loss value, gradient update the M neurons in the last layer of the pre-trained multimodal semantic coding model with high to low importance to obtain a fine-tuned multimodal semantic coding model, and deploy the fine-tuned multimodal semantic coding model in the edge device.

[0115] In this embodiment, based on the first loss value, the M neurons in the last layer of the pre-trained multimodal semantic coding model with high to low importance can be gradient updated to obtain a fine-tuned multimodal semantic coding model, and then the fine-tuned multimodal semantic coding model can be deployed in the edge device.

[0116] In this embodiment, the importance of the last layer of neurons in the pre-trained multimodal semantic coding model is continuously counted during the fine-tuning training process. When the multimodal semantic coding model is deployed on the edge device, the number of neuron updates M determined by the computing power of the edge device is used to select M neurons with high to low importance for gradient updates. This allows the multimodal semantic coding model of the edge device to quickly update its own parameters and perform adaptive computing power optimization under resource-constrained hardware conditions.

[0117] For example, the last layer of the pre-trained multimodal semantic coding model has a total of 10 neurons. When the resources of the side device are sufficient, the number of updated neurons M obtained based on the resource information of the side device is 7. Then, based on the first loss value, the 7 neurons in the last layer of the pre-trained multimodal semantic coding model with high to low importance can be gradient updated to obtain the fine-tuned multimodal semantic coding model; when the resources of the side device are limited, the number of updated neurons M obtained based on the resource information of the side device is 2. Then, based on the first loss value, the 2 neurons in the last layer of the pre-trained multimodal semantic coding model with high to low importance can be gradient updated to obtain the fine-tuned multimodal semantic coding model.

[0118] In this embodiment, a lightweight fine-tuning method for an edge-deployed multimodal semantic coding model is proposed. The effect of "removing" neurons is achieved by setting the parameter w of the neurons to 0. The impact of removing neurons on the loss is used as the importance of the neurons. During the joint fine-tuning training with the large cloud model, the importance of the neurons in the last layer of the edge model is counted. When deploying the application, the neurons with the highest importance are selected for gradient update based on the computing power situation, so that the edge model can quickly update its own parameters under resource-constrained hardware conditions.

[0119] In one embodiment, if Figure 7 As shown, Figure 7 FIG is a schematic diagram of a lightweight fine-tuning method for a multimodal semantic coding model according to an embodiment of the present invention. Figure 7In this method, the edge model is a pre-trained multimodal semantic encoding model to be deployed on the edge device. The multimodal information representation output by the edge model and the representation of the smart home device cluster in the edge device can be input into the cloud model (i.e., the pre-trained multimodal large model) to obtain sample instructions. Based on the sample instructions and user instructions, the first loss of all neurons in the last layer of the edge model and the second loss of each neuron removed from the last layer are calculated. The importance of each neuron in the last layer is determined based on the difference between the second loss and the first loss. Based on the computing power of the edge device and the importance of each neuron, all neurons in the last layer of the edge model are divided into important neurons (neurons to be updated with gradients) and ordinary neurons (neurons that do not undergo gradient updates). If the first loss representation generates an incorrect instruction or cannot be processed (for example, there is a large difference between the sample instruction and the user instruction), no gradient update is performed. If the first loss representation generates a correct and executable instruction (for example, there is a small difference between the sample instruction and the user instruction), the gradients of the important neurons in the last layer of the edge model are updated to achieve fine-tuning of the edge model.

[0120] In one embodiment, if Figure 8 As shown, Figure 8 This is a schematic diagram of a method for implementing semantic communication of a whole-house intelligent central control robot according to an embodiment of the present invention. Figure 8 In this method, a terminal device, a side device (central control robot), and a cloud server are mainly implemented. The central control robot is deployed with a pre-trained multimodal semantic encoding model, and the cloud server is deployed with a pre-trained multimodal large model. The central control robot extracts semantic information from the multimodal information collected by the terminal device and / or related instructions issued by the user, and transmits it to the multimodal large model in the cloud. Because the multimodal large model already has a strong understanding and reasoning ability for text data, it can align the semantic information such as device group representation encoding information, audio, images, and videos sent by the central control robot with the text representation. This allows the multimodal large model to simultaneously receive input forms of multiple different modal combinations and output control instructions that can be executed by all smart devices in the house. The central control robot receives the control instructions generated by the cloud large model and receives the lightweight fine-tuning parameters of the side model from the cloud large model.

[0121] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0122] Based on the same inventive concept, an embodiment of the present invention provides a multimodal large model cloud-edge collaborative reasoning device based on semantic coding. Figure 9 , Figure 9 This is a structural block diagram of a multimodal large model cloud-edge collaborative reasoning device based on semantic coding provided by one embodiment of the present invention. Figure 9 As shown, the semantically encoded multimodal large model cloud-edge collaborative reasoning device of this embodiment may include: The semantic encoding module is deployed on the edge device and is used to receive multimodal data sent by the terminal device, semantically encode the multimodal data through the pre-trained multimodal semantic encoding model, obtain multimodal semantic information, and send it to the cloud server; An inference module, deployed on the cloud server, is used to process the multimodal semantics based on a pre-trained multimodal large model, obtain control instructions, and send them to the edge device; The instruction control module is deployed on the edge device and is used to send the control instruction to the terminal device so that the terminal device performs the corresponding instruction operation based on the control instruction.

[0123] Optionally, the pre-trained multimodal semantic coding model includes at least: a video semantic coding model, and the video semantic coding model includes at least: a video semantic coding module and a mask adaptation module; The semantic encoding module includes: a first mask module, configured to, when the multimodal data includes video data, perform mask processing on a current video frame in the video data to obtain a current image mask feature; A first encoding module, configured to perform semantic encoding on the current image mask feature through the video semantic encoding module to obtain current video semantic information; A first sorting module is configured to sort the current video semantic information by importance through the mask adaptation module, and divide the current video frame into a first semantic layer, a second semantic layer, and a third semantic layer according to the importance of the current video semantic information carried; a second encoding module, configured to perform semantic encoding on the first semantic layer, the second semantic layer, and the third semantic layer through the video semantic encoding module to obtain first video semantic information, second video semantic information, and third video semantic information, respectively; the first video semantic information has a higher importance than the second video semantic information, and the third video semantic information represents the private information in the current video frame; A transmission module is used to send the first video semantic information, the second video semantic information and the third video semantic information to the cloud server through the mask adaptation module when the current network state is not higher than the first state threshold; and send the first video semantic information to the cloud server when the current network state is higher than the first state threshold.

[0124] Optionally, the first mask module includes: a first adjustment module, configured to adjust the initial mask area according to the current network state through the mask adaptation module to obtain a first mask area; a second mask module, configured to perform mask processing on redundant pixel blocks in the first mask area of the current video frame to obtain a first image feature; The splicing module is used to splice the first image feature with the preset pixel block feature to obtain the current image mask feature.

[0125] Optionally, the video semantic coding model is obtained by training based on an initial video semantic coding model, wherein the initial video semantic coding model includes at least an initial video semantic coding module and an initial mask adaptation module; the apparatus further includes a first training module for training the initial video semantic coding model, wherein the first training module includes: a second adjustment module, configured to adjust the sample initial mask region according to the sample network state through the initial mask adaptation module to obtain a sample first mask region; a third mask module, configured to perform mask processing on redundant pixel blocks in a sample first mask region of a sample current video frame in the sample video to obtain a sample first image feature and to concatenate the masked features with the sample preset pixel block feature to obtain a sample current image mask feature; A third encoding module, configured to input the sample current image mask feature into the initial video semantic encoding module to obtain sample video semantic information; a second sorting module, configured to sort the sample video semantic information by importance using the initial mask adaptive module, and divide the sample current video frame into a first sample semantic layer, a second sample semantic layer, and a third sample semantic layer according to the importance of the sample video semantic information carried; a fourth encoding module, configured to perform semantic encoding on the first sample semantic layer, the second sample semantic layer, and the third sample semantic layer through the initial video semantic encoding module to obtain first sample video semantic information, second sample video semantic information, and third sample video semantic information, respectively; a reconstruction module, configured to input the first sample video semantic information, the second sample video semantic information, and the third sample video semantic information into a decoder to obtain a current reconstructed image; The second training module is used to train the initial video semantic coding module and the initial mask adaptation module based on the current reconstructed image and the sample current video frame to obtain the video semantic coding module and the mask adaptation module.

[0126] Optionally, the pre-trained multimodal semantic encoding model includes at least: a text semantic encoder, an image semantic encoder, and an audio semantic encoder; the pre-trained multimodal large model is obtained by training based on the initial multimodal large model, and the text semantic encoder, the image semantic encoder, and the audio semantic encoder are respectively trained based on the initial text semantic encoder, the initial image semantic encoder, and the initial audio semantic encoder. The device also includes: a third training module for training the initial text semantic encoder, the initial image semantic encoder, and the initial audio semantic encoder, and the third training module includes: a fifth encoding module, configured to input the first sample text data, the first sample image data, and the first sample audio data into the initial text semantic encoder, the initial image semantic encoder, and the initial audio semantic encoder, respectively, to obtain first sample text semantic information, first sample image semantic information, and first sample audio semantic information; a model processing module, configured to input the sample device description encoding information corresponding to the sample terminal device, the first sample text semantic information, the first sample image semantic information, and the first sample audio semantic information into the initial multimodal large model, align the sample device description encoding information, the first sample image semantic information, and the first sample audio semantic information with the first sample text semantic information, and obtain a first sample instruction, a second sample instruction, and a third sample instruction, respectively; a fourth training module, configured to train the initial text semantic encoder based on the first sample instruction and the label instruction corresponding to the first sample text data to obtain the text semantic encoder; a fifth training module, configured to train the initial image semantic encoder based on the second sample instruction and the label instruction corresponding to the first sample image data to obtain the image semantic encoder; The sixth training module is used to train the initial audio semantic encoder based on the third sample instruction and the label instruction corresponding to the first sample audio data to obtain the audio semantic encoder.

[0127] Optionally, the pre-trained multimodal large model is obtained based on the training of the initial multimodal large model, and the initial multimodal large model at least includes: a pre-trained large language model, a device description linear layer to be trained, and a linear layer to be trained corresponding to each modality; the apparatus further includes: a seventh training module for training the initial multimodal large model, and the seventh training module includes: A first input module, configured to input first sample multimodal data into the pre-trained multimodal semantic encoding model to obtain first sample multimodal semantic information; A second input module is configured to input the multimodal semantic information of the first sample into the corresponding linear layers to be trained, thereby obtaining multiple sample modal representations respectively; A third input module is used to input the sample device description encoding information corresponding to the sample terminal device into the device description linear layer to be trained to obtain the sample device cluster representation; a fourth input module, configured to input the plurality of sample modal representations and the sample device cluster representation into the pre-trained large language model to obtain a plurality of first sample instructions corresponding to the first sample multimodal data; a linear layer training module, configured to train the device description linear layer to be trained and the linear layers to be trained corresponding to the respective modalities based on the plurality of first sample instructions and the plurality of label instructions corresponding to the first sample multimodal data, to obtain trained device description linear layers and trained linear layers corresponding to the respective modalities; A model determination module is used to obtain the pre-trained multimodal large model based on the pre-trained large language model, the trained device description linear layer and the trained linear layers corresponding to each modality.

[0128] Optionally, the device further comprises: A quantity determination module, configured to determine the number M of neuron updates based on resource information of the edge device; A first processing module is configured to perform semantic encoding on the second sample multimodal data using the pre-trained multimodal semantic encoding model to obtain the second sample multimodal semantic information and send the obtained information to the pre-trained multimodal large model; A second processing module is configured to process the second sample multimodal semantics using the pre-trained multimodal large model to obtain a plurality of second sample instructions corresponding to the second sample multimodal data; a first calculation module, configured to perform loss calculation based on the plurality of second sample instructions and the plurality of label instructions corresponding to the second sample multimodal data, to obtain a first loss value; A second calculation module is configured to set a parameter of a neuron in the last layer of the pre-trained multimodal semantic encoding model to 0, and then perform loss calculation based on the multiple second sample instructions and the multiple label instructions corresponding to the second sample multimodal data to obtain a second loss value; a degree determination module, configured to determine the importance of neurons whose parameters are set to 0 based on the difference between the first loss value and the second loss value, until the importance of all neurons in the last layer of the pre-trained multimodal semantic encoding model is obtained; A gradient update module is configured to perform a gradient update on M neurons in the last layer of the pre-trained multimodal semantic coding model, ranked from high to low importance, based on the first loss value, to obtain a fine-tuned multimodal semantic coding model, and deploy the fine-tuned multimodal semantic coding model in the edge device; Semantic encoding module, including: The semantic coding submodule is deployed on the edge device and is used to receive the multimodal data sent by the terminal device, semantically encode the multimodal data through the fine-tuned multimodal semantic coding model, obtain multimodal semantic information and send it to the cloud server.

[0129] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in the multimodal large model cloud-edge collaborative reasoning method based on semantic encoding as described in any of the above embodiments of the present invention.

[0130] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes, it implements the steps of the semantic coding-based multimodal large model cloud-edge collaborative reasoning method described in any of the above embodiments of the present invention.

[0131] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to in detail.

[0132] The above is a detailed introduction to the semantic coding-based multimodal large-model cloud-edge collaborative reasoning method, device, equipment and medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A multimodal large model cloud-edge collaborative reasoning method based on semantic encoding, characterized by: The method comprises: The edge device receives the multimodal data sent by the terminal device, semantically encodes the multimodal data through a pre-trained multimodal semantic encoding model, obtains multimodal semantic information and sends it to the cloud server; The cloud server processes the multimodal semantics based on a pre-trained multimodal large model, obtains a control instruction, and sends it to the edge device; The edge device sends the control instruction to the terminal device, so that the terminal device performs a corresponding instruction operation based on the control instruction.

2. The multimodal large model cloud-edge collaborative reasoning method based on semantic encoding according to claim 1 is characterized in that: The pre-trained multimodal semantic coding model at least includes: a video semantic coding model, and the video semantic coding model at least includes: a video semantic coding module and a mask adaptation module; Semantically encoding the multimodal data using a pre-trained multimodal semantic encoding model to obtain multimodal semantic information and send it to a cloud server, including: In a case where the multimodal data includes video data, performing mask processing on a current video frame in the video data to obtain a current image mask feature; Performing semantic encoding on the current image mask feature by the video semantic encoding module to obtain current video semantic information; sorting the current video semantic information by importance through the mask adaptive module, and dividing the current video frame into a first semantic layer, a second semantic layer, and a third semantic layer according to the importance of the current video semantic information carried; semantically encoding the first semantic layer, the second semantic layer, and the third semantic layer by the video semantic encoding module to obtain first video semantic information, second video semantic information, and third video semantic information, respectively; the first video semantic information has a higher importance than the second video semantic information, and the third video semantic information represents the private information in the current video frame; Through the mask adaptive module, when the current network status is not higher than the first status threshold, the first video semantic information, the second video semantic information and the third video semantic information are sent to the cloud server; when the current network status is higher than the first status threshold, the first video semantic information is sent to the cloud server.

3. The multimodal large model cloud-edge collaborative reasoning method based on semantic encoding according to claim 2 is characterized in that: Performing mask processing on the current video frame in the video data to obtain current image mask features, including: By means of the mask adaptation module, the initial mask area is adjusted according to the current network state to obtain a first mask area; Performing mask processing on redundant pixel blocks in a first mask area of the current video frame to obtain a first image feature; The first image feature is spliced with the preset pixel block feature to obtain the current image mask feature.

4. The multimodal large model cloud-edge collaborative reasoning method based on semantic coding according to claim 2 is characterized in that: The video semantic coding model is obtained by training based on the initial video semantic coding model, and the initial video semantic coding model at least includes: an initial video semantic coding module and an initial mask adaptation module; the training steps of the initial video semantic coding model include: By means of the initial mask adaptation module, the initial mask region of the sample is adjusted according to the network state of the sample to obtain a first mask region of the sample; Masking the redundant pixel blocks in the sample first mask area of the sample current video frame in the sample video to obtain the sample first image feature and splicing it with the sample preset pixel block feature to obtain the sample current image mask feature; Inputting the sample current image mask feature into the initial video semantic encoding module to obtain sample video semantic information; sorting the sample video semantic information by importance through the initial mask adaptive module, and dividing the sample current video frame into a first sample semantic layer, a second sample semantic layer, and a third sample semantic layer according to the importance of the sample video semantic information carried; Performing semantic encoding on the first sample semantic layer, the second sample semantic layer, and the third sample semantic layer by the initial video semantic encoding module to obtain first sample video semantic information, second sample video semantic information, and third sample video semantic information, respectively; Inputting the first sample video semantic information, the second sample video semantic information, and the third sample video semantic information into a decoder to obtain a current reconstructed image; The initial video semantic coding module and the initial mask adaptation module are trained based on the current reconstructed image and the sample current video frame to obtain the video semantic coding module and the mask adaptation module.

5. The multimodal large model cloud-edge collaborative reasoning method based on semantic encoding according to claim 1 is characterized in that: The pre-trained multimodal semantic encoding model includes at least: a text semantic encoder, an image semantic encoder, and an audio semantic encoder; the pre-trained multimodal large model is obtained by training based on the initial multimodal large model, the text semantic encoder, the image semantic encoder, and the audio semantic encoder are respectively trained based on the initial text semantic encoder, the initial image semantic encoder, and the initial audio semantic encoder, and the training steps of the initial text semantic encoder, the initial image semantic encoder, and the initial audio semantic encoder include: Inputting the first sample text data, the first sample image data and the first sample audio data into the initial text semantic encoder, the initial image semantic encoder and the initial audio semantic encoder respectively to obtain first sample text semantic information, first sample image semantic information and first sample audio semantic information; Inputting the sample device description encoding information corresponding to the sample terminal device, the first sample text semantic information, the first sample image semantic information, and the first sample audio semantic information into the initial multimodal large model, and aligning the sample device description encoding information, the first sample image semantic information, and the first sample audio semantic information with the first sample text semantic information to obtain a first sample instruction, a second sample instruction, and a third sample instruction, respectively; Training the initial text semantic encoder based on the first sample instruction and the label instruction corresponding to the first sample text data to obtain the text semantic encoder; Training the initial image semantic encoder based on the second sample instruction and the label instruction corresponding to the first sample image data to obtain the image semantic encoder; The initial audio semantic encoder is trained based on the third sample instruction and the label instruction corresponding to the first sample audio data to obtain the audio semantic encoder.

6. The multimodal large model cloud-edge collaborative reasoning method based on semantic encoding according to claim 1 is characterized in that: The pre-trained multimodal large model is obtained by training based on the initial multimodal large model. The initial multimodal large model at least includes: a pre-trained large language model, a device description linear layer to be trained, and linear layers to be trained corresponding to each modality. The training steps of the initial multimodal large model include: Inputting the first sample multimodal data into the pre-trained multimodal semantic encoding model to obtain first sample multimodal semantic information; Inputting the first sample multimodal semantic information into the corresponding linear layers to be trained to obtain multiple sample modal representations respectively; Inputting the sample device description encoding information corresponding to the sample terminal device into the device description linear layer to be trained to obtain a sample device cluster representation; Inputting the multiple sample modal representations and the sample device cluster representation into the pre-trained large language model to obtain multiple first sample instructions corresponding to the first sample multimodal data; Based on the multiple first sample instructions and the multiple label instructions corresponding to the first sample multimodal data, the device description linear layer to be trained and the linear layers to be trained corresponding to the respective modalities are trained to obtain the trained device description linear layer and the trained linear layers corresponding to the respective modalities; The pre-trained multimodal large model is obtained based on the pre-trained large language model, the trained device description linear layer and the trained linear layers corresponding to each modality.

7. The multimodal large model cloud-edge collaborative reasoning method based on semantic coding according to any one of claims 1 to 6, characterized in that: The method further comprises: Determine the number of neurons to update M based on the resource information of the edge device; Performing semantic encoding on the second sample multimodal data using the pre-trained multimodal semantic encoding model to obtain the second sample multimodal semantic information and sending it to the pre-trained multimodal large model; Processing the second sample multimodal semantics by using the pre-trained multimodal large model to obtain a plurality of second sample instructions corresponding to the second sample multimodal data; Performing loss calculation based on the plurality of second sample instructions and a plurality of label instructions corresponding to the second sample multimodal data to obtain a first loss value; After setting a parameter of a neuron in the last layer of the pre-trained multimodal semantic encoding model to 0, performing loss calculation again based on the multiple second sample instructions and the multiple label instructions corresponding to the second sample multimodal data to obtain a second loss value; Determining the importance of neurons whose parameters are set to 0 based on the difference between the first loss value and the second loss value, until the importance of all neurons in the last layer of the pre-trained multimodal semantic encoding model is obtained; Based on the first loss value, gradient update the M neurons in the last layer of the pre-trained multimodal semantic coding model, from high to low importance, to obtain a fine-tuned multimodal semantic coding model, and deploy the fine-tuned multimodal semantic coding model on the side device; The edge device receives the multimodal data sent by the terminal device, semantically encodes the multimodal data using a pre-trained multimodal semantic encoding model, obtains multimodal semantic information, and sends it to the cloud server, including: The edge device receives the multimodal data sent by the terminal device, semantically encodes the multimodal data through the fine-tuned multimodal semantic encoding model, obtains multimodal semantic information and sends it to the cloud server.

8. A multimodal large model cloud-edge collaborative reasoning device based on semantic coding, characterized by: The device comprises: The semantic encoding module is deployed on the edge device and is used to receive multimodal data sent by the terminal device, semantically encode the multimodal data through the pre-trained multimodal semantic encoding model, obtain multimodal semantic information, and send it to the cloud server; An inference module, deployed on the cloud server, is used to process the multimodal semantics based on a pre-trained multimodal large model, obtain control instructions, and send them to the edge device; The instruction control module is deployed on the edge device and is used to send the control instruction to the terminal device so that the terminal device performs the corresponding instruction operation based on the control instruction.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the multimodal large model cloud-edge collaborative reasoning method based on semantic encoding is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the semantic coding-based multimodal large model cloud-edge collaborative reasoning method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Data transmission method and device of cloud computer, electronic equipment and storage medium

    CN120979819A