Data transmission method and device of cloud computer, electronic equipment and storage medium

By using a multimodal feature extraction model to identify and encrypt sensitive data in cloud computer sessions, the problem of inaccurate multimodal data identification in existing technologies is solved, thereby improving the security and efficiency of data transmission.

CN120979819APending Publication Date: 2025-11-18SHENZHEN WANCHENG IOT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511423795.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing cloud computer data transmission methods cannot accurately identify key information in multimodal data (such as voice, images, and video), leading to the exposure of sensitive content to risks and affecting the security and efficiency of data transmission.

Method used

The trained multimodal feature extraction model is used to extract semantics from cloud computer session data, identify and encrypt sensitive data, generate target transmission data, and transmit it to the cloud.

Benefits of technology

It improves the accuracy of identifying key information in multimodal data, ensuring the security and efficiency of sensitive content during transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979819A_ABST
    Figure CN120979819A_ABST
Patent Text Reader

Abstract

The invention provides a data transmission method of a cloud computer. The method comprises the following steps: acquiring to-be-transmitted data generated in a session process of the cloud computer; performing semantic extraction processing on the to-be-transmitted data through the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data; identifying target sensitive data in the to-be-transmitted data based on the target semantic vector; based on the target sensitive data, performing encryption coding processing on the to-be-transmitted data to generate target transmission data; and transmitting the target transmission data to the cloud based on the target transmission data. According to the invention, the problem that the security and efficiency of data transmission are affected due to the fact that sensitive contents in the cloud computer session process are exposed to risks because key information of common multi-modal data (such as voice, images, videos and the like) in the session cannot be accurately identified and the key information cannot be encrypted and transmitted in time in the existing data transmission method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cloud computers, and in particular to a data transmission method and device for a cloud computer, an electronic device and a storage medium. BACKGROUND

[0002] With the popularization of cloud computing technology, cloud computers, as a kind of virtual desktop service based on the cloud, are widely used in enterprise office, remote collaboration and other fields. At present, the data transmission of cloud computers mainly selects key information for encrypted transmission based on theme recognition and semantic difference analysis of text conversation data, which to some extent improves the security and processing efficiency of text data. However, the existing data transmission method often ignores the common multi-modal data (such as voice, image, video, etc.) in the conversation. In actual business scenarios, conversations often contain information in various forms, and processing only text data may result in missing context information, affecting the accuracy of key information identification. Therefore, the existing data transmission method cannot accurately identify the key information of common multi-modal data (such as voice, image, video, etc.) in the conversation, cannot timely encrypt and transmit the key information, and leads to the exposure of sensitive content in the cloud computer conversation process, affecting the security and efficiency of data transmission. SUMMARY

[0003] The present application provides a data transmission method for a cloud computer, which aims to solve the problem that the existing data transmission method cannot accurately identify the key information of common multi-modal data (such as voice, image, video, etc.) in the conversation, cannot timely encrypt and transmit the key information, and leads to the exposure of sensitive content in the cloud computer conversation process, affecting the security and efficiency of data transmission. By training a multi-modal feature extraction model, the semantic extraction processing of the generated data to be transmitted in the cloud computer conversation process is performed, and the target semantic vector corresponding to the data to be transmitted is obtained. According to the target semantic vector, the target sensitive data in the data to be transmitted is identified, and the data to be transmitted is encrypted and encoded by the target sensitive data to generate target transmission data, and the target transmission data is transmitted to the cloud, solving the problem that the existing data transmission method cannot accurately identify the key information of common multi-modal data (such as voice, image, video, etc.) in the conversation, cannot timely encrypt and transmit the key information, and leads to the exposure of sensitive content in the cloud computer conversation process, affecting the security and efficiency of data transmission.

[0004] In a first aspect, the present application provides a data transmission method for a cloud computer, which includes the following steps:

[0005] Obtaining the data to be transmitted generated in the cloud computer conversation process;

[0006] The trained multi-modal feature extraction model is used for semantic extraction processing on the to-be-transmitted data, to obtain a target semantic vector corresponding to the to-be-transmitted data.

[0007] Based on the target semantic vector, target sensitive data in the to-be-transmitted data is identified.

[0008] Based on the target sensitive data, the to-be-transmitted data is encrypted and coded to generate target transmission data.

[0009] Based on the target transmission data, the target transmission data is transmitted to the cloud.

[0010] Optionally, the to-be-transmitted data includes first text data, and the trained multi-modal feature extraction model is used for semantic extraction processing on the to-be-transmitted data to obtain a target semantic vector corresponding to the to-be-transmitted data.

[0011] The trained multi-modal feature extraction model is used for first semantic extraction on the first text data to obtain a first semantic vector corresponding to the first text data.

[0012] Optionally, the to-be-transmitted data includes a screen frame sequence and second text data corresponding to the screen frame sequence, and the trained multi-modal feature extraction model is used for semantic extraction processing on the to-be-transmitted data to obtain a target semantic vector corresponding to the to-be-transmitted data.

[0013] The trained multi-modal feature extraction model is used for second semantic extraction processing on the screen frame sequence to obtain a second semantic vector corresponding to the screen frame sequence.

[0014] The second text data is subjected to third semantic extraction processing to obtain a third semantic vector corresponding to the second text data.

[0015] Based on the second semantic vector and the third semantic vector, a target semantic vector corresponding to the screen frame sequence is obtained.

[0016] Optionally, the to-be-transmitted data includes audio data, and the trained multi-modal feature extraction model is used for semantic extraction processing on the to-be-transmitted data to obtain a target semantic vector corresponding to the to-be-transmitted data.

[0017] The audio data is subjected to data preprocessing to obtain preprocessed audio data.

[0018] The trained multi-modal feature extraction model is used for semantic extraction processing on the preprocessed audio data to obtain a target semantic vector corresponding to the audio data.

[0019] Optionally, before the semantic extraction processing of the to-be-transmitted data by the trained multi-modal feature extraction model to obtain the target semantic vector corresponding to the to-be-transmitted data, the method further comprises:

[0020] Obtain a text training data set and a pre-trained multi-modal feature extraction model, the text training data set comprising sample text data and semantic annotation data corresponding to the sample text data;

[0021] Train the pre-trained multi-modal feature extraction model through the text training data set. During the training process, the pre-trained multi-modal feature extraction model is adjusted in parameters through a minimum loss function. After the training is completed, a trained first multi-modal feature extraction model is obtained.

[0022] Alternatively, obtain an audio training data set and a pre-trained multi-modal feature extraction model, the audio training data set comprising sample audio data and semantic annotation data corresponding to the sample audio data;

[0023] Train the pre-trained multi-modal feature extraction model through the audio training data set. During the training process, the pre-trained multi-modal feature extraction model is adjusted in parameters through a minimum loss function. After the training is completed, a trained second multi-modal feature extraction model is obtained.

[0024] Alternatively, obtain a screen frame sequence training data set and a pre-trained multi-modal feature extraction model, the screen frame sequence data training set comprising sample screen frame sequences, image semantic annotation data of image data corresponding to the screen frame sequences, and text semantic annotation data of image data;

[0025] Train the pre-trained multi-modal feature extraction model through the screen frame sequence training data set. During the training process, the pre-trained multi-modal feature extraction model is adjusted in parameters through a minimum loss function. After the training is completed, a trained third multi-modal feature extraction model is obtained.

[0026] Alternatively, obtain a text training data set, an audio training data set, a screen frame sequence training data set, and a pre-trained multi-modal feature extraction model, the text training data set comprising sample text data and semantic annotation data corresponding to the sample text data, the audio training data set comprising sample audio data and semantic annotation data corresponding to the sample audio data, and the screen frame sequence data training set comprising sample screen frame sequences, image semantic annotation data of image data corresponding to the screen frame sequences, and text semantic annotation data of image data;

[0027] The pre-trained multi-modal feature extraction model is trained through the text training data set, the audio training data set and the screen frame sequence training data set, in the training process, the pre-trained multi-modal feature extraction model is adjusted through a minimum loss function, and after the training is completed, a trained fourth multi-modal feature extraction model is obtained.

[0028] Optionally, the encryption and coding processing of the to-be-transmitted data based on the target sensitive data includes:

[0029] Sensitive level identification processing is performed on the target sensitive data to obtain a target sensitive level corresponding to the target sensitive data.

[0030] The encryption and coding processing of the to-be-transmitted data based on the target sensitive level generates target transmission data.

[0031] Optionally, the encryption and coding processing of the to-be-transmitted data based on the target sensitive level includes:

[0032] Based on the correspondence between the sensitive level and the encryption algorithm, the target sensitive level is determined in the correspondence between the sensitive level and the encryption algorithm to determine the corresponding target encryption algorithm, and different sensitive levels correspond to different encryption algorithms.

[0033] The encryption and coding processing of the to-be-transmitted data based on the target encryption algorithm obtains target transmission data.

[0034] In a second aspect, the embodiment of the present application further provides a data transmission device of a cloud computer, and the data transmission device of the cloud computer includes:

[0035] An acquisition module is configured to acquire to-be-transmitted data generated in a cloud computer session process.

[0036] A first processing module is configured to perform semantic extraction processing on the to-be-transmitted data through the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data.

[0037] An identification module is configured to identify target sensitive data in the to-be-transmitted data based on the target semantic vector.

[0038] A second processing module is configured to perform encryption and coding processing on the to-be-transmitted data based on the target sensitive data to generate target transmission data.

[0039] A transmission module is configured to transmit the target transmission data to the cloud based on the target transmission data.

[0040] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the data transmission method of the cloud computer when executing the computer program

[0041] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps in the data transmission method of the cloud computer.

[0042] In the embodiment of the present application, the data to be transmitted generated in the cloud computer session process is obtained, the multi-modal feature extraction model is trained, the semantic extraction processing is performed on the data to be transmitted, the target semantic vector corresponding to the data to be transmitted is obtained, the target sensitive data in the data to be transmitted is identified based on the target semantic vector, the target transmission data is generated by performing encryption coding processing on the data to be transmitted based on the target sensitive data, and the target transmission data is transmitted to the cloud based on the target transmission data. In the embodiment of the present application, the multi-modal feature extraction model is trained, the semantic extraction processing is performed on the data to be transmitted generated in the cloud computer session process, the target semantic vector corresponding to the data to be transmitted is obtained, the target sensitive data in the data to be transmitted is identified based on the target semantic vector, the target transmission data is generated by performing encryption coding processing on the data to be transmitted based on the target sensitive data, and the target transmission data is transmitted to the cloud, thereby solving the problem that the existing data transmission method cannot accurately identify the key information of the common multi-modal data (such as voice, image, video, etc.) in the session, cannot timely encrypt and transmit the key information, and the sensitive content in the cloud computer session process is exposed to risk, thereby affecting the data transmission safety and efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0044] Figure 1 is a flowchart of a data transmission method of a cloud computer provided by an embodiment of the present application;

[0045] Figure 2 is a structural schematic diagram of a data transmission device of a cloud computer provided by an embodiment of the present application;

[0046] Figure 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0048] As Figure 1 shown, Figure 1 is a flowchart of a data transmission method of a cloud computer provided by an embodiment of the present application. The data transmission method of the cloud computer comprises the following steps:

[0049] 101. Obtain to-be-transmitted data generated in a cloud computer session process.

[0050] In the embodiment of the present application, the data transmission method of the cloud computer can be applied to the cloud computer. The cloud computer is a virtual desktop service based on cloud computing technology, which migrates the computing, storage and running environment of a traditional personal computer to a cloud server. A user can access a cloud computer with complete functions by connecting a lightweight terminal device (such as a notebook computer, a tablet computer or a mobile phone) to a network.

[0051] The cloud computer session process can be understood as a process in which a user interacts with the cloud computer.

[0052] The to-be-transmitted data can be understood as data that needs to be sent to the cloud in the cloud computer session process.

[0053] It should be noted that the to-be-transmitted data can be multi-modal to-be-transmitted data, can be text data, can be audio data, or can be video data, etc.

[0054] 102. Perform semantic extraction processing on the to-be-transmitted data by using the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data.

[0055] In the embodiments of the present application, the trained multi-modal feature extraction model described above can be a multi-modal feature extraction model constructed based on deep learning or machine learning, such as CLIP, LLM, etc. The CLIP (Contrastive Language-Image Pretraining) described above is a multi-modal model that realizes cross-modal semantic alignment of images and texts through contrastive learning. The core idea of CLIP is to embed images and texts into the same latent space, so that matching image-text pairs are closer in the space, and non-matching pairs are farther apart. The core mechanism of the LLM (Large Language Model) described above is based on the Transformer architecture, which realizes language processing by analyzing long-distance dependencies in text.

[0056] The trained multi-modal feature extraction model described above can identify the semantic vector of the data to be transmitted.

[0057] The semantic extraction process described above can be understood as a process of analyzing the data to be transmitted by the trained multi-modal feature extraction model and extracting the semantic features in the data to be transmitted.

[0058] The semantic vector described above is used to represent the semantic features of the data to be transmitted.

[0059] The target semantic vector described above is a key feature extracted from the data to be transmitted, which can represent the subject and content of the entire data to be transmitted. The target semantic vector is used to capture the complete meaning expressed by the data to be transmitted, and the target semantic vector includes all information and concepts contained in the data to be transmitted.

[0060] In a possible embodiment, when the data to be transmitted is text data, the trained multi-modal feature extraction model can be used to perform semantic extraction on the text data, and output the semantic vector corresponding to the text data. When the data to be transmitted is audio data, the multi-modal feature extraction model can be used to perform semantic extraction on the audio data, and output the semantic vector corresponding to the audio data. When the data to be transmitted is video data, the multi-modal feature extraction model can be used to perform semantic extraction on the video data, and output the semantic vector corresponding to the video data.

[0061] 103、Based on the target semantic vector, the target sensitive data in the data to be transmitted is identified.

[0062] In the embodiments of the present application, the target sensitive data in the data to be transmitted can be identified through the target semantic vector corresponding to the data to be transmitted.

[0063] The sensitive data can be understood as data that may cause serious harm to society or individuals after being leaked, and the sensitive data includes personal privacy information and enterprise confidential information.

[0064] The target sensitive data can be sensitive data in the to-be-transmitted data.

[0065] In a possible embodiment, the target semantic vector includes keywords such as "name", "telephone", "bank account number", and "password", which can be identified in the to-be-transmitted data, and the data information of "name", "telephone", "bank account number", and "password" in the to-be-transmitted data is identified, and the data information of "name", "telephone", "bank account number", and "password" is determined as the target sensitive data. Further, the sensitive information can be encrypted and coded to ensure that it is not leaked during transmission.

[0066] 104. Based on the target sensitive data, the to-be-transmitted data is encrypted and coded to generate target transmission data.

[0067] In the embodiment of the application, the encryption and coding process can be understood as a process of converting the target sensitive data in the to-be-transmitted data into a format that cannot be directly identified through a specific rule, so as to protect the data from unauthorized access.

[0068] The target transmission data can be transmission data after encryption and coding.

[0069] In a possible embodiment, the target sensitive data in the to-be-transmitted data can be encrypted using an encryption algorithm to generate encrypted target transmission data. It can be understood that, during data transmission, even if other users intercept the target transmission data, they cannot know the original content of the target transmission data. The encryption algorithm can be understood as a mathematical function that converts plaintext into ciphertext through a specific rule, and needs to be used with a key. Only authorized parties can decrypt and restore the information. The encryption algorithm can be AES, DES, or the like.

[0070] 105. Based on the target transmission data, the target transmission data is transmitted to the cloud.

[0071] In the embodiment of the application, after obtaining the target transmission data, the target data can be transmitted to the cloud.

[0072] In the embodiment of the present application, the data to be transmitted generated in the cloud computer session process is obtained; the trained multi-modal feature extraction model is used for semantic extraction processing of the data to be transmitted to obtain the target semantic vector corresponding to the data to be transmitted; based on the target semantic vector, the target sensitive data in the data to be transmitted is identified; based on the target sensitive data, the data to be transmitted is encrypted and encoded to generate target transmission data; and based on the target transmission data, the target transmission data is transmitted to the cloud. Through the trained multi-modal feature extraction model, the data to be transmitted generated in the cloud computer session process is subjected to semantic extraction processing to obtain the target semantic vector corresponding to the data to be transmitted. According to the target semantic vector, the target sensitive data in the data to be transmitted is identified, and the data to be transmitted is encrypted and encoded based on the target sensitive data to generate target transmission data, and the target transmission data is transmitted to the cloud. The problem that the existing data transmission method cannot accurately identify the key information of the common multi-modal data (such as voice, image, video, etc.) in the session, cannot timely encrypt and transmit the key information, and causes the sensitive content in the cloud computer session process to be exposed to risk, affecting the data transmission safety and efficiency is solved.

[0073] It can be understood that in the specific embodiments of the present application, data related to transmission data, feature data, semantic data, sensitive data, user data, etc. are involved. When the embodiments in the present application are applied to specific products or technologies, the user's permission or consent needs to be obtained, and the collection, use and processing of related data, as well as the training, deployment and calling of algorithm models, need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0074] Optionally, the data to be transmitted includes first text data. In the step of obtaining the target semantic vector corresponding to the data to be transmitted by using the trained multi-modal feature extraction model to perform semantic extraction processing on the data to be transmitted, the first text data can be subjected to first semantic extraction by using the trained multi-modal feature extraction model to obtain a first semantic vector corresponding to the first text data.

[0075] In the embodiment of the present application, the first text data can be text data to be transmitted.

[0076] The trained multi-modal feature extraction model can be a multi-modal feature extraction model constructed based on deep learning or machine learning, such as BERT, LLM, etc. The above-mentioned BERT (Bidirectional Encoder Representations from Transformers) is a bidirectional encoder model based on the Transformer architecture. BERT is stacked by multiple Transformer encoder layers, adopts a bidirectional self-attention mechanism, and can capture left and right context information of words at the same time. The above-mentioned LLM (Large Language Model) is used to understand and generate natural language. The core mechanism of LLM is based on the Transformer architecture, which realizes language processing by analyzing long-distance dependencies in text.

[0077] The trained multi-modal feature extraction model is obtained by training a pre-trained multi-modal feature extraction model based on a text training data set. The text training data set includes sample text data and semantic annotation data corresponding to the sample text data. The annotation data is a process of converting raw data into a form understandable by machine learning models. By adding semantic labels or structured information to the data, the machine can learn and perform classification, detection, and other tasks. The pre-trained multi-modal feature can be a multi-modal feature extraction model constructed based on deep learning or machine learning, such as BERT, LLM, etc. The training can be supervised training. Supervised training is to train a model using a set of data with known labels. By optimizing model parameters, the model can predict the labels of new data or make decisions based on the characteristics of existing data. During the training process, the parameters of the model can be adjusted using a minimum loss function to minimize the difference between the output labels of the model and the input data. The loss function is used to measure the difference between the model's prediction and the true result. The purpose is to minimize the loss function value by adjusting the model parameters, thereby improving the prediction accuracy. The loss function can be a mean square error loss function, a cross-entropy loss function, etc. The parameter adjustment refers to the process of adjusting the weights and biases of the model to optimize the performance of the model. During the training process, the parameters of the model are optimized based on the annotation data to make the model have better prediction or decision-making ability.

[0078] The trained multi-modal feature extraction model can identify semantic vectors in text data.

[0079] The first semantic extraction can be understood as a process of analyzing the first text data by the trained multi-modal feature extraction model and extracting semantic features in the first text data.

[0080] The first semantic vector can be a semantic vector corresponding to the first text data.

[0081] It should be noted that the first text data can be subjected to first semantic extraction by the trained multi-modal feature extraction model to obtain a first semantic vector corresponding to the first text data.

[0082] Optionally, the to-be-transmitted data includes a screen frame sequence and second text data corresponding to the screen frame sequence, and in the step of performing semantic extraction on the to-be-transmitted data by the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data, the trained multi-modal feature extraction model can be used to perform second semantic extraction on the screen frame sequence to obtain a second semantic vector corresponding to the screen frame sequence, perform third semantic extraction on the second text data to obtain a third semantic vector corresponding to the second text data, and obtain the target semantic vector corresponding to the screen frame sequence based on the second semantic vector and the third semantic vector.

[0083] In the embodiments of the present application, the screen frame sequence can be understood as a sequence arranged in time sequence by continuous image frames, wherein each frame is a picture, and they are arranged in time sequence to form a video or animation.

[0084] The second text data can be understood as text description or label information corresponding to each frame image in the screen frame sequence.

[0085] The trained multi-modal feature extraction model can be a multi-modal feature extraction model constructed based on deep learning or machine learning, such as CLIP, VIT, etc. The CLIP (Contrastive Language-Image Pretraining) is a multi-modal model that realizes cross-modal semantic alignment of images and texts through contrastive learning. The core idea of CLIP is to embed images and texts into the same latent space, so that matching image-text pairs are closer in the space, and non-matching pairs are farther apart. The core of VIT (Vision Transformer) is to divide the image into fixed-size patches, treat each patch as an independent element in the sequence, and process global dependency through a Transformer encoder.

[0086] The trained multi-modal feature extraction model is obtained by training a pre-trained multi-modal feature extraction model based on a training data set. The training data set includes a sample screen frame sequence, image semantic annotation data corresponding to image data of the screen frame sequence, and text semantic annotation data of the image data. The pre-trained multi-modal feature extraction model can be a multi-modal feature extraction model based on deep learning or machine learning, such as CLIP, VIT, etc. The training can be supervised training, which uses a set of data with known labels to train the model, optimizes the model parameters to enable the model to predict the labels of new data or make decisions based on the characteristics of existing data. During the training process, the parameters of the model can be adjusted using a minimum loss function to minimize the difference between the output labels of the model and the input data.

[0087] The trained multi-modal feature extraction model can identify the semantic vector corresponding to the screen frame sequence.

[0088] The second semantic extraction process can be understood as a process of analyzing the screen frame sequence by the trained multi-modal feature extraction model and extracting the semantic features.

[0089] The second semantic vector can be a semantic vector corresponding to the screen frame sequence extracted by analyzing the screen frame sequence by the trained multi-modal feature extraction model.

[0090] The third semantic extraction process can be understood as a process of analyzing the second text data by the trained multi-modal feature extraction model and extracting the semantic features in the second text data.

[0091] The third semantic vector can be a semantic vector corresponding to the second text data extracted by analyzing the second text data by the trained multi-modal feature extraction model.

[0092] Further, the second semantic vector and the third semantic vector can be fused to obtain a target semantic vector corresponding to the screen frame sequence. The fusion process can be understood as a process of fusing the second semantic vector and the third semantic vector to obtain a fused feature vector.

[0093] Optionally, the data to be transmitted includes audio data. In the step of performing semantic extraction processing on the data to be transmitted by the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the data to be transmitted, the audio data can be preprocessed to obtain preprocessed audio data; and the preprocessed audio data is subjected to semantic extraction processing by the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the audio data.

[0094] In the embodiments of the present application, the data preprocessing described above can be a processing process of noise reduction, normalization and the like on the audio data. The purpose of the data preprocessing described above is to improve the data quality.

[0095] The trained multi-modal feature extraction model described above can be a multi-modal feature extraction model constructed based on deep learning or machine learning, such as CLIP, RNN, etc. The RNN described above can mine the time sequence information and context dependency in the data through a loop structure.

[0096] The trained multi-modal feature extraction model described above is obtained by training a pre-trained multi-modal feature extraction model using a training data set. The training data set includes sample audio data and semantic annotation data corresponding to the sample audio data. The pre-trained multi-modal feature extraction model can be a multi-modal feature extraction model constructed based on deep learning or machine learning, such as CLIP, RNN, etc. The training can be supervised training. Supervised training is to train a model using a set of data with known labels. By optimizing the model parameters, the model can predict the labels of new data or make decisions based on the characteristics of existing data. In the training process, the parameters of the model can be adjusted using a minimum loss function to minimize the difference between the output labels of the model and the input data.

[0097] The semantic extraction process described above can be understood as a process of analyzing the preprocessed audio data by the trained multi-modal feature extraction model and extracting the semantic features in the preprocessed audio data.

[0098] The target semantic vector described above can be obtained by performing semantic extraction processing on the preprocessed audio data by the trained multi-modal feature extraction model to obtain the target semantic vector corresponding to the audio data.

[0099] Optionally, before the step of performing semantic extraction on the to-be-transmitted data by the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data, a text training data set and a pre-trained multi-modal feature extraction model can also be obtained; the pre-trained multi-modal feature extraction model is trained through the text training data set, and in the training process, the pre-trained multi-modal feature extraction model is adjusted in parameters through a minimum loss function, and after the training is completed, a trained first multi-modal feature extraction model is obtained; or, an audio training data set and a pre-trained multi-modal feature extraction model are obtained; the pre-trained multi-modal feature extraction model is trained through the audio training data set, and in the training process, the pre-trained multi-modal feature extraction model is adjusted in parameters through a minimum loss function, and after the training is completed, a trained second multi-modal feature extraction model is obtained; or, a screen frame sequence training data set and a pre-trained multi-modal feature extraction model are obtained; the pre-trained multi-modal feature extraction model is trained through the screen frame sequence training data set, and in the training process, the pre-trained multi-modal feature extraction model is adjusted in parameters through a minimum loss function, and after the training is completed, a trained third multi-modal feature extraction model is obtained; or, a text training data set, an audio training data set, a screen frame sequence training data set and a pre-trained multi-modal feature extraction model are obtained; the pre-trained multi-modal feature extraction model is trained through the text training data set, the audio training data set and the screen frame sequence training data set, and in the training process, the pre-trained multi-modal feature extraction model is adjusted in parameters through a minimum loss function, and after the training is completed, a trained fourth multi-modal feature extraction model is obtained.

[0100] In the embodiment of the application, the text training data set includes sample text data and semantic annotation data corresponding to the sample text data. The annotation data is a process of converting original data into a form understandable by a machine learning model, enabling the machine to learn and perform classification, detection and other tasks by adding semantic labels or structured information to the data.

[0101] The audio training data set includes sample audio data and semantic annotation data corresponding to the sample audio data.

[0102] The screen frame sequence data training set includes sample screen frame sequences, image semantic annotation data of image data corresponding to the screen frame sequences, and text semantic annotation data of the image data.

[0103] The pre-trained multi-modal feature extraction model can be a multi-modal feature extraction model constructed based on deep learning or machine learning, such as CLIP, LLM, etc. The above-mentioned CLIP (Contrastive Language-Image Pretraining) is a multi-modal model that realizes cross-modal semantic alignment of images and texts through contrastive learning. The core idea of CLIP is to embed images and texts into the same latent space, so that matching image-text pairs are closer in the space, and non-matching pairs are farther apart. The core mechanism of the above-mentioned LLM (Large Language Model) is based on the Transformer architecture, which realizes language processing by analyzing long-distance dependencies in text.

[0104] The training can be supervised training, which is to train the model using a set of data with known labels, and to optimize the model parameters so that the model can predict the labels of new data or make decisions based on the characteristics of existing data. During the training process, the parameters of the model can be adjusted using a minimum loss function to minimize the difference between the output label of the model and the input data. The loss function is used to measure the difference between the model's prediction and the true result, and the purpose is to minimize the loss function value by adjusting the model parameters, thereby improving the prediction accuracy. The loss function can be a mean square error loss function, a cross-entropy loss function, etc. The parameter adjustment refers to the process of adjusting the weights and biases of the model to optimize the performance of the model. During the training process, the parameters of the model are optimized by labeled data to make the model have better prediction or decision-making ability.

[0105] Specifically, during the training process, the minimum loss function can be used as the optimization objective, and the model parameters can be adjusted by the backpropagation algorithm. The adjustment process of the model parameters is iterated until the error loss is less than a preset value, or the number of iterations reaches a preset number, the training process ends, and a trained behavior prediction model is obtained. The backpropagation algorithm is a supervised learning algorithm that updates weights by calculating the gradient of the loss function to minimize the error between the predicted output and the true value.

[0106] The trained first multi-modal feature extraction model is obtained by training the pre-trained multi-modal feature extraction model using a text training data set. The trained first multi-modal feature extraction model can identify the text semantic vector of the text data.

[0107] The trained second multi-modal feature extraction model is obtained by training the pre-trained multi-modal feature extraction model using an audio training data set. The trained second multi-modal feature extraction model can identify the semantic vector of the audio data.

[0108] The trained third multi-modal feature extraction model is obtained by training a pre-trained multi-modal feature extraction model based on the screen frame sequence training data set.

[0109] The trained fourth multi-modal feature extraction model is obtained by training a pre-trained multi-modal feature extraction model based on the text training data set, the audio training data set, and the screen frame sequence training data set. The trained fourth multi-modal feature extraction model can identify the text semantic vector of the text data, the semantic vector of the audio data, and the semantic vector of the screen frame sequence.

[0110] Optionally, in the step of performing encryption and encoding processing on the to-be-transmitted data based on the target sensitive data to generate target transmission data, sensitive level identification processing can be performed on the target sensitive data to obtain a target sensitive level corresponding to the target sensitive data; and the to-be-transmitted data is processed based on the target sensitive level to generate the target transmission data.

[0111] In the embodiments of the present application, the target sensitive data can be sensitive data in the to-be-transmitted data.

[0112] The sensitive level identification processing can be understood as a processing process of identifying the sensitive level of the target sensitive data.

[0113] The encryption and encoding processing can be understood as a processing process of converting the target sensitive data of the to-be-transmitted data into a format that cannot be directly identified according to the target sensitive level and a specific rule, so as to protect the data from unauthorized access.

[0114] The target transmission data can be transmission data after encryption and encoding processing.

[0115] It should be noted that different encryption algorithms can be used to encrypt and encode sensitive data of different sensitive levels.

[0116] Optionally, in the step of performing encryption and encoding processing on the to-be-transmitted data based on the target sensitive level to generate target transmission data, the target sensitive level can be determined in the corresponding relationship between the sensitive level and the encryption algorithm to determine the corresponding target encryption algorithm based on the corresponding relationship between the sensitive level and the encryption algorithm, different sensitive levels correspond to different encryption algorithms; and the to-be-transmitted data is processed based on the target encryption algorithm to obtain the target transmission data.

[0117] In the embodiments of the present application, the correspondence between the above-mentioned sensitivity levels and encryption algorithms can be a correspondence set in advance by the system, different sensitivity levels corresponding to different encryption algorithms, such as, for example, a high-level sensitivity level corresponding to an AES algorithm; a middle-level sensitivity level corresponding to a DES algorithm; a low-level sensitivity level corresponding to an XOR encryption algorithm, and the like. The above-mentioned AES algorithm (Advanced Encryption Standard) uses a substitution-permutation network (SPN) structure to implement encryption, and the decryption process is the inverse operation of encryption. The above-mentioned DES algorithm (Data Encryption Standard) uses a Feistel network structure, and implements data encryption through 16 rounds of iteration encryption, wherein, in each round, 64-bit plaintext is divided into left and right parts, the left data is encrypted using a sub-key, and then the next round of encryption is performed after the right data is exchanged, and finally the ciphertext is generated. It should be noted that the key length is 64 bits, but only 56 bits are actually involved in the operation, and the remaining bits are check bits. The above-mentioned XOR encryption algorithm is a basic encryption algorithm that is implemented by comparing binary codes bit by bit with a key.

[0118] The above-mentioned target encryption algorithm can be understood as an encryption algorithm corresponding to the target sensitivity level.

[0119] The above-mentioned encryption and encoding processing can be understood as a processing process of encrypting and encoding the target sensitive data in the to-be-transmitted data by using the target encryption algorithm.

[0120] The above-mentioned target transmission data can be transmission data after encryption and encoding processing.

[0121] In a possible embodiment, for example, in the correspondence between the sensitivity levels and the encryption algorithms, a high-level sensitivity level corresponds to an AES algorithm, a middle-level sensitivity level corresponds to a DES algorithm, and a low-level sensitivity level corresponds to an XOR encryption algorithm. When the target sensitivity level is identified as a high-level sensitivity level, the target encryption algorithm corresponding to the target sensitivity level can be determined as the AES algorithm according to the correspondence between the sensitivity levels and the encryption algorithms, and the to-be-transmitted data can be encrypted and encoded by using the AES algorithm to obtain the target transmission data; when the target sensitivity level is identified as a middle-level sensitivity level, the target encryption algorithm corresponding to the target sensitivity level can be determined as the DES algorithm according to the correspondence between the sensitivity levels and the encryption algorithms, and the to-be-transmitted data can be encrypted and encoded by using the DES algorithm to obtain the target transmission data, and the like.

[0122] It should be noted that the target sensitive level can be matched in the correspondence between the sensitive levels and the encryption algorithms according to the correspondence between the sensitive levels and the encryption algorithms, the target encryption algorithm corresponding to the target sensitive level is determined, and the target transmission data is obtained by performing encryption coding processing on the to-be-transmitted data through the target encryption algorithm. The application can use different strength encryption algorithms according to the sensitivity of the data, and improve the security of the data.

[0123] As shown in Figure 2 The cloud computer data transmission device provided by the embodiment of the application comprises:

[0124] The acquisition module 201 is configured to acquire to-be-transmitted data generated in a cloud computer session process.

[0125] The first processing module 202 is configured to perform semantic extraction processing on the to-be-transmitted data through a trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data.

[0126] The identification module 203 is configured to identify target sensitive data in the to-be-transmitted data based on the target semantic vector.

[0127] The second processing module 204 is configured to perform encryption coding processing on the to-be-transmitted data based on the target sensitive data to generate target transmission data.

[0128] The transmission module 205 is configured to transmit the target transmission data to the cloud based on the target transmission data.

[0129] Optionally, the first processing module 202 is further configured to perform first semantic extraction on the first text data through the trained multi-modal feature extraction model to obtain a first semantic vector corresponding to the first text data.

[0130] Optionally, the first processing module 202 is further configured to perform second semantic extraction processing on the screen frame sequence through the trained multi-modal feature extraction model to obtain a second semantic vector corresponding to the screen frame sequence, perform third semantic extraction processing on the second text data to obtain a third semantic vector corresponding to the second text data, and obtain a target semantic vector corresponding to the screen frame sequence based on the second semantic vector and the third semantic vector.

[0131] Optionally, the first processing module 202 is further configured to perform data preprocessing on the audio data to obtain preprocessed audio data, and perform semantic extraction processing on the preprocessed audio data through the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the audio data.

[0132] Optionally, the device is further configured to obtain a text training data set and a pre-trained multi-modal feature extraction model, the text training data set comprising sample text data and semantic annotation data corresponding to the sample text data; train the pre-trained multi-modal feature extraction model using the text training data set, wherein, during the training, the pre-trained multi-modal feature extraction model is adjusted in parameters using a minimum loss function, and, after the training, a first trained multi-modal feature extraction model is obtained; or, obtain an audio training data set and a pre-trained multi-modal feature extraction model, the audio training data set comprising sample audio data and semantic annotation data corresponding to the sample audio data; train the pre-trained multi-modal feature extraction model using the audio training data set, wherein, during the training, the pre-trained multi-modal feature extraction model is adjusted in parameters using a minimum loss function, and, after the training, a second trained multi-modal feature extraction model is obtained; or, obtain a screen frame sequence training data set and a pre-trained multi-modal feature extraction model, the screen frame sequence training data set comprising sample screen frame sequences, image semantic annotation data of image data corresponding to the screen frame sequences, and text semantic annotation data of the image data; train the pre-trained multi-modal feature extraction model using the screen frame sequence training data set, wherein, during the training, the pre-trained multi-modal feature extraction model is adjusted in parameters using a minimum loss function, and, after the training, a third trained multi-modal feature extraction model is obtained; or, obtain a text training data set, an audio training data set, a screen frame sequence training data set, and a pre-trained multi-modal feature extraction model, the text training data set comprising sample text data and semantic annotation data corresponding to the sample text data, the audio training data set comprising sample audio data and semantic annotation data corresponding to the sample audio data, and the screen frame sequence training data set comprising sample screen frame sequences, image semantic annotation data of image data corresponding to the screen frame sequences, and text semantic annotation data of the image data; train the pre-trained multi-modal feature extraction model using the text training data set, the audio training data set, and the screen frame sequence training data set, wherein, during the training, the pre-trained multi-modal feature extraction model is adjusted in parameters using a minimum loss function, and, after the training, a fourth trained multi-modal feature extraction model is obtained.

[0133] Optionally, the second processing module 204 is further configured to perform sensitive level identification processing on the target sensitive data to obtain a target sensitive level corresponding to the target sensitive data; and perform encryption and coding processing on the to-be-transmitted data based on the target sensitive level to generate target transmission data.

[0134] Optionally, the second processing module 204 is further configured to determine a target encryption algorithm corresponding to the target sensitive level in the correspondence between sensitive levels and encryption algorithms based on the correspondence between sensitive levels and encryption algorithms, different sensitive levels corresponding to different encryption algorithms; and perform encryption coding processing on the to-be-transmitted data based on the target encryption algorithm to obtain target transmission data.

[0135] As shown in Figure 3 the embodiment of the present application further provides an electronic device, comprising a processor, and the processor can execute the data transmission method of the cloud computer.

[0136] Specifically, the electronic device comprises a processor 301 and a memory 302, and a computer program for executing the data transmission method of the cloud computer stored in the memory 302 and capable of running on the processor 301, wherein:

[0137] The processor 301 runs the computer program of the data transmission method of the cloud computer stored in the memory 302 to execute the following steps:

[0138] obtaining to-be-transmitted data generated in a cloud computer session process;

[0139] performing semantic extraction processing on the to-be-transmitted data by using the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data;

[0140] identifying target sensitive data in the to-be-transmitted data based on the target semantic vector;

[0141] performing encryption coding processing on the to-be-transmitted data based on the target sensitive data to generate target transmission data;

[0142] transmitting the target transmission data to the cloud based on the target transmission data.

[0143] Optionally, the to-be-transmitted data comprises first text data, and the processor 301 performs the semantic extraction processing on the to-be-transmitted data by using the trained multi-modal feature extraction model to obtain the target semantic vector corresponding to the to-be-transmitted data, comprising:

[0144] performing first semantic extraction on the first text data by using the trained multi-modal feature extraction model to obtain a first semantic vector corresponding to the first text data.

[0145] Optionally, the to-be-transmitted data comprises a screen frame sequence and second text data corresponding to the screen frame sequence, and the processor 301 performs the semantic extraction processing on the to-be-transmitted data by using the trained multi-modal feature extraction model to obtain the target semantic vector corresponding to the to-be-transmitted data, comprising:

[0146] performing second semantic extraction processing on the screen frame sequence by using the trained multi-modal feature extraction model to obtain a second semantic vector corresponding to the screen frame sequence;

[0147] performing third semantic extraction processing on the second text data to obtain a third semantic vector corresponding to the second text data;

[0148] obtaining a target semantic vector corresponding to the screen frame sequence based on the second semantic vector and the third semantic vector.

[0149] Optionally, the to-be-transmitted data includes audio data, and the performing, by the processor 301, of the semantic extraction processing on the to-be-transmitted data by using the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data includes:

[0150] performing data preprocessing on the audio data to obtain preprocessed audio data;

[0151] performing semantic extraction processing on the preprocessed audio data by using the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the audio data.

[0152] Optionally, before the performing, by the processor 301, of the semantic extraction processing on the to-be-transmitted data by using the trained multi-modal feature extraction model to obtain a target semantic vector corresponding to the to-be-transmitted data, the method further includes:

[0153] obtaining a text training data set and a pre-trained multi-modal feature extraction model, the text training data set including sample text data and semantic annotation data corresponding to the sample text data;

[0154] training the pre-trained multi-modal feature extraction model by using the text training data set, in the training process, adjusting parameters of the pre-trained multi-modal feature extraction model by using a minimum loss function, and obtaining a trained first multi-modal feature extraction model after the training is completed;

[0155] or, obtaining an audio training data set and a pre-trained multi-modal feature extraction model, the audio training data set including sample audio data and semantic annotation data corresponding to the sample audio data;

[0156] training the pre-trained multi-modal feature extraction model by using the audio training data set, in the training process, adjusting parameters of the pre-trained multi-modal feature extraction model by using a minimum loss function, and obtaining a trained second multi-modal feature extraction model after the training is completed;

[0157] Or, obtain a screen frame sequence training data set and a pre-trained multi-modal feature extraction model, the screen frame sequence data training set includes a sample screen frame sequence, and image semantic annotation data and text semantic annotation data of image data corresponding to the screen frame sequence;

[0158] Train the pre-trained multi-modal feature extraction model through the screen frame sequence training data set, in the training process, parameters of the pre-trained multi-modal feature extraction model are adjusted through a minimum loss function, and after the training is completed, a trained third multi-modal feature extraction model is obtained.

[0159] Or, obtain a text training data set, an audio training data set, a screen frame sequence training data set and a pre-trained multi-modal feature extraction model, the text training data set includes sample text data and semantic annotation data corresponding to the sample text data, the audio training data set includes sample audio data and semantic annotation data corresponding to the sample audio data, and the screen frame sequence data training set includes a sample screen frame sequence and image semantic annotation data and text semantic annotation data of image data corresponding to the screen frame sequence.

[0160] Train the pre-trained multi-modal feature extraction model through the text training data set, the audio training data set and the screen frame sequence training data set, in the training process, parameters of the pre-trained multi-modal feature extraction model are adjusted through a minimum loss function, and after the training is completed, a trained fourth multi-modal feature extraction model is obtained.

[0161] Optionally, the processor 301 performs the encryption coding processing on the to-be-transmitted data based on the target sensitive data to generate target transmission data, including:

[0162] Sensitive level identification processing is performed on the target sensitive data to obtain a target sensitive level corresponding to the target sensitive data.

[0163] Based on the target sensitive level, the encryption coding processing is performed on the to-be-transmitted data to generate target transmission data.

[0164] Optionally, the processor 301 performs the encryption coding processing on the to-be-transmitted data based on the target sensitive level to generate target transmission data, including:

[0165] Based on the correspondence between the sensitive level and the encryption algorithm, the target sensitive level is determined in the correspondence between the sensitive level and the encryption algorithm to determine a corresponding target encryption algorithm, and different sensitive levels correspond to different encryption algorithms.

[0166] Based on the target encryption algorithm, the to-be-transmitted data is encrypted and coded to obtain target transmission data.

[0167] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program.

[0168] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiment methods can be included. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM) and the like.

[0169] The above only describes the preferred embodiments of the present application, and of course cannot limit the scope of the present application, so equivalent changes made according to the claims of the present application are still within the scope of the present application.

Claims

1. A data transmission method for a cloud computer, characterized in that, The method includes: Retrieve data to be transmitted generated during a cloud computer session; The target semantic vector corresponding to the data to be transmitted is obtained by performing semantic extraction processing on the data to be transmitted through a trained multimodal feature extraction model. Based on the target semantic vector, the target sensitive data in the data to be transmitted is identified; Based on the target sensitive data, the data to be transmitted is encrypted and encoded to generate the target transmission data; Based on the target transmission data, the target transmission data is transmitted to the cloud.

2. The method according to claim 1, characterized in that, The data to be transmitted includes first text data. The step of performing semantic extraction processing on the data to be transmitted using a trained multimodal feature extraction model to obtain the target semantic vector corresponding to the data to be transmitted includes: The first semantic vector corresponding to the first text data is obtained by performing first semantic extraction on the first text data using a trained multimodal feature extraction model.

3. The method according to claim 1, characterized in that, The data to be transmitted includes a screen frame sequence and corresponding second text data. The step of performing semantic extraction processing on the data to be transmitted using a trained multimodal feature extraction model to obtain a target semantic vector corresponding to the data to be transmitted includes: The screen frame sequence is processed by a trained multimodal feature extraction model to extract the second semantics, thereby obtaining the second semantic vector corresponding to the screen frame sequence. The second text data is subjected to third semantic extraction processing to obtain the third semantic vector corresponding to the second text data; Based on the second semantic vector and the third semantic vector, the target semantic vector corresponding to the screen frame sequence is obtained.

4. The method according to claim 1, characterized in that, The data to be transmitted includes audio data. The step of performing semantic extraction processing on the data to be transmitted using a trained multimodal feature extraction model to obtain the target semantic vector corresponding to the data to be transmitted includes: The audio data is preprocessed to obtain preprocessed audio data; The preprocessed audio data is semantically extracted using a trained multimodal feature extraction model to obtain the target semantic vector corresponding to the audio data.

5. The method according to claim 4, characterized in that, Before performing semantic extraction processing on the data to be transmitted using a trained multimodal feature extraction model to obtain the target semantic vector corresponding to the data to be transmitted, the method further includes: Obtain a text training dataset and a pre-trained multimodal feature extraction model. The text training dataset includes sample text data and semantic annotation data corresponding to the sample text data. The pre-trained multimodal feature extraction model is trained using the text training dataset. During the training process, the parameters of the pre-trained multimodal feature extraction model are adjusted using the minimum loss function. Once the training is complete, the trained first multimodal feature extraction model is obtained. Alternatively, obtain an audio training dataset and a pre-trained multimodal feature extraction model, wherein the audio training dataset includes sample audio data and semantic annotation data corresponding to the sample audio data; The pre-trained multimodal feature extraction model is trained using the audio training dataset. During the training process, the parameters of the pre-trained multimodal feature extraction model are adjusted using the minimum loss function. Once training is complete, a trained second multimodal feature extraction model is obtained. Alternatively, obtain a screen frame sequence training dataset and a pre-trained multimodal feature extraction model. The screen frame sequence training dataset includes sample screen frame sequences, image semantic annotation data of the image data corresponding to the screen frame sequences, and text semantic annotation data of the image data. The pre-trained multimodal feature extraction model is trained using the screen frame sequence training dataset. During the training process, the parameters of the pre-trained multimodal feature extraction model are adjusted using the minimum loss function. Once training is complete, a trained third multimodal feature extraction model is obtained. Alternatively, obtain a text training dataset, an audio training dataset, a screen frame sequence training dataset, and a pre-trained multimodal feature extraction model. The text training dataset includes sample text data and semantic annotation data corresponding to the sample text data. The audio training dataset includes sample audio data and semantic annotation data corresponding to the sample audio data. The screen frame sequence training dataset includes sample screen frame sequences, image semantic annotation data of the image data corresponding to the screen frame sequences, and text semantic annotation data of the image data. The pre-trained multimodal feature extraction model is trained using the text training dataset, the audio training dataset, and the screen frame sequence training dataset. During training, the adoption number of the pre-trained multimodal feature extraction model is adjusted using the minimum loss function. After training is completed, a trained fourth multimodal feature extraction model is obtained.

6. The method according to claim 1, characterized in that, The step of encrypting and encoding the data to be transmitted based on the target sensitive data to generate target transmission data includes: The target sensitive data is subjected to sensitivity level identification processing to obtain the target sensitivity level corresponding to the target sensitive data; Based on the target sensitivity level, the data to be transmitted is encrypted and encoded to generate the target transmission data.

7. The method according to claim 6, characterized in that, The step of encrypting and encoding the data to be transmitted based on the target sensitivity level to generate target transmission data includes: Based on the correspondence between sensitivity levels and encryption algorithms, the target sensitivity level is determined to correspond to a target encryption algorithm in the correspondence between sensitivity levels and encryption algorithms, and different sensitivity levels correspond to different encryption algorithms; Based on the target encryption algorithm, the data to be transmitted is encrypted and encoded to obtain the target transmission data.

8. A data transmission device for a cloud computer, characterized in that, The data transmission device of the cloud computer includes: The acquisition module is used to acquire the data to be transmitted generated during the cloud computer session; The first processing module is used to perform semantic extraction processing on the data to be transmitted using a trained multimodal feature extraction model to obtain the target semantic vector corresponding to the data to be transmitted. The identification module is used to identify target sensitive data in the data to be transmitted based on the target semantic vector; The second processing module is used to encrypt and encode the data to be transmitted based on the target sensitive data to generate target transmission data; The transmission module is used to transmit the target transmission data to the cloud based on the target transmission data.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the data transmission method of the cloud computer as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the data transmission method of the cloud computer as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Http encryption transmission method and device, computer equipment and storage medium

    CN112511514A

  • Data asset safety monitoring system and method based on cloud computing

    CN118278048A

  • Data transmission method and system of cloud computer

    CN118427859A

  • Multi-cloud data processing control method and system

    CN118611948A

  • Enterprise-level intelligent customer service interaction control system under multi-round dialogue scene

    CN120372680A