Computer transmission data security protection method and system
By classifying and analyzing the importance of user historical data, using a hierarchical encryption strategy, selecting the number of key bits of the RSA algorithm based on the category and user coefficients, the problem of waste of resources and poor user experience caused by the RSA algorithm using the same key for all data is solved, and efficient and secure data transmission is achieved.
Patent Information
- Application Number
- CN202510592061.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the prior art, the RSA algorithm encrypts all data using the same long key, resulting in a long encryption and decryption time, wasting network resources and reducing user experience.
By analyzing the user's historical transmission data, classifying users and obtaining category coefficients and user coefficients, calculating the encryption degree of the data to be encrypted based on the weighting, encrypting using the RSA algorithm, and selecting the appropriate number of key bits.
Effective utilization of network resources and improved user experience. Through a hierarchical encryption strategy, the appropriate number of key bits is selected according to the importance of data for encryption, improving data transmission efficiency and security.
Smart Images

Figure CN120145424B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a computer transmission data security protection method and system. Background Art
[0002] In today's internet age, social media sites attract a large number of users. To protect the privacy and security of user data during transmission over the internet, these sites generally use the SSL / TLS (Secure Sockets Layer / Transport Layer Security) protocol. This protocol combines multiple encryption algorithms to establish a secure encrypted channel between the client and server, ensuring the confidentiality and integrity of data during transmission and preventing it from being intercepted or tampered with by unauthorized third parties. The RSA (Rivest-Shamir-Adleman) algorithm is a typical asymmetric encryption algorithm whose security and reliability are based on the difficulty of factoring extremely large integers. In the RSA algorithm, the public key and private key are a pair of related functions generated by large prime numbers, typically 100 to 200 decimal numbers or even larger. The difficulty of recovering the original plaintext from the public key and ciphertext is comparable to factoring the product of two large prime numbers, making cracking RSA encryption extremely difficult.
[0003] Since the complexity of encryption and decryption directly affects the efficiency of data transmission, using larger prime numbers for encryption improves security, but also leads to a decrease in data transmission speed. Moreover, with the continuous growth of the number of social media users and the increasing amount of invalid content, if the same high-difficulty encryption algorithm is used for all data, it will not only waste network resources but also greatly reduce the user experience. That is, using the same long key to encrypt all data will waste resources and affect transmission efficiency. Summary of the Invention
[0004] In order to solve the technical problem that when the existing technology uses the RSA algorithm to encrypt user data, the same long key is used to encrypt all data, resulting in long encryption and decryption time, waste of network resources and poor user experience, the purpose of the present invention is to provide a computer transmission data security protection method, the technical solution adopted is as follows:
[0005] Obtain historical transmission data of several users from any social networking site;
[0006] Process historical transmission data to classify users into categories and obtain the category coefficient for each category; analyze the corresponding features of historical transmission data, calculate the data importance of historical transmission data, and determine the user coefficients of all users;
[0007] The transmission data to be encrypted is weighted according to the category coefficient and the user coefficient to obtain the encryption degree of the transmission data to be encrypted;
[0008] Based on the encryption level, the number of key bits corresponding to the transmitted data to be encrypted is obtained and encrypted using the RSA algorithm to achieve complete protection.
[0009] Preferably, the transmission data includes first text data, second text data and third text data.
[0010] Preferably, processing historical transmission data to classify users into categories and obtaining category coefficients for each category includes:
[0011] Determining the data volume of the first text data and the second text data respectively, and calculating the activity of all users;
[0012] Analyze the similarity between any two users, cluster all users based on similarity to obtain several cluster categories, and obtain the average activity of each cluster category according to the activity of each user in the cluster category, which is recorded as the category coefficient.
[0013] Preferably, the activity of all users is calculated, and the corresponding calculation formula is:
[0014]
[0015] in, Indicates the activity level of any user; Indicates the data amount of the first text data; indicating the data amount of the second text data; represents the average number of comments in the first text data;
[0016] Analyze the similarity between any two users, and the corresponding calculation formula is:
[0017]
[0018] in, Indicates the users and Similarity of users; Indicates the The activity level of each user; Indicates the The activity of a user.
[0019] Preferably, analyzing the corresponding features of the historical transmission data, calculating the data importance of the historical transmission data, and determining the user coefficients of all users include:
[0020] Segmenting the first text data to obtain a plurality of words of the first text data, and determining a word entry from the first text data;
[0021] Matching the vocabulary with a professional vocabulary dictionary to obtain professional vocabulary and non-professional vocabulary respectively, and calculating the depth and breadth of any first text data in turn to determine the importance of the first text data;
[0022] matching the second text data and the third text data to obtain related words, and calculating the importance of the second text data;
[0023] Obtaining the importance of all first text data and the importance of all second text data, respectively obtaining a first text data importance sequence and a second text data importance sequence, performing analysis to respectively determine a frequency of change in the importance of the first text data and a frequency of change in the importance of the second text data, and calculating a user coefficient of the user;
[0024] determining a corresponding depth level according to a product between the validity of the first text data and the content depth;
[0025] The calculation formula of the effectiveness is:
[0026]
[0027] in, Indicates the vocabulary category of the first text data; The number of words representing the first text data; Indicates the first text data The first word is words, that is, the proportion of adjacent words appearing in the first text data; Indicates the first text data The word after words, that is, the proportion of adjacent words appearing in the first text data; represents the normalization function;
[0028] The calculation formula of the content depth is:
[0029]
[0030] in, represents the number of professional words in the first text data; Indicates the number of words in the professional vocabulary dictionary.
[0031] Preferably, the process of obtaining the breadth includes:
[0032] Determining a corresponding breadth degree according to a product of term popularity and discussion degree of the first text data;
[0033] The calculation formula for the term popularity is:
[0034]
[0035] in, Indicates the popularity of a term in the first text data; The number of words representing all entries in the first text data; Represents the first word of all the words in the first text data The number of times a word appears in all the first text data; represents the number of terms in the first text data; Indicates the first text data The number of words in each entry; Indicates the first text data The first The number of times a word appears in all entries;
[0036] The calculation formula of the discussion degree is:
[0037]
[0038] in, Indicates the discussion degree of the first text data; Indicates the first text data among all the entries The number of times a word appears in the first text data.
[0039] Preferably, the process of calculating the importance of the second text data includes:
[0040]
[0041] determining the importance of the second text data based on a product of the content relevance of the second text data and the comment repetition;
[0042] The calculation formula for content relevance includes:
[0043]
[0044] in, is the content relevance of the second text data; The number of words representing the second text data; Indicates the first The number of related words for each word;
[0045]
[0046] in, The repetitiveness of comments for the second text data; Indicates the first The number of times a word appears in all second text data of the user currently being analyzed; Represents an exponential function with a natural constant as the base. Preferably, the user coefficient of the user is calculated, and the corresponding calculation formula is:
[0047]
[0048] in, represents the user coefficient of any user; represents the average importance of all first text data of the user currently being analyzed; Indicates the frequency of changes in the importance of the first text data; represents the average importance of all second text data of the currently analyzed user; Indicates the frequency of changes in the importance of the second text data; Represents the normalization function.
[0049] Preferably, weighting the transmission data to be encrypted according to the category coefficient and the user coefficient to obtain the encryption degree of the transmission data to be encrypted includes:
[0050] Define the first text data as , the second text data is , the corresponding calculation formula is:
[0051]
[0052] in, Indicates the encryption level of the transmitted data to be encrypted; Indicates the importance of the first text data to be encrypted; Indicates the importance of the second text data to be encrypted; Indicates the user coefficient of the user corresponding to the transmission data to be encrypted; Indicates the category coefficient of the user corresponding to the transmission data to be encrypted.
[0053] To solve the above problems, the present application also provides a computer transmission data security protection system, which stores program data, and when the program data is executed, implements a computer transmission data security protection method as described in any of the above items.
[0054] The present invention has the following beneficial effects:
[0055] 1. Analyze the user's historical transmission data, classify the users, and obtain the category coefficient of each classified user; analyze the historical transmission data, determine the importance of each historical transmission data, and obtain the user coefficient of each user; comprehensively weight the category coefficient and the user coefficient to determine the encryption degree of the transmission data to be encrypted, and obtain the corresponding key bit number, so that more important transmission data to be encrypted uses a larger key, that is, analyze the historical transmission data, classify the importance of the first text data and the second text data, and use different encryption processing according to different classifications and different key bit numbers, which is beneficial to the effective use of network resources and the user's website usage experience.
[0056] 2. The computer transmission data security protection system provided by the present invention has the same beneficial effects as the computer transmission data security protection method provided by the present invention, and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0058] Figure 1 A flowchart of a method for protecting computer data transmission security provided by one embodiment of the present invention;
[0059] Figure 2 A flowchart of an RSA encryption algorithm for a computer transmission data security protection method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0060] To further illustrate the technical means and effectiveness of the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail a computer transmission data security protection method and system proposed in accordance with the present invention, including its specific implementation, structure, features, and effectiveness. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0061] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0062] The following describes in detail a computer transmission data security protection method and system provided by the present invention with reference to the accompanying drawings.
[0063] Since the existing RSA algorithm uses the same high-difficulty encryption algorithm for all data, it is not only a waste of network resources but also greatly reduces the user experience. That is, the same long key is used to encrypt all data, which wastes resources and affects transmission efficiency. In one embodiment of the present invention, a computer transmission data security protection method is provided. By analyzing the user's historical transmission data, the user is classified and the category coefficient of each classified user is obtained; the historical transmission data is analyzed to determine the importance of each historical transmission data to obtain the user coefficient of each user; the comprehensive category coefficient and user coefficient are weighted to determine the encryption degree of the transmission data to be encrypted, and the corresponding key bit number is obtained for encryption. In order to realize a computer transmission data security protection method, a computer transmission data security protection system is provided. The system is essentially a software system, which is composed of modules that realize corresponding functions. The specific steps in the method are now introduced in detail.
[0064] See also Figure 1 , which shows a flowchart of a method for protecting computer transmission data security provided by one embodiment of the present invention, the method comprising:
[0065] Step S1: Obtain historical transmission data of several users from any social networking site;
[0066] Step S2: Process the historical transmission data to classify users into categories and obtain the category coefficient of each category; analyze the corresponding features of the historical transmission data, calculate the data importance of the historical transmission data, and determine the user coefficients of all users;
[0067] Step S3: Weighting the transmission data to be encrypted according to the category coefficient and the user coefficient to obtain the encryption degree of the transmission data to be encrypted;
[0068] Step S4: Based on the encryption level, the number of key bits corresponding to the transmission data to be encrypted is obtained and encrypted using the RSA algorithm to achieve complete protection.
[0069] For better explanation, this application is optimized based on the RSA algorithm, where the RSA algorithm is a public key encryption algorithm based on a simple number theory fact: it is easy to multiply two large prime numbers, but it is very difficult to decompose the product of the two back to the original prime numbers. The core lies in the key generation process, which involves two keys: a public key and a private key; the public key is used to encrypt data, while the private key is used for decryption. When generating a key pair, two large prime numbers are first randomly selected, and the product is calculated as the modulus; then the Euler function is calculated, which is the number of positive integers less than the modulus that are coprime to the modulus, a small integer coprime to the Euler function is selected as the public key exponent, and the modular inverse of the public key exponent with respect to the Euler function is calculated as the private key exponent.
[0070] In this embodiment, for a social media website with a large number of users, as the user base continues to grow and invalid content increases, using the same high-difficulty encryption for all data is not only a waste of network resources but also greatly reduces the user experience. Therefore, to address this problem, the importance of the data is analyzed and a hierarchical encryption strategy is adopted. The user's comprehensive confidentiality level is determined, and different data is encrypted with corresponding key bits based on the importance of the content, thereby reducing the use of system resources while maintaining data transmission efficiency.
[0071] Specifically, historical transmission data of several clients of any social networking site within one month is collected through a server interface or service; the historical transmission data includes: client IP (Internet Protocol) address, user ID (Identifier), text data and creation time, wherein one ID may correspond to multiple IP addresses, that is, one user ID has logged in on different devices or locations; in this embodiment, several client IP addresses of the same user ID are obtained, and the text data under each client IP address is used as the historical transmission data of the current analyzed user, and is sorted by creation time to better understand the time series of the data and user behavior patterns.
[0072] Furthermore, in step S1 , the transmission data includes first text data, second text data and third text data.
[0073] It can be explained that the first text data refers to the content of posts posted by users on social networking sites; the second text data is the content of comments made by users under content posted by others; the third text data refers to the post described in the comment, that is, the comment content posted by the commentator belongs to the corresponding post content; for example, assuming that user A posts content, and users B, C, and D comment on the content of the post, which are recorded as B1, C1, and D1 respectively, at this time, the content of the post is the first text data, B1, C1, and D1 are the second text data, and the content of the post to which comment B1 belongs is the third text data; it can be understood that not all users will choose to post content, and a user's comment may also be associated with multiple different posts, so for the same user, the first text data and the third text data may be the same or different, depending on whether the user only comments on the content of the posts he or she posts.
[0074] In addition, all comments published under the same post content, that is, the second text data, belong to the same comment type of text data. At this time, the post content to which the second text data belongs is all third text data, indicating that the second text data and the third text data correspond one-to-one.
[0075] Understandably, different users have different usage habits on social networking sites. For example, some users may simply browse information without leaving any comments or posting content; some users may frequently participate in comments but rarely or never post content; or some users post content that can stimulate a large amount of interaction and feedback. Therefore, before transmitting data, users can be differentiated by the number and proportion of comments and posts. Those user groups that are more active and whose content can generate more discussion are assigned a higher category coefficient. The category coefficient reflects the user's behavioral pattern and participation in the social networking site. Then, based on the historical transmission data on the social networking site, analysis is conducted. The data volume of historical transmission data can only reflect the user's activity, but the content of historical transmission data can reflect the validity and importance of the data. For example, between two users in the same category, the user who publishes more professional content and keeps up with current hot topics and trends should be given higher importance because such content not only provides value to other users but also promotes discussion and knowledge sharing within the social networking site, thereby improving the content quality and user engagement of the entire social networking site.
[0076] Furthermore, in step S2, historical transmission data is processed to classify users into categories, and category coefficients for each category are obtained, including:
[0077] Step S211: determining the data volume of the first text data and the second text data respectively, and calculating the activity of all users.
[0078] It is explained that for the content of a post published by any user or the content of a feedback comment, the data volume of the first text data and the second text data is determined to calculate the activity of any user, wherein activity refers to the frequency of user participation in interaction on a social networking site, that is, the user's activity is quantified by counting the number of posts and comments published by the user. Users with high activity indicate high participation and interest in the social networking site.
[0079] Furthermore, in step S211, the activity of all users is calculated, and the corresponding calculation formula is:
[0080]
[0081] in, Indicates the activity level of any user; Indicates the data amount of the first text data; indicating the data amount of the second text data; represents the average number of comments in the first text data.
[0082] Step S212: Analyze the similarity between any two users, cluster all users based on the similarity to obtain several cluster categories, and obtain the average activity of each cluster category according to the activity of each user in the cluster category, which is recorded as the category coefficient.
[0083] It is explained that the closer the activity levels of two users are, the higher the similarity between the two users is.
[0084] Analyze the similarity between any two users, and the corresponding calculation formula is:
[0085]
[0086] in, Indicates the users and Similarity of users; Indicates the The activity level of each user; Indicates the The activity of a user.
[0087] To illustrate, when the difference in activity between two users is smaller, that is, the activity of two users is similar, the closer the difference is to zero, the similarity The closer it is to 1, the higher the similarity between the two users.
[0088] Specifically, in this embodiment, all users are clustered using the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm with similarity as the distance. The DBSCAN algorithm is an algorithm for data clustering that can divide areas with sufficiently high density into clusters and can find cluster categories of arbitrary shapes in a spatial database with noise. The algorithm parameter Eps (Epsilon) takes a distance greater than 0.7, that is, when the similarity is greater than 0.7, the activity of two users is similar. MinPts (MinimumPoints) takes 5, that is, the minimum number of points required to form a dense area is specified to be 5, indicating that each cluster category contains at least 5 users, avoiding errors in data analysis. Several cluster categories are obtained, and each cluster category represents users of the same type. The average activity of all users in any cluster category is then recorded as the category coefficient of the cluster category. , that is, the higher the average activity of users, the higher the category coefficient, that is, the user's activity and category coefficient are positively correlated.
[0089] It can be understood that a social networking site is a platform for content production and exchange based on relationships between users. It not only provides a space for sharing various post contents, but also allows users to establish and maintain their own social networks to effectively disseminate information. For users who publish relevant professional post contents in specific vertical fields, such as technology, art or education, although the scope of dissemination may be relatively small, based on the professionalism and depth of the post content, it is still of considerable importance. In addition, when the disseminated content has a high timeliness and receives a lot of attention, it also indicates a high importance. Therefore, the importance of the first text data is calculated by analyzing the depth and breadth of the historical transmission data.
[0090] In terms of comment content, that is, the second text data is generally used more frequently than the post content. This means that although some users do not publish a large amount of post content, they often participate in comments. Therefore, the importance of the comment text is reflected by its relevance to the post to which it belongs. That is, the importance of the second text data is reflected by the third text data. Users who post highly relevant comments under multiple post contents are given greater importance.
[0091] Furthermore, in step S2, corresponding features of historical transmission data are analyzed, data importance of historical transmission data is calculated, and user coefficients of all users are determined, including:
[0092] Step S221: Segment the first text data to obtain a plurality of words of the first text data, and determine entries from the first text data.
[0093] Specifically, the first text data is segmented using a bidirectional maximum matching method to obtain several words of the first text data, wherein the bidirectional maximum matching method is an efficient algorithm for string matching, which improves matching efficiency by matching from both ends of the string at the same time and can better handle ambiguity, that is, when encountering ambiguity, it can perform more accurate word segmentation based on context information; vocabulary refers to the basic elements that constitute the first text data, such as Chinese characters, words or phrases.
[0094] Search the pound sign key from the first text data to determine the entry, and the number of characters in the entry is less than 50. For example, the first text data is: "#Earthquake News# The earthquake network officially determined: A 4.1-magnitude earthquake occurred in XXX, with a focal depth of 18 kilometers." Among them, the entry is Earthquake News.
[0095] Step S222: Match the vocabulary with the professional vocabulary dictionary to obtain professional vocabulary and non-professional vocabulary respectively, calculate the depth and breadth of any first text data in turn, and determine the importance of the first text data.
[0096] To clarify, a professional vocabulary dictionary refers to a professional academic dictionary, which contains a large number of terms and expressions in professional fields, has specific meanings and usages, and is widely accepted and used in specific fields; that is, dividing vocabulary into professional vocabulary and non-professional vocabulary can more accurately understand and analyze the information in the first text data.
[0097] Furthermore, in step S222, the depth and breadth of any first text data are calculated in sequence to determine the importance of the first text data, including:
[0098] It can be explained that the depth refers to the content depth and complexity of the vocabulary in the first text data, reflecting the depth of the first text data in the professional field; the breadth refers to the scope of fields covered by the first text data and the diversity of non-professional vocabulary, reflecting the breadth and diversity of the first text data.
[0099] determining a corresponding depth level according to a product between the validity of the first text data and the content depth;
[0100] The calculation formula of the effectiveness is:
[0101]
[0102] in, Indicates the vocabulary category of the first text data; The number of words representing the first text data; Indicates the first text data The first word is words, that is, the proportion of adjacent words appearing in the first text data; Indicates the first text data The word after words, that is, the proportion of adjacent words appearing in the first text data; represents the normalization function;
[0103] The calculation formula of the content depth is:
[0104]
[0105] in, represents the number of professional words in the first text data; Indicates the number of words in a professional vocabulary dictionary; makes a description, The validity of the first text data can be demonstrated, wherein and To reflect the proportion of the same consecutive words appearing together, that is Indicates the frequency with which a word is used together with the preceding and following words, reflecting the effectiveness of the word. If the text content is not repeated but the word is not used regularly, the effectiveness will also be low. It indicates the degree of non-repetition of words in the first text data. The closer it is to 1, the lower the repetition of the post content and the higher the validity. Therefore, when analyzing the depth of the first text data, the validity of the first text data is considered first. The higher the validity, the better the depth foundation of the first text data. It is analyzed from the appearance of professional vocabulary. Indicates the proportion of professional vocabulary in the first text data, It indicates the concentration of professional vocabulary in the professional vocabulary dictionary. The larger the proportion of professional vocabulary, the more concentrated it is, which indicates the depth of the first text data, that is, the deeper the content.
[0106] It should be noted that in order to ensure that the calculation results are meaningful, when performing fractional operations in the embodiments of the present invention, when encountering a situation where the denominator is 0, it is necessary to add a parameter adjustment factor greater than 0 to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer according to actual conditions, and this application does not impose any special restrictions.
[0107] Furthermore, the process of obtaining the breadth of the first text data includes:
[0108] Determining a corresponding breadth degree according to a product of term popularity and discussion degree of the first text data;
[0109] The calculation formula for the term popularity is:
[0110]
[0111] in, Indicates the popularity of a term in the first text data; The number of words representing all entries in the first text data; Represents the first word of all the words in the first text data The number of times a word appears in all the first text data; represents the number of terms in the first text data; Indicates the first text data The number of words in each entry; Indicates the first text data The first The number of times a word appears in all entries;
[0112] The calculation formula of the discussion degree is:
[0113]
[0114] in, Indicates the discussion degree of the first text data; Indicates the first text data among all the entries The number of times a word appears in the first text data. can reflect the popularity of the term in the first text data, wherein popularity refers to the popularity or popularity of the term in the first text data; The occurrence of all words in the entry in all the first text data is summarized. The more times it appears, the more widely the word is discussed, and the greater the breadth of the first text data; Indicates the relevance between terms, filters out the duplication of words between terms, and avoids the situation where the first text data contains multiple irrelevant terms; It indicates the relevance between the term and the first text data. If there is only the term but the post content is irrelevant to what the term reflects, the importance of the first text data will be reduced. The larger the value is, the more popular the first text data is. On this basis, the stronger the relevance between the term and the post content, the greater the breadth of the first text data.
[0115] The importance of the first text data is determined according to the depth and breadth, and the corresponding calculation formula is:
[0116]
[0117] in, Indicates the importance of any first text data; Represents the depth of the corresponding first text data; Represents The breadth of any corresponding first text data.
[0118] An explanation is made by comprehensively considering the depth and breadth of the first text data to obtain the importance of the first text data, that is, the first text data is scored by the core degree of the terms in the first text data and the universality and popularity of the terms. The higher the score, the greater the value of the first text data in the historical transmission data.
[0119] Step S223: Match the second text data and the third text data to obtain related words, and calculate the importance of the second text data.
[0120] Specifically, any second text data and third text data are matched, and the same words between the two text data are screened out and recorded as related words.
[0121] Furthermore, in step S223, the process of calculating the importance of the second text data includes:
[0122] determining the importance of the second text data based on a product of the content relevance of the second text data and the comment repetition;
[0123] The calculation formula for content relevance includes:
[0124]
[0125] in, is the content relevance of the second text data; The number of words representing the second text data; Indicates the first The number of related words for each word;
[0126]
[0127] in, The repetitiveness of comments for the second text data; Indicates the first The number of times a word appears in all second text data of the user currently being analyzed; Represents an exponential function with a natural constant as its base.
[0128] Make an explanation, It can reflect the correlation between the second text data and the third text data, that is, the correlation between the comment content and the content of the post to which the comment belongs. The more words in the text data of the two, the more likely the words in the comment content appear in the corresponding post content, and the higher the importance of the second text data. The number of times that the words in the second text data, i.e., the comment content, appear in all the second text data of the currently analyzed user can reflect the repetitiveness of the user's comments. The stronger the repetitiveness, the lower the importance, i.e., the lower the validity of the second text data.
[0129] Step S224: Obtain the importance of all first text data and the importance of all second text data, obtain the first text data importance sequence and the second text data importance sequence respectively, and perform analysis to determine the frequency of change of the first text data importance and the frequency of change of the second text data importance respectively, and calculate the user coefficient of the user.
[0130] Specifically, the importance of all first text data and second text data is calculated, and the importance of all first text data and second text data is sorted based on the time sequence to obtain the first text data importance sequence and the second text data importance sequence respectively, and the two sequences are subjected to first-order backward difference to obtain the first text data importance difference sequence and the second text data importance difference sequence respectively; wherein, the first-order backward difference refers to the difference between the current data point and the previous data point in the data sequence, which is used to calculate the rate of change of the data sequence.
[0131] The number of times the data points in the first text data importance differential sequence and the second text data importance differential sequence change in positive and negative signs are counted respectively, and recorded as the first text data importance change frequency and the second text data importance change frequency respectively. When counting the number of changes in positive and negative signs, if there are continuous positive signs or continuous negative signs without positive and negative transitions, the number is recorded as 0. During the collection days of the currently analyzed user, the higher the importance of the historical transmission data, the higher the user coefficient of the user.
[0132] Furthermore, in step S224, the user coefficient of the user is calculated, and the corresponding calculation formula is:
[0133]
[0134] in, represents the user coefficient of any user; represents the average importance of all first text data of the user currently being analyzed; Indicates the frequency of changes in the importance of the first text data; represents the average importance of all second text data of the currently analyzed user; Indicates the frequency of changes in the importance of the second text data; Represents the normalization function.
[0135] It is explained that the user coefficient is used to evaluate the stability and importance of users in the data transmission process, and can reflect the reliability of data transmission and the degree of influence of users on the entire social networking site.
[0136] Furthermore, in step S3, the transmission data to be encrypted is weighted according to the category coefficient and the user coefficient to obtain the encryption degree of the transmission data to be encrypted, including:
[0137] Define the first text data as , the second text data is , the corresponding calculation formula is:
[0138]
[0139] in, Indicates the encryption level of the transmitted data to be encrypted; Indicates the importance of the first text data to be encrypted; Indicates the importance of the second text data to be encrypted; Indicates the user coefficient of the user corresponding to the transmission data to be encrypted; Indicates the category coefficient of the user corresponding to the transmission data to be encrypted.
[0140] It can be explained that the transmission data to be encrypted is weighted according to the category coefficient and the user coefficient, that is, the historical transmission data of the user is analyzed, and the importance of the first text data and the second text data in the predicted transmission data to be encrypted, the user coefficient and the category coefficient are obtained in turn, so as to assign corresponding weights to the first text data and the second text data to be encrypted, and obtain the corresponding encryption degree of the transmission data to be encrypted.
[0141] See also Figure 2 , which shows a flow chart of an RSA encryption algorithm of a computer transmission data security protection method provided by an embodiment of the present invention. First, a pair of prime numbers are randomly selected, the public modulus is calculated, and then the Euler function is calculated to generate a public key and a private key respectively. The transmission data to be encrypted is encrypted based on the public key to obtain a ciphertext; and the ciphertext is decrypted using the private key to obtain a plaintext; wherein the public key can be shared publicly and used to encrypt data; and the private key needs to be kept confidential and used to decrypt data.
[0142] To illustrate, in step S4, the number of key bits corresponding to the transmission data to be encrypted is obtained based on the encryption level and encrypted using the RSA algorithm to achieve complete protection.
[0143] Specifically, the security level of the RSA encryption algorithm is closely related to the number of bits of the key prime number used. Generally speaking, the larger the number of bits of the prime number, the more difficult it is to crack and the longer it takes to decrypt. Therefore, the higher the encryption level, the larger the number of bits of the selected prime number. In actual applications, in order to ensure data security, commonly used prime number bits include 1024 bits, 2048 bits, and 4096 bits to provide different levels of security protection, which are used to meet encryption requirements of different strengths and effectively resist various potential attack methods.
[0144] Based on the encryption degree, the key number corresponding to the transmission data to be encrypted is obtained and encrypted through the RSA algorithm, that is, the prime number is selected according to the size of the encryption degree. When , choose a 1024-bit prime key; When selecting a 2048-bit prime key; Select a 4096-bit prime key; that is, through the mapping relationship between the preset encryption degree range and the number of prime numbers, the system can automatically select the most appropriate key bit number for encryption operation to ensure the security of data transmission under different security requirements; this selection mechanism is more flexible and can be adjusted and optimized according to the needs of actual application scenarios to adapt to the ever-changing security environment.
[0145] It can be explained that for the transmission data to be encrypted, RSA encryption is performed after selecting the number of key bits to achieve security protection and ensure the security of the user's transmission data during the transmission process; and this security protection method, even if the data is intercepted during the transmission process, a third party without the private key cannot decrypt the data, effectively protecting the security of the user's transmission data during the transmission process.
[0146] Preferably, during the security protection process, after the encrypted transmission data is encrypted, an encryption identifier is generated, which is associated with the encrypted transmission data; and the entire system will also record log information of the encryption operation, including key information such as the encryption time, the identity of the encryptor, and the number of key bits used, so as to facilitate safe storage and facilitate subsequent information leakage or troubleshooting.
[0147] It can be understood that the user's historical transmission data is analyzed, the users are classified, and the category coefficient of each classified user is obtained; the historical transmission data is analyzed, the importance of each historical transmission data is determined, so as to obtain the user coefficient of each user; the comprehensive category coefficient and user coefficient are weighted to determine the encryption degree of the transmission data to be encrypted, and the corresponding key bit number is obtained, so that the more important transmission data to be encrypted uses a larger key, that is, the historical transmission data is analyzed, the first text data and the second text data are graded in importance, and according to different grades, the key bit number is different, and corresponding encryption processing is used, which is conducive to the effective use of network resources and the user's website usage experience.
[0148] An embodiment of the present invention proposes a computer transmission data security protection system, which stores program data. When the program data is executed, a computer transmission data security protection method as described in the aforementioned embodiment is implemented; this system has the same beneficial effects as the computer transmission data security protection method provided above, and will not be repeated here.
[0149] It can be understood that when a module of a computer transmission data security protection system is in operation, it is necessary to utilize a computer transmission data security protection method provided by the aforementioned embodiment. Therefore, whether the method is integrated with program data or different hardware is configured to produce functions similar to the effects achieved by the present invention, it falls within the scope of protection of the present invention.
[0150] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0151] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A computer transmission data security protection method, characterized in that: The method comprises: Obtain historical transmission data of several users from any social networking site; Process historical transmission data to classify users into categories and obtain the category coefficient for each category; analyze the corresponding features of historical transmission data, calculate the data importance of historical transmission data, and determine the user coefficients of all users; The transmission data to be encrypted is weighted according to the category coefficient and the user coefficient to obtain the encryption degree of the transmission data to be encrypted; Based on the encryption level, the key bits corresponding to the transmission data to be encrypted are obtained and encrypted using the RSA algorithm to achieve complete protection; The transmission data includes first text data, second text data and third text data; Process historical transmission data to classify users into different categories and obtain the category coefficient for each category, including: Determining the data volume of the first text data and the second text data respectively, and calculating the activity of all users; Analyze the similarity between any two users, cluster all users based on similarity to obtain several cluster categories, and obtain the average activity of each cluster category according to the activity of each user in the cluster category, which is recorded as the category coefficient; Analyze the corresponding characteristics of historical transmission data, calculate the data importance of historical transmission data, and determine the user coefficients of all users, including: Segmenting the first text data to obtain a plurality of words of the first text data, and determining a word entry from the first text data; Matching the vocabulary with a professional vocabulary dictionary to obtain professional vocabulary and non-professional vocabulary respectively, and calculating the depth and breadth of any first text data in turn to determine the importance of the first text data; matching the second text data and the third text data to obtain related words, and calculating the importance of the second text data; Obtaining the importance of all first text data and the importance of all second text data, respectively obtaining a first text data importance sequence and a second text data importance sequence, performing analysis to respectively determine a frequency of change in the importance of the first text data and a frequency of change in the importance of the second text data, and calculating a user coefficient of the user; determining a corresponding depth level according to a product between the validity of the first text data and the content depth; The calculation formula of the effectiveness is: in, Indicates the vocabulary category of the first text data; The number of words representing the first text data; Indicates the first text data The first word is words, that is, the proportion of adjacent words appearing in the first text data; Indicates the first text data The word after words, that is, the proportion of adjacent words appearing in the first text data; represents the normalization function; The calculation formula of the content depth is: in, represents the number of professional words in the first text data; Indicates the number of words in the professional vocabulary dictionary; The process of obtaining the breadth degree includes: Determining a corresponding breadth degree according to a product of term popularity and discussion degree of the first text data; The calculation formula for the term popularity is: in, Indicates the popularity of a term in the first text data; The number of words representing all entries in the first text data; Represents the first word of all the words in the first text data The number of times a word appears in all the first text data; represents the number of terms in the first text data; Indicates the first text data The number of words in each entry; Indicates the first text data The first The number of times a word appears in all entries; The calculation formula of the discussion degree is: in, Indicates the discussion degree of the first text data; Indicates the first text data among all the entries The number of times a word appears in the first text data.
2. A computer transmission data security protection method according to claim 1, characterized in that: Calculate the activity of all users. The corresponding calculation formula is: in, Indicates the activity level of any user; Indicates the data amount of the first text data; indicating the data amount of the second text data; represents the average number of comments in the first text data; represents the normalization function; Analyze the similarity between any two users, and the corresponding calculation formula is: in, Indicates the users and Similarity of users; Indicates the The activity level of each user; Indicates the The activity level of each user; Represents an exponential function with a natural constant as its base.
3. A computer transmission data security protection method according to claim 1, characterized in that: The process of calculating the importance of the second text data includes: determining the importance of the second text data based on a product of the content relevance of the second text data and the comment repetition; The calculation formula for content relevance includes: in, is the content relevance of the second text data; The number of words representing the second text data; Indicates the first The number of related words for each word; in, The repetitiveness of comments for the second text data; Indicates the first The number of times a word appears in all second text data of the user currently being analyzed; Represents an exponential function with a natural constant as its base.
4. A computer transmission data security protection method according to claim 1, characterized in that: Calculate the user coefficient of the user. The corresponding calculation formula is: in, represents the user coefficient of any user; represents the average importance of all first text data of the user currently being analyzed; Indicates the frequency of changes in the importance of the first text data; represents the average importance of all second text data of the currently analyzed user; Indicates the frequency of changes in the importance of the second text data; Represents the normalization function.
5. A computer transmission data security protection method according to claim 1, characterized in that: The encryption degree of the transmission data to be encrypted is obtained by weighting the transmission data to be encrypted according to the category coefficient and the user coefficient, including: Define the first text data as , the second text data is , the corresponding calculation formula is: in, Indicates the encryption level of the transmitted data to be encrypted; Indicates the importance of the first text data to be encrypted; Indicates the importance of the second text data to be encrypted; Indicates the user coefficient of the user corresponding to the transmission data to be encrypted; Indicates the category coefficient of the user corresponding to the transmission data to be encrypted.
6. A computer transmission data security protection system, characterized in that: The system stores program data, and when the program data is executed, a computer transmission data security protection method as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data protection method and protection system in cloud computing environment
CN119249452A