Data publishing method and system, and electronic device
By obtaining user permission levels and splitting target statements into word sequences, calculating information entropy and gain, and using neural network algorithms to assign risk weights, the problem of unsafe and unreliable data release risk assessment in existing technologies is solved, achieving highly secure and applicable data release and a simplified system architecture.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QINGDAO PENGHAI SOFT CO LTD
- Filing Date
- 2022-10-26
- Publication Date
- 2026-04-24
AI Technical Summary
The current risk assessment process before data release lacks a highly secure, reliable, and applicable risk assessment scheme. Different users need to establish different complex algorithms for different data requests, and there are problems with insufficient security and reliability due to deviations in algorithm and parameter design.
By obtaining user permission levels, splitting the target statement into a sequence of terms, calculating the information entropy and information gain of the terms, assigning risk weights using a neural network algorithm, adjusting the risk weight sequence, judging the security of data release, and the server deciding whether to release the data based on the output results.
It achieves highly secure, reliable, and applicable data publishing, simplifies system architecture, reduces production, installation, and maintenance costs, and improves risk analysis efficiency and data security.
Smart Images

Figure CN115576790B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data publishing methods, and specifically discloses a data publishing method, system, and electronic device thereof. Background Technology
[0002] With the rapid development and widespread adoption of mobile technology and 5G networks, massive amounts of data are constantly being generated across various industries. As this data accumulates and becomes more refined, its research and utilization value becomes increasingly apparent. Therefore, risk assessment before data release is crucial. However, current risk assessment processes often require different complex algorithms for different user data requests. Furthermore, the risk assessment process itself may lack security and reliability due to design flaws in the algorithms and parameters. Consequently, current risk assessment methods lack highly secure, reliable, and applicable solutions.
[0003] Therefore, the existing technology still needs further development and improvement. Summary of the Invention
[0004] To address the shortcomings of existing technologies and solve the aforementioned problems, this invention proposes a highly secure, highly reliable, and highly applicable data publishing method, system, and electronic device. The invention provides the following technical solution:
[0005] According to a first aspect of the present invention, a data publishing method is provided, the method comprising:
[0006] Obtain the user's login information and parse it to determine the user's permission level;
[0007] Retrieve the set of target statements that retrieve data requested by the user;
[0008] Each target statement in the target statement set is split to obtain a word sequence corresponding to each target statement, and the word sequence includes at least one word obtained after splitting the target statement;
[0009] The information entropy of each term in the term sequence of each target statement is determined. The information entropy of each term is processed according to a preset method to obtain the factor value of each term. By calculating the conditional entropy and information gain, the factor value of each term in the term sequence of each target statement is assigned a risk weight according to the information gain of each term and a preset risk weight sequence. The factor value of each term in each target statement and its corresponding risk weight are respectively input into the neural network algorithm to calculate the output result of each target statement. The output result of each target statement is compared with a preset threshold. When the output result of each target statement is less than the preset threshold, the risk weights in the preset risk weight sequence are adjusted according to a first preset ratio for a first preset number of times within the first adjustment period. After each adjustment, the factor value of each term in each target statement and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result of each target statement.
[0010] When the number of times each risk weight in the preset risk weight sequence is adjusted reaches the first preset number, if the output result of each target statement in this neural network is still less than the preset threshold, then the target statement set is determined to meet the publishing requirements, and the server publishes the data.
[0011] Furthermore, the method also includes:
[0012] When the output of each target statement is less than the preset threshold, the risk weights in the preset risk weight sequence are increased by a first preset number of times within the first detection period according to a first preset ratio. After each adjustment, the factor value of each term in each target statement and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. If, during the process of increasing the risk weights in the preset risk weight sequence, the output of any target statement is greater than the preset threshold, it is determined that there is a significant risk in publishing the target statement set, and the server prohibits the publishing of data and sends a data transmission failure prompt to the terminal.
[0013] Furthermore, the preset method includes:
[0014] The factor value of each term is obtained by rounding down the information entropy of each term.
[0015] Furthermore, the step of assigning risk weights to the factor values of each term in the term sequence according to the information gain of each term through a preset risk weight sequence by calculating conditional entropy and information gain includes:
[0016] Calculate the conditional entropy and information gain of each term in each term sequence of each target statement. Based on the information gain of each term in each term sequence of each target statement, assign the factor values of each term in each term sequence to the risk weights in the preset risk weight sequence in descending order.
[0017] Furthermore, the risk weights in the preset risk weight sequence are arranged in descending order.
[0018] Furthermore, the method also includes:
[0019] If, after assigning risk weights to each term in the term sequence from high to low according to the information gain of each term in each target statement, there are still terms whose factor values have not been assigned risk weights, then the factor values of all terms that have not been assigned risk weights will not be assigned risk weights and will not participate in the calculation of the neural network algorithm.
[0020] Furthermore, the preset threshold includes a first preset threshold, a second preset threshold, a third preset threshold, and a fourth preset threshold. The user's permission level is one of four: data owner, data manager, data producer, and data user. When the user's permission is data owner, the preset threshold is set to the first preset threshold; when the user's permission is data manager, the preset threshold is set to the second preset threshold; when the user's permission is data producer, the preset threshold is set to the third preset threshold; and when the user's permission is data user, the preset threshold is set to the fourth preset threshold. The first preset threshold is greater than the second preset threshold, the second preset threshold is greater than the third preset threshold, and the third preset threshold is greater than the fourth preset threshold.
[0021] Furthermore, the first adjustment cycle is 5 seconds, the first preset number of times is 10, and the first preset ratio is 1%.
[0022] According to a second aspect of the present invention, a data publishing system is provided, the data publishing system comprising:
[0023] The acquisition module is used to acquire user login information and parse it to obtain the user's permission level; or acquire the target statement set of data called by the user; or split each target statement in the target statement set to obtain the word sequence corresponding to each target statement and send the word sequence to the analysis module; the word sequence includes at least one word obtained after splitting the target statement;
[0024] The analysis module is used to determine the information entropy of each term in the term sequence of each target statement. It processes the information entropy of each term in the target statement according to a preset method to obtain the factor value of each term. It determines the risk level of each term in the term sequence of each target statement using conditional entropy and information gain, and assigns risk weights to the factor values of each term according to a preset risk weight sequence. It inputs the factor values and corresponding risk weights of each term in the target statement into a neural network algorithm to calculate the output of each target statement. It compares the output of each target statement with a preset threshold and sends the comparison result to the server. If the output of each target statement is less than the preset threshold, it adjusts the risk weights in the preset risk weight sequence a first preset number of times within the first detection period according to a first preset ratio. After each adjustment, it inputs the factor values and corresponding adjusted risk weights of each term in the target statement into the neural network algorithm to calculate the output of each target statement and compares it with the preset threshold, sending the comparison result to the server.
[0025] The server is used to uniformly manage the control programs of all modules based on the comparison results, or to control the risk weight adjustment of the analysis module, or to control or prohibit data release and send a data release failure prompt to the terminal, or to determine that the data meets the release requirements and release the data if the number of times each risk weight in the preset risk weight sequence is equal to the first preset number and the output result of each target statement of the next neural network is still less than the preset threshold.
[0026] According to a third aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0027] A memory; and a processor, the memory storing computer-readable instructions which, when executed by the processor, implement the data publishing method according to any one of claims 1 to 8.
[0028] The beneficial effects of this invention are:
[0029] 1. This invention can evaluate the output of the neural network algorithm by using different preset thresholds when the user's permission level is different. The higher the user's permission level, the higher the corresponding preset threshold. That is, the user can obtain data with a higher risk level through this data publishing system, which greatly improves the intelligence level of this invention and greatly expands the application scenarios of this invention.
[0030] 2. This invention can analyze information gain and assign risk weights to the factor values of each term in the term sequence in real time according to the magnitude of the information gain of each term through a preset risk weight sequence. It can assign the factor value of the term with the largest information gain the largest risk weight. The same applies to the remaining terms in the target statement, that is, assign the factor value of the term with the greatest impact on risk assessment in the target statement to the largest risk weight. The remaining terms in the target statement are also assigned risk weights in the preset risk weight sequence according to the magnitude of the information gain of each term. This can achieve a highly sensitive, reliable, and applicable risk assessment of the target statement, which greatly improves the intelligence and reliability of this invention and further broadens the application scenarios of this invention.
[0031] 3. This invention sets a first adjustment period. When the output results of all target statements are less than a preset threshold, the risk weights in the preset risk weight sequence are adjusted according to a first preset ratio for a first preset number of times within the first adjustment period. After each adjustment, the factor value of each term and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. Only when the adjustment of each risk weight in the preset risk weight sequence reaches the first preset number of times and the output results of each target statement in the neural network are still less than the preset threshold, is the data determined to meet the publication requirements, and the server publishes the data. This can greatly improve the security of information, improve the reliability of risk analysis, prevent system misjudgment, and further improve the reliability of the risk assessment results of this invention.
[0032] 4. This invention can achieve highly reliable and applicable risk assessment and analysis of the data used for the call through algorithms alone, enabling timely and efficient risk assessment before data release. It does not require complex systems and modeling calculations, greatly simplifies the system architecture while ensuring the accuracy of analysis, significantly reduces the system production, installation and maintenance costs, and greatly improves enterprise efficiency.
[0033] 5. This invention, through ingenious design, sets the first adjustment cycle to 5 seconds, the first preset number of times to 10, and the first preset ratio to 1%. It optimizes the design of various circuit parameters, avoids the influence of various factors on the system risk judgment results, optimizes the program architecture, greatly improves the efficiency of risk analysis, saves a lot of time and manpower costs, and realizes the low-power and high-reliability operation of the data publishing system. Attached Figure Description
[0034] Figure 1 This is a flowchart of the data publishing method in a specific embodiment of the present invention;
[0035] Figure 2 This is a schematic diagram of a data publishing method in another specific embodiment of the present invention;
[0036] Figure 3 This is a schematic diagram of the data publishing system in a specific embodiment of the present invention. Detailed Implementation
[0037] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Based on the embodiments in this application, other similar embodiments obtained by those skilled in the art without creative effort should all fall within the scope of protection of this application. Furthermore, directional terms mentioned in the following embodiments, such as "up," "down," "left," and "right," are only for reference to the directions in the accompanying drawings; therefore, the directional terms used are for illustrative purposes and not for limiting the invention.
[0038] The present invention will be further described below with reference to the accompanying drawings and preferred embodiments.
[0039] Please see Figure 1 This invention provides a data publishing method, comprising:
[0040] S100: Obtain the user's login information and parse it to obtain the user's permission level.
[0041] It should be noted that the data publishing system of the present invention is pre-set with an acquisition module, which is used to acquire the user's login information and parse the user's permission level; or acquire the target statement set of the data called by the user; or split each target statement in the target statement set to obtain the word sequence corresponding to each target statement and send the word sequence to the analysis module; the word sequence includes at least one word obtained after splitting the target statement.
[0042] The steps preceding step S100 include: setting a first adjustment period, a first preset number of times, and a first preset ratio. Preferably, the first detection period is 5 seconds, the first preset ratio is 1%, and the first preset number of times is 10.
[0043] S200, Obtain the set of target statements for the data invoked by the user.
[0044] It should be noted here that the server retrieves the target set of data according to the user's needs, providing a data foundation for risk analysis of the target set of data.
[0045] S300. Each target statement in the target statement set is split to obtain a word sequence corresponding to each target statement. The word sequence includes at least one word obtained after splitting the target statement.
[0046] It should be noted that each target statement in the target statement set is broken down for subsequent risk assessment. If the risk assessment result of the target statement set is that it meets the data release requirements, the server releases the data. This greatly improves the reliability of the risk assessment results of the target statement set in this invention.
[0047] S400. Determine the information entropy of each term in the term sequence. Process the information entropy of each term according to a preset method to obtain the factor value of each term. Calculate the conditional entropy and information gain. Assign risk weights to the factor values of each term in the term sequence according to the information gain of each term through a preset risk weight sequence. Input the factor value of each term and its corresponding risk weight into a neural network algorithm to calculate the output result. Compare the output result with a preset threshold. If the output result is less than the preset threshold, adjust the risk weights in the preset risk weight sequence for a first preset number of times according to a first preset ratio within the first adjustment period. After each adjustment, input the factor value of each term and its corresponding adjusted risk weight into the neural network algorithm to calculate the output result.
[0048] It should be noted that the preset thresholds include a first preset threshold, a second preset threshold, a third preset threshold, and a fourth preset threshold. The user's permission level is one of four: data owner, data manager, data producer, and data user. When the user's permission is data owner, the preset threshold is set to the first preset threshold; when the user's permission is data manager, the preset threshold is set to the second preset threshold; when the user's permission is data producer, the preset threshold is set to the third preset threshold; and when the user's permission is data user, the preset threshold is set to the fourth preset threshold. The first preset threshold is greater than the second preset threshold, the second preset threshold is greater than the third preset threshold, and the third preset threshold is greater than the fourth preset threshold. This setting enables intelligent switching of preset thresholds, greatly improving the reliability of analysis results and further enhancing the intelligence of the invention. It allows the evaluation of the neural network algorithm's output results using different preset thresholds when the user's permission level is different. The higher the user's permission level, the higher the corresponding preset threshold, meaning that users can obtain data with higher risk levels through this data publishing system. This greatly improves the intelligence level of the invention and significantly expands its application scenarios. The module acquires the user's login information and parses it to obtain the user's permission level; it acquires the target statement set of the data called by the user, and rounds the information entropy of each term to obtain the factor value of each term. This setting can realize the use of the information entropy of each term as input to the neural network algorithm, and can realize the analysis of the impact of each term on the overall data system, which greatly improves the analytical reliability and usability of the present invention, and greatly expands the application scenarios of the present invention. The conditional entropy and information gain of each term in each term sequence of each target statement are calculated respectively. According to the information gain of each term in each term sequence of each target statement, the factor values of each term in each term sequence are assigned to the risk weights in the preset risk weight sequence from high to low. This invention enables real-time risk weight assignment based on the information gain of each term in a term sequence through information gain analysis and a preset risk weight sequence. It assigns the highest risk weight to the term with the highest information gain, and similarly assigns the highest risk weight to the remaining terms in the target statement. The remaining terms in the target statement are also assigned risk weights according to their information gain, thus achieving highly sensitive, reliable, and applicable risk assessment of the target statement. This significantly improves the intelligence and reliability of the invention and further expands its application scenarios.This invention, through ingenious design, sets the first adjustment cycle to 5 seconds, the first preset number of adjustments to 10, and the first preset ratio to 1%. By optimizing various circuit parameters, it avoids the influence of various factors on the system's risk assessment results, optimizes the program architecture, significantly improves risk analysis efficiency, and greatly saves time and manpower costs, achieving low-power, high-reliability operation of the data publishing system. The output results of each target statement are calculated by multiplying the factor value of each term by its corresponding risk weight and then summing the results.
[0049] Specifically, the preset risk weight sequence includes a first risk weight, a second risk weight, a third risk weight, a fourth risk weight, and a fifth risk weight. The risk weights in the preset risk weight sequence are arranged in descending order. If, after assigning risk weights to each term in the term sequence according to its information gain, from highest to lowest, based on the preset risk weight sequence, there are still terms whose factor values have not been assigned risk weights, then all term values without risk weight assignments will not be assigned risk weights and will not participate in the neural network algorithm's calculations. This design can ensure the reliability of risk analysis results, greatly simplify the algorithm and system architecture, and achieve timely and efficient risk assessment before data release. It eliminates the need for complex systems and modeling calculations, significantly simplifying the system architecture while ensuring analytical accuracy, greatly reducing system production, installation, and maintenance costs, and significantly improving enterprise efficiency. The first risk weight is 50, the second risk weight is 40, the third risk weight is 30, the fourth risk weight is 20, and the fifth risk weight is 10. Setting the first risk weight to 50, the second risk weight to 40, the third risk weight to 30, the fourth risk weight to 20, and the fifth risk weight to 10 was determined by the inventors through extensive experiments. Since the set of target statements invoked by the user is uncertain, the neural network algorithm of this invention analyzes the information entropy and information gain of each term in each target statement through an analysis module. This allows for the assessment of the information content and publication risk of the invoked target statement set, achieving a high-sensitivity, high-reliability, and high-applicability risk assessment of the target statements. This significantly improves the intelligence and reliability of this invention and further broadens its application scenarios.Setting the first risk weight to 50, the second risk weight to 40, the third risk weight to 30, the fourth risk weight to 20, and the fifth risk weight to 10 ensures that the factor value corresponding to the term with the greatest impact on the data system is always the highest risk weight. This setting significantly improves the high sensitivity and intelligent analysis of the terms that have the greatest impact on the release risk assessment in this invention. This setting can complete the release risk assessment of the called target statement set while greatly simplifying the system architecture and reducing system complexity, significantly reducing the production and maintenance costs of this system, greatly improving the usability of this invention, and greatly expanding the application scenarios of this invention. This calculation method can process the information entropy and information gain of each term in the target statement called by the user into the input of a neural network algorithm and calculate the output result, realizing highly sensitive intelligent integration monitoring of the set of target statements called by the user. The analysis module compares the output results of each target statement with a preset threshold. If the analysis module determines that the output result of a target statement is greater than the preset threshold, it indicates that the release of the target statement set poses a significant risk, and the server does not release the data. By setting a first adjustment period, when the output results of all target statements are less than the preset threshold, the risk weights in the preset risk weight sequence are adjusted according to a first preset ratio for a first preset number of times within the first adjustment period. After each adjustment, the factor value of each term and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. Only when the risk weights in the preset risk weight sequence are adjusted to the first preset number of times and the output results of each target statement in that neural network are still less than the preset threshold, is the data deemed to meet the release requirements, and the server releases the data. This greatly improves information security, enhances the reliability of risk analysis, and prevents system misjudgments, further improving the reliability of the risk assessment results of this invention. By analyzing only the information entropy and information gain of each term in each target statement, the analysis of the release risk of the target statement set can be completed, greatly simplifying the analysis process, significantly improving analysis efficiency, reducing system complexity, significantly reducing the costs of system production, installation, operation, and maintenance, and significantly improving enterprise efficiency.
[0050] Specifically, when the output results of all target statements are less than the preset threshold, it indicates that the preliminary risk assessment of the target statement set meets the release requirements. At this point, further release risk assessment is required. Within the first adjustment cycle, the risk weights in the preset risk weight sequence are adjusted according to the first preset ratio for the first preset number of times. After each adjustment, the factor value of each term and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. Setting the first preset ratio to 1% can achieve good traversal adjustment of risk weights. By adjusting the risk weights, the risk stability of the target statement set is assessed, preventing the release of high-risk information caused by unreasonable manual setting of risk weights, further improving system security and the reliability of analysis results. Setting the first adjustment cycle to 5 seconds and the first preset number of times to 10 times can reduce the computational complexity of the system program while ensuring adjustment accuracy. The program design is ingenious, reasonable, simple, and efficient, saving user time and improving user experience. It solves the problem that the existing technology of manually adjusting the neural network model is very complex and cumbersome, greatly improving adjustment efficiency and enterprise benefits. If, during the risk weight adjustment process, the output of any target statement exceeds a preset threshold, the server determines that the publication of the target statement set poses a significant risk. The server then prohibits data publication and sends a data publication failure message to the terminal, thereby further improving the data security and reliability of the risk assessment results of this system.
[0051] S500: When the number of times each risk weight in the preset risk weight sequence is adjusted reaches the first preset number, if the output results of all target statements in this neural network are still less than the preset threshold, then the data is deemed to meet the release requirements, and the server releases the data. This greatly improves the data security and reliability of the risk assessment results of this system.
[0052] Please see Figure 2 This invention proposes a specific embodiment, providing a data publishing method, the method comprising:
[0053] P0: Start.
[0054] P1: Obtain the user's login information and parse it to obtain the user's permission level.
[0055] It should be noted that the data publishing system of this invention is pre-configured with an acquisition module, which is used to acquire user login information and parse it to obtain the user's permission level; or acquire the target statement set of data called by the user; or split each target statement in the target statement set to obtain the word sequence corresponding to each target statement and send the word sequence to the analysis module; the word sequence includes at least one word obtained after splitting the target statement.
[0056] Before step P1, the steps include setting a first adjustment period, a first preset number of times, and a first preset ratio. Preferably, the first detection period is 5 seconds, the first preset ratio is 1%, and the first preset number of times is 10.
[0057] P2. Obtain the set of target statements for the data invoked by the user.
[0058] It should be noted here that the server retrieves the target set of data according to the user's needs, providing a data foundation for risk analysis of the target set of data.
[0059] P3. Each target statement in the target statement set is split to obtain the word sequence corresponding to each target statement.
[0060] It should be noted that each target statement in the target statement set is broken down for subsequent risk assessment. If the risk assessment result of the target statement set is that it meets the data release requirements, the server releases the data. This greatly improves the reliability of the risk assessment results of the target statement set in this invention.
[0061] P4. Determine the information entropy of each term in the term sequence of each target statement, and process the information entropy of each term according to the preset method to obtain the factor value of each term.
[0062] It should be noted that the preset thresholds include a first preset threshold, a second preset threshold, a third preset threshold, and a fourth preset threshold. The user's permission level is one of four: data owner, data manager, data producer, and data user. When the user's permission is data owner, the preset threshold is set to the first preset threshold; when the user's permission is data manager, the preset threshold is set to the second preset threshold; when the user's permission is data producer, the preset threshold is set to the third preset threshold; and when the user's permission is data user, the preset threshold is set to the fourth preset threshold. The first preset threshold is greater than the second preset threshold, the second preset threshold is greater than the third preset threshold, and the third preset threshold is greater than the fourth preset threshold. This setting enables intelligent switching of preset thresholds, greatly improving the reliability of analysis results and further enhancing the intelligence of the invention. It allows the evaluation of the neural network algorithm's output results using different preset thresholds when the user's permission level is different. The higher the user's permission level, the higher the corresponding preset threshold, meaning that users can obtain data with higher risk levels through this data publishing system. This greatly improves the intelligence level of the invention and significantly expands its application scenarios. The module acquires the user's login information and parses it to obtain the user's permission level; it acquires the target statement set of the data called by the user, and rounds the information entropy of each term to obtain the factor value of each term. This setting can realize the use of the information entropy of each term as input to the neural network algorithm, and can realize the analysis of the impact of each term on the overall data system, which greatly improves the analytical reliability and usability of the present invention, and greatly expands the application scenarios of the present invention. The conditional entropy and information gain of each term in each term sequence of each target statement are calculated respectively. According to the information gain of each term in each term sequence of each target statement, the factor values of each term in each term sequence are assigned to the risk weights in the preset risk weight sequence from high to low. This invention enables real-time risk weight assignment based on the information gain of each term in a term sequence through information gain analysis and a preset risk weight sequence. It assigns the highest risk weight to the term with the highest information gain, and similarly assigns the highest risk weight to the remaining terms in the target statement. The remaining terms in the target statement are also assigned risk weights according to their information gain, thus achieving highly sensitive, reliable, and applicable risk assessment of the target statement. This significantly improves the intelligence and reliability of the invention and further expands its application scenarios.This invention, through ingenious design, sets the first adjustment cycle to 5 seconds, the first preset number of adjustments to 10, and the first preset ratio to 1%. By optimizing various circuit parameters, it avoids the influence of various factors on the system's risk assessment results, optimizes the program architecture, significantly improves risk analysis efficiency, and greatly saves time and manpower costs, achieving low-power, high-reliability operation of the data publishing system. The output results of each target statement are calculated by multiplying the factor value of each term by its corresponding risk weight and then summing the results.
[0063] P5. By calculating conditional entropy and information gain, the factor values of each term in the term sequence of each target statement are assigned risk weights according to the information gain of each term and the preset risk weight sequence.
[0064] Specifically, the preset risk weight sequence includes a first risk weight, a second risk weight, a third risk weight, a fourth risk weight, and a fifth risk weight. The risk weights in the preset risk weight sequence are arranged in descending order. If, after assigning risk weights to each term in the term sequence according to its information gain, from highest to lowest, based on the preset risk weight sequence, there are still terms whose factor values have not been assigned risk weights, then all term values without risk weight assignments will not be assigned risk weights and will not participate in the neural network algorithm's calculations. This design can ensure the reliability of risk analysis results, greatly simplify the algorithm and system architecture, and achieve timely and efficient risk assessment before data release. It eliminates the need for complex systems and modeling calculations, significantly simplifying the system architecture while ensuring analytical accuracy, greatly reducing system production, installation, and maintenance costs, and significantly improving enterprise efficiency. The first risk weight is 50, the second risk weight is 40, the third risk weight is 30, the fourth risk weight is 20, and the fifth risk weight is 10. Setting the first risk weight to 50, the second risk weight to 40, the third risk weight to 30, the fourth risk weight to 20, and the fifth risk weight to 10 was determined by the inventors through extensive experiments. Since the set of target statements invoked by the user is uncertain, the neural network algorithm of this invention analyzes the information entropy and information gain of each term in each target statement through an analysis module. This allows for the assessment of the information content and publication risk of the invoked target statement set, achieving a high-sensitivity, high-reliability, and high-applicability risk assessment of the target statements. This significantly improves the intelligence and reliability of this invention and further broadens its application scenarios.Setting the first risk weight to 50, the second risk weight to 40, the third risk weight to 30, the fourth risk weight to 20, and the fifth risk weight to 10 ensures that the factor value corresponding to the term with the greatest impact on the data system is always the highest risk weight. This setting significantly improves the high sensitivity and intelligent analysis of the terms that have the greatest impact on the release risk assessment in this invention. This setting can complete the release risk assessment of the called target statement set while greatly simplifying the system architecture and reducing system complexity, significantly reducing the production and maintenance costs of this system, greatly improving the usability of this invention, and greatly expanding the application scenarios of this invention. This calculation method can process the information entropy and information gain of each term in the target statement called by the user into the input of a neural network algorithm and calculate the output result, realizing highly sensitive intelligent integration monitoring of the set of target statements called by the user. The analysis module compares the output results of each target statement with a preset threshold. If the analysis module determines that the output result of a target statement is greater than the preset threshold, it indicates that the release of the target statement set poses a significant risk, and the server does not release the data. By setting a first adjustment period, when the output results of all target statements are less than the preset threshold, the risk weights in the preset risk weight sequence are adjusted according to a first preset ratio for a first preset number of times within the first adjustment period. After each adjustment, the factor value of each term and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. Only when the risk weights in the preset risk weight sequence are adjusted to the first preset number of times and the output results of each target statement in that neural network are still less than the preset threshold, is the data deemed to meet the release requirements, and the server releases the data. This greatly improves information security, enhances the reliability of risk analysis, and prevents system misjudgments, further improving the reliability of the risk assessment results of this invention. By analyzing only the information entropy and information gain of each term in each target statement, the analysis of the release risk of the target statement set can be completed, greatly simplifying the analysis process, significantly improving analysis efficiency, reducing system complexity, significantly reducing the costs of system production, installation, operation, and maintenance, and significantly improving enterprise efficiency.
[0065] P6. Input the factor values of each term in each target statement and its corresponding risk weight into the neural network algorithm to calculate the output of each target statement.
[0066] It should be noted that the output results of each target statement are calculated by multiplying the factor value of each term by its corresponding risk weight and then summing the results.
[0067] P7. Compare the output results of each target statement. Are there any items that exceed the preset threshold? If yes, proceed to step P12; otherwise, proceed to step P8.
[0068] Specifically, when the output results of all target statements are less than the preset threshold, it indicates that the preliminary risk assessment of the target statement set meets the release requirements. At this point, further release risk assessment is required. Within the first adjustment cycle, the risk weights in the preset risk weight sequence are adjusted according to the first preset ratio for the first preset number of times. After each adjustment, the factor value of each term and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. Setting the first preset ratio to 1% can achieve good traversal adjustment of risk weights. By adjusting the risk weights, the risk stability of the target statement set is assessed, preventing the release of high-risk information caused by unreasonable manual setting of risk weights, further improving system security and the reliability of analysis results. Setting the first adjustment cycle to 5 seconds and the first preset number of times to 10 times can reduce the computational complexity of the system program while ensuring adjustment accuracy. The program design is ingenious, reasonable, simple, and efficient, saving user time and improving user experience. It solves the problem that the existing technology of manually adjusting the neural network model is very complex and cumbersome, greatly improving adjustment efficiency and enterprise benefits. If, during the risk weight adjustment process, the output of any target statement exceeds a preset threshold, the server determines that the publication of the target statement set poses a significant risk. The server then prohibits data publication and sends a data publication failure message to the terminal, thereby further improving the data security and reliability of the risk assessment results of this system.
[0069] P8. During the first adjustment period, increase the risk weights in the preset risk weight sequence by the first preset ratio, and after each adjustment, input the factor values of each term in each target statement and its corresponding adjusted risk weights into the neural network algorithm to calculate the output of each target statement.
[0070] Specifically, when the output results of all target statements are less than the preset threshold, it indicates that the preliminary risk assessment of the target statement set meets the release requirements. At this point, further release risk assessment is required. Within the first adjustment cycle, the risk weights in the preset risk weight sequence are adjusted according to the first preset ratio for the first preset number of times. After each adjustment, the factor value of each term and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. Setting the first preset ratio to 1% can achieve good traversal adjustment of risk weights. By adjusting the risk weights, the risk stability of the target statement set is assessed, preventing the release of high-risk information caused by unreasonable manual setting of risk weights, further improving system security and the reliability of analysis results. Setting the first adjustment cycle to 5 seconds and the first preset number of times to 10 times can reduce the computational complexity of the system program while ensuring adjustment accuracy. The program design is ingenious, reasonable, simple, and efficient, saving user time and improving user experience. It solves the problem that the existing technology of manually adjusting the neural network model is very complex and cumbersome, greatly improving adjustment efficiency and enterprise benefits.
[0071] P9. Compare the output results of each target statement. Are there any items that exceed the preset threshold? If yes, proceed to step P12; otherwise, proceed to step P10.
[0072] P10. Has the number of weight adjustments reached the first preset number? If yes, proceed to step P11; otherwise, return to step P8.
[0073] It should be noted that during the risk weight adjustment process, if the output of any target statement exceeds the preset threshold, the server determines that the publication of the target statement set poses a significant risk. The server then prohibits data publication and sends a data publication failure message to the terminal, further improving the data security and reliability of the risk assessment results of this system.
[0074] P11. The server determines that the target statement set meets the publishing requirements and then publishes the data.
[0075] P12. If the server determines that the set of target statements to be published poses a significant risk, the server will prohibit the publication of data and send a data transmission failure message to the terminal.
[0076] P13, End.
[0077] Please see Figure 3 The present invention provides another embodiment, which provides a data publishing system, the data publishing system comprising:
[0078] It includes acquisition module 1, analysis module 2, and server 3;
[0079] The acquisition module 1 is used to acquire the user's login information and parse it to obtain the user's permission level; or acquire the target statement set of data called by the user; or split each target statement in the target statement set to obtain the word sequence corresponding to each target statement and send the word sequence to the analysis module; the word sequence includes at least one word obtained after splitting the target statement;
[0080] Analysis module 2 is used to determine the information entropy of each term in the term sequence. It processes the information entropy of each term according to a preset method to obtain the factor value of each term. It determines the risk level of each term in the term sequence using conditional entropy and information gain, and assigns risk weights to the factor values of each term according to a preset risk weight sequence. It inputs the factor values of each term and their corresponding risk weights into a neural network algorithm to calculate the output result, compares the output result with a preset threshold, and sends the comparison result to server 3. Alternatively, if the output result is less than the preset threshold, it adjusts the risk weights in the preset risk weight sequence a first preset number of times within the first detection period according to a first preset ratio. After each adjustment, it inputs the factor values of each term and their corresponding adjusted risk weights into the neural network algorithm to calculate the output result, compares it with the preset threshold, and sends the comparison result to server 3.
[0081] Server 3 is used to uniformly manage the control programs of all modules based on the comparison results, or to control the risk weight adjustment of the analysis module, or to control or prohibit data release and send a data release failure prompt to the terminal, or to determine that the data meets the release requirements and release the data if the number of times each risk weight in the preset risk weight sequence is adjusted equal to the first preset number and the output result of the next neural network is still less than the preset threshold.
[0082] In a preferred embodiment, this application also provides an electronic device, the electronic device comprising:
[0083] The computer device includes a memory and a processor, wherein the memory stores computer-readable instructions that, when executed by the processor, implement the data publishing method described herein. The computer device can be broadly categorized as a server, terminal, or any other electronic device with the necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the computer device can be used to provide the necessary computing, processing, and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface and communication interface of the computer device can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the method of the present invention.
[0084] This invention can be implemented as a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the steps of the methods of embodiments of the invention to be performed. In one embodiment, the computer program is distributed across multiple network-coupled computer devices or processors, such that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be executed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be executed by one or more computer devices or processors, and one or more other method steps / operations may be executed by one or more other computer devices or processors. One or more computer devices or processors may execute a single method step / operation, or execute two or more method steps / operations.
[0085] Those skilled in the art will understand that the method steps of this invention can be performed by a computer program instructing related hardware, such as a computer device or processor, to perform the steps of this invention when executed. Depending on the context, any references herein to memory, storage, databases, or other media may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.
[0086] The technical features described above can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification, provided that such combination does not contain contradictions.
[0087] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A data publishing method, characterized in that, Obtain the user's login information and parse it to determine the user's permission level; Retrieve the set of target statements that retrieve data requested by the user; Each target statement in the target statement set is split to obtain a word sequence corresponding to each target statement, and the word sequence includes at least one word obtained after splitting the target statement; The information entropy of each term in the term sequence of each target statement is determined. The information entropy of each term is processed according to a preset method to obtain the factor value of each term. By calculating conditional entropy and information gain, the factor value of each term in the term sequence of each target statement is assigned a risk weight according to the information gain of each term and a preset risk weight sequence. The factor value of each term in each target statement and its corresponding risk weight are respectively input into a neural network algorithm to calculate the output result of each target statement. The output result of each target statement is compared with a preset threshold. When the output result of each target statement is less than the preset threshold, the risk weights in the preset risk weight sequence are adjusted according to a first preset ratio for a first preset number of times within the first adjustment period. After each adjustment, the factor value of each term in each target statement and its corresponding adjusted risk weight are input into a neural network algorithm to calculate the output result of each target statement. When the number of times each risk weight in the preset risk weight sequence is adjusted reaches the first preset number, and the output result of each target statement of the judgment neural network is still less than the preset threshold, then the target statement set is judged to meet the publishing requirements, and the server publishes the data.
2. The data publishing method according to claim 1, characterized in that, The method further includes: When the output of each target statement is less than the preset threshold, the risk weights in the preset risk weight sequence are increased by a first preset number of times within the first detection period according to a first preset ratio. After each adjustment, the factor value of each term in each target statement and its corresponding adjusted risk weight are input into the neural network algorithm to calculate the output result. If, during the process of increasing the risk weights in the preset risk weight sequence, the output of any target statement is greater than the preset threshold, it is determined that there is a significant risk in publishing the target statement set, and the server prohibits the publishing of data and sends a data transmission failure prompt to the terminal.
3. The data publishing method according to claim 1, characterized in that, The preset method includes: The factor value of each term is obtained by rounding down the information entropy of each term.
4. The data publishing method according to claim 1, characterized in that, The step of assigning risk weights to the factor values of each term in the term sequence according to the information gain of each term based on the calculated conditional entropy and information gain, according to a preset risk weight sequence, includes: Calculate the conditional entropy and information gain of each term in each term sequence of each target statement. Based on the information gain of each term in each term sequence of each target statement, assign the factor values of each term in each term sequence to the risk weights in the preset risk weight sequence in descending order.
5. A data publishing method according to claim 4, characterized in that, The risk weights in the preset risk weight sequence are arranged in descending order.
6. The data publishing method according to claim 1, characterized in that, The method further includes: If, after assigning risk weights to each term in the term sequence from high to low according to the information gain of each term in each target statement, there are still terms whose factor values have not been assigned risk weights, then the factor values of all terms that have not been assigned risk weights will not be assigned risk weights and will not participate in the calculation of the neural network algorithm.
7. A data publishing method according to claim 5, characterized in that, The preset thresholds include a first preset threshold, a second preset threshold, a third preset threshold, and a fourth preset threshold. The user's permission level is one of four: data owner, data manager, data producer, and data user. When the user's permission is data owner, the preset threshold is set to the first preset threshold; when the user's permission is data manager, the preset threshold is set to the second preset threshold; when the user's permission is data producer, the preset threshold is set to the third preset threshold; and when the user's permission is data user, the preset threshold is set to the fourth preset threshold. The first preset threshold is greater than the second preset threshold, the second preset threshold is greater than the third preset threshold, and the third preset threshold is greater than the fourth preset threshold.
8. A data publishing method according to claim 7, characterized in that, The first adjustment cycle is 5 seconds, the first preset number of times is 10, and the first preset ratio is 1%.
9. A data publishing system, characterized in that, include: The acquisition module is used to obtain the user's login information and parse it to obtain the user's permission level; Or obtain the target set of statements that retrieve the data called by the user; Alternatively, each target statement in the target statement set can be split to obtain a word sequence corresponding to each target statement, and the word sequence can be sent to the analysis module; The term sequence includes at least one term obtained after splitting the target statement; The analysis module is used to determine the information entropy of each term in the term sequence of each target statement. It processes the information entropy of each term in the target statement according to a preset method to obtain the factor value of each term. It determines the risk level of each term in the term sequence of each target statement using conditional entropy and information gain, and assigns risk weights to the factor values of each term according to a preset risk weight sequence. It inputs the factor values and corresponding risk weights of each term in the target statement into a neural network algorithm to calculate the output of each target statement. It compares the output of each target statement with a preset threshold and sends the comparison result to the server. If the output of each target statement is less than the preset threshold, it adjusts the risk weights in the preset risk weight sequence a first preset number of times within the first detection period according to a first preset ratio. After each adjustment, it inputs the factor values and corresponding adjusted risk weights of each term in the target statement into the neural network algorithm to calculate the output of each target statement and compares it with the preset threshold, sending the comparison result to the server. The server is used to uniformly manage the control programs of all modules based on the comparison results, or to control the risk weight adjustment of the analysis module, or to control or prohibit data release and send a data release failure prompt to the terminal, or to determine that the data meets the release requirements and release the data if the number of times each risk weight in the preset risk weight sequence is equal to the first preset number and the output results of each target statement of the next neural network are still less than the preset threshold.
10. An electronic device, characterized in that, include: Memory; The processor, wherein the memory stores computer-readable instructions that, when executed by the processor, implement the data publishing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
A method and a device for training an item weight calculation model
CN108959263A
Information publishing method, device, equipment and storage medium
CN113761110A