Data acquisition, display and analysis system

By designing a data acquisition, display and analysis system, and using convolutional neural networks and semantic feature tables to identify and optimize data uploaded by users, the problem of how to effectively identify and optimize risk information in data information is solved, and the effect of reducing the risk of data information is achieved.

CN120104897APending Publication Date: 2025-06-06INNER MONGOLIA SHANSHAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510161532.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

How to effectively identify and optimize the risk information contained in the data uploaded by users to reduce the risk of the data information finally displayed.

Method used

Design a data acquisition, display and analysis system, including a data acquisition module, a data processing module, a data analysis module and a data display module. The system communicates with the user through the monitoring center, collects and processes user data, extracts data features using a convolutional neural network, and generates data tags through a semantic feature table, and optimizes and displays data information based on the tag.

Benefits of technology

It realizes the risk identification and optimization of the data uploaded by users, reduces the probability of risk content in the data information, and improves the security of data display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104897A_ABST
    Figure CN120104897A_ABST
Patent Text Reader

Abstract

The invention discloses a data acquisition, display and analysis system, and relates to the technical field of data information processing, normalization processing is performed on various types of data uploaded by a user, and data features in normalized data obtained after normalization processing are extracted through a trained convolutional neural network model, so that the data acquisition, display and analysis efficiency is improved. The extracted data features are endowed with different types of data labels through the constructed semantic feature table, whether risks exist in the extracted data features or not is judged according to the endowed labels, then whether risks exist in the corresponding data information or not is judged, and when the risks exist in the data information, the data information is extracted from the extracted data features. According to the method, the data information with risks is replaced by the semantic features without risks, so that optimization of the data information uploaded by the user is realized, and the probability that risk content appears in the data information content uploaded by the user is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data information processing, and in particular to a data collection, display and analysis system. Background Art

[0002] The number of users on major Internet social platforms has increased dramatically, and the platform content is constantly updated. Users can post a large amount of text information anytime and anywhere, and more and more information is mixed with various suspected bad information, and these information may contain a large number of risky words;

[0003] How to effectively identify the risk information contained in the data information uploaded by users, and after identification, to perform corresponding optimization, so as to reduce the risk of the data information finally displayed, is the problem we need to solve. To this end, a data collection, display and analysis system is now provided. Summary of the invention

[0004] The purpose of the present invention is to provide a data collection, display and analysis system.

[0005] The object of the present invention can be achieved by the following technical solutions: A data acquisition, display and analysis system comprises a monitoring center, wherein the monitoring center is communicatively connected with a data acquisition module, a data processing module, a data analysis module and a data display module;

[0006] The data collection module is used to collect data information uploaded by users;

[0007] The data processing module is used to normalize the data information uploaded by the user, obtain normalized data information, and extract data features in the normalized data information;

[0008] The data analysis module is used to analyze the extracted data features and generate corresponding data labels according to the analysis results;

[0009] The data display module is used to optimize the corresponding part of the data information uploaded by the user according to the generated data tags, and display the optimization results in data.

[0010] Furthermore, the process of collecting the data information uploaded by the user includes:

[0011] Set up a user data upload port. After the user registers his / her identity information, build a corresponding user terminal according to the user's identity registration information, and upload the data information to be uploaded through the user data upload port.

[0012] Constructing a temporary data storage space in the user terminal, and importing the data information uploaded by the user into the temporary data storage space;

[0013] The data information stored in the temporary data storage space is processed by the data processing module.

[0014] Furthermore, the process of normalizing the data information uploaded by the user includes:

[0015] Traversing the data information in the temporary data storage space, obtaining the data type of each piece of data information, generating a corresponding data type label according to the data type corresponding to each piece of data information obtained, and associating the generated data type label with the corresponding data information; wherein the data type includes text type, audio type, video type and image type;

[0016] Aggregate the same type of data information to generate corresponding data information sets;

[0017] According to the data type label in the data information set, the data information in the data information set is preprocessed accordingly;

[0018] The data information in each data information set after preprocessing is spliced ​​to obtain normalized data information.

[0019] Furthermore, the process of preprocessing the data information in the data information set includes:

[0020] The preprocessing process for the data information in the text type data information set is:

[0021] Use regular expressions to identify punctuation marks and non-text special symbols in text-type data information;

[0022] Then divide the data information into corresponding words and phrases;

[0023] After removing stop words from the segmented words and phrases, the words and phrases are morphologically restored;

[0024] The preprocessing process of the data information in the audio type data information set is:

[0025] The data information will be de-noised, and after the de-noising of the data information is completed, the data information will be amplitude normalized and mean-variance normalized;

[0026] Finally, the DC component in the data information is removed;

[0027] The preprocessing process of the data information in the video type data information set includes:

[0028] Converting data information into a number of video frames, and performing rasterization processing on each video frame;

[0029] The rasterized video frame is gray-scaled to obtain a corresponding gray-scale image.

[0030] Furthermore, the process of splicing the data information in each data information set after preprocessing to obtain normalized data information includes:

[0031] Convert each data information in the same data information set into a corresponding data stream segment, generate a type string according to the corresponding data type, splice each data stream segment in the same data information set, obtain an initial data stream of the corresponding data type, and bind and associate the initial data stream with the corresponding type string;

[0032] The initial data streams corresponding to the various data information sets are combined to obtain the corresponding normalized data information.

[0033] Further, the process of extracting data features in the normalized data set:

[0034] Construct a convolutional neural network model and initialize the parameters of the constructed convolutional neural network model;

[0035] Input sample data of different data types into the initialized convolutional neural network model to train the convolutional neural network model, and test the accuracy of the trained convolutional neural network model through the test set until the accuracy of the convolutional neural network model converges, thus completing the training of the convolutional neural network model;

[0036] The normalized data information is input into the trained convolutional neural network model, the convolutional neural network model is used to extract features of the input normalized data information, and all the extracted data features are output, and all the output data features are summarized to obtain the corresponding data feature set.

[0037] Furthermore, the extracted data features are analyzed, and the process of generating corresponding data labels according to the analysis results includes:

[0038] Constructing a semantic feature table, wherein the semantic feature table includes semantic features and label features, and each semantic feature corresponds to at least one label feature;

[0039] A corresponding risk level is set for each label feature, and the label features are divided into risk label features, sub-risk label features, and safety label features according to the risk level;

[0040] When the label feature corresponding to the semantic feature is a risk label feature, the semantic feature is associated with at least one semantic feature whose label feature is a safety label feature;

[0041] Match the data features in the obtained data feature set with each semantic feature in the semantic feature table. If the match is successful, associate the corresponding data feature with the corresponding semantic feature, and obtain the label feature corresponding to the semantic feature;

[0042] If the match is unsuccessful, the corresponding data features are uploaded to the monitoring center, which assigns label features to the data features. After the label features are assigned, the data features and the assigned label features are imported into the semantic feature table to complete the update of the semantic feature table.

[0043] The corresponding risk level is obtained according to the label characteristics, and then the corresponding data label is obtained.

[0044] Furthermore, according to the generated data tags, the corresponding part of the data information uploaded by the user is optimized, and the process of displaying the optimization result as data includes:

[0045] When the data tag is a security tag feature, it means that the corresponding data information is normal and no operation is performed on the data information;

[0046] When the data label is a secondary risk label feature, the corresponding data information is uploaded to the monitoring center, which manually judges the data information and reassigns the safety label feature or risk label feature based on the judgment result;

[0047] When the data label is a risk label feature, it means that the corresponding data information is abnormal, and the data label associated with the semantic feature corresponding to the data information is replaced with a semantic feature of a security label feature, thereby completing the optimization of the data information;

[0048] Visualize the optimized data information.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] The various types of data uploaded by users are normalized, and the data features within the normalized data obtained after the normalization processing are extracted through the trained convolutional neural network model. Different types of data labels are then assigned to the extracted data features through the constructed semantic feature table. According to the assigned labels, it is determined whether the extracted data features are risky, and then whether the corresponding data information is risky. When the data information is risky, the risky data information is replaced by semantic features that do not carry risks, thereby optimizing the data information uploaded by users and reducing the probability of risky content appearing in the data information uploaded by users. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0052] Figure 1 It is a schematic diagram of the present invention. DETAILED DESCRIPTION

[0053] like Figure 1 As shown, a data acquisition, display and analysis system includes a monitoring center, wherein the monitoring center is communicatively connected to a data acquisition module, a data processing module, a data analysis module and a data display module;

[0054] The data collection module is used to collect data information uploaded by users;

[0055] The data processing module is used to normalize the data information uploaded by the user, obtain normalized data information, and extract data features in the normalized data information;

[0056] The data analysis module is used to analyze the extracted data features and generate corresponding data labels according to the analysis results;

[0057] The data display module is used to optimize the corresponding part of the data information uploaded by the user according to the generated data tags, and display the optimization results in data.

[0058] It should be further explained that, in the specific implementation process, the process of the data collection module collecting the data information uploaded by the user includes:

[0059] Set up a user data upload port. After the user registers his / her identity information, build a corresponding user terminal according to the user's identity registration information, and upload the data information to be uploaded through the user data upload port.

[0060] Constructing a temporary data storage space in the user terminal, and importing the data information uploaded by the user into the temporary data storage space;

[0061] The data information stored in the temporary data storage space is processed by the data processing module.

[0062] It should be further explained that, in the specific implementation process, the process of user identity information registration includes:

[0063] The user data upload port is provided with a registration unit and a login unit. The registration unit is used for the user to register identity information. The user uploads personal basic information through the registration unit, and the personal basic information uploaded by the user is sent to the monitoring center, which reviews the personal basic information uploaded by the user. It should be further explained that in the specific implementation process, the personal basic information includes name, gender, age, ID card number and real-name authenticated mobile phone number;

[0064] After the monitoring center has reviewed and approved the personal basic information uploaded by the user, it will generate a corresponding login account and login password, send the generated login account and login password to the user, and upload them to the monitoring center for storage;

[0065] The user enters the received login account and login password into the login unit, and the login account and login password entered by the user into the login unit are uploaded to the monitoring center for verification. After the verification is passed, the account login is completed;

[0066] It should be further explained that, in the specific implementation process, after completing the login, the user can make corresponding changes to the login password in the user terminal, and update the modified login password to the monitoring center.

[0067] It should be further explained that, in the specific implementation process, the process of the data processing module normalizing the data information uploaded by the user includes:

[0068] Traversing the data information in the temporary data storage space, obtaining the data type of each piece of data information, generating a corresponding data type label according to the data type corresponding to each piece of data information obtained, and associating the generated data type label with the corresponding data information; wherein the data type includes text type, audio type, video type and image type;

[0069] Aggregate the same type of data information to generate corresponding data information sets;

[0070] According to the data type label in the data information set, the data information in the data information set is preprocessed accordingly;

[0071] The data information in each data information set after preprocessing is spliced ​​to obtain normalized data information.

[0072] It should be further explained that, in the specific implementation process, the process of preprocessing the data information in the data information set includes:

[0073] The preprocessing process for the data information in the text type data information set is:

[0074] For the punctuation marks and non - text special symbols in the text - type data information through regular expressions; the punctuation marks include full stops, commas, question marks, exclamation marks, quotation marks, brackets, etc., and the special symbols include emoji, etc.;

[0075] Then split the data information into corresponding words and phrases;

[0076] After removing the stop words in the split words and phrases, perform lemmatization on the words and phrases, where the stop words include "de", "le", etc.;

[0077] The pre - processing process for the data information in the audio - type data information set is as follows:

[0078] Remove the noise from the data information, and after completing the noise removal of the data information, perform amplitude normalization and mean - variance normalization on the data information;

[0079] Finally, remove the DC component in the data information;

[0080] The pre - processing process for the data information in the video - type data information set includes:

[0081] Convert the data information into a number of video frames, and perform rasterization processing on each video frame;

[0082] Perform grayscale processing on the rasterized video frames to obtain the corresponding grayscale images.

[0083] It should be further noted that in the specific implementation process, the process of splicing the data information in each pre - processed data information set to obtain the normalized data information includes:

[0084] Convert each data information in the same data information set into corresponding data stream segments, generate a type string according to the corresponding data type, splice each data stream segment in the same data information set to obtain the initial data stream of the corresponding data type, and bind and associate the initial data stream with the corresponding type string;

[0085] Combine the initial data streams corresponding to each data information set to obtain the corresponding normalized data information.

[0086] The process by which the data processing module extracts data features from the normalized data set:

[0087] Construct a convolutional neural network model and perform parameter initialization on the constructed convolutional neural network model;

[0088] Input sample data of different data types into the initialized convolutional neural network model to train the convolutional neural network model, and test the accuracy of the trained convolutional neural network model through the test set until the accuracy of the convolutional neural network model converges, thus completing the training of the convolutional neural network model;

[0089] The normalized data information is input into the trained convolutional neural network model, the convolutional neural network model is used to extract features of the input normalized data information, and all the extracted data features are output, and all the output data features are summarized to obtain the corresponding data feature set.

[0090] The data analysis module analyzes the extracted data features, and the process of generating corresponding data labels according to the analysis results includes:

[0091] Constructing a semantic feature table, wherein the semantic feature table includes semantic features and label features, and each semantic feature corresponds to at least one label feature;

[0092] A corresponding risk level is set for each label feature, and the label features are divided into risk label features, sub-risk label features, and safety label features according to the risk level;

[0093] When the label feature corresponding to the semantic feature is a risk label feature, the semantic feature is associated with at least one semantic feature whose label feature is a safety label feature;

[0094] Match the data features in the obtained data feature set with each semantic feature in the semantic feature table. If the match is successful, associate the corresponding data feature with the corresponding semantic feature, and obtain the label feature corresponding to the semantic feature;

[0095] If the match is unsuccessful, the corresponding data features are uploaded to the monitoring center, which assigns label features to the data features. After the label features are assigned, the data features and the assigned label features are imported into the semantic feature table to complete the update of the semantic feature table.

[0096] The corresponding risk level is obtained according to the label characteristics, and then the corresponding data label is obtained.

[0097] The data display module optimizes the corresponding part of the data information uploaded by the user according to the generated data tags, and displays the optimization results in data, including:

[0098] When the data tag is a security tag feature, it means that the corresponding data information is normal and no operation is performed on the data information;

[0099] When the data label is a secondary risk label feature, the corresponding data information is uploaded to the monitoring center, which manually judges the data information and reassigns the safety label feature or risk label feature based on the judgment result;

[0100] When the data label is a risk label feature, it means that the corresponding data information is abnormal, and the data label associated with the semantic feature corresponding to the data information is replaced with a semantic feature of a security label feature, thereby completing the optimization of the data information;

[0101] Visualize the optimized data information.

[0102] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any modification or equivalent replacement of the above embodiments made according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A data collection, display and analysis system, including a monitoring center, characterized in that: The monitoring center is communicatively connected to a data acquisition module, a data processing module, a data analysis module and a data display module; The data collection module is used to collect data information uploaded by users; The data processing module is used to normalize the data information uploaded by the user, obtain normalized data information, and extract data features in the normalized data information; The data analysis module is used to analyze the extracted data features and generate corresponding data labels according to the analysis results; The data display module is used to optimize the corresponding part of the data information uploaded by the user according to the generated data tags, and display the optimization results in data.

2. A data collection, display and analysis system according to claim 1, characterized in that: The process of collecting data information uploaded by users includes: Set up a user data upload port. After the user registers his / her identity information, build a corresponding user terminal according to the user's identity registration information, and upload the data information to be uploaded through the user data upload port. Constructing a temporary data storage space in the user terminal, and importing the data information uploaded by the user into the temporary data storage space; The data information stored in the temporary data storage space is processed by the data processing module.

3. A data collection, display and analysis system according to claim 2, characterized in that: The process of normalizing the data information uploaded by the user includes: Traversing the data information in the temporary data storage space, obtaining the data type of each piece of data information, generating a corresponding data type label according to the data type corresponding to each piece of data information obtained, and associating the generated data type label with the corresponding data information; wherein the data type includes text type, audio type, video type and image type; Aggregate the same type of data information to generate corresponding data information sets; According to the data type label in the data information set, the data information in the data information set is preprocessed accordingly; The data information in each data information set after preprocessing is spliced ​​to obtain normalized data information.

4. A data collection, display and analysis system according to claim 3, characterized in that: The process of preprocessing the data information in the data information set includes: The preprocessing process for the data information in the text type data information set is: Use regular expressions to identify punctuation marks and non-text special symbols in text-type data information; Then divide the data information into corresponding words and phrases; After removing stop words from the segmented words and phrases, the words and phrases are morphologically restored; The preprocessing process of the data information in the audio type data information set is: The data information will be de-noised, and after the de-noising of the data information is completed, the data information will be amplitude normalized and mean-variance normalized; Finally, the DC component in the data information is removed; The preprocessing process of the data information in the video type data information set includes: Converting data information into a number of video frames, and performing rasterization processing on each video frame; The rasterized video frame is gray-scaled to obtain a corresponding gray-scale image.

5. A data collection, display and analysis system according to claim 4, characterized in that: The process of splicing the data information in each data information set after preprocessing to obtain normalized data information includes: Convert each data information in the same data information set into a corresponding data stream segment, generate a type string according to the corresponding data type, splice each data stream segment in the same data information set, obtain an initial data stream of the corresponding data type, and bind and associate the initial data stream with the corresponding type string; The initial data streams corresponding to the various data information sets are combined to obtain the corresponding normalized data information.

6. A data collection, display and analysis system according to claim 5, characterized in that: The process of extracting data features within a normalized dataset: Construct a convolutional neural network model and initialize the parameters of the constructed convolutional neural network model; Input sample data of different data types into the initialized convolutional neural network model to train the convolutional neural network model, and test the accuracy of the trained convolutional neural network model through the test set until the accuracy of the convolutional neural network model converges, thus completing the training of the convolutional neural network model; The normalized data information is input into the trained convolutional neural network model, the convolutional neural network model is used to extract features of the input normalized data information, and all the extracted data features are output, and all the output data features are summarized to obtain the corresponding data feature set.

7. A data collection, display and analysis system according to claim 6, characterized in that: The process of analyzing the extracted data features and generating corresponding data labels based on the analysis results includes: Constructing a semantic feature table, wherein the semantic feature table includes semantic features and label features, and each semantic feature corresponds to at least one label feature; A corresponding risk level is set for each label feature, and the label features are divided into risk label features, sub-risk label features, and safety label features according to the risk level; When the label feature corresponding to the semantic feature is a risk label feature, the semantic feature is associated with at least one semantic feature whose label feature is a safety label feature; Match the data features in the obtained data feature set with each semantic feature in the semantic feature table. If the match is successful, associate the corresponding data feature with the corresponding semantic feature, and obtain the label feature corresponding to the semantic feature; If the match is unsuccessful, the corresponding data features are uploaded to the monitoring center, which assigns label features to the data features. After the label features are assigned, the data features and the assigned label features are imported into the semantic feature table to complete the update of the semantic feature table. The corresponding risk level is obtained according to the label characteristics, and then the corresponding data label is obtained.

8. A data collection, display and analysis system according to claim 7, characterized in that: The process of optimizing the corresponding part of the data information uploaded by the user according to the generated data tags and displaying the optimization results in data includes: When the data tag is a security tag feature, it means that the corresponding data information is normal and no operation is performed on the data information; When the data label is a secondary risk label feature, the corresponding data information is uploaded to the monitoring center, which manually judges the data information and reassigns the safety label feature or risk label feature based on the judgment result; When the data label is a risk label feature, it means that the corresponding data information is abnormal, and the data label associated with the semantic feature corresponding to the data information is replaced with a semantic feature of a security label feature, thereby completing the optimization of the data information; Visualize the optimized data information.