Material preprocessing method and system based on sensitive content detection
By adopting a combination method of real-time preprocessing and sensitive information identification model in the sensitive content detection system, the problem of low sensitivity identification accuracy in the prior art is solved, efficient identification and removal of sensitive information is achieved, and the security and reliability of the system are improved.
Patent Information
- Application Number
- CN202411140195.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-08-20
AI Technical Summary
In the detection of sensitive content, the prior art has low sensitivity recognition accuracy due to inadequate processing of material data in sensitive content, making it difficult to effectively identify and eliminate sensitive information.
The material preprocessing method and system based on sensitive content detection is adopted to receive and preprocess material data in real time, and combine sensitive information identification models to identify and eliminate, including special processing of text, image and video data.
It effectively reduces the risk of dissemination of sensitive information in the system, improves the security of content, ensures that the system receives material data from high-quality data sources, and enhances the reliability and processing efficiency of the system.
Smart Images

Figure CN119046613B_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a method and system for preprocessing materials based on sensitive content detection, belonging to the technical field of material data processing. Background Art
[0002] Social networks are deeply loved by the majority of Internet users for their convenient and flexible information release and fast and efficient network dissemination methods. However, they also provide room for the spread of sensitive content. In order to create a clean and upright network space and cultivate a positive, healthy, upward and good network culture, it is necessary to use advanced technical means to detect and filter social network content. Sensitive information publishers often distort sensitive words in various ways to avoid detection. Due to the inadequate processing of material data by traditional sensitive content detection methods, there is a problem of low accuracy in identifying the sensitivity of subsequent materials. Summary of the Invention
[0003] The present invention provides a method and system for preprocessing materials based on sensitive content detection to solve the above technical problems in the prior art. The technical solutions adopted are as follows:
[0004] A method for preprocessing materials based on sensitive content detection, the method for preprocessing materials based on sensitive content detection includes:
[0005] Receiving material data from a data source in real time, and preprocessing the material data according to the type of the material data to obtain preprocessed material data;
[0006] Identifying sensitive information in the material data through a sensitive information recognition model, and removing the material data with sensitive information; wherein, the sensitive information recognition model includes a sensitive information recognition model for image recognition and a sensitive information recognition model for text recognition;
[0007] Evaluating the material quality of the data source according to the removal status of the material data, and performing material reception shielding processing on the data source whose material quality does not meet the evaluation requirements.
[0008] Further, receiving material data from a data source in real time, and preprocessing the material data according to the type of the material data to obtain preprocessed material data, includes:
[0009] Receiving material data from a data source in real time;
[0010] Classifying the material data to obtain different types of data information, wherein the different types of data information include text data information, image data information and video data information;
[0011] Perform data processing on the text data information to obtain the text material data after data processing corresponding to the text data information;
[0012] Perform data processing on the image data information to obtain the image material data after data processing corresponding to the image data information;
[0013] Perform data processing on the video data information to obtain the frame image material data after data processing corresponding to the video data information.
[0014] Furthermore, performing data processing on the text data information to obtain the text material data after data processing corresponding to the text data information includes:
[0015] Perform cleaning data processing on the text data information to remove the irrelevant information in the text data information and obtain the cleaned text data information, where the irrelevant information includes HTML tags, special characters, URL links, etc.;
[0016] Use a word segmentation tool to segment the cleaned text data information to obtain the independent words and phrases corresponding to the cleaned text data information;
[0017] Perform stop word removal and unified format conversion on the independent words and phrases corresponding to the cleaned text data information to obtain the text data information after format conversion, and use the text data information after format conversion as the target text data information;
[0018] Among them, the target text data information is the text material data after data processing corresponding to the text data information.
[0019] Furthermore, performing data processing on the image data information to obtain the image material data after data processing corresponding to the image data information includes:
[0020] Perform noise reduction processing on the image data information to obtain the image data information after noise reduction processing;
[0021] Perform gray scale adjustment on the image data information after noise reduction processing to obtain the image data information after gray scale adjustment;
[0022] Perform image segmentation on the image data information after gray scale adjustment to obtain the image data blocks corresponding to different regions;
[0023] Perform feature extraction on the image data blocks to obtain the key features corresponding to each image data block, where the key features include color features, shape features, texture features, etc.;
[0024] Among them, the key feature corresponding to each image data block is the image material data after data processing corresponding to the image data information.
[0025] Further, perform gray-scale adjustment on the image data information after the noise reduction process to obtain the image data information after gray-scale adjustment, including:
[0026] Extract the gray value corresponding to the target pixel block;
[0027] Extract each pixel block included in the image data information after the noise reduction process, and use each pixel block as the target pixel block;
[0028] Extract the pixel blocks adjacent to each pixel block, and use the pixel blocks adjacent to each pixel block as the first selected pixel blocks;
[0029] Extract the gray values of the first selected pixel blocks, and obtain the first gray coefficient by using the gray values of the first selected pixel blocks; among them, the first gray coefficient is obtained through the following formula:
[0030]
[0031] Among them, h 01 represents the first gray coefficient; H m represents the gray value corresponding to the target pixel block; n represents the number of pixel blocks adjacent to the target pixel block; H 01i represents the gray value corresponding to the i-th pixel block adjacent to the target pixel block; exp represents the exponential function symbol with e as the base;
[0032] Extract the pixel blocks adjacent to each first selected pixel block, and exclude the target pixel block from the pixel blocks adjacent to the first selected pixel block. At the same time, use the pixel blocks adjacent to each first selected pixel block excluding the target pixel block as the second selected pixel blocks;
[0033] Extract the gray values of the second selected pixel blocks, and obtain the second gray coefficient by using the gray values of the second selected pixel blocks; among them, the second gray coefficient is obtained through the following formula:
[0034]
[0035] Among them, h 02 represents the second gray coefficient; m represents the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i represents the gray value corresponding to the i-th second selected pixel block; K represents the first adjustment coefficient; among them, the first adjustment coefficient is obtained through the following formula:
[0036]
[0037] where m represents the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i represents the gray value corresponding to the i-th second selected pixel block; H x represents the gray value of the first selected pixel block corresponding to the second selected pixel block;
[0038] Performing gray-scale adjustment on the target pixel block by combining the gray value corresponding to the target pixel block with a first gray-scale coefficient and a second gray-scale coefficient;
[0039] Traversing all target pixel blocks and performing gray-scale adjustment on each target pixel block to obtain image data information after gray-scale adjustment.
[0040] Furthermore, performing gray-scale adjustment on the target pixel block by combining the gray value corresponding to the target pixel block with a first gray-scale coefficient and a second gray-scale coefficient includes:
[0041] Extracting the gray value corresponding to the target pixel block;
[0042] Extracting a first gray-scale coefficient and a second gray-scale coefficient;
[0043] Obtaining a target gray value corresponding to the target pixel block by using the gray value corresponding to the target pixel block, a first gray-scale coefficient, and a second gray-scale coefficient; wherein, the target gray value is obtained through the following formula:
[0044]
[0045] where H f represents the target gray value; h 01 represents the first gray-scale coefficient; h 02 represents the second gray-scale coefficient; H m represents the gray value corresponding to the target pixel block;
[0046] Adjusting the gray value of the target pixel block according to the target gray value to obtain a target pixel block with an adjusted gray value.
[0047] Furthermore, performing image segmentation on the image data information after gray-scale adjustment to obtain image data blocks corresponding to different regions includes:
[0048] Extracting the image data information after gray-scale adjustment;
[0049] Obtaining a first gray-scale threshold and a second gray-scale threshold according to the gray values included in the image data information after gray-scale adjustment; wherein, the first gray-scale threshold and the second gray-scale threshold are obtained through the following formula:
[0050]
[0051] Among them, H y01 represents the first grayscale threshold; H y02 represents the second grayscale threshold; H 01 represents the initial first grayscale threshold; H 02 represents the initial second grayscale threshold; k represents the number of pixel blocks whose target grayscale value is lower than the grayscale value before grayscale adjustment; r represents the number of pixel blocks whose target grayscale value is higher than the grayscale value before grayscale adjustment; H mj represents the grayscale value of the j-th pixel block whose target grayscale value is lower than the grayscale value before grayscale adjustment; H fj represents the target grayscale value of the j-th pixel block whose target grayscale value is lower than the grayscale value before grayscale adjustment; H mt represents the grayscale value of the t-th pixel block whose target grayscale value is higher than the grayscale value before grayscale adjustment; H ft represents the target grayscale value of the t-th pixel block whose target grayscale value is higher than the grayscale value before grayscale adjustment;
[0052] Compare the grayscale value of each pixel block in the image data information after the grayscale adjustment with the first grayscale threshold and the second grayscale threshold;
[0053] Form one or more first image regions with pixel blocks whose grayscale values are lower than the first grayscale threshold;
[0054] Form one or more second image regions with pixel blocks whose grayscale values are not lower than the first grayscale threshold but lower than the second grayscale threshold;
[0055] Form one or more third image regions with pixel blocks whose grayscale values are not lower than the second grayscale threshold;
[0056] Perform image segmentation according to the enclosed regions of the first image region, the second image region and the third image region to obtain image data blocks corresponding to different regions.
[0057] Furthermore, perform data processing on the video data information to obtain the frame image material data corresponding to the video data information, including:
[0058] Perform frame processing on the video data information to obtain the frame image data set corresponding to the video data information;
[0059] Obtain the frame interval number range for the frame image data; among them, the upper limit value and the lower limit value corresponding to the frame interval number range are obtained through the following formula:
[0060]
[0061] Among them, M up and Mdown represents the upper limit value and the lower limit value corresponding to the range of the number of frame intervals; x represents the number of unit times corresponding to the reception of video data, and the value of the unit time is 2s - 5s; M c represents a preset reference value of the number of frames; M p represents the average number of frames corresponding to the unit time; M i represents the number of frames of the frame data included in the video data information received corresponding to the i-th unit time; M d0 and M u0 represents the upper limit value and the lower limit value of the number corresponding to the preset initial frame interval range; s represents a second adjustment coefficient, and the second adjustment coefficient is obtained through the following formula:
[0062]
[0063] where s represents the second adjustment coefficient; M z represents the median of the number of frames included in the received video data corresponding to n unit times; M max and M min represent the maximum value and the minimum value of the number of frames included in the received video data corresponding to the unit time;
[0064] Randomly set the number of frame intervals that conform to the range of the number of frame intervals according to the range of the number of frame intervals, and extract frame images from the frame image dataset according to the number of frame intervals to obtain target frame image data;
[0065] Perform noise reduction processing on the target frame image data to obtain the target frame image data after noise reduction processing;
[0066] Perform gray level adjustment on the target frame image data after noise reduction processing to obtain the target frame image data after gray level adjustment;
[0067] Perform image segmentation on the target frame image data after gray level adjustment to obtain image data blocks corresponding to different regions of the target frame image data;
[0068] Extract features for the image data blocks corresponding to the target frame image data to obtain the key features included in the image data blocks corresponding to each target frame image data, where the key features include color features, shape features, texture features, etc.;
[0069] Among them, the key features corresponding to the image data blocks corresponding to each target frame image data are the image material data after data processing corresponding to the image data information.
[0070] Further, evaluate the material quality of the data source according to the material data rejection status, and perform material reception shielding processing on the data source whose material quality does not meet the evaluation requirements, including:
[0071] Extract the rejection status information of the material data of each data source;
[0072] Obtain the material data quality evaluation parameter corresponding to each data source according to the rejection status information of the material data; wherein, the material data quality evaluation parameter is obtained by the following formula:
[0073]
[0074] where Q represents the material data quality evaluation parameter; p represents the number of transmission unit times experienced by the data source to send the material data, and the transmission unit time is 1s; P i represents the data proportion of the sensitive information data in the total transmitted data in the i-th transmission unit time; P c represents the preset data proportion threshold; z represents the number of transmission unit times in which the data proportion of the sensitive information data in the total transmitted data in p unit times does not exceed the preset data proportion threshold; y represents the number of transmission unit times in which the data proportion of the sensitive information data in the total transmitted data in p unit times exceeds the preset data proportion threshold; P j represents the data proportion of the sensitive information data corresponding to the transmission unit time in which the data proportion of the j-th sensitive information data in the total transmitted data does not exceed the preset data proportion threshold; P g represents the data proportion of the sensitive information data corresponding to the transmission unit time in which the data proportion of the g-th sensitive information data in the total transmitted data exceeds the preset data proportion threshold;
[0075] Compare the material data quality evaluation parameter with the preset parameter threshold;
[0076] When the material data quality evaluation parameter is lower than the preset parameter threshold, it is determined that the material quality of the data source does not meet the evaluation requirements;
[0077] Perform material reception shielding processing on the data source whose material quality does not meet the evaluation requirements.
[0078] A material preprocessing system based on sensitive content detection, the material preprocessing system based on sensitive content detection includes:
[0079] A material data receiving module, configured to receive material data from a data source in real time, and preprocess the material data according to the type of the material data to obtain preprocessed material data;
[0080] A sensitive information recognition module is used to recognize sensitive information in the material data through a sensitive information recognition model and eliminate the material data with sensitive information; among them, the sensitive information recognition model includes a sensitive information recognition model for image recognition and a sensitive information recognition model for text recognition;
[0081] A material quality evaluation module is used to evaluate the quality of the data source according to the elimination status of the material data and perform material reception shielding processing on the data source whose material quality does not meet the evaluation requirements.
[0082] Advantages of the present invention:
[0083] The material preprocessing method and system based on sensitive content detection proposed by the present invention effectively reduce the risk of the spread of sensitive information in the system and improve the content security by receiving and preprocessing the material data in real time and combining with a sensitive information recognition model for accurate recognition and elimination. The data source quality evaluation and shielding mechanism can ensure that the system receives material data from high-quality data sources, reduce the risk of problems caused by the system receiving low-quality data, and enhance the reliability of the system. Adopting corresponding preprocessing methods and sensitive information recognition models for different types of material data enables the system to process a large amount of material data more quickly and accurately, improving the processing efficiency. This technical solution is based on modular design and can flexibly adjust and optimize the functions and performances of each module according to actual needs, with strong scalability. At the same time, with the continuous development of technology, the system can continuously introduce new preprocessing methods and sensitive information recognition models to adapt to the changing network environment and user needs. Description of the drawings
[0084] Figure 1 is a flowchart of the method of the present invention;
[0085] Figure 2 is a system block diagram of the system of the present invention. Detailed implementation manners
[0086] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0087] The embodiment of the present invention proposes a material preprocessing method based on sensitive content detection, as Figure 1 shown, the material preprocessing method based on sensitive content detection includes:
[0088] S1. Receive material data from a data source in real time and preprocess the material data according to the type of the material data to obtain preprocessed material data;
[0089] S2. Use the sensitive information recognition model to identify sensitive information in the material data and remove the material data with sensitive information. Among them, the sensitive information recognition model includes a sensitive information recognition model for image recognition and a sensitive information recognition model for text recognition.
[0090] S3. Evaluate the material quality of the data source according to the status of material data removal, and perform material reception shielding processing on the data source whose material quality does not meet the evaluation requirements.
[0091] The working principle of the above technical solution is as follows: The system receives material data from the data source in real time. The above data may include various forms such as images, videos, and texts. According to the type of material data, the system performs corresponding preprocessing on the received data. For images, it may include operations such as scaling, grayscaling, and denoising; for texts, it may include word segmentation, stop word removal, and part-of-speech tagging. The above preprocessing steps aim to improve the efficiency and accuracy of subsequent sensitive information recognition.
[0092] The preprocessed material data is sent to the sensitive information recognition model for further analysis. The model includes two sensitive information recognition models for images and texts. The image recognition model may adopt deep learning techniques, such as convolutional neural networks (CNNs), to detect sensitive content in images, such as porn and violence.
[0093] The text recognition model may adopt natural language processing techniques, such as word segmentation and part-of-speech tagging, combined with machine learning or deep learning algorithms, to identify sensitive words or sensitive topics in the text. Once sensitive information is detected, the system will remove the above material data with sensitive information to ensure that the subsequent processed materials are safe and compliant.
[0094] The system evaluates the quality of the data source according to the status of material data removal. Specifically, the proportion of sensitive information materials removed from the materials sent by each data source can be counted as an evaluation index for the quality of the data source. For the data source whose quality does not meet the evaluation requirements, the system will take shielding measures to suspend or terminate receiving material data from the data source. This helps to reduce the risk of sensitive information entering the system and improve the overall security and compliance of the system.
[0095] The effects of the above technical solution are as follows: By receiving and preprocessing material data in real time, and combining with a sensitive information recognition model for accurate recognition and elimination, the risk of sensitive information spreading in the system is effectively reduced, and the security of the content is improved. The data source quality assessment and shielding mechanism can ensure that the system receives material data from high-quality data sources, reducing the risk of problems caused by the system receiving low-quality data and enhancing the reliability of the system. Appropriate preprocessing methods and sensitive information recognition models are adopted for different types of material data, enabling the system to process a large amount of material data more quickly and accurately, and improving the processing efficiency. This technical solution is based on modular design, and the functions and performance of each module can be flexibly adjusted and optimized according to actual needs, with strong scalability. At the same time, with the continuous development of technology, new preprocessing methods and sensitive information recognition models can be continuously introduced into the system to adapt to the changing network environment and user requirements.
[0096] In one embodiment of the present invention, material data is received in real time from a data source, and the material data is preprocessed according to the type of the material data to obtain preprocessed material data, including:
[0097] S101. Receive material data in real time from a data source;
[0098] S102. Classify the material data to obtain different types of data information, where the different types of data information include text data information, image data information, and video data information;
[0099] S103. Perform data processing on the text data information to obtain text material data after data processing corresponding to the text data information;
[0100] S104. Perform data processing on the image data information to obtain image material data after data processing corresponding to the image data information;
[0101] S105. Perform data processing on the video data information to obtain frame image material data after data processing corresponding to the video data information.
[0102] The working principle of the above technical solution is as follows: The system or software receives material data in real time from a specified data source (such as a database, API interface, file system, etc.). Real-time reception ensures the timeliness and freshness of the data, enabling subsequent processing to be based on the latest data. The received material data is classified into different types at this step, including text data information, image data information, and video data information. Data classification provides a basis for subsequent specialized processing of different types of data.
[0103] For text data information, the system performs specific data processing operations, such as denoising, word segmentation, part-of-speech tagging, named entity recognition, sentiment analysis, etc. The above processing helps to extract key information from the text or change its format to meet the needs of subsequent applications.
[0104] For image data information, the system may perform image enhancement, denoising, feature extraction, target detection, image segmentation, etc. The above processing can improve image quality, extract key features, or convert the image into a form more suitable for subsequent analysis or application.
[0105] For video data information, processing usually involves video frame extraction, key frame selection, video transcoding, video feature extraction, etc. Video data is usually converted into frame image material data for frame-by-frame analysis or processing.
[0106] The effect of the above technical solution is: through data cleaning, denoising and other processing, the accuracy and reliability of the data are improved, providing a better foundation for subsequent data analysis or application. Data classification and specialized processing enable different types of data to be used more effectively. For example, text data can be used for natural language processing tasks, image data can be used for image recognition or analysis, and video data can be used for video analysis and processing. Real-time reception and processing of data ensures the timeliness and rapid response of data. In addition, specialized processing for different types of data can improve processing efficiency and reduce unnecessary consumption of computing resources. The technical solution is capable of processing multiple types of data, including text, images and videos, and can therefore support a wide range of application scenarios, such as natural language processing, image recognition, video analysis, and the like. The various steps in the technical solution can be used independently or in combination, and can be flexibly configured according to specific needs. In addition, with the emergence of new data processing technologies, the solution can also be easily expanded to support new data types or processing requirements.
[0107] In one embodiment of the present invention, data processing is performed on text data information to obtain text material data after data processing corresponding to the text data information, including:
[0108] S1031, performing data cleaning processing on the text data information to remove irrelevant information in the text data information and obtain cleaned text data information, wherein the irrelevant information includes HTML tags, special characters and URL links, etc.;
[0109] S1032, using a word segmentation tool to segment the cleaned text data information to obtain independent words and phrases corresponding to the cleaned text data information;
[0110] S1033. Remove stop words from and perform unified format conversion on the independent words and phrases corresponding to the text data information after cleaning, obtain the text data information after format conversion, and use the text data information after format conversion as the target text data information;
[0111] Among them, the target text data information is the text material data after data processing corresponding to the text data information.
[0112] The working principle of the above technical solution is as follows: The system will clean the received text data information, aiming to remove the irrelevant information therein, such as HTML tags, special characters, and URL links, etc. The above irrelevant information may cause interference or is unnecessary in subsequent text processing. Therefore, cleaning can improve the quality of text data.
[0113] The text data information after cleaning will be sent to a word segmentation tool for text segmentation. The word segmentation tool will segment the continuous text into independent words and phrases. This step is a basic step in natural language processing and provides a basis for subsequent tasks such as part-of-speech tagging and named entity recognition.
[0114] Next, the system will remove stop words from the segmented words and phrases. Stop words usually refer to those words with high occurrence frequencies but little contribution to the text meaning, such as "de" (of), "shi" (is), etc. Removing the above words can reduce the redundancy of text data.
[0115] At the same time, the system will also perform unified format conversion on the words and phrases, such as converting all words to lowercase, removing punctuation marks, etc. The purpose of this step is to unify the format of text data and facilitate subsequent analysis and processing.
[0116] The text data obtained after the above steps of processing, that is, the text data after cleaning, segmentation, stop word removal, and format conversion, is called the target text data information. This is the text material data after data processing corresponding to the text data information and can be used for subsequent natural language processing tasks or applications.
[0117] The effects of the above technical solution are as follows: By cleaning and format conversion, irrelevant information and redundant parts in the text are removed, making the text data cleaner and more standardized, and improving the quality of the text data. After word segmentation and stop word removal, the text data is segmented into independent words and phrases, and words that contribute little to the meaning of the text are removed, making subsequent natural language processing tasks more efficient and accurate. The processed text data can be used as the input for natural language processing tasks such as text classification, sentiment analysis, and keyword extraction, supporting multiple application scenarios. Operations such as cleaning, word segmentation, stop word removal, and format conversion are common steps in text processing. The above operations can improve processing efficiency through algorithm optimization and parallel computing, making text data processing faster. This technical solution can be flexibly configured according to specific requirements, such as adding new cleaning rules and replacing word segmentation tools. At the same time, with the continuous development of natural language processing technology, this solution can also be easily extended to support more text processing tasks and application scenarios.
[0118] In one embodiment of the present invention, data processing is performed on image data information to obtain image material data after data processing corresponding to the image data information, including:
[0119] S1041. Perform noise reduction processing on the image data information to obtain the image data information after noise reduction processing;
[0120] S1042. Perform gray level adjustment on the image data information after noise reduction processing to obtain the image data information after gray level adjustment;
[0121] S1043. Perform image segmentation on the image data information after gray level adjustment to obtain image data blocks corresponding to different regions;
[0122] S1044. Perform feature extraction on the image data blocks to obtain key features corresponding to each image data block, where the key features include color features, shape features, texture features, etc.;
[0123] Among them, the key features corresponding to each image data block are the image material data after data processing corresponding to the image data information.
[0124] The working principle of the above technical solution is as follows: Remove the noise in the image, that is, the interference information that is not desired to exist in the image, such as random pixel points, color fluctuations, etc. Noise reduction processing is usually implemented through digital filters or algorithms, such as Gaussian filtering, median filtering, etc., to improve the clarity and quality of the image.
[0125] Gray level adjustment is to adjust the brightness or contrast of the image, making the image clearer visually or meeting specific requirements. This is usually achieved by stretching the image histogram, applying gamma correction, etc., to improve the visual effect of the image.
[0126] Image segmentation is the process of dividing an image into multiple regions or objects. This is usually achieved based on the differences in certain features (such as color, brightness, texture, etc.) in the image. The result of image segmentation is to divide the image into multiple regions with similar characteristics, and each region corresponds to a block of image data.
[0127] For each block of image data, the system performs feature extraction operations to obtain the key features of that region. The above key features usually include color features (such as color histograms, color moments, etc.), shape features (such as edges, contours, region areas, etc.), and texture features (such as gray-level co-occurrence matrices, wavelet transforms, etc.). The above features can describe the visual characteristics of the image data block and provide a basis for subsequent applications (such as image recognition, object detection, etc.).
[0128] The effects of the above technical solutions are as follows: Through noise reduction processing and gray-scale adjustment, the clarity and visual effects of the image are improved, providing a better basis for subsequent image analysis and processing. Image segmentation and feature extraction operations can extract the key regions and features in the image, and the above regions and features are crucial for subsequent tasks such as image recognition and object detection. The processed image material data can be used in various application scenarios, such as face recognition, license plate recognition, medical image analysis, industrial automation, etc. Operations such as noise reduction, gray-scale adjustment, image segmentation, and feature extraction can usually improve the processing efficiency through algorithm optimization and parallel computing, making the image data processing faster and more efficient. This technical solution can be flexibly configured according to specific requirements, such as adding new noise reduction methods, adjusting gray-scale adjustment parameters, improving image segmentation algorithms, etc. At the same time, with the continuous development of image processing technology, this solution can also be easily extended to support more image processing and analysis tasks.
[0129] In one embodiment of the present invention, gray-scale adjustment is performed on the image data information after the noise reduction processing to obtain the image data information after gray-scale adjustment, including:
[0130] Step A1: Extract the gray value corresponding to the target pixel block;
[0131] Step A2: Extract each pixel block included in the image data information after the noise reduction processing, and use each pixel block as the target pixel block;
[0132] Step A3: Extract the pixel blocks adjacent to each pixel block, and use the pixel blocks adjacent to each pixel block as the first selected pixel blocks;
[0133] Step A4: Extract the gray value of the first selected pixel block, and obtain the first gray coefficient by using the gray value of the first selected pixel block; wherein, the first gray coefficient is obtained through the following formula:
[0134]
[0135] Among them, h 01 represents the first gray coefficient; H m represents the gray value corresponding to the target pixel block; n represents the number of pixel blocks adjacent to the target pixel block; H 01i represents the gray value corresponding to the i-th pixel block adjacent to the target pixel block; exp represents the exponential function symbol with base e;
[0136] Step A5: Extract the pixel blocks adjacent to each first selected pixel block, and exclude the target pixel block from the pixel blocks adjacent to the first selected pixel block. At the same time, use the pixel blocks adjacent to each first selected pixel block excluding the target pixel block as the second selected pixel blocks;
[0137] Step A6: Extract the gray values of the second selected pixel blocks, and obtain the second gray coefficient by using the gray values of the second selected pixel blocks; among them, the second gray coefficient is obtained through the following formula:
[0138]
[0139] Among them, h 02 represents the second gray coefficient; m represents the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i represents the gray value corresponding to the i-th second selected pixel block; K represents the first adjustment coefficient; among them, the first adjustment coefficient is obtained through the following formula:
[0140]
[0141] Among them, m represents the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i represents the gray value corresponding to the i-th second selected pixel block; H x represents the gray value of the first selected pixel block corresponding to the second selected pixel block;
[0142] Step A7: Perform gray adjustment on the target pixel block by using the gray value corresponding to the target pixel block in combination with the first gray coefficient and the second gray coefficient;
[0143] Step A8: Traverse all target pixel blocks, and perform gray adjustment on each target pixel block to obtain the image data information after gray adjustment.
[0144] The working principle of the above technical solution is as follows: First, extract the gray value of each pixel block (or a single pixel if a single pixel is being processed) from the image after noise reduction processing.
[0145] Next, each pixel block or pixel in the image is taken as the target pixel block and is prepared for grayscale adjustment.
[0146] For each target pixel block, the pixel blocks adjacent to it (i.e., neighbor pixel blocks) are extracted, and a certain relationship (through an exponential function) between the grayscale values of the above neighbor pixel blocks and the grayscale value of the target pixel block is calculated. This relationship is used to calculate the first grayscale coefficient, which reflects the grayscale relationship between the target pixel block and its direct neighbors.
[0147] Then, the pixel blocks adjacent to each neighbor pixel block (i.e., second-level neighbors) are further considered, excluding the target pixel block itself. The grayscale values of the above second-level neighbors are used to calculate the second grayscale coefficient, which reflects the grayscale relationship between the target pixel block and its indirect neighbors. A first adjustment coefficient is also involved in this process, which takes into account the relationship between the grayscale values of the second-level neighbors and the direct neighbors.
[0148] Using the original grayscale value of the target pixel block, combined with the first grayscale coefficient and the second grayscale coefficient, the grayscale adjustment of the target pixel block is performed. The above process takes into account the grayscale relationship between the target pixel block and its direct and indirect neighbors, thereby obtaining an adjusted grayscale value.
[0149] Finally, all pixel blocks or pixels in the image are traversed, and the above grayscale adjustment process is repeated for each pixel block or pixel, thereby obtaining the grayscale adjustment result of the entire image.
[0150] The effects of the above technical solution are as follows: By considering the grayscale relationship between pixel blocks and their neighbors, this technology can more finely adjust the grayscale values of the image, enabling the image to have a better visual effect while maintaining details. Since the grayscale adjustment is based on the grayscale values of pixel blocks (or pixels) and their neighbors, this technology has local adaptability and can perform different adjustments according to the grayscale characteristics of different regions. Parameters such as the first adjustment coefficient involved in the technology can be adjusted according to specific requirements to adapt to different image processing needs. This grayscale adjustment technology can be further extended, for example, by introducing more neighbor pixel block levels, using more complex grayscale relationship calculation methods, etc., to improve the accuracy and effect of grayscale adjustment. Grayscale adjustment is one of the basic operations in image processing, and this technology can be applied to various scenarios that require improving the grayscale distribution of images, such as image enhancement, feature extraction, target detection, etc.
[0151] Meanwhile, by calculating the grayscale coefficient in two steps, the grayscale value of the target pixel block can be adjusted more precisely, making the grayscale of the final image more in line with expectations. This step-by-step adjustment method helps to improve the meticulousness and accuracy of image processing. Adjusting using the grayscale information of adjacent pixel blocks can ensure the grayscale consistency within the local area of the image, reducing the unnaturalness caused by grayscale mutations. This method helps to improve the overall quality and visual effect of the image. By combining the first grayscale coefficient and the second grayscale coefficient, the effects of the two adjustments can be effectively utilized, further improving the grayscale adjustment accuracy and consistency of the image. This superimposed effect can optimize the quality of the final image. At the same time, the ability of dynamic adjustment is provided, enabling this method to be flexibly adjusted according to the characteristics of specific images, thus adapting to the processing requirements of different types of images.
[0152] Through the calculation of the above formula, the grayscale value of each pixel block can be accurately obtained. The calculations such as the average value and difference involved in the formula ensure the accuracy of the result. An adjustment coefficient is introduced in the second grayscale adjustment formula. By dynamically adjusting the first grayscale adjustment result, the second grayscale adjustment can more flexibly adapt to the characteristics of different images, thus achieving a better adjustment effect. At the same time, the weighted average and difference calculations of the grayscale information of adjacent pixel blocks in the above formula make the adjustment result not only consider the grayscale of the current pixel block but also comprehensively consider the grayscale information of the surrounding pixel blocks. This method can improve the overall consistency and naturalness of the image, reducing the unnaturalness caused by grayscale mutations.
[0153] Meanwhile, by calculating the grayscale coefficient in two steps, each adjustment gradually refines the calculation of the grayscale value, making the final result more accurate and natural. The first adjustment provides a good foundation for the second adjustment. The way of layer-by-layer fine adjustment can gradually optimize the image quality. On the other hand, during the calculation process, by averaging and weighting the grayscale values, random noise in the image can be effectively suppressed, enhancing the smoothness and quality of the image. This provides an effective solution to the problem of noise interference in image processing. The exponential function and average value calculation involved in the formula ensure the stability and robustness of the adjustment result. Even when processing different types of images, the formula can still maintain a good adjustment effect.
[0154] Generally speaking, the above formula can achieve high-precision grayscale adjustment through accurate calculation and step-by-step adjustment methods, improve the overall quality and consistency of the image, and have strong dynamic adaptation ability and noise suppression effect.
[0155] In an embodiment of the present invention, the grayscale adjustment of the target pixel block is performed by using the grayscale value corresponding to the target pixel block in combination with the first grayscale coefficient and the second grayscale coefficient, including:
[0156] Step A701: Extract the grayscale value corresponding to the target pixel block;
[0157] Step A702: Extract the first grayscale coefficient and the second grayscale coefficient;
[0158] Step A703: Obtain the target grayscale value corresponding to the target pixel block by using the grayscale value corresponding to the target pixel block, the first grayscale coefficient, and the second grayscale coefficient; wherein, the target grayscale value is obtained through the following formula:
[0159]
[0160] wherein, H f represents the target grayscale value; h 01 represents the first grayscale coefficient; h 02 represents the second grayscale coefficient; H m represents the grayscale value corresponding to the target pixel block;
[0161] Step A704: Adjust the grayscale value of the target pixel block according to the target grayscale value to obtain the target pixel block with the adjusted grayscale value.
[0162] The working principle of the above technical solution is as follows: First, extract the original grayscale value (Hm) corresponding to the target pixel block from the image data.
[0163] Then, according to the results calculated in the previous steps, extract the first grayscale coefficient (h01) and the second grayscale coefficient (h02).
[0164] Use a specific formula to combine the original grayscale value (Hm) of the target pixel block, the first grayscale coefficient (h01), and the second grayscale coefficient (h02) to calculate the target grayscale value (Hf). The above formula considers the grayscale relationship between the target pixel block and its direct neighbors and indirect neighbors, and adjusts the original grayscale value by weighting these two coefficients.
[0165] Use the calculated target grayscale value (Hf) to update the grayscale value of the target pixel block. This usually means setting the grayscale value of the target pixel block to the newly calculated target grayscale value, thus completing the grayscale adjustment.
[0166] The effects of the above technical solution are as follows: By considering the gray - level relationship between the target pixel block and its neighbors, this technology can enhance the local contrast of the image. If the gray - level value difference between the target pixel block and its neighbors is large, then the adjusted gray - level value may be more obvious, thus improving the visual quality of the image. Since this technology calculates based on the gray - level values of the pixel block (or pixel) and its neighbors, it can retain the detail information of the image to a certain extent. This is particularly useful for application scenarios that require preserving tiny details or textures in the image. The calculation methods of the first gray - level coefficient and the second gray - level coefficient can be adjusted according to specific requirements to adapt to different image - processing needs. In addition, since this technology calculates based on the gray - level values of the pixel block (or pixel), it can be flexibly applied to images of different resolutions and sizes. Through gray - level adjustment, the visual effect of the image can be significantly improved. For example, in some cases, the brightness and contrast of the image may be more uniform, making the image look clearer and easier to observe. Gray - level adjustment is a basic operation in image processing, and this technology can be applied to various scenarios, such as image enhancement, image restoration, object detection, etc. By adjusting the gray - level value, the visual effect of the image can be improved, providing a better basis for subsequent processing and analysis.
[0167] Meanwhile, by introducing the first gray - level coefficient (h01) and the second gray - level coefficient (h02), this technical solution allows for flexible adjustment of the gray - level value of the target pixel block. These two coefficients can be set according to specific application scenarios and requirements, thereby achieving fine control of the gray - level distribution of the image. This flexibility enables this technical solution to adapt to various image - processing needs, such as enhancing image contrast, improving image brightness, etc. By adjusting the gray - level coefficients, the gray - level dynamic range of the image can be effectively expanded or compressed. This technical solution performs gray - level adjustment on the target pixel block, which means it can perform fine processing on local regions of the image. This local - processing ability is very useful for enhancing or weakening specific regions (such as dark regions, bright regions, or edge regions) in the image, contributing to improving the overall quality of the image. At the same time, the basic framework of this technical solution (i.e., extracting gray - level values, gray - level coefficients, and applying formulas to calculate the target gray - level value) provides the possibility for further expansion and customization. For example, more gray - level coefficients can be introduced to implement more complex gray - level adjustment strategies, or this method can be combined with other image - processing technologies to achieve more advanced image - processing effects.
[0168] In one embodiment of the present invention, image segmentation is performed on the image data information after gray - level adjustment to obtain image data blocks corresponding to different regions, including:
[0169] Step B1: Extract the image data information after gray - level adjustment;
[0170] Step B2. Obtain a first gray-scale threshold and a second gray-scale threshold according to the gray-scale values included in the image data information after gray-scale adjustment; wherein, the first gray-scale threshold and the second gray-scale threshold are obtained through the following formula:
[0171]
[0172] wherein, H y01 represents the first gray-scale threshold; H y02 represents the second gray-scale threshold; H 01 represents the initial first gray-scale threshold; H 02 represents the initial second gray-scale threshold; k represents the number of pixel blocks whose target gray-scale values are lower than the gray-scale values before gray-scale adjustment; r represents the number of pixel blocks whose target gray-scale values are higher than the gray-scale values before gray-scale adjustment; H mj represents the gray-scale value of the jth pixel block whose target gray-scale value is lower than the gray-scale value before gray-scale adjustment; H fj represents the target gray-scale value of the jth pixel block whose target gray-scale value is lower than the gray-scale value before gray-scale adjustment; H mt represents the gray-scale value of the tth pixel block whose target gray-scale value is higher than the gray-scale value before gray-scale adjustment; H ft represents the target gray-scale value of the tth pixel block whose target gray-scale value is higher than the gray-scale value before gray-scale adjustment;
[0173] Step B3. Compare the gray-scale value of each pixel block in the image data information after gray-scale adjustment with the first gray-scale threshold and the second gray-scale threshold;
[0174] Step B4. Form one or more first image regions with the pixel blocks whose gray-scale values are lower than the first gray-scale threshold;
[0175] Step B5. Form one or more second image regions with the pixel blocks whose gray-scale values are not lower than the first gray-scale threshold but lower than the second gray-scale threshold;
[0176] Step B6. Form one or more third image regions with the pixel blocks whose gray-scale values are not lower than the second gray-scale threshold;
[0177] Step B7. Perform image segmentation according to the regions enclosed by the first image region, the second image region, and the third image region to obtain image data blocks corresponding to different regions.
[0178] The working principle of the above technical solution is as follows: First, extract image data information from the image that has been subjected to gray-scale adjustment, and the above information includes the gray-scale value of each pixel block.
[0179] According to the gray value distribution of the image data information after gray scale adjustment, two gray scale thresholds are calculated: the first gray scale threshold (Hy01) and the second gray scale threshold (Hy02). The calculation of these two thresholds takes into account the change of the target gray value relative to the gray value before gray scale adjustment and the distribution of the above changes. The initial thresholds (H01 and H02) can be set according to the characteristics or experience of the image, and then adjusted by iterative or adaptive methods.
[0180] The gray value of each pixel block in the image after gray scale adjustment is compared with the two thresholds to determine the gray scale range to which each pixel block belongs.
[0181] According to the result of the gray value comparison, the pixel blocks with gray values lower than the first gray scale threshold are combined into one or more first image regions.
[0182] The pixel blocks with gray values between the first gray scale threshold and the second gray scale threshold are combined into one or more second image regions.
[0183] The pixel blocks with gray values higher than the second gray scale threshold are combined into one or more third image regions.
[0184] Based on the image regions with different gray scale ranges formed above, the image is segmented to obtain image data blocks corresponding to different regions. The above image data blocks represent the regions in the image with different gray scale characteristics.
[0185] The effects of the above technical solutions are as follows: The image segmentation method based on gray scale thresholds usually has high computational efficiency and can quickly divide the image into different regions. Since a clear gray scale threshold is used for segmentation, the segmentation result has strong interpretability and is convenient for subsequent analysis and processing. The calculation of the gray scale threshold takes into account the change of the gray value of the pixel block before and after gray scale adjustment, so this segmentation method can adapt to images with different gray scale distribution characteristics. The initial gray scale threshold can be adjusted according to the specific application scenario to meet different segmentation requirements. In addition, the segmentation process can be further refined as needed, for example, by introducing more gray scale thresholds to form more image regions. This image segmentation technology can be applied to a variety of image processing and analysis tasks, such as object detection, feature extraction, image classification, etc. By segmenting the image regions with different gray scale characteristics, more useful information can be provided for subsequent tasks.
[0186] On the other hand, by dynamically calculating the first grayscale threshold (Hy01) and the second grayscale threshold (Hy02), this technical solution can adaptively adjust the threshold according to the image data information after grayscale adjustment. This adaptability enables the threshold to more accurately reflect the actual changes in the image content, thereby improving the accuracy and robustness of image segmentation. When calculating the threshold, this technical solution not only considers the initial grayscale thresholds (H01 and H02), but also considers the changes in the grayscale values of pixel blocks before and after grayscale adjustment (such as Hmj and Hfj, Hmt and Hft). This consideration makes the determination of the threshold more in line with the characteristics of the image after grayscale adjustment, contributing to more precisely segmenting different regions in the image.
[0187] By setting two grayscale thresholds, this technical solution can divide the image into three different regions (the first image region, the second image region, and the third image region), thus achieving more refined image segmentation. This refined segmentation is helpful for subsequent image analysis and processing tasks, such as object detection, feature extraction, etc. Since the determination of the threshold is based on the overall image data information after grayscale adjustment, this technical solution has a certain robustness to noise and local outliers in the image. This means that even if there are some undesirable grayscale value changes in the image, this technical solution can still relatively accurately segment different regions in the image.
[0188] At the same time, the basic framework of this technical solution (i.e., grayscale adjustment, threshold calculation, pixel block comparison, and region division) provides the possibility for further expansion and optimization. For example, more complex threshold calculation methods can be introduced, combined with other image features for segmentation, etc., to meet different image processing requirements.
[0189] In one embodiment of the present invention, data processing is performed on video data information to obtain frame image material data corresponding to the video data information, including:
[0190] Step C1: Perform frame processing on the video data information to obtain a frame image data set corresponding to the video data information;
[0191] Step C2: Obtain the frame interval number range for the frame image data; wherein, the upper limit value and the lower limit value corresponding to the frame interval number range are obtained through the following formula:
[0192]
[0193] wherein, M up and M down represent the upper limit value and the lower limit value corresponding to the frame interval number range; x represents the number of unit times corresponding to the reception of video data, and the value of the unit time is 2s - 5s; M c represents a preset frame number reference value; Mp represents the average number of frames corresponding to a unit time; M i represents the number of frames of the frame data included in the video data information received corresponding to the i-th unit time; M d0 and M u0 represent the upper limit value and the lower limit value corresponding to the preset initial frame interval quantity range; s represents the second adjustment coefficient, and the second adjustment coefficient is obtained through the following formula:
[0194]
[0195] wherein, s represents the second adjustment coefficient; M z represents the median of the number of frames contained in the received video data corresponding to n unit times; M max and M min represent the maximum value and the minimum value of the number of frames contained in the received video data corresponding to a unit time;
[0196] Step C3: Randomly set the frame interval quantity that conforms to the frame interval quantity range according to the frame interval quantity range, and extract frame images from the frame image data set according to the frame interval quantity to obtain target frame image data;
[0197] Step C4: Perform noise reduction processing on the target frame image data to obtain the target frame image data after noise reduction processing;
[0198] Step C5: Perform gray scale adjustment on the target frame image data after noise reduction processing to obtain the target frame image data after gray scale adjustment;
[0199] Step C6: Perform image segmentation on the target frame image data after gray scale adjustment to obtain image data blocks corresponding to different regions of the target frame image data;
[0200] Step C7: Extract features for the image data blocks corresponding to the target frame image data to obtain the key features included in the image data blocks corresponding to each target frame image data, wherein the key features include color features, shape features, texture features, etc.;
[0201] wherein, the key features corresponding to the image data blocks corresponding to each target frame image data are the image material data after data processing corresponding to the image data information.
[0202] The working principle of the above technical solution is as follows: First, perform frame processing on the video data information, that is, split the continuous video data stream into independent frame images to form a frame image data set.
[0203] Next, according to the number of frames within the unit time of video data reception, a reasonable range of frame interval numbers is calculated. The upper and lower limit values of the above range are obtained through a series of complex calculations, considering factors such as the average number of frames of the video, the median, maximum, and minimum number of frames within the unit time, as well as the initial range of frame interval numbers, and are dynamically adjusted using the second adjustment coefficient s.
[0204] After determining the range of frame interval numbers, a frame interval number that meets the above range is randomly selected, and the corresponding number of frame images is extracted from the frame image dataset to form the target frame image data.
[0205] Noise reduction processing is performed on the extracted target frame image data to reduce noise interference in the image and improve the image quality.
[0206] Gray level adjustment is performed on the target frame image data after noise reduction processing, which may include enhancing the contrast, brightness, etc. of the image to further improve the visualization effect of the image.
[0207] The target frame image data after gray level adjustment is subjected to image segmentation, and it is divided into different regions, with each region corresponding to an image data block.
[0208] Feature extraction is performed on each image data block to obtain key features including color features, shape features, and texture features, etc. The above key features will be used as the image material data after image processing for subsequent analysis or applications.
[0209] The effects of the above technical solution are as follows: Through frame processing and frame image extraction, key frame images can be quickly extracted from video data, improving the efficiency of data processing. Through steps such as noise reduction processing, gray level adjustment, and image segmentation, the quality of the image is improved, and the accuracy of feature extraction is increased. The determination of the range of frame interval numbers considers various factors and uses a dynamic adjustment method, enabling this technical solution to adapt to video data of different qualities and contents. By extracting various features including color, shape, and texture, rich image material data is obtained, providing more possibilities for subsequent analysis or applications. This technical solution can be extended and optimized according to specific requirements, such as adding new feature extraction methods, adjusting the determination method of the range of frame interval numbers, etc., to meet different application requirements.
[0210] On the other hand, by dynamically calculating the range of the number of frame intervals (Mup and Mdown), this technical solution can flexibly adjust the number of frame intervals according to the actual situation of video data information (such as the number of frames received per unit time, the maximum and minimum values of the number of frames, etc.), so as to achieve effective extraction of video frames. This flexibility enables this technical solution to adapt to the processing requirements of video data with different frame rates and different scenarios. By setting a reasonable range of the number of frame intervals and randomly extracting frame images within this range, this technical solution can effectively reduce the number of frame images to be processed while ensuring data representativeness, thereby optimizing data processing efficiency. This is particularly important for large-scale video data processing tasks.
[0211] The noise reduction processing and gray-scale adjustment steps help to improve the quality of the target frame image data. The noise reduction processing can reduce the noise interference in the image and make the image clearer; the gray-scale adjustment can adjust parameters such as the brightness and contrast of the image according to needs, making the image more in line with the requirements of subsequent processing or analysis. The image segmentation step divides the target frame image data into different regions and extracts the key features of each region. This fine image segmentation helps subsequent feature extraction and image analysis tasks and can more accurately capture the useful information in the image. This technical solution not only extracts basic image features such as color features and shape features, but also considers more advanced features such as texture features. This comprehensive feature extraction can provide rich information support for subsequent image recognition, classification and other tasks. The basic framework of this technical solution (i.e., frame processing, calculation of the range of the number of frame intervals, extraction of frame images, noise reduction processing, gray-scale adjustment, image segmentation and feature extraction) provides the possibility for further expansion and customization. For example, more complex noise reduction algorithms, gray-scale adjustment strategies or feature extraction methods can be introduced to adapt to different application scenarios and requirements. By considering multiple factors such as the average number of frames per unit time, the maximum and minimum values of the number of frames to calculate the range of the number of frame intervals, this technical solution has a certain robustness to outliers or noise in the video data. This helps to ensure that relatively reliable data processing results can be obtained even when the quality of the video data is not high.
[0212] In one embodiment of the present invention, the material quality of the data source is evaluated according to the material data rejection status, and the data source whose material quality does not meet the evaluation requirements is subjected to material reception shielding processing, including:
[0213] S301. Extract the rejection status information of the material data of each data source;
[0214] S302. Obtain the material data quality evaluation parameters corresponding to each data source according to the rejection status information of the material data; wherein, the material data quality evaluation parameters are obtained through the following formula:
[0215]
[0216] Among them, Q represents the evaluation parameter of the quality of material data; p represents the number of transmission unit time periods experienced by the data source when sending material data, and the transmission unit time is 1 s; P i represents the data proportion of sensitive information data in the total transmitted data in the i-th transmission unit time period; P c represents the preset data proportion threshold; z represents the number of transmission unit time periods in which the data proportion of sensitive information data in p unit time periods does not exceed the preset data proportion threshold; y represents the number of transmission unit time periods in which the data proportion of sensitive information data in p unit time periods exceeds the preset data proportion threshold; P j represents the data proportion of sensitive information data corresponding to the transmission unit time period in which the data proportion of the j-th sensitive information data does not exceed the preset data proportion threshold; P g represents the data proportion of sensitive information data corresponding to the transmission unit time period in which the data proportion of the g-th sensitive information data exceeds the preset data proportion threshold;
[0217] S303. Compare the material data quality evaluation parameter with a preset parameter threshold;
[0218] S304. When the material data quality evaluation parameter is lower than the preset parameter threshold, it is determined that the material quality of the data source does not meet the evaluation requirements;
[0219] S305. Perform material reception shielding processing on the data source whose material quality does not meet the evaluation requirements.
[0220] The working principle of the above technical solution is: Extract the rejection status information of the material data from the data source. The rejection status information usually reflects whether the material data contains sensitive information or non-compliant content during transmission or processing, and the rejection situation of the above content.
[0221] Based on the rejection status information, calculate the material data quality evaluation parameter (Q) corresponding to each data source. The calculation of this parameter takes into account the proportion of sensitive information data in the transmission unit time period (such as 1 second) experienced by the data source when sending material data, and compares it with the preset data proportion threshold (Pc). Each term in the formula reflects the different impacts of the proportion of sensitive information data in different time periods.
[0222] Compare the calculated material data quality evaluation parameter (Q) with the preset parameter threshold. The above parameter threshold represents the lower limit of acceptable data quality.
[0223] If the quality evaluation parameter (Q) of the material data is lower than the preset parameter threshold, it is determined that the quality of the material from the data source does not meet the evaluation requirements. This means that the material data sent by the data source contains too much sensitive information or content that does not meet the requirements, and needs to be processed.
[0224] For data sources whose material quality does not meet the evaluation requirements, material reception shielding processing is performed. This usually means that the system or platform will no longer receive the material data sent by the data source, or perform special processing (such as filtering, marking, etc.) on the above data to ensure the quality of the data used subsequently.
[0225] The effects of the above technical solution are as follows: By evaluating the quality of the material data sent by the data source, it can be ensured that the quality of the data received by the system or platform reaches a certain standard, thereby improving the accuracy and reliability of subsequent processing. By calculating the proportion of sensitive information data and setting corresponding thresholds, the risk of sensitive information leakage can be effectively reduced. For data sources containing too much sensitive information, shielding processing can avoid potential security threats posed by the above sensitive information to the system or platform. Shielding data sources that do not meet the evaluation requirements can reduce the amount of invalid data processed by the system or platform, thereby improving the efficiency and accuracy of data processing. The preset parameter threshold and data ratio threshold can be adjusted according to actual needs to adapt to different application scenarios and data quality requirements. This makes the technical solution highly configurable and flexible.
[0226] On the other hand, by comprehensively considering the proportion of sensitive information data in the process of data sources sending material data and whether these proportions exceed a preset threshold, this technical solution can relatively accurately evaluate the quality of material data of each data source. This accuracy helps to accurately identify those data sources that frequently send excessive amounts of sensitive information, thus ensuring the reliability of subsequent data processing and applications. This technical solution uses a dynamic time window (i.e., the sending unit time) to evaluate the quality of material data of data sources in different time periods. This dynamic adaptability enables the evaluation results to better reflect the actual performance of data sources, especially in cases where the performance of data sources is unstable or the data quality fluctuates greatly. The formulas involved in the calculation process are relatively simple, mainly based on proportion and counting operations, which makes the entire evaluation process have high computational efficiency. At the same time, by quickly comparing with the preset parameter thresholds, it can be quickly determined which data sources' material quality does not meet the evaluation requirements, and then shielding measures can be taken to improve the processing efficiency. The parameters in the technical solution (such as the preset data proportion threshold, parameter threshold, etc.) can be adjusted according to actual needs to adapt to different evaluation criteria and application scenarios. This flexibility enables this technical solution to be widely applied to various material data quality evaluation and data source management scenarios. By comprehensively considering the proportion of sensitive information data in multiple sending unit times, this technical solution has a certain robustness to individual outliers or noises. Even if the quality of the material data sent by a data source is poor in a certain time period, but if the overall performance still meets the evaluation criteria, then this data source will not be wrongly shielded. The entire evaluation process realizes automation and intelligence, and can automatically complete the evaluation of material data quality and the shielding process of data sources without manual intervention. This reduces the cost and error rate of manual operations and improves work efficiency and accuracy.
[0227] In summary, this technical solution has technical effects such as accuracy, dynamic adaptability, high efficiency, flexibility, robustness, and automation and intelligence in material quality evaluation and data source management, and can effectively improve the quality and efficiency of data processing.
[0228] An embodiment of the present invention proposes a material preprocessing system based on sensitive content detection, such as Figure 2 shown, the material preprocessing system based on sensitive content detection includes:
[0229] A material data receiving module, configured to receive material data from a data source in real time, and preprocess the material data according to the type of the material data to obtain preprocessed material data;
[0230] A sensitive information recognition module, which is used to recognize sensitive information in the material data through a sensitive information recognition model and eliminate the material data with sensitive information; among them, the sensitive information recognition model includes a sensitive information recognition model for image recognition and a sensitive information recognition model for text recognition;
[0231] A material quality evaluation module, which is used to evaluate the material quality of the data source according to the elimination status of the material data, and perform material reception shielding processing on the data source whose material quality does not meet the evaluation requirements.
[0232] The working principle of the above technical solution is as follows: The system receives material data from the data source in real time, and the above data may include various forms such as images, videos, and texts. According to the type of the material data, the system performs corresponding preprocessing on the received data. For images, it may include operations such as scaling, grayscale conversion, and denoising; for texts, it may include word segmentation, stop word removal, and part-of-speech tagging. The above preprocessing steps are aimed at improving the efficiency and accuracy of subsequent sensitive information recognition.
[0233] The preprocessed material data is sent to the sensitive information recognition model for further analysis. The model includes two sensitive information recognition models for images and texts. The image recognition model may adopt deep learning technologies, such as convolutional neural networks (CNNs), to detect sensitive content in images, such as porn and violence.
[0234] The text recognition model may adopt natural language processing technologies, such as word segmentation and part-of-speech tagging, combined with machine learning or deep learning algorithms, to identify sensitive words or sensitive topics in the text. Once sensitive information is detected, the system will eliminate the above material data with sensitive information to ensure that the materials for subsequent processing are safe and compliant.
[0235] The system evaluates the quality of the data source according to the elimination status of the material data. Specifically, the proportion of the sensitive information materials eliminated in the materials sent by each data source can be counted, and this is used as an evaluation index for the quality of the data source. For the data source whose quality does not meet the evaluation requirements, the system will take shielding measures to suspend or terminate receiving material data from this data source. This helps to reduce the risk of sensitive information entering the system and improve the overall security and compliance of the system.
[0236] The effects of the above technical solution are as follows: By receiving and preprocessing material data in real time and combining with a sensitive information recognition model for accurate recognition and elimination, the risk of sensitive information spreading in the system is effectively reduced, and the security of the content is improved. The data source quality assessment and shielding mechanism can ensure that the system receives material data from high-quality data sources, reducing the risk of problems caused by the system receiving low-quality data and enhancing the reliability of the system. Adopting corresponding preprocessing methods and sensitive information recognition models for different types of material data enables the system to process a large amount of material data more quickly and accurately, improving the processing efficiency. This technical solution is based on modular design and can flexibly adjust and optimize the functions and performance of each module according to actual needs, with strong scalability. At the same time, with the continuous development of technology, the system can continuously introduce new preprocessing methods and sensitive information recognition models to adapt to the changing network environment and user needs.
[0237] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if the above modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include the above modifications and variations.
Claims
1. A material preprocessing method based on sensitive content detection, characterized in that: The material preprocessing method based on sensitive content detection includes: Material data is received from a data source in real time, and the material data is preprocessed according to the type of the material data to obtain the preprocessed material data; wherein the type of the material data includes text data information, image data information and video data information; and the data processing for the image data information includes: obtaining a first grayscale coefficient by using the grayscale value of a first selected pixel block connected to the pixel block after the image data information is segmented; obtaining a second grayscale coefficient by using the grayscale value of a second selected pixel block connected to the first selected pixel block excluding the target pixel block; obtaining a target grayscale value of the target pixel block by using the first grayscale coefficient and the second grayscale coefficient; wherein the first grayscale coefficient is obtained by the following formula: in, h 01 represents the first grayscale coefficient; H m Indicates the grayscale value corresponding to the target pixel block; n Indicates the number of pixel blocks connected to the target pixel block; H 01i Indicates i The grayscale value corresponding to the pixel block connected to the target pixel block; exp represents the grayscale value corresponding to the pixel block connected to the target pixel block; e The symbol of the exponential function with base ; And, the second grayscale coefficient is obtained by the following formula: in, h 02 represents the second grayscale coefficient; m Indicates the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i Indicates i A grayscale value corresponding to a second selected pixel block; K represents the first adjustment coefficient; wherein the first adjustment coefficient is obtained by the following formula: in, m Indicates the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i Indicates i A grayscale value corresponding to a second selected pixel block; H x Indicates the grayscale value of the first selected pixel block corresponding to the second selected pixel block; At the same time, the target grayscale value is obtained by the following formula: in, H f Indicates the target gray value; h 01 represents the first grayscale coefficient; h 02 represents the second grayscale coefficient; H m Indicates the grayscale value corresponding to the target pixel block; Performing sensitive information identification on the material data through a sensitive information identification model, and removing the material data with sensitive information; wherein the sensitive information identification model includes a sensitive information identification model for image identification and a sensitive information identification model for text identification; The material quality of the data source is evaluated according to the material data rejection status, and the data source whose material quality does not meet the evaluation requirements is subjected to material reception shielding processing.
2. The material preprocessing method based on sensitive content detection according to claim 1 is characterized in that: Receiving material data from a data source in real time, and preprocessing the material data according to the type of the material data to obtain the preprocessed material data, including: Receive material data from data source in real time; Classifying the material data to obtain different types of data information, wherein the different types of data information include text data information, image data information and video data information; Performing data processing on the text data information to obtain text material data after data processing corresponding to the text data information; Performing data processing on the image data information to obtain image material data after data processing corresponding to the image data information; Data processing is performed on the video data information to obtain frame image material data after data processing corresponding to the video data information.
3. The material preprocessing method based on sensitive content detection according to claim 2 is characterized in that: Data processing is performed on the text data information to obtain text material data corresponding to the text data information after data processing, including: Performing data cleaning processing on the text data information to remove irrelevant information in the text data information and obtain cleaned text data information, wherein the irrelevant information includes HTML tags, special characters and URL links; Using a word segmentation tool to segment the cleaned text data information to obtain independent words and phrases corresponding to the cleaned text data information; Performing stop word extraction and unified format conversion on the independent words and phrases corresponding to the cleaned text data information, obtaining the text data information after the format conversion, and using the text data information after the format conversion as the target text data information; The target text data information is the text material data after data processing corresponding to the text data information.
4. The material preprocessing method based on sensitive content detection according to claim 2 is characterized in that: Performing data processing on the image data information to obtain image material data corresponding to the image data information after data processing includes: Performing noise reduction processing on the image data information to obtain the image data information after the noise reduction processing; Performing grayscale adjustment on the image data information after the noise reduction processing to obtain the image data information after the grayscale adjustment; Performing image segmentation on the grayscale-adjusted image data information to obtain image data blocks corresponding to different areas; Performing feature extraction on the image data blocks to obtain key features corresponding to each image data block, wherein the key features include color features, shape features, and texture features; The key feature corresponding to each image data block is the image material data after data processing corresponding to the image data information.
5. The material preprocessing method based on sensitive content detection according to claim 4 is characterized in that: Performing grayscale adjustment on the image data information after the noise reduction process to obtain the image data information after the grayscale adjustment includes: Extracting the grayscale value corresponding to the target pixel block; Extracting each pixel block contained in the image data information after the noise reduction process, and taking each pixel block as a target pixel block; Extracting a pixel block connected to each pixel block, and taking the pixel block connected to each pixel block as a first selected pixel block; Extracting the grayscale value of the first selected pixel block, and obtaining a first grayscale coefficient using the grayscale value of the first selected pixel block; Extracting pixel blocks connected to each first selected pixel block, excluding the target pixel block from the pixel blocks connected to the first selected pixel blocks, and taking the pixel blocks connected to each first selected pixel block excluding the target pixel block as second selected pixel blocks; Extracting the grayscale value of the second selected pixel block, and obtaining a second grayscale coefficient using the grayscale value of the second selected pixel block; Performing grayscale adjustment on the target pixel block by using the grayscale value corresponding to the target pixel block in combination with the first grayscale coefficient and the second grayscale coefficient; All target pixel blocks are traversed, and grayscale adjustment is performed on each target pixel block to obtain image data information after grayscale adjustment.
6. The material preprocessing method based on sensitive content detection according to claim 5 is characterized in that: The grayscale value corresponding to the target pixel block is combined with the first grayscale coefficient and the second grayscale coefficient to adjust the grayscale of the target pixel block, including: Extracting the grayscale value corresponding to the target pixel block; extracting a first grayscale coefficient and a second grayscale coefficient; Obtaining a target grayscale value corresponding to the target pixel block by using the grayscale value corresponding to the target pixel block and the first grayscale coefficient and the second grayscale coefficient; The grayscale value of the target pixel block is adjusted according to the target grayscale value to obtain the target pixel block after the grayscale value is adjusted.
7. The material preprocessing method based on sensitive content detection according to claim 4 is characterized in that: Performing image segmentation on the grayscale-adjusted image data information to obtain image data blocks corresponding to different regions includes: Extracting image data information after grayscale adjustment; According to the grayscale value contained in the grayscale-adjusted image data information, a first grayscale threshold and a second grayscale threshold are obtained; wherein the first grayscale threshold and the second grayscale threshold are obtained by the following formula: in, H y01 represents the first grayscale threshold; H y02 represents the second grayscale threshold; H 01 represents the initial first grayscale threshold; H 02 represents the initial second grayscale threshold; k Indicates the number of pixel blocks whose target grayscale value is lower than the grayscale value before grayscale adjustment; r Indicates the number of pixel blocks whose target grayscale value is higher than the grayscale value before grayscale adjustment; H mj Indicates j The grayscale value of a pixel block whose target grayscale value is lower than the grayscale value before grayscale adjustment; H fj Indicates j The target grayscale value of the pixel block whose target grayscale value is lower than the grayscale value before grayscale adjustment; H mt Indicates t The grayscale value of a pixel block whose target grayscale value is higher than the grayscale value before grayscale adjustment; H ft Indicates t The target grayscale value of the pixel block whose target grayscale value is higher than the grayscale value before grayscale adjustment; Comparing the grayscale value of each pixel block in the grayscale-adjusted image data information with the first grayscale threshold and the second grayscale threshold; Pixel blocks with grayscale values lower than the first grayscale threshold form one or more first image areas; Pixel blocks whose grayscale values are not lower than the first grayscale threshold but lower than the second grayscale threshold form one or more second image regions; The pixel blocks whose grayscale values are not lower than the second grayscale threshold form one or more third image areas; Image segmentation is performed according to the areas enclosed by the first image area, the second image area and the third image area to obtain image data blocks corresponding to different areas.
8. The material preprocessing method based on sensitive content detection according to claim 2, characterized in that: Performing data processing on the video data information to obtain frame image material data corresponding to the video data information after data processing includes: Performing frame processing on the video data information to obtain a frame image data set corresponding to the video data information; A frame interval quantity range is obtained for the frame image data; wherein the upper limit quantity value and the lower limit quantity value corresponding to the frame interval quantity range are obtained by the following formula: in, M up and M down Indicates the upper limit and lower limit of the frame interval range. x Indicates the number of unit times corresponding to the reception of video data, and the value of the unit time is 2s-5s; M c Indicates the preset frame number reference value; M p Indicates the average number of frames per unit time; M i Indicates i The number of frames of frame data contained in the received video data information corresponding to each unit time; M d0 and M u0 represents the upper limit value and the lower limit value of the preset initial frame interval quantity range; s represents the second adjustment coefficient, and the second adjustment coefficient is obtained by the following formula: in, s represents the second adjustment factor; M z express n The median value of the number of frames contained in the received video data corresponding to a unit time; M max and M min Indicates the maximum and minimum number of frames contained in the received video data corresponding to a unit time; Randomly setting a frame interval number that meets the frame interval number range according to the frame interval number range, and extracting frame images from the frame image data set according to the frame interval number to obtain target frame image data; Performing noise reduction processing on the target frame image data to obtain the target frame image data after the noise reduction processing; Performing grayscale adjustment on the target frame image data after the noise reduction processing to obtain the target frame image data after the grayscale adjustment; Performing image segmentation on the grayscale-adjusted target frame image data to obtain image data blocks corresponding to the target frame image data corresponding to different regions; Performing feature extraction on the image data blocks corresponding to the target frame image data to obtain key features contained in the image data blocks corresponding to each target frame image data, wherein the key features include color features, shape features and texture features; The key feature corresponding to the image data block corresponding to each target frame image data is the image material data after data processing corresponding to the image data information.
9. The material preprocessing method based on sensitive content detection according to claim 1, characterized in that: The data source is evaluated for material quality according to the material data rejection status, and a material receiving shielding process is performed on the data source whose material quality does not meet the evaluation requirements, including: Extract the rejection status information of the material data of each data source; The material data quality evaluation parameter corresponding to each data source is obtained according to the rejection status information of the material data; wherein the material data quality evaluation parameter is obtained by the following formula: in, Q Indicates material data quality evaluation parameters; p Indicates the number of sending unit times that the data source takes to send material data, and the sending unit time is 1s; P i Indicates i The proportion of sensitive information data to the total data sent in a sending unit time; P c Indicates the preset data ratio threshold; z express p The number of sending unit times in which the proportion of sensitive information data in the total data sent does not exceed a preset data proportion threshold; y express p The number of sending unit times in which the proportion of sensitive information data in the total sent data exceeds a preset data proportion threshold; P j Indicates j The data ratio of sensitive information data to the total data sent does not exceed the preset data ratio threshold value for the sending unit time corresponding to the sensitive information data to the total data sent; P g Indicates g The data ratio of sensitive information data to the total data sent exceeds the preset data ratio threshold value during the sending unit time; Comparing the material data quality evaluation parameter with a preset parameter threshold; When the material data quality evaluation parameter is lower than a preset parameter threshold, it is determined that the material quality of the data source does not meet the evaluation requirements; Data sources whose material quality does not meet the evaluation requirements are blocked from receiving materials.
10. A material preprocessing system based on sensitive content detection, characterized in that: The material preprocessing system based on sensitive content detection includes: The material data receiving module is used to receive material data from a data source in real time, and pre-process the material data according to the type of the material data to obtain the pre-processed material data; wherein the type of the material data includes text data information, image data information and video data information; and the data processing for the image data information includes: obtaining a first grayscale coefficient by using the grayscale value of the first selected pixel block connected to the pixel block after the image data information is segmented; obtaining a second grayscale coefficient by using the grayscale value of the second selected pixel block connected to the first selected pixel block excluding the target pixel block; obtaining a target grayscale value of the target pixel block by using the first grayscale coefficient and the second grayscale coefficient; wherein the first grayscale coefficient is obtained by the following formula: in, h 01 represents the first grayscale coefficient; H m Indicates the grayscale value corresponding to the target pixel block; n Indicates the number of pixel blocks connected to the target pixel block; H 01i Indicates i The grayscale value corresponding to the pixel block connected to the target pixel block; exp represents the grayscale value corresponding to the pixel block connected to the target pixel block; e The symbol of the exponential function with base ; And, the second grayscale coefficient is obtained by the following formula: in, h 02 represents the second grayscale coefficient; m Indicates the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i Indicates i A grayscale value corresponding to a second selected pixel block; K represents the first adjustment coefficient; wherein the first adjustment coefficient is obtained by the following formula: in, m Indicates the number of second selected pixel blocks corresponding to each first selected pixel block; H 02i Indicates i A grayscale value corresponding to a second selected pixel block; H x Indicates the grayscale value of the first selected pixel block corresponding to the second selected pixel block; At the same time, the target grayscale value is obtained by the following formula: in, H f Indicates the target gray value; h 01 represents the first grayscale coefficient; h 02 represents the second grayscale coefficient; H m Indicates the grayscale value corresponding to the target pixel block; A sensitive information identification module, used to identify sensitive information of the material data through a sensitive information identification model, and remove the material data with sensitive information; wherein the sensitive information identification model includes a sensitive information identification model for image identification and a sensitive information identification model for text identification; The material quality assessment module is used to assess the material quality of the data source according to the material data rejection status, and to perform material reception shielding processing on the data source whose material quality does not meet the assessment requirements.
Citation Information
Patent Citations
Network information detection method and system based on big data
CN118300851A