Data cleaning and transcription method, device, electronic device and storage medium
Through a multi-step data cleaning method, including format cleaning, preset template comparison and text content recognition, the problem of cleaning sensitive information and malicious content in the data is solved, data security and compliance are achieved, and data leakage is prevented.
Patent Information
- Application Number
- CN202510616414.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Existing technologies are difficult to effectively clean sensitive information and malicious content embedded in data, especially when facing 0-day attacks and APT attacks. Traditional methods cannot completely remove malicious information in extended attributes, resulting in a high risk of data leakage.
Through multi-step cleaning methods such as format cleaning, comparison between preset outlines and chapter templates, text content cleaning, OCR and ASR model recognition, namespace and business logic tag tree structure inspection, combined with virtual display and frequency domain processing, data security and compliance are ensured.
It effectively removes potential risk information from the data, ensures that the file structure and content meet security standards, prevents data leakage and the spread of malicious content, and maintains the integrity and availability of the files required for business.
Smart Images

Figure CN120124107B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data cleaning and transcription method, device, electronic device and storage medium. Background Art
[0002] With the continuous development of network technology, network security protection has become increasingly important. Traditional network attacks have gradually transformed from traditional network attacks to 0-day attacks or other APT attacks, and the form of attacks has also changed from traffic attacks to data attacks. Data-based attacks mainly focus on data implantation, and the output data may also be leaked through entrainment.
[0003] Data cleaning can effectively avoid the security risks brought by implanted data. Therefore, how to perform data cleaning more effectively has become an urgent problem to be solved in the industry. Summary of the Invention
[0004] The present invention provides a data cleaning and transcription method, device, electronic device and storage medium, which are used to solve the defects of the prior art in how to more effectively perform data cleaning.
[0005] The present invention provides a data cleaning and transcription method, comprising:
[0006] When receiving a file to be cleaned in a document format, performing format cleaning on the file to be cleaned to obtain a first file to be cleaned after format cleaning; wherein the format cleaning includes: cleaning metadata and extended attribute data of the file to be cleaned;
[0007] Performing outline and chapter cleaning on the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned;
[0008] The second file to be cleaned is subjected to text content cleaning to obtain a cleaned file in a document format.
[0009] According to a data cleaning and transcription method provided by the present invention, the first file to be cleaned is cleaned in terms of outline and chapters according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned, including:
[0010] Comparing the content outline of the first file to be cleaned with the preset outline template, deleting the outline and outline reference content in the content outline of the first file to be cleaned that are inconsistent with the preset outline template, to obtain a third file to be cleaned;
[0011] Compare the content chapters of the third file to be cleaned with the chapter template, delete the chapters and chapter references in the chapter template of the third file to be cleaned that are inconsistent with the chapter template, and obtain a cleaned second file to be cleaned.
[0012] According to a data cleaning and transcription method provided by the present invention, the text content of the second file to be cleaned is cleaned to obtain a cleaned file in a document format, including:
[0013] After capturing the virtual display content of the second file to be cleaned by using a simulation display technology, the virtual display content of the second file to be cleaned is converted into first YUV data;
[0014] Through a preset optical character recognition model, text recognition is performed on the first YUV data to obtain a cleaned file in a document format.
[0015] According to a data cleaning and transcription method provided by the present invention, the text content of the second file to be cleaned is cleaned to obtain a cleaned file in a document format, including:
[0016] Capturing audio data corresponding to the text in the second file to be cleaned by using virtual playback technology;
[0017] The audio data is subjected to speech recognition through an automatic speech recognition model, and a cleaned file in a document format is obtained according to the speech recognition content.
[0018] According to a data cleaning and transcription method provided by the present invention, the method further comprises:
[0019] When a fourth file to be cleaned in XML format is received, obtaining a first namespace of each element in the fourth file to be cleaned;
[0020] Removing elements in each of the first namespaces that do not conform to a preset namespace template to obtain a cleaned fifth file to be cleaned;
[0021] The fifth file to be cleaned is cleaned of contents that are inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned.
[0022] According to a data cleaning and transcription method provided by the present invention, after the step of cleaning the fifth file to be cleaned for content that is inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned, the method further includes:
[0023] Cleaning the attribute value, attribute field, and attribute relationship of each tag in the sixth file to be cleaned according to a preset tag attribute template to obtain a seventh file to be cleaned;
[0024] After capturing the virtual display content of the seventh file to be cleaned through virtual display technology, the virtual display content of the seventh file to be cleaned is converted into second YUV data; through a preset optical character recognition model, text recognition is performed on the second YUV data to obtain a cleaned file in XML format.
[0025] According to a data cleaning and transcription method provided by the present invention, the method further includes:
[0026] When an eighth file to be cleaned in a picture format is received, the least significant position of each pixel in the eighth file to be cleaned is set to zero to obtain a ninth file to be cleaned;
[0027] superimposing preset spatial noise onto the pixel matrix of the ninth file to be cleaned to obtain a tenth file to be cleaned;
[0028] Using discrete cosine transform or fast Fourier transform, converting the tenth file to be cleaned from the spatial domain to the frequency domain to obtain frequency domain data of the tenth file to be cleaned;
[0029] The preset noise is superimposed on the frequency domain data, and the superimposed frequency domain data is restored back to the spatial domain to obtain a cleaned file in an image format.
[0030] The present invention also provides a data cleaning and transcription device, comprising:
[0031] A first cleaning module is configured to, upon receiving a file to be cleaned in a document format, perform format cleaning on the file to be cleaned to obtain a first file to be cleaned after the format cleaning; wherein the format cleaning includes: cleaning metadata and extended attribute data of the file to be cleaned;
[0032] A second cleaning module is configured to clean the outline and chapters of the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned;
[0033] The third cleaning module is used to clean the text content of the second file to be cleaned to obtain a cleaned file in a document format.
[0034] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0035] Comparing the content outline of the first file to be cleaned with the preset outline template, deleting the outline and outline reference content in the content outline of the first file to be cleaned that are inconsistent with the preset outline template, to obtain a third file to be cleaned;
[0036] Compare the content chapters of the third file to be cleaned with the chapter template, delete the chapters and chapter references in the chapter template of the third file to be cleaned that are inconsistent with the chapter template, and obtain a cleaned second file to be cleaned.
[0037] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0038] After capturing the virtual display content of the second file to be cleaned by using a simulation display technology, the virtual display content of the second file to be cleaned is converted into first YUV data;
[0039] Through a preset optical character recognition model, text recognition is performed on the first YUV data to obtain a cleaned file in a document format.
[0040] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0041] Capturing audio data corresponding to the text in the second file to be cleaned by using virtual playback technology;
[0042] The audio data is subjected to speech recognition through an automatic speech recognition model, and a cleaned file in a document format is obtained according to the speech recognition content.
[0043] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0044] When a fourth file to be cleaned in XML format is received, obtaining a first namespace of each element in the fourth file to be cleaned;
[0045] Removing elements in each of the first namespaces that do not conform to a preset namespace template to obtain a cleaned fifth file to be cleaned;
[0046] The fifth file to be cleaned is cleaned of contents that are inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned.
[0047] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0048] Cleaning the attribute value, attribute field, and attribute relationship of each tag in the sixth file to be cleaned according to a preset tag attribute template to obtain a seventh file to be cleaned;
[0049] After capturing the virtual display content of the seventh file to be cleaned through virtual display technology, the virtual display content of the seventh file to be cleaned is converted into second YUV data; through a preset optical character recognition model, text recognition is performed on the second YUV data to obtain a cleaned file in XML format.
[0050] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0051] When an eighth file to be cleaned in a picture format is received, the least significant position of each pixel in the eighth file to be cleaned is set to zero to obtain a ninth file to be cleaned;
[0052] superimposing preset spatial noise onto the pixel matrix of the ninth file to be cleaned to obtain a tenth file to be cleaned;
[0053] Using discrete cosine transform or fast Fourier transform, converting the tenth file to be cleaned from the spatial domain to the frequency domain to obtain frequency domain data of the tenth file to be cleaned;
[0054] The preset noise is superimposed on the frequency domain data, and the superimposed frequency domain data is restored back to the spatial domain to obtain a cleaned file in an image format.
[0055] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any of the above-described data cleaning and transcription methods is implemented.
[0056] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described data cleaning and transcription methods.
[0057] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned data cleaning and transcription methods.
[0058] The data cleaning and transcription method, device, electronic device and storage medium provided by the present invention receive document format files uploaded by users, parse the metadata of the files, remove or reset unnecessary metadata items, check and clear the extended attribute data of the files, and ensure the purity of the file format. The cleaned file is compared with the preset outline template, and the content that does not conform to the outline structure is deleted. Each chapter is checked in detail, and the parts that do not conform to the chapter template requirements are deleted or modified, so as to achieve comprehensive cleaning of the document file and ensure the security and compliance of the data. Format cleaning removes external impurities of the file, outline and chapter cleaning ensures the standardization of the file structure, and text content cleaning goes deep into the core content of the file and removes potential risk information. The final cleaned and transcribed file meets security standards while maintaining the integrity and availability required by the business, effectively preventing data leakage and the spread of malicious content. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0060] Figure 1 Schematic diagram of data entrainment leakage in related technologies;
[0061] Figure 2 It is a flow chart of the data cleaning and transcription method provided by the present invention;
[0062] Figure 3 A schematic diagram illustrating the namespace template provided by the present invention;
[0063] Figure 4 This is an illustration of the tag tree template provided by the present invention;
[0064] Figure 5 This is a schematic diagram illustrating the tag attribute template provided by the present invention;
[0065] Figure 6 This is a schematic diagram of image format data cleaning provided by the present invention;
[0066] Figure 7 This is one of the file cleaning schematics provided by the present invention;
[0067] Figure 8 The second schematic diagram of file cleaning provided by the present invention;
[0068] Figure 9 The third schematic diagram of file cleaning provided by the present invention;
[0069] Figure 10 This is the fourth file cleaning diagram provided by the present invention;
[0070] Figure 11 This is the fifth file cleaning diagram provided by the present invention;
[0071] Figure 12 The template data transcription and cleaning principle diagram provided by the present invention;
[0072] Figure 13 A schematic diagram of the structure of the data cleaning and transcription device provided by the present invention;
[0073] Figure 14 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0074] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0075] In related technologies, malware detection and killing software based on features and behaviors cannot detect or completely remove extended attributes and hidden unknown malicious information, and cannot block APT's long-term data-based attack links.
[0076] Information leakage prevention tools based on DLP can search the internal content of output files to prevent the leakage of sensitive content. However, various implants in extended attributes, especially when the data is specifically converted or encrypted, cannot be detected through conventional keyword searches.
[0077] With the continuous development of AI technology, AI-based prompt word attacks are gradually emerging. This type of attack through the inclusion of content in text is difficult to detect through traditional data security checks.
[0078] Figure 1 Schematic diagram of data entrainment leakage in related technologies, such as Figure 1 As shown in the figure, hidden entrained data 1 and entrained data 2 are embedded in two places in the original data, namely the data content and the extended attributes. Entrained data 1 is embedded in the main content area of the file, while entrained data 2 is embedded in the extended attributes of the file. Using specific extraction methods, entrained data 1 and entrained data 2 can be completely extracted from the embedded file.
[0079] Normal file A is maliciously implanted with sensitive file B during the outflow process, generating file C after implantation. File C is no different from file A during use. Then, file C is transferred to user 1, and user 1 extracts sensitive file B through specific means. This process does not cause any abnormality in the use of the user or organization, but sensitive file B has been extracted by user 1. This process will neither match the marked features nor can the DLC's information leakage prevention system scan the entrained sensitive files, which causes serious data leakage.
[0080] Figure 2 Schematic diagram of the data cleaning and transcription method provided by the present invention. Figure 2 As shown, the method includes the following:
[0081] Step 210: When a file to be cleaned in a document format is received, the file to be cleaned is format cleaned to obtain a first file to be cleaned after format cleansing; wherein the format cleansing includes: cleaning metadata and extended attribute data of the file to be cleaned;
[0082] In the present invention, the document format may refer to doc format, docx format, txt format, rtf format, wps format, html format, etc.
[0083] In the present invention, metadata refers to information about the creation, modification, author, etc. of a file. Cleaning metadata involves removing or resetting this information to prevent it from being maliciously exploited or leaked.
[0084] Extended attribute data refers to additional attributes in a file beyond the main content, such as custom attributes and summary information. Cleaning extended attribute data involves removing or standardizing these attributes to avoid potential security risks.
[0085] In this invention, format cleaning is the process of cleaning the metadata and extended attribute data of a document file. The purpose of this step is to remove or reset sensitive information and unnecessary attributes that may be contained in the file, ensuring the security and compliance of the file in subsequent processing.
[0086] Specifically, in the present invention, a document processing library (such as Python-docx) is used to read the metadata of the document file, including the author, creation time, modification time, etc. Unnecessary metadata items are set to empty or standardized values, for example, the author information is set to "anonymous".
[0087] More specifically, the extended attributes of the file, such as custom attributes, summary information, etc., are read, and unnecessary extended attributes are deleted or set to default values.
[0088] After format cleaning is completed, a new document file is generated, which is the first file to be cleaned. This file has been stripped of potential risks in metadata and extended attributes, making it ready for subsequent cleaning steps.
[0089] In the present invention, the security and compliance of document files in subsequent processing are ensured by format cleaning. By cleaning metadata and extended attribute data, potential sensitive information and unnecessary attributes are removed, effectively ensuring data security.
[0090] For example, the original file to be cleaned is transcribed into a predefined file template, and then the file content is parsed and compared with the standard template for content verification; by reading the file metadata (version, editing history, etc.), it is determined whether it meets the requirements, and non-compliant ones are deleted; by checking the extended attributes (macros, embedded objects, etc.), unnecessary content is removed, thereby completing the format cleaning and obtaining the first file to be cleaned after format cleaning.
[0091] Step 220: Clean the outline and chapters of the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned;
[0092] In this method, the outline structure of the first document to be cleaned is first analyzed, all headings and subheadings are extracted, and then compared with a preset outline template. The outline template clearly defines the expected heading hierarchy and structure of the document. Through this comparison, redundant or incorrect heading structures that do not meet the template requirements can be accurately identified.
[0093] For sections that do not conform to the outline template requirements, necessary adjustments will be made. This may include deleting unnecessary headings, merging related sections, or rearranging the order of sections to ensure that the overall structure of the document conforms to the preset business logic and specifications.
[0094] After the outline structure is adjusted, we will delve into the content of each chapter and conduct a detailed review based on the preset chapter template. The chapter template specifies the content type, format requirements, and keywords that each chapter should include. Through comparison, we can identify content fragments that do not meet the requirements.
[0095] Content that does not conform to the chapter template requirements will be cleaned up. This may involve removing irrelevant content, correcting formatting errors, or supplementing missing key information.
[0096] After a comprehensive cleansing of the outline and chapters, the processed document content is saved as a new file, the second cleansing file. This file is more standardized and secure in both structure and content, making it fully prepared for the subsequent text content cleaning steps.
[0097] In the present invention, by strictly following the preset outline and chapter templates, it is possible to ensure that the structure and content of the document comply with established business specifications and security standards, effectively avoiding the risks caused by inconsistent document formats or content violations, and by cleaning out content and structures that do not meet the requirements, it is possible to effectively eliminate sensitive information or malicious content that may be hidden in the document, thereby enhancing the security of the document and preventing potential security threats.
[0098] Step 230: Clean the text content of the second file to be cleaned to obtain a cleaned file in a document format.
[0099] Optionally, performing text content cleaning on the second file to be cleaned to obtain a cleaned file in a document format includes:
[0100] After capturing the virtual display content of the second file to be cleaned by using a simulation display technology, the virtual display content of the second file to be cleaned is converted into first YUV data;
[0101] Through a preset optical character recognition model, text recognition is performed on the first YUV data to obtain a cleaned file in a document format.
[0102] In the present invention, the content of the second file to be cleaned is rendered into a virtual display environment using simulated display technology without actually displaying it on a physical screen. In this way, the visual content of the file can be accurately captured without being interfered with by non-visual elements (such as metadata, extended attributes, etc.).
[0103] Then, the captured virtual display content is converted into a YUV data format to obtain first YUV data. YUV is a color space representation commonly used in image and video processing that can effectively separate brightness and chrominance information, facilitating subsequent image processing and analysis.
[0104] The first YUV data is input into the preset optical character recognition (OCR) model. The OCR model accurately identifies the text content in the image and converts it into an editable text format. The recognized text content is reorganized and transcribed into a new document format, ensuring the integrity and accuracy of the text content.
[0105] In the present invention, the transcribed document format file is saved as the final cleaned file. This file contains only verified and processed text content, removing all potential sensitive information, malicious content, and unnecessary formats and attributes, ensuring the security and compliance of the file.
[0106] In this invention, the OCR model accurately identifies text content, ensuring that important information from the original document is not lost or misinterpreted during the conversion process, improving the accuracy and reliability of text extraction. By combining analog display technology with the OCR model, it effectively removes potentially malicious embedded content and sensitive information from files, preventing data leaks and malicious attacks.
[0107] Optionally, performing text content cleaning on the second file to be cleaned to obtain a cleaned file in a document format includes:
[0108] Capturing audio data corresponding to the text in the second file to be cleaned by using virtual playback technology;
[0109] The audio data is subjected to speech recognition through an automatic speech recognition model, and a cleaned file in a document format is obtained according to the speech recognition content.
[0110] In the present invention, the text content in the second file to be cleaned is converted into audio data using virtual playback technology. This process is carried out in a virtual environment without actually playing the audio, thus avoiding dependence on physical devices and potential security risks.
[0111] The captured audio data is then fed into a pre-set automatic speech recognition (ASR) model, which accurately converts the audio signal into text and identifies the text in the document.
[0112] The text content recognized by the ASR model will be extracted to obtain a cleaned file in document format, removing all potential sensitive information, malicious content, and unnecessary formats and attributes, ensuring the security and compliance of the file.
[0113] In the present invention, by receiving a document format file uploaded by a user, parsing the metadata of the file, removing or resetting unnecessary metadata items, checking and clearing the extended attribute data of the file, the purity of the file format is ensured. The cleaned file is compared with the preset outline template, the content that does not conform to the outline structure is deleted, each chapter is checked in detail, and the parts that do not conform to the chapter template requirements are deleted or modified, so that the document file can be fully cleaned to ensure data security and compliance. Format cleaning removes external impurities from the file, outline and chapter cleaning ensures the standardization of the file structure, and text content cleaning goes deep into the core content of the file and removes potential risk information. The final cleaned file not only meets security standards but also maintains the integrity and availability required by the business, effectively preventing data leakage and the spread of malicious content.
[0114] Optionally, the step of performing outline and chapter cleaning on the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned includes:
[0115] Comparing the content outline of the first file to be cleaned with the preset outline template, deleting the outline and outline reference content in the content outline of the first file to be cleaned that are inconsistent with the preset outline template, to obtain a third file to be cleaned;
[0116] Compare the content chapters of the third file to be cleaned with the chapter template, delete the chapters and chapter references in the chapter template of the third file to be cleaned that are inconsistent with the chapter template, and obtain a cleaned second file to be cleaned.
[0117] In the present invention, a first file to be cleaned is parsed to extract its content outline, including structural information such as titles and subtitles, to form an outline structure tree. Simultaneously, a preset outline template also exists in the form of an outline structure tree, specifying the titles and subtitles that the document should contain and their hierarchical relationships. The outline structure tree of the first file to be cleaned is compared layer by layer with the outline structure tree of the preset outline template to check whether each node matches.
[0118] Any titles and subtitles in the first cleaned file that do not conform to the pre-set outline template will be deleted. At the same time, the document content will be checked for references to these titles and subtitles and deleted as well. This ensures that the document structure strictly adheres to the pre-set outline template and avoids non-compliant chapter structures.
[0119] Parse the third file to be cleaned, extract its chapter content, and generate a chapter content list. The preset chapter template specifies the content requirements for each chapter, such as the allowed content types, formats, and keywords. Compare the chapter content list of the third file to the chapter template one by one to check whether the content of each chapter meets the template requirements.
[0120] In the present invention, for the chapter contents in the third file to be cleaned that are inconsistent with the chapter template, they will be deleted. At the same time, the references related to these chapters in the document will be checked and deleted. This ensures that the chapter contents of the document strictly comply with the preset chapter template requirements and avoids the occurrence of non-compliant content.
[0121] In the present invention, the first file to be cleaned can be subjected to in-depth outline and chapter cleaning to ensure that the structure and content of the document fully comply with the preset specifications and security requirements. The resulting second file to be cleaned is a strictly cleaned and processed document, effectively ensuring data security.
[0122] Optionally, the method further comprises:
[0123] When a fourth file to be cleaned in XML format is received, obtaining a first namespace of each element in the fourth file to be cleaned;
[0124] Removing elements in each of the first namespaces that do not conform to a preset namespace template to obtain a cleaned fifth file to be cleaned;
[0125] The fifth file to be cleaned is cleaned of contents that are inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned.
[0126] In the present invention, a namespace is a mechanism for resolving conflicts between element and attribute names in Extensible Markup Language (XML). It allows elements and attributes with the same name to be used in the same XML document, but distinguished by different namespaces. A namespace is defined by a unique identifier (URI). This URI is typically a URL, but does not necessarily point to an actual resource; it is simply a unique string.
[0127] When processing an XML document, the parser identifies the namespace to which elements and attributes belong based on the namespace declaration. This means that even if the elements have the same name, the parser will treat them as different elements as long as their namespaces are different.
[0128] In the present invention, the received fourth file to be cleaned in XML format is first parsed to extract the first namespace of each element in the file. Namespaces are used in XML to distinguish elements with the same name, ensuring their uniqueness and accuracy. For example, in an XML document, different namespaces can be used to distinguish elements from different sources, even if they have the same name.
[0129] The first namespace of each extracted element is compared with the preset namespace template. The preset namespace template specifies the namespaces that can exist. Elements that do not conform to the preset namespace template are deleted to ensure that the remaining elements conform to the namespace requirements, resulting in the fifth file to be cleaned.
[0130] Figure 3 The schematic diagram of the namespace template provided by the present invention is as follows: Figure 3 As shown, the namespace of the XML file is checked, elements that do not conform to the namespace are removed, and the security of the namespace is checked.
[0131] In this invention, the pre-defined business logic tag tree structure refers to the tag hierarchy and semantic rules pre-defined in XML documents based on business requirements and specifications. This structure specifies which tags are allowed, as well as their parent-child relationships and nesting rules. This ensures that the structure of the XML document conforms to the specified business logic, thereby improving data accuracy and usability.
[0132] Assume that the preset namespace template specifies the following tags: a, b, c, d, a1, b1, c1, d1. Furthermore, based on these tags, the following hierarchical relationship is further specified: the parent tag of a1 must be a; the parent tag of b1 must be b.
[0133] a, b, c, d: These are first-level tags that can be subtags of the root tag or other first-level tags. a1, b1, c1, d1: These are second-level tags that must be nested under a specific first-level tag.
[0134] These hierarchical relationships ensure that the structure of the XML document meets business requirements. For example, the content of a1 might be a detailed description of item a, while the content of b1 might be a detailed description of item b. This structure ensures that the data is organized in a manner that meets business specifications, facilitating subsequent data processing and analysis.
[0135] In the present invention, during the data cleaning process, it is checked whether the tags in the XML file conform to the preset business logic tag tree structure.
[0136] First, parse the XML file and extract the tags and their hierarchical relationships; compare the extracted tag tree structure with the preset business logic tag tree structure to check whether there are any tags that do not comply with the rules; for tags and content that do not comply with the preset business logic tag tree structure, they will be deleted or corrected to ensure that the final XML file complies with business specifications.
[0137] Figure 4 The label tree template provided by the present invention is illustrated as follows: Figure 4 As shown, the left side is the original XML file content, which contains the following tags and attributes:
[0138]
[0139] A-label
[0140] <a1 apropallow="val1" apropban="val2">< / a1>
[0141]
[0142] B Tag< >
[0143] <a1 apropallow="val1" apropban="val2"> a1 tag< / a1>
[0144] <b1> B1 Label< / b1>
[0145] The middle part is the business-specific tag tree template, which specifies the following:
[0146] Specify the hierarchical relationship of a, a1, b, and b1.
[0147] The parent tag of a1 must be a.
[0148] The parent tag of b1 must be b.
[0149] First read the contents of the original XML file into memory.
[0150] Parse the XML file to form a tag tree structure for subsequent cleaning operations.
[0151] Check whether each tag complies with the specified hierarchical relationship based on the tag tree template specified by the business.
[0152] Tags that do not conform to the hierarchical relationship will be deleted.
[0153] The cleaned XML content is written to a new file.
[0154] The right side shows the cleaned XML file content, which contains the following tags and attributes:
[0155]
[0156] A-label
[0157] <a1 apropallow="val1" apropban="val2">< / a1>
[0158]
[0159] B Tag< >
[0160] In this invention, by reading the original file, forming a tag tree, and deleting tags that do not conform to the hierarchical relationship, a cleaned XML file that conforms to the business logic is finally generated. This process ensures that the structure and content of the XML file conform to the preset business specifications and security requirements.
[0161] In this invention, the preset business logic tag tree structure is an important mechanism to ensure that the XML file structure meets business requirements. By defining the allowed tags and their hierarchical relationships, it is possible to effectively clean up content that does not meet the standards and improve the quality and security of data.
[0162] Optionally, after the step of cleaning the fifth file to be cleaned for content that is inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned, the method further includes:
[0163] Cleaning the attribute value, attribute field, and attribute relationship of each tag in the sixth file to be cleaned according to a preset tag attribute template to obtain a seventh file to be cleaned;
[0164] After capturing the virtual display content of the seventh file to be cleaned through virtual display technology, the virtual display content of the seventh file to be cleaned is converted into second YUV data; through a preset optical character recognition model, text recognition is performed on the second YUV data to obtain a cleaned file in XML format.
[0165] In this method, an XML file is parsed and each tag and its attributes are extracted. Each tag's attribute values, attribute fields, and attribute relationships are checked to see if they conform to the requirements of a pre-set template. Any attribute values, fields, or relationships that do not conform to the requirements are corrected or deleted to ensure attribute compliance.
[0166] Figure 5 This is a schematic diagram illustrating the tag attribute template provided by the present invention, such as Figure 5 As shown, the left side is the original XML file content, which contains the following tags and attributes:
[0167]
[0168] A-label
[0169] <a1 apropallow="val1" apropban="val2">< / a1>
[0170]
[0171] B Tag< >
[0172] The middle part is the business-specific tag attribute template, which specifies the following:
[0173] Tags a and a1 only allow property aPropAllow. Tag b only allows property bPropAllow.
[0174] First read the contents of the original XML file into memory.
[0175] Check the attributes of each tag according to the tag attribute template specified by the business, and delete the attributes that do not meet the requirements.
[0176] For tags a and a1, delete the property aPropBan. For tag b, delete the property bPropBan.
[0177] The right side shows the cleaned XML file content, which contains the following tags and attributes:
[0178]
[0179] A-label
[0180] <a1 apropallow="val1">< / a1>
[0181]
[0182] B Tag< >
[0183] In the present invention, by reading the original file and deleting the attributes that do not meet the requirements, a cleaned XML file that meets the business specifications is finally generated. This process ensures that the attributes of the XML file meet the preset business specifications and security requirements.
[0184] More specifically, using virtual display technology, the contents of the seventh file to be cleaned are rendered into a virtual display environment without actually displaying it on a physical screen. This allows the file's visual content to be accurately captured without interference from non-visual elements such as metadata and extended attributes. The captured virtual display content is then converted into a YUV data format. YUV is a color space representation commonly used in image and video processing, effectively separating luminance and chrominance information, facilitating subsequent image processing and analysis.
[0185] Input the YUV data into the preset optical character recognition (OCR) model. The OCR model accurately recognizes the text content in the image and converts it into an editable text format. The recognized text content is reorganized and generated into a new XML file.
[0186] Optionally, the method further includes:
[0187] When an eighth file to be cleaned in a picture format is received, the least significant position of each pixel in the eighth file to be cleaned is set to zero to obtain a ninth file to be cleaned;
[0188] superimposing preset spatial noise onto the pixel matrix of the ninth file to be cleaned to obtain a tenth file to be cleaned;
[0189] Using discrete cosine transform or fast Fourier transform, converting the tenth file to be cleaned from the spatial domain to the frequency domain to obtain frequency domain data of the tenth file to be cleaned;
[0190] The preset noise is superimposed on the frequency domain data, and the superimposed frequency domain data is restored back to the spatial domain to obtain a cleaned file in an image format.
[0191] In the present invention, when receiving the eighth file to be cleaned in image format, the least significant bit of each pixel in the eighth file is first set to zero, resulting in the ninth file to be cleaned. This step aims to remove any hidden steganographic information that may be hidden in the least significant bit of a pixel, preventing data leakage. The least significant bit (LSB) is the least significant bit in a pixel value and is often used to hide information because the human eye is insensitive to such small changes. By setting the LSB to zero, this type of hidden information can be effectively removed.
[0192] Next, a preset spatial noise is superimposed on the pixel matrix of the ninth file to be cleaned, resulting in the tenth file to be cleaned. This step destroys the steganographic information hidden in the image by adding random noise to the pixel matrix. Spatial noise refers to random noise added directly to the pixel space of an image. It can disrupt the regularity of the steganographic information, thereby improving the security of the image.
[0193] More specifically, the present invention can also use a discrete cosine transform (DCT) or a fast Fourier transform (FFT) to convert the tenth file to be cleaned from the spatial domain to the frequency domain, thereby obtaining frequency domain data for the tenth file to be cleaned. DCT and FFT are two commonly used mathematical transformation methods for converting signals from the spatial domain to the frequency domain. DCT is commonly used in image compression and processing, while FFT is widely used in signal processing and image analysis. By converting images to the frequency domain, certain types of steganographic information can be processed more efficiently, as this information may exhibit specific patterns in the frequency domain.
[0194] Finally, the pre-set noise is superimposed on the frequency domain data, and the superimposed frequency domain data is then restored back to the spatial domain to produce a cleaned image file. This step further destroys any frequency domain steganographic information by adding the pre-set noise to the frequency domain. Afterwards, an inverse transform is used to convert the frequency domain data back to the spatial domain to produce the final cleaned image file. This step ensures that the image's primary visual information is not affected, while also enhancing image security.
[0195] Figure 6 This is a schematic diagram of image format data cleaning provided by the present invention, such as Figure 6 As shown in the figure, during the initial stage of image cleaning, each pixel in the image is processed by setting the least significant bit (LSB) of each pixel value to zero. This step aims to remove any hidden steganographic information that may be hidden in the least significant bit of the pixel, as the LSB is often used to hide information, and the human eye is insensitive to such small changes. This step effectively prevents data leakage.
[0196] After the LSBs are zeroed, a preset spatial noise is added to the image's pixel matrix. Spatial noise refers to random noise added directly to the image's pixel space. This step introduces randomness, destroying any other steganographic information that may exist, further enhancing the image's security.
[0197] Next, the image is converted from the spatial domain to the frequency domain using either a discrete cosine transform (DCT) or a fast Fourier transform (FFT). DCT and FFT are two common mathematical transformations that convert an image's pixel values into frequency components. This step allows for more efficient processing in the frequency domain, as some types of steganographic information are easier to detect and remove in the frequency domain.
[0198] In the frequency domain, a preset noise level is superimposed on the frequency domain data. This step, by adding noise to the frequency domain, further destroys any possible frequency domain steganographic information. The addition of noise to the frequency domain disrupts the regularity of the hidden information, making it difficult to recover in subsequent processing.
[0199] After frequency domain processing, the frequency domain data is converted back to the spatial domain through an inverse transform (such as inverse DCT or inverse FFT). This step restores the processed data to the pixel data of the image, generating the final cleaned image.
[0200] The cleaned image is the output of the entire process. After the above steps, the image has been processed to remove any potential hidden information and other security risks, while retaining the main visual information, ensuring the security and integrity of the image.
[0201] The entire process is designed to fully remove potential hidden information and other security threats from images while maintaining their usability and visual quality. This process is crucial for protecting data privacy and preventing information leakage.
[0202] Figure 7 This is one of the file cleaning schematics provided by the present invention. Figure 8 The second file cleaning diagram provided by the present invention is: Figure 9 The third file cleaning diagram provided by the present invention is: Figure 10 The fourth file cleaning diagram provided by the present invention is: Figure 11 The fifth file cleaning diagram provided by the present invention is as follows: Figure 7-11 As shown,
[0203] First, File Transcription Template Lv1 reads the original file content according to the specified file format, transcribes it into a predefined file template, and parses the file content to verify it against the standard template. During this process, it verifies compliance with requirements by reading file metadata (such as version and edit history), deleting any non-compliant metadata. It also checks the file's extended attributes (such as macros and embedded objects) and removes unnecessary content, resulting in a preliminarily cleaned file.
[0204] Next, the Content Summary Template Lv2 further processes the initially cleaned file. It examines and identifies the file's content, extracts it according to the template specifications, and compares it with predefined templates. During this process, irrelevant sections, content, and extensions are filtered out, and irrelevant titles and text are deleted to ensure that the file better meets business requirements.
[0205] The Visual Transcription Template Lv3 then visualizes the file processed by the Content Summarization Template. For document files, virtual display technology is used to capture the document content, convert it into YUV data, and perform operations such as stretching and adding white noise. After these operations, the document is regenerated using the trained OCR model. For video files or video streams, they are converted into YUV data and similar operations such as stretching and adding white noise are performed to regenerate the video stream or video file.
[0206] Finally, the speech-based transcription template Lv4 performs speech-based transcription on the file. For document files, TTS text-to-speech technology is used to capture the document content, convert it into PCM data, and perform operations such as adding white noise. The document is then regenerated using the trained ASR model. For audio files or audio streams, they are converted into PCM data, and operations such as adding white noise are performed to regenerate the audio stream or audio file.
[0207] More specifically, the file transcription template Lv1 reads the original file content according to the specified file format and transcribes it into a predefined file template. It then parses the file content and compares it with the standard template for content verification. It determines whether it meets the requirements by reading the file metadata (version, editing history, etc.), and deletes non-compliant content. It removes unnecessary content by checking extended attributes (macros, embedded objects, etc.).
[0208] Content summary template Lv2, this template checks and identifies the content of the file, extracts the file content according to the template specifications, compares it with the predefined template, filters out irrelevant chapters, content, extension elements, etc., deletes irrelevant titles and their text, and ensures that the file meets business requirements.
[0209] Visual transcription template Lv3, when processing document file data, this template uses virtual display technology to capture the content of the document file, converts it into YUV data, and performs stretching transformation, etc., adds white noise and other operations, and then regenerates the document through the trained OCR model; when processing video files or video streams, this template converts them into YUV data, and performs stretching transformation, etc., and then adds white noise and other operations to regenerate the video stream or video file.
[0210] Voice transcription template Lv4: When processing document file data, this template uses TTS text-to-speech technology to capture the document file content, convert it into PCM data, add white noise and other operations, and then regenerate the document through the trained ASR model. When processing audio files or audio streams, this template converts them into PCM data, adds white noise and other operations, and then regenerates the audio stream or audio file.
[0211] Figure 12 The present invention provides a schematic diagram of the template data transcription and cleaning principle, such as Figure 12 As shown, the entire data cleansing process starts with the original data and covers the processing of three types of files: documents, videos, and audio. After the original data is input, any non-compliant portions are discarded. Document files enter the document template processing module. After visualization using virtual display technology, they are converted into speech using TTS text-to-speech technology. They then undergo image and audio template processing before being re-formatted. Video files and video streams enter the image template processing module. After data type identification and format parsing, they are converted into YUV / RGB data, undergo steganography detection and removal, and finally re-formatted. Audio files and audio streams enter the audio template processing module. After data type identification and format parsing, they are converted into PCM data, undergo steganography detection and removal, and finally re-formatted. The entire process is designed to ensure data security, compliance, and integrity, while preserving the data's essential content and value, adapting to different file types, and effectively removing potential risks and unnecessary information.
[0212] More specifically, the original document data is passed to the document-based template. The document template will transcribe the data into YUV or PCM data for virtual display or text-to-speech conversion as required. The data is then processed by the image template and audio template and repackaged into a document to generate cleaned data. The original data such as video files and video streams are passed to the image-based template. The image template will first detect the data type, then parse the format, convert its content into standardized YUV / RGB data, and remove steganography in the images by slightly stretching and adjusting the images and adding white noise. The processed data will be repackaged and repackaged according to business needs to generate cleaned data.
[0213] Pass the audio file and audio stream data to the audio-based template. The audio template first detects the data type, then parses the format, converts its content into standardized PCM data, adds white noise and other operations to the data, and then repackages and repacks it according to business needs to generate cleaned data.
[0214] The data cleaning and transcription device provided by the present invention is described below. The data cleaning and transcription device described below and the data cleaning and transcription method described above can be referenced to each other.
[0215] Figure 13 The data cleaning and transcription device provided by the present invention is shown in FIG. Figure 13 Shown, including:
[0216] The first cleaning module 1310 is configured to, upon receiving a file to be cleaned in a document format, perform format cleaning on the file to be cleaned to obtain a first file to be cleaned after the format cleaning; wherein the format cleaning includes cleaning metadata and extended attribute data of the file to be cleaned;
[0217] The second cleaning module 1320 is used to clean the outline and chapters of the first file to be cleaned according to the preset outline template and chapter template to obtain a cleaned second file to be cleaned;
[0218] The third cleaning module 1330 is used to clean the text content of the second file to be cleaned to obtain a cleaned file in a document format.
[0219] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0220] Comparing the content outline of the first file to be cleaned with the preset outline template, deleting the outline and outline reference content in the content outline of the first file to be cleaned that are inconsistent with the preset outline template, to obtain a third file to be cleaned;
[0221] Compare the content chapters of the third file to be cleaned with the chapter template, delete the chapters and chapter references in the chapter template of the third file to be cleaned that are inconsistent with the chapter template, and obtain a cleaned second file to be cleaned.
[0222] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0223] After capturing the virtual display content of the second file to be cleaned by using a simulation display technology, the virtual display content of the second file to be cleaned is converted into first YUV data;
[0224] Through a preset optical character recognition model, text recognition is performed on the first YUV data to obtain a cleaned file in a document format.
[0225] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0226] Capturing audio data corresponding to the text in the second file to be cleaned by using virtual playback technology;
[0227] The audio data is subjected to speech recognition through an automatic speech recognition model, and a cleaned file in a document format is obtained according to the speech recognition content.
[0228] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0229] When a fourth file to be cleaned in XML format is received, obtaining a first namespace of each element in the fourth file to be cleaned;
[0230] Removing elements in each of the first namespaces that do not conform to a preset namespace template to obtain a cleaned fifth file to be cleaned;
[0231] The fifth file to be cleaned is cleaned of contents that are inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned.
[0232] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0233] Cleaning the attribute value, attribute field, and attribute relationship of each tag in the sixth file to be cleaned according to a preset tag attribute template to obtain a seventh file to be cleaned;
[0234] After capturing the virtual display content of the seventh file to be cleaned through virtual display technology, the virtual display content of the seventh file to be cleaned is converted into second YUV data; through a preset optical character recognition model, text recognition is performed on the second YUV data to obtain a cleaned file in XML format.
[0235] According to a data cleaning and transcription device provided by the present invention, the device is further used for:
[0236] When an eighth file to be cleaned in a picture format is received, the least significant position of each pixel in the eighth file to be cleaned is set to zero to obtain a ninth file to be cleaned;
[0237] superimposing preset spatial noise onto the pixel matrix of the ninth file to be cleaned to obtain a tenth file to be cleaned;
[0238] Using discrete cosine transform or fast Fourier transform, converting the tenth file to be cleaned from the spatial domain to the frequency domain to obtain frequency domain data of the tenth file to be cleaned;
[0239] The preset noise is superimposed on the frequency domain data, and the superimposed frequency domain data is restored back to the spatial domain to obtain a cleaned file in an image format.
[0240] In the present invention, by receiving a document format file uploaded by a user, parsing the metadata of the file, removing or resetting unnecessary metadata items, checking and clearing the extended attribute data of the file, the purity of the file format is ensured. The cleaned file is compared with the preset outline template, the content that does not conform to the outline structure is deleted, each chapter is checked in detail, and the parts that do not conform to the chapter template requirements are deleted or modified, so that the document file can be fully cleaned to ensure data security and compliance. Format cleaning removes external impurities from the file, outline and chapter cleaning ensures the standardization of the file structure, and text content cleaning goes deep into the core content of the file and removes potential risk information. The final cleaned file not only meets security standards but also maintains the integrity and availability required by the business, effectively preventing data leakage and the spread of malicious content.
[0241] Figure 14 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 14 As shown, the electronic device may include: a processor 1410, a communications interface 1420, a memory 1430, and a communication bus 1440, wherein the processor 1410, the communications interface 1420, and the memory 1430 communicate with each other via the communication bus 1440. The processor 1410 may call logic instructions in the memory 1430 to execute a data cleaning and transcription method, which includes: upon receiving a file to be cleaned in a document format, performing format cleaning on the file to be cleaned to obtain a first file to be cleaned after format cleaning; wherein the format cleaning includes: cleaning metadata and extended attribute data of the file to be cleaned;
[0242] Performing outline and chapter cleaning on the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned;
[0243] The second file to be cleaned is subjected to text content cleaning to obtain a cleaned file in a document format.
[0244] Furthermore, the logic instructions in the aforementioned memory 1430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0245] On the other hand, the present invention further provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data cleaning and transcription method provided by the above methods, the method comprising: upon receiving a file to be cleaned in a document format, performing format cleaning on the file to be cleaned to obtain a first file to be cleaned after format cleaning; wherein the format cleaning includes: cleaning metadata and extended attribute data of the file to be cleaned;
[0246] Performing outline and chapter cleaning on the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned;
[0247] The second file to be cleaned is subjected to text content cleaning to obtain a cleaned file in a document format.
[0248] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program is implemented to perform the data cleaning and transcription method provided by the above methods, the method comprising: upon receiving a file to be cleaned in a document format, performing format cleaning on the file to be cleaned to obtain a first file to be cleaned after format cleaning; wherein the format cleaning comprises: cleaning metadata and extended attribute data of the file to be cleaned;
[0249] Performing outline and chapter cleaning on the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned;
[0250] The second file to be cleaned is subjected to text content cleaning to obtain a cleaned file in a document format.
[0251] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0252] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A data cleaning and transcription method, characterized in that: include: When receiving a file to be cleaned in a document format, performing format cleaning on the file to be cleaned to obtain a first file to be cleaned after format cleaning; wherein the format cleaning includes: cleaning metadata and extended attribute data of the file to be cleaned; Performing outline and chapter cleaning on the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned; Cleaning the text content of the second file to be cleaned to obtain a cleaned file in a document format; The second file to be cleaned is subjected to text content cleaning to obtain a cleaned file in a document format, including: After capturing the virtual display content of the second file to be cleaned by using a simulation display technology, the virtual display content of the second file to be cleaned is converted into first YUV data; Through a preset optical character recognition model, text recognition is performed on the first YUV data to obtain a cleaned file in a document format.
2. The data cleaning and transcription method according to claim 1, characterized in that: The step of performing outline and chapter cleaning on the first file to be cleaned according to the preset outline template and chapter template to obtain a cleaned second file to be cleaned includes: Comparing the content outline of the first file to be cleaned with the preset outline template, deleting the outline and outline reference content in the content outline of the first file to be cleaned that are inconsistent with the preset outline template, to obtain a third file to be cleaned; Compare the content chapters of the third file to be cleaned with the chapter template, delete the chapters and chapter references in the chapter template of the third file to be cleaned that are inconsistent with the chapter template, and obtain a cleaned second file to be cleaned.
3. The data cleaning and transcription method according to claim 1, characterized in that: Cleaning the text content of the second file to be cleaned to obtain a cleaned file in a document format, including: Capturing audio data corresponding to the text in the second file to be cleaned by using virtual playback technology; The audio data is subjected to speech recognition through an automatic speech recognition model, and a cleaned file in a document format is obtained according to the speech recognition content.
4. The data cleaning and transcription method according to claim 1, characterized in that: The method further comprises: When a fourth file to be cleaned in XML format is received, obtaining a first namespace of each element in the fourth file to be cleaned; Removing elements in each of the first namespaces that do not conform to a preset namespace template to obtain a cleaned fifth file to be cleaned; The fifth file to be cleaned is cleaned of contents that are inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned.
5. The data cleaning and transcription method according to claim 4, characterized in that: After the step of cleaning the fifth file to be cleaned for content that is inconsistent with the business logic of the preset business logic tag tree structure to obtain a sixth file to be cleaned, the method further includes: Cleaning the attribute value, attribute field, and attribute relationship of each tag in the sixth file to be cleaned according to a preset tag attribute template to obtain a seventh file to be cleaned; After capturing the virtual display content of the seventh file to be cleaned through virtual display technology, the virtual display content of the seventh file to be cleaned is converted into second YUV data; through a preset optical character recognition model, text recognition is performed on the second YUV data to obtain a cleaned file in XML format.
6. The data cleaning and transcription method according to claim 1, characterized in that: The method further comprises: When an eighth file to be cleaned in a picture format is received, the least significant position of each pixel in the eighth file to be cleaned is set to zero to obtain a ninth file to be cleaned; superimposing preset spatial noise onto the pixel matrix of the ninth file to be cleaned to obtain a tenth file to be cleaned; Using discrete cosine transform or fast Fourier transform, converting the tenth file to be cleaned from the spatial domain to the frequency domain to obtain frequency domain data of the tenth file to be cleaned; The preset noise is superimposed on the frequency domain data, and the superimposed frequency domain data is restored back to the spatial domain to obtain a cleaned file in an image format.
7. A data cleaning and transcription device, characterized in that: include: A first cleaning module is configured to, upon receiving a file to be cleaned in a document format, perform format cleaning on the file to be cleaned to obtain a first file to be cleaned after the format cleaning; wherein the format cleaning includes: cleaning metadata and extended attribute data of the file to be cleaned; A second cleaning module is configured to clean the outline and chapters of the first file to be cleaned according to a preset outline template and chapter template to obtain a cleaned second file to be cleaned; a third cleaning module, configured to clean the text content of the second file to be cleaned to obtain a cleaned file in a document format; Wherein, the device is also used for: After capturing the virtual display content of the second file to be cleaned by using a simulation display technology, the virtual display content of the second file to be cleaned is converted into first YUV data; Through a preset optical character recognition model, text recognition is performed on the first YUV data to obtain a cleaned file in a document format.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the data cleaning and transcription method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data cleaning and transcription method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
File reconstruction method and device, transmission equipment, electronic equipment, program product and medium
CN113852602A
Document-based business processing method and device, equipment, storage medium and program product
CN119862883A