Method for automatically annotating unstructured data based on neural network

Through the automated annotation method based on neural network, the problem of reducing annotation efficiency caused by the diverse data formats and data sources in unstructured data is solved, and efficient annotation of unstructured data and unified standards are established.

CN120197598AActive Publication Date: 2025-06-24SHENZHEN HAOYUAN TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510685350.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

When processing unstructured data, it is difficult to establish unified standards for cross-type data when dealing with unstructured data formats and data sources, resulting in reduced annotation efficiency and user information reception efficiency.

Method used

An automated annotation method based on neural network is adopted, and a preprocessing annotation method is established through NLP technology and CV technology, and the unstructured data is preprocessed, and the preferred neural network is obtained through neural network training to achieve automated annotation of unannotated data.

Benefits of technology

It improves the annotation efficiency of unstructured data, realizes unified standards for cross-type data, reduces user information processing time and improves user data reception efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197598A_ABST
    Figure CN120197598A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically annotating unstructured data based on a neural network, and relates to the technical field of unstructured data, and the method comprises the steps: building a preprocessing annotation method based on an NLP technology and a CV technology; using a preprocessing annotation method to obtain an annotation processing interval; establishing a neural network, and establishing a deep annotation method; training the neural network to obtain an optimal neural network; performing automatic annotation on the unannotated data based on the optimized neural network; the method is used for solving the problems that in an existing unstructured data annotation method, when data formats in unstructured data are many and data sources are many, each data source needs to be independently analyzed, so that a unified standard of cross-type data is difficult to establish, and the unstructured data is difficult to annotate. The annotation efficiency is reduced; and the user information receiving efficiency is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unstructured data, and specifically to a method for automatically annotating unstructured data based on a neural network. Background Art

[0002] Unstructured data is data with irregular or incomplete data structures, without a predefined data model, and is inconvenient to be represented by a two-dimensional logical table of a database; it includes office documents, texts, pictures, HTML, various reports, images, audio, video information, etc. in all formats; the formats of unstructured data are very diverse, and the standards are also diverse. Moreover, unstructured information is more difficult to standardize and understand technically than structured information. Therefore, more intelligent IT technologies are required for storage, retrieval, publishing, and utilization; the annotation of unstructured data is mainly through methods such as data cleaning, data standardization, and image processing.

[0003] Existing methods for annotating unstructured data usually analyze the unstructured data through an analysis tool after receiving it, and display the summary information to the user. The user modifies the summary information and interacts with the unstructured data analysis system through the annotation layer, so as to achieve the annotation of unstructured data. Although this improved method can assist the user in fully understanding the unstructured data, when there are many data formats in the unstructured data, summarizing the unstructured data only through the analysis tool will make it difficult to establish a unified standard for cross-type data due to the large number of data types. As a result, when there are many data sources, each data source needs to be analyzed independently, which leads to a decrease in annotation efficiency and a decrease in the user information reception efficiency. For example, in the patent application with the publication number CN107368506A, an unstructured data analysis system and method are disclosed. This solution is to apply one or more analysis tools to the unstructured data, and display the summary information to one or more users, allowing one or more users to modify the granularity of the summary information and interact with the unstructured data analysis system through the annotation layer at the same time. Other improvements in annotating unstructured data are usually based on improvements in data recognition efficiency during transmission, and still cannot solve the problem that when there are many data formats and many data sources in the unstructured data, each data source needs to be analyzed independently, resulting in difficulty in establishing a unified standard for cross-type data, thereby causing a decrease in annotation efficiency and a decrease in the user information reception efficiency. In view of this, it is necessary to improve the existing methods for annotating unstructured data. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems in the prior art to some extent. By proposing a method for automatically annotating unstructured data based on a neural network, it is used to solve the problem in the existing methods for annotating unstructured data that when there are many data formats and data sources in unstructured data, each data source needs to be analyzed independently, resulting in difficulty in establishing a unified standard for cross-type data, thus reducing the annotation efficiency and the user information reception efficiency.

[0005] To achieve the above object, the present application provides a method for automatically annotating unstructured data based on a neural network, including the following steps: Obtain unannotated unstructured data, denoted as unannotated data; establish a preprocessing annotation method based on NLP technology and CV technology; use the preprocessing annotation method to preprocess the unannotated data, and obtain an annotation processing interval based on the preprocessing result; Obtain annotated unstructured data, denoted as annotated data; establish a neural network, introduce the preprocessing annotation method and the annotation processing interval into the neural network, and establish an in-depth annotation method in the neural network; train the neural network based on the unannotated data and the annotated data, and obtain an optimized neural network based on the training result; Step S3, automatically annotate the unannotated data based on the optimized neural network.

[0006] Furthermore, the preprocessing annotation method includes: Obtain unannotated data, extract the text in the unannotated data based on NLP technology, and sequentially denote the obtained independent text segments as data text SW1 to data text SW n , where n is the number of independent text segments; Extract the images, audio, and video in the unannotated data based on CV technology, and denote the extracted pictures, audio, and video as data images, data audio, and data video respectively.

[0007] Furthermore, the preprocessing annotation method further includes: For any data image: Denote the data image with data text SW as a text image, and denote the data image without data text SW as a regular image; For any data audio: Perform text conversion on the data audio based on NLP technology, and denote the obtained text as audio text; Obtain the background noise in the data audio based on AI recognition, and denote the data audio with background noise greater than the standard recognition decibel in the data audio as noisy audio; Denote the audio text in the noisy audio as interfering audio text, and denote the audio text not in the noisy audio as regular audio text.

[0008] Furthermore, the preprocessing annotation method further includes: For any data video: Extract all the pictures in the data video based on CV technology, and denote the obtained pictures as video pictures; Based on NLP technology, perform text conversion on the audio of the data video and extract text from the video pictures respectively, and denote the obtained texts as video audio text and video picture text respectively; Denote the video pictures with video audio text as video text images, and denote the video pictures without video picture text as video standard images; Obtain the background noise in the data video based on AI recognition, and denote the data videos with background noise greater than the standard recognition decibels in the data video as noise videos; Denote the video audio text in the noise video in the video audio text as interference video text, and denote the video audio text not in the noisy audio in the video audio text as regular video text.

[0009] Furthermore, use the preprocessing annotation method to preprocess the unannotated data, and obtain the annotation processing interval based on the preprocessing result, including: Denote a set of unannotated data as single test data, use the preprocessing annotation method to preprocess the single test data, and denote the obtained data text SW, regular audio text, and regular video text as clear processing text; Denote the interference audio text and interference video text as fuzzy processing text; Denote the text images and video text images as text processing images; Denote the regular images and video standard images as regular processing images.

[0010] Furthermore, use the preprocessing annotation method to preprocess the unannotated data, and obtain the annotation processing interval based on the preprocessing result, which further includes: Denote multiple sets of unannotated data as parameter test data; For any parameter test data: Obtain the clear processing text, fuzzy processing text, text processing images, and regular processing images of the parameter test data respectively based on the preprocessing annotation method, and denote the obtained times as clear text time, fuzzy text time, text image time, and regular image time respectively, and denote the memory space occupied by the parameter test data as test occupied memory; Obtain the clear text time, fuzzy text time, text image time, regular image time, and test occupied memory corresponding to all parameter test data, and denote the values corresponding to the minimum test occupied memory and the maximum test occupied memory as XX1 and XX2 respectively.

[0011] Furthermore, use the preprocessing annotation method to preprocess the unannotated data, and obtain the annotation processing interval based on the preprocessing result, which further includes: Establish a plane rectangular coordinate system, denoted as the parameter analysis coordinate system. Among them, the unit of the X-axis of the parameter analysis coordinate system is MB, and the unit of the Y-axis is ms; for any parameter test data: in the parameter analysis coordinate system, use the test occupancy of the parameter test data as the abscissa, and use the clear text time, blurred text time, text image time, and conventional image time of the parameter test data as the ordinates respectively for punctuation, and denote them as the clear text point, blurred text point, text image point, and conventional image point respectively; The curve obtained by fitting all the clear text points is denoted as curve QW; the curve obtained by fitting all the blurred text points is denoted as curve MW; the curve obtained by fitting all the text image points is denoted as curve WT; the curve obtained by fitting all the conventional image points is denoted as curve CT; For any abscissa X0 from XX1 to XX2 on the X-axis, the point with the smallest ordinate among the points on curve QW, curve MW, curve WT, and curve CT with abscissa X0 is denoted as the low parameter point; the curve formed by all the low parameter points in the parameter analysis coordinate system is denoted as the parameter analysis curve.

[0012] Furthermore, using the preprocessing annotation method to preprocess the unannotated data, and obtaining the annotation processing interval based on the preprocessing result also includes: The interval formed by the abscissas of the region where the parameter analysis curve coincides with curve QW is denoted as the QW preferred interval; the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve MW is denoted as the MW preferred interval; The interval formed by the abscissas of the region where the parameter analysis curve coincides with curve WT is denoted as the WT preferred interval; The interval formed by the abscissas of the region where the parameter analysis curve coincides with curve CT is denoted as the CT preferred interval; The QW preferred interval, MW preferred interval, WT preferred interval, and CT preferred interval are denoted as the annotation processing interval.

[0013] Furthermore, the in-depth annotation method includes: The neurons in the input layer of the neural network receive data, and the memory space occupied by the received data is denoted as the annotation occupancy; the interval where the annotation occupancy is located in the annotation processing interval is denoted as the preferred analysis interval, and the data type corresponding to the preferred analysis interval is denoted as the preferred screening type, where the data types include clear processed text, blurred processed text, text processed image, and conventional processed image; Based on the preprocessing annotation method, preferentially obtain the preferred screening type of the data, and obtain the data types other than the preferred screening type; The data type obtained by the preprocessing annotation method is annotated based on the target detection box, named entity recognition and speaker separation technology, and the annotated data is output by the neurons of the output layer of the neural network.

[0014] Further, training a neural network based on the unannotated data and the annotated data, and obtaining a preferred neural network based on the training result includes: For any value X0 from XX1 to XX2, the unannotated data with the annotation X0 and the annotated data are respectively input into the neural network and annotated, wherein the method of annotating the annotated data is recorded as the annotated method, and when the annotated data is input into the neural network, the neural network re-annotates the annotated data based on the annotated method; When the time that the unannotated data is annotated by the neural network is less than or equal to the time that the annotated data is annotated by the neural network, X0 is recorded as optimized memory and the annotated method corresponding to X0 is removed from the neural network; When the time that the unannotated data is annotated by the neural network is greater than the time that the annotated data is annotated by the neural network, X0 is recorded as the memory to be learned and the annotated method corresponding to X0 is recorded as the preferred annotation method of X0; When all values ​​in XX1 to XX2 are recorded as optimized memory, or recorded as memory to be learned and the corresponding preferred annotation method is entered into the neural network, the neural network is recorded as the preferred neural network; For unannotated data annotated by a neural network: when the annotation share is an optimized share, the unannotated data is annotated based on the preprocessing annotation method and the annotation processing interval; when the annotation share is a share to be learned, the unannotated data is annotated based on the preferred annotation method of the share to be learned.

[0015] Beneficial effects of the present invention: The present application first obtains unannotated unstructured data, recorded as unannotated data; establishes a preprocessing annotation method based on NLP technology and CV technology; uses the preprocessing annotation method to preprocess the unannotated data, and obtains the annotation processing interval based on the preprocessing result. The advantage of this is that by establishing the preprocessing annotation method based on NLP technology and CV technology, when there are many formats and data sources in the unstructured data, the text, pictures, audio and video in the unstructured data can be effectively preprocessed by the preprocessing annotation method to make them pure text or pictures, thereby achieving the establishment of a unified standard for cross-type data, thereby improving the subsequent annotation efficiency; and by obtaining the annotation processing interval, the type of data extraction with the highest efficiency when using the preprocessing annotation method can be obtained based on the size of the memory occupied by the unstructured data, which helps to preprocess the unstructured data with the fastest preprocessing efficiency when facing unstructured data of different sizes; The present application also obtains the unstructured data that has been annotated, denoted as the annotated data; establishes a neural network, introduces the preprocessing annotation method and the annotation processing interval into the neural network, and establishes an in-depth annotation method in the neural network; trains the neural network based on the unannotated data and the annotated data, and obtains an optimized neural network based on the training results; finally, automatically annotates the unannotated data based on the optimized neural network. The advantage of this is that after establishing the neural network and training it, using the optimized neural network for automatic annotation can effectively improve the annotation efficiency of unstructured data when there are many data formats and data sources in the unstructured data, thereby reducing the user information processing time while improving the user's data reception efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of the steps of the method of the present invention; Figure 2 is a schematic diagram of the parameter analysis coordinate system of the present invention; Figure 3 is a schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment 1, please refer to Figure 1 As shown, the present application provides a method for automatically annotating unstructured data based on a neural network, including the following steps: Step S1, obtain the unannotated unstructured data, denoted as the unannotated data; establish a preprocessing annotation method based on NLP technology and CV technology; use the preprocessing annotation method to preprocess the unannotated data, and obtain an annotation processing interval based on the preprocessing result; Step S1 includes: Step S101, the preprocessing annotation method includes: Step S1011, obtain the unannotated data, extract the text in the unannotated data based on NLP technology, and sequentially denote the obtained independent text segments as data text SW1 to data text SW n , where n is the number of independent text segments; In the specific implementation process, the data analysis SW may include the text in the picture in the unannotated data and the text outside the picture in the unannotated data; Step S1012: Extract images, audio, and video from the unannotated data based on CV technology, and denote the extracted pictures, audio, and video as data images, data audio, and data video respectively; Step S1013: For any data image: Denote the data image with the data text SW as a text image, and denote the data image without the data text SW as a regular image; Step S1014: For any data audio: Perform speech-to-text processing on the data audio based on NLP technology, and denote the obtained text as audio text; Obtain the background noise in the data audio based on AI recognition, and denote the data audio with background noise greater than the standard recognition decibel as noisy audio; In the specific implementation process, in this embodiment, the standard recognition decibel is set to 40 decibels, and the standard recognition decibel can be adjusted according to the actual size of the background noise of the audio in the unstructured data and the actual environment where the audio is located during actual analysis. For example, when the actual environment where the audio is located is a relatively quiet place such as an auditorium or a classroom, the standard recognition decibel can be reduced; By obtaining the background noise and marking the noisy audio, it is possible to distinguish between the audio in which the text can be clearly recognized and the audio in which the text cannot be clearly recognized in the data audio, which helps to distinguish between clear text and fuzzy text during subsequent analysis and annotation; Step S1015: Denote the audio text in the noisy audio as interfering audio text, and denote the audio text not in the noisy audio as regular audio text; Step S1016: For any data video: Extract all the pictures in the data video based on CV technology, and denote the obtained pictures as video pictures; Step S1017: Perform speech-to-text processing on the audio of the data video and extract text from the video pictures based on NLP technology respectively, and denote the obtained texts as video audio text and video picture text; Step S1018: Denote the video picture with video audio text as a video text image, and denote the video picture without video picture text as a video standard image; Obtain the background noise in the data video based on AI recognition, and denote the data video with background noise greater than the standard recognition decibel as a noisy video; Step S1019: Denote the video audio text in the noisy video as interfering video text, and denote the video audio text not in the noisy audio as regular video text; In the specific implementation process, by processing the text, images, audio, and video in the unannotated data, and obtaining the data text SW, regular audio text, interfering audio text, regular video text, interfering video text, text images, video text images, regular images, and video standard images, and unifying them into clear processed text, blurred processed text, text processed images, and regular processed images during subsequent analysis, it is possible to achieve that when there are many data formats and data sources in unstructured data, the text, pictures, audio, and video in the unstructured data become pure text or pictures, so as to establish a unified standard for cross-type data, and further improve the subsequent annotation efficiency; Step S102: Denote a set of unannotated data as single test data, preprocess the single test data using the preprocessing annotation method, and denote the obtained data text SW, regular audio text, and regular video text as clear processed text; Step S1021: Denote the interfering audio text and interfering video text as blurred processed text; Step S1022: Denote the text images and video text images as text processed images; Denote the regular images and video standard images as regular processed images.

[0019] Step S1 further includes: Step S103: Denote multiple sets of unannotated data as parameter test data; For any parameter test data: Based on the preprocessing annotation method, respectively obtain the clear processed text, blurred processed text, text processed images, and regular processed images of the parameter test data, and denote the acquisition times as clear text time, blurred text time, text image time, and regular image time respectively, and denote the memory space occupied by the parameter test data as test occupied memory; Step S104: Obtain the clear text time, blurred text time, text image time, regular image time, and test occupied memory corresponding to all parameter test data, and denote the values corresponding to the minimum test occupied memory and the maximum test occupied memory as XX1 and XX2 respectively; In the specific implementation process, for example, during a data analysis, the test occupied memories of all parameter test data obtained are 2MB, 231MB, 322MB, 500MB, and 12MB respectively. Then, through analysis, it can be obtained that XX1 and XX2 are 2MB and 500MB respectively; Step S105, establish a plane rectangular coordinate system, denoted as the parameter analysis coordinate system. Among them, the unit of the X-axis of the parameter analysis coordinate system is MB, and the unit of the Y-axis is ms; for any parameter test data: in the parameter analysis coordinate system, use the test memory occupancy of the parameter test data as the abscissa, and use the clear text time, blurred text time, text image time, and conventional image time of the parameter test data as the ordinates respectively to mark points, and denote them as the clear text point, blurred text point, text image point, and conventional image point respectively; Step S106, denote the curve obtained by fitting all clear text points as curve QW; denote the curve obtained by fitting all blurred text points as curve MW; denote the curve obtained by fitting all text image points as curve WT; denote the curve obtained by fitting all conventional image points as curve CT; Step S107, for any abscissa X0 from XX1 to XX2 on the X-axis, denote the point with the smallest ordinate among the points on curve QW, curve MW, curve WT, and curve CT with abscissa X0 as the low parameter point; denote the curve formed by all low parameter points in the parameter analysis coordinate system as the parameter analysis curve; In the specific implementation process, for example, during a data analysis, the obtained parameter analysis coordinate system is as Figure 2 shown in PP1. Through analysis, it can be obtained that the curve CF in PP2 is the parameter analysis curve; by obtaining the parameter analysis curve, it is possible to obtain the type of data extraction with the fastest efficiency when using the preprocessing annotation method based on the size of the memory occupied by unstructured data, which helps to preprocess unstructured data with the fastest preprocessing efficiency when facing unstructured data of different sizes. For example, for the point CO1 on curve CF, through analysis, it can be obtained that point CO1 is on curve WT, indicating that when the memory occupied by unstructured data is NC, the text image time is the smallest, that is, the processing efficiency of the text processing image is the fastest. Therefore, when the memory occupied by unstructured data is CO1, the text processing image can be preferentially extracted to achieve the purpose of improving the annotation efficiency; Step S108, denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve QW as the QW preferred interval; denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve MW as the MW preferred interval; Step S109, denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve WT as the WT preferred interval; Denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve CT as the CT preferred interval; Denote the QW preferred interval, MW preferred interval, WT preferred interval, and CT preferred interval as the annotation processing intervals; In the specific implementation process, through Figure 2It can be obtained that the QW preferred interval, MW preferred interval, WT preferred interval and CT preferred interval are [C1, C2], [C2, C3], [C3, C4] and [C4, C5] respectively; for the memory that is in two intervals at the same time, such as C2 that is in both the QW preferred interval and the MW preferred interval, in the subsequent analysis, when the space occupied by the unstructured data is C2, the QW preferred interval or the MW preferred interval can be selected for analysis.

[0020] Step S2, obtaining annotated unstructured data, recorded as annotated data; establishing a neural network, introducing the preprocessing annotation method and the annotation processing interval into the neural network, and establishing a deep annotation method in the neural network; training the neural network based on the unannotated data and the annotated data, and obtaining a preferred neural network based on the training results; Step S2 has the following sub-steps: Step S201, the in-depth annotation method includes: Step S2011, receiving data by neurons in the input layer of the neural network, and recording the memory space occupied by the received data as the annotation memory; recording the interval where the annotation memory is located in the annotation processing interval as the preferred analysis interval, and recording the data type corresponding to the preferred analysis interval as the preferred screening type, wherein the data type includes clear processing text, fuzzy processing text, text processing image and conventional processing image; Step S2012, based on the preprocessing annotation method, preferentially obtaining the preferred screening type of data, and obtaining data types other than the preferred screening type; Step S2013, annotating the data type obtained by the preprocessing annotation method based on the target detection frame, named entity recognition and speaker separation technology, and outputting the annotated data through the neurons of the output layer of the neural network; In the specific implementation process, the data types obtained by the preprocessing annotation method can be annotated according to the annotation tools that can be applied in the actual analysis, so as to ensure that the annotation is more comprehensive and efficient.

[0021] Step S2 also includes: step S202, for any value X0 from XX1 to XX2, inputting the unannotated data with the annotation X0 and the annotated data into the neural network respectively, and performing annotation processing, wherein the method of annotating the annotated data is recorded as the annotated method, and when the annotated data is input into the neural network, the neural network re-annotates the annotated data based on the annotated method; Step S203, when the time for which the unannotated data is annotated by the neural network is less than or equal to the time for which the annotated data is annotated by the neural network, X0 is recorded as optimized memory and the annotated method corresponding to X0 is removed from the neural network; Step S204, when the time for the unannotated data to be annotated by the neural network is greater than the time for the annotated data to be annotated by the neural network, record X0 as the storage to be learned and record the corresponding annotated method of X0 as the preferred annotation method of X0; In the specific implementation process, for example, during a data analysis, the time for the unannotated data to be annotated by the neural network is 1 minute, and the time for the annotated data to be annotated by the neural network is 0.8 minute. This indicates that the annotation efficiency of the annotated method for the annotated data is better than that of the neural network. Therefore, the neural network can be improved by using the annotated method for learning, and in subsequent annotations, based on the preferred annotation method learned by the neural network, the preprocessing annotation method stored in the neural network, and the annotation processing interval, the unannotated data can be annotated quickly and accurately; Step S205, when all values from XX1 to XX2 are recorded as the optimized storage, or are recorded as the storage to be learned and the corresponding preferred annotation method is entered into the neural network, record the neural network as the preferred neural network; Step S206, for the unannotated data annotated by the neural network: when the annotation storage is the optimized storage, annotate the unannotated data based on the preprocessing annotation method and the annotation processing interval; when the annotation storage is the storage to be learned, annotate the unannotated data based on the preferred annotation method of the storage to be learned.

[0022] Step S3, perform automated annotation on the unannotated data based on the preferred neural network.

[0023] Embodiment 2, please refer to Figure 3 as shown Figure 3 illustrates a schematic structural diagram of an electronic device. The electronic device may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the method for automatically annotating unstructured data based on a neural network are run to achieve the following functions: First, obtain the unannotated unstructured data, recorded as unannotated data; establish a preprocessing annotation method based on NLP technology and CV technology; use the preprocessing annotation method to preprocess the unannotated data, and obtain an annotation processing interval based on the preprocessing result; then obtain the annotated unstructured data, recorded as annotated data; establish a neural network, introduce the preprocessing annotation method and the annotation processing interval into the neural network, and establish an in-depth annotation method in the neural network; train the neural network based on the unannotated data and the annotated data, and obtain a preferred neural network based on the training result; finally, perform automated annotation on the unannotated data based on the preferred neural network.

[0024] In addition, when the logic instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0025] Embodiment 3, this application also provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the method for automatically annotating unstructured data based on a neural network provided by the above-mentioned various methods. This method includes: First, obtain unannotated unstructured data, denoted as unannotated data; establish a preprocessing annotation method based on NLP technology and CV technology; use the preprocessing annotation method to preprocess the unannotated data, and obtain an annotation processing interval based on the preprocessing result; then obtain annotated unstructured data, denoted as annotated data; establish a neural network, introduce the preprocessing annotation method and the annotation processing interval into the neural network, and establish a deep annotation method in the neural network; train the neural network based on the unannotated data and the annotated data, and obtain an optimized neural network based on the training result; finally, automatically annotate the unannotated data based on the optimized neural network.

[0026] Example 4. The present application also provides a computer-readable storage medium. The present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method for automatically annotating unstructured data based on a neural network are run to implement the following functions: First, obtain unannotated unstructured data, denoted as unannotated data; establish a preprocessing annotation method based on NLP technology and CV technology; use the preprocessing annotation method to preprocess the unannotated data, and obtain an annotation processing interval based on the preprocessing result; then obtain annotated unstructured data, denoted as annotated data; establish a neural network, introduce the preprocessing annotation method and the annotation processing interval into the neural network, and establish a deep annotation method in the neural network; train the neural network based on the unannotated data and the annotated data, and obtain an optimized neural network based on the training result; finally, automatically annotate the unannotated data based on the optimized neural network.

[0027] Through the description of the above embodiments, the embodiments of the present invention can be provided as a method, a system or a computer program product. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0028] In the embodiments provided by the present application, it should be understood that the disclosed system or method can be implemented in other ways. The above-described embodiments are merely illustrative. For example, the division of modules or units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the system, module and unit can be in an electrical, mechanical or other form.

[0029] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for automatically annotating unstructured data based on a neural network, characterized in that, It includes the following steps: Obtain unannotated unstructured data, denoted as unannotated data; establish a preprocessing annotation method based on NLP technology and CV technology; Preprocess the unannotated data using the preprocessing annotation method, and obtain an annotation processing interval based on the preprocessing result; Obtain annotated unstructured data, denoted as annotated data; establish a neural network, introduce the preprocessing annotation method and the annotation processing interval into the neural network, and establish an in-depth annotation method in the neural network; Train the neural network based on the unannotated data and the annotated data, and obtain an optimized neural network based on the training result; Automatically annotate the unannotated data based on the optimized neural network.

2. The method for automatically annotating unstructured data based on a neural network according to claim 1, characterized in that, The preprocessing annotation method includes: Obtain unannotated data, extract the text in the unannotated data based on NLP technology, and sequentially record the obtained independent text segments as data text SW1 to data text SW according to the extraction order of the independent text segments in the extraction result n , where n is the number of independent text segments; Extract images, audio, and video in the unannotated data based on CV technology, and denote the extracted pictures, audio, and video as data images, data audio, and data videos respectively.

3. The method for automatically annotating unstructured data based on a neural network according to claim 2, wherein The preprocessing annotation method also includes: For any data image: Denote the data image with data text SW as a text image, and denote the data image without data text SW as a regular image; For any data audio: Perform speech-to-text processing on the data audio based on NLP technology, and denote the obtained text as audio text; Obtain the background noise in the data audio based on AI recognition, and denote the data audio with background noise greater than the standard recognition decibel in the data audio as noisy audio; Denote the audio text in the noisy audio as interfering audio text, and denote the audio text not in the noisy audio as regular audio text.

4. The method for automatically annotating unstructured data based on a neural network according to claim 3, wherein The preprocessing annotation method also includes: For any data video: Extract all pictures in the data video based on CV technology, and denote the obtained pictures as video pictures; Perform speech-to-text processing on the audio of the data video and extract text from the video pictures respectively based on NLP technology, and denote the obtained texts as video audio text and video picture text respectively; Denote the video pictures with video audio text as video text images, and denote the video pictures without video picture text as video standard images; Obtain the background noise in the data video based on AI recognition, and denote the data video with background noise greater than the standard recognition decibel in the data video as noisy video; Denote the video audio text in the noisy video as interfering video text, and denote the video audio text not in the noisy audio as regular video text.

5. The method for automatically annotating unstructured data based on a neural network according to claim 4, characterized in that Preprocessing the unannotated data using the preprocessing annotation method and obtaining an annotation processing interval based on the preprocessing result includes: Denote a set of unannotated data as single test data, preprocess the single test data using the preprocessing annotation method, and denote the obtained data text SW, regular audio text, and regular video text as clearly processed text; Denote the interfering audio text and the interfering video text as blurred processed text; Denote the text images and the video text images as text processed images; Denote the regular images and the video standard images as regular processed images.

6. The method for automatically annotating unstructured data based on a neural network according to claim 5, wherein Preprocessing the unannotated data using the preprocessing annotation method and obtaining an annotation processing interval based on the preprocessing result also includes: Record multiple groups of unannotated data as parameter test data; for any parameter test data: respectively obtain the clear processed text, blurred processed text, text processed image, and conventional processed image of the parameter test data based on the preprocessing annotation method, and record the acquisition times as the clear text time, blurred text time, text image time, and conventional image time respectively, and record the memory space occupied by the parameter test data as the test memory occupancy; Obtain the clear text time, blurred text time, text image time, conventional image time, and test memory occupancy corresponding to all parameter test data, and record the values corresponding to the minimum test memory occupancy and the maximum test memory occupancy as XX1 and XX2 respectively.

7. The method for automatically annotating unstructured data based on a neural network according to claim 6, wherein Preprocess the unannotated data using the preprocessing annotation method, and the annotation processing interval obtained based on the preprocessing result further includes: Establish a plane rectangular coordinate system, denoted as the parameter analysis coordinate system, where the unit of the X-axis of the parameter analysis coordinate system is MB and the unit of the Y-axis is ms; for any parameter test data: in the parameter analysis coordinate system, use the test memory occupancy of the parameter test data as the abscissa, and use the clear text time, blurred text time, text image time, and conventional image time of the parameter test data as the ordinates respectively for punctuation, and denote them as the clear text point, blurred text point, text image point, and conventional image point respectively; Denote the curve obtained by fitting all clear text points as curve QW; denote the curve obtained by fitting all blurred text points as curve MW; denote the curve obtained by fitting all text image points as curve WT; denote the curve obtained by fitting all conventional image points as curve CT; For any abscissa X0 from XX1 to XX2 on the X-axis, denote the point with the minimum ordinate among the points on curve QW, curve MW, curve WT, and curve CT with abscissa X0 as the low parameter point; denote the curve formed by all low parameter points in the parameter analysis coordinate system as the parameter analysis curve.

8. The method for automatically annotating unstructured data based on a neural network according to claim 7, wherein Preprocess the unannotated data using the preprocessing annotation method, and the annotation processing interval obtained based on the preprocessing result further includes: Denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve QW as the QW preferred interval; denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve MW as the MW preferred interval; Denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve WT as the WT preferred interval; Denote the interval formed by the abscissas of the region where the parameter analysis curve coincides with curve CT as the CT preferred interval; Denote the QW preferred interval, MW preferred interval, WT preferred interval, and CT preferred interval as the annotation processing interval.

9. The method for automatically annotating unstructured data based on a neural network according to claim 8, wherein The in-depth annotation method includes: The neurons in the input layer of the neural network receive data, and record the memory space occupied by the received data as the annotation memory occupancy; denote the interval where the annotation memory occupancy is located in the annotation processing interval as the preferred analysis interval, and denote the data type corresponding to the preferred analysis interval as the preferred screening type, where the data types include clear processed text, blurred processed text, text processed image, and conventional processed image; Based on the preprocessing annotation method, the preferred screening type of data is preferentially obtained, and the data types other than the preferred screening type are obtained; Based on object detection frames, named entity recognition, and speaker separation technologies, the data types obtained by the preprocessing annotation method are annotated, and the annotated data is output by the neurons of the output layer of the neural network.

10. The method for automatically annotating unstructured data based on a neural network according to claim 9, characterized in that The neural network is trained based on the unannotated data and the annotated data, and the preferred neural network is obtained based on the training results, including: For any value X0 from XX1 to XX2, the unannotated data and the annotated data with the annotation occupancy of X0 are respectively input into the neural network and annotated. Among them, the method of annotating the annotated data is recorded as the annotated method. When the annotated data is input into the neural network, the neural network re-annotates the annotated data based on the annotated method; When the time for the unannotated data to be annotated by the neural network is less than or equal to the time for the annotated data to be annotated by the neural network, X0 is recorded as the optimized occupancy, and the annotated method corresponding to X0 is removed from the neural network; When the time for the unannotated data to be annotated by the neural network is greater than the time for the annotated data to be annotated by the neural network, X0 is recorded as the to-be-learned occupancy, and the annotated method corresponding to X0 is recorded as the preferred annotation method of X0; When all values from XX1 to XX2 are recorded as optimized occupancy, or are recorded as to-be-learned occupancy and the corresponding preferred annotation methods are entered into the neural network, the neural network is recorded as the preferred neural network; For the unannotated data annotated by the neural network: when the annotation occupancy is optimized occupancy, the unannotated data is annotated based on the preprocessing annotation method and the annotation processing interval; when the annotation occupancy is to-be-learned occupancy, the unannotated data is annotated based on the preferred annotation method of the to-be-learned occupancy.

Citation Information

Patent Citations

  • Unstructured data analytics systems and methods

    CN107368506A

  • Method and system for automated column type annotation

    EP4296865A1

  • Annotation pipeline for machine learning algorithm training and optimization

    US20210034920A1

  • Systems and methods for automatic context-based annotation

    US20220051009A1

  • Annotation creation, storage, and visualization methods for knowledge management in external contexts

    US20220129615A1