Data identification method and device

By extracting and clustering data feature vectors, identifying data sample libraries with high propagation and automatically updating them, the problems of inefficiency and high cost of traditional data recognition methods are solved, and wider data coverage and lower labor costs are achieved.

CN119939270APending Publication Date: 2025-05-06SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510090550.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional specific types of data identification methods are difficult to fully cover existing specific types of data forms, and identification is inefficient and requires a lot of labor costs.

Method used

By obtaining reported data, extracting its feature vectors, clustering based on text similarity, determining that the data in the group whose propagation exceeds the threshold is data of the data sample library, and automatically update the data sample library.

Benefits of technology

It improves the coverage of target type data identification, reduces labor costs, and actively perceives the latest target type data cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939270A_ABST
    Figure CN119939270A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data identification method and device, computer equipment, a medium and a program product. The data identification method comprises the following steps: acquiring reported data; extracting a plurality of feature vectors corresponding to the reported data; clustering the plurality of feature vectors on the basis of file similarity to obtain different groups; under the condition that the propagation degree of any reported data in the same group is greater than a preset threshold value, determining the feature vector of each piece of reported data in the same group as data of a data sample library; the propagation degree of the reported data is obtained according to the number of the data division areas associated with the reported data; based on the data of the data sample library, updating the data sample library; wherein the data sample library is used for identifying target type data. According to the technical scheme provided by the embodiment of the invention, the latest target type data case can be actively sensed to automatically update the data sample library, so that the coverage range of target type data identification is expanded, and meanwhile, the labor cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of information security technology, and in particular, to a data identification method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] In today's Internet environment, with the rapid development of various online businesses such as social platforms, e-commerce, and games.

[0003] The technological demands in the field of data recognition are growing rapidly, but traditional methods for identifying specific types of data are difficult to fully cover existing specific types of data forms, have low recognition efficiency and require a lot of manpower costs.

[0004] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the invention

[0005] The embodiments of the present application provide a data identification method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the technical problems raised above.

[0006] One aspect of an embodiment of the present application provides a data identification method, the method comprising: Get reported data; Extracting multiple feature vectors corresponding to the reported data; Based on the text similarity, clustering the plurality of feature vectors to obtain different groups; When the propagation degree of any reported data in the same group is greater than a preset threshold, the feature vectors of each reported data in the same group are determined as data sample library data; the propagation degree of the reported data is obtained according to the number of data partition areas associated with the reported data; Based on the data in the data sample library, the data sample library is updated; wherein the data sample library is used to identify target type data.

[0007] Optionally, the method further comprises: Obtain target detection data; Extracting a target feature vector corresponding to the target detection data; The target feature vector is compared with the vector in the data sample library, and whether the target detection data is the target type data is determined according to the comparison result.

[0008] Optionally, updating the data sample library based on the data in the data sample library includes: Create multiple processes based on online service instances; The multiple processes are updated to update the data sample library and respond to real-time requests from requesting clients.

[0009] Optionally, updating the data sample library based on the data in the data sample library includes: Packing the data sample library data by overloading the service to create a data sample library and set the data sample library version; The reload service is bound to the online service so that the online service can perceive the data sample library version in real time; wherein the online service is deployed in a plurality of online service instances.

[0010] Optionally, the updating of the data sample library based on the process includes: Recording the update status of the online service instance and the current version of the data sample library; In the case where the current version is inconsistent with the running version of the online service instance, querying the update status of the online service instance; and When the update status of the online service instance is an updateable status, the data sample library is updated online.

[0011] Optionally, when the update status of the online service instance is updateable, updating the data sample library online includes: In the process of updating the data sample library, setting the update state of the online service instance to a non-updatable state; The non-updatable state is used to indicate that only one process among multiple online service instances is updating the data sample library.

[0012] Optionally, the method further comprises: In a preset time window, when the number of times that the detection data is identified as the target type data reaches a threshold, the target type data identification is stopped in the current time window.

[0013] Another aspect of an embodiment of the present application provides a data identification device, the device comprising: Acquisition module, used to obtain reported data; An extraction module, used to extract multiple feature vectors corresponding to the reported data; A clustering module, used for clustering the plurality of feature vectors based on text similarity to obtain different groups; A determination module, configured to determine the characteristic vectors of each reported data in the same group as data sample library data when the propagation degree of any reported data in the same group is greater than a preset threshold; the propagation degree of the reported data is obtained according to the number of data partition areas associated with the reported data; An updating module is used to update a data sample library based on the data in the data sample library; wherein the data sample library is used to identify target type data.

[0014] Another aspect of an embodiment of the present application provides a computer device, including: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein: the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0015] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.

[0016] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the method described above when executed by a processor.

[0017] The embodiments of the present application using the above technical solution may include the following advantages: obtaining reported data and extracting corresponding multiple feature vectors, and grouping multiple feature vectors by clustering. When the propagation degree of any reported data in a group exceeds a preset threshold, the feature vector of the group is included in the data sample library, and the data sample library is automatically updated accordingly for subsequent data identification. This method can actively perceive the latest target type data cases to automatically update the data sample library, thereby increasing the coverage of target type data identification and reducing labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0019] Figure 1 A flowchart of a data identification method according to Embodiment 1 of the present application is schematically shown; Figure 2 Schematically shows Figure 1 Flow chart of sub-steps in step S208; Figure 3 Schematically shows Figure 1 Flow chart of sub-steps in step S208; Figure 4 Schematically shows Figure 1Flow chart of sub-steps in step S208; Figure 5 The following schematically shows a newly added flow chart of the data identification method according to the first embodiment of the present application; Figure 6 A block diagram schematically shows a data identification device according to the second embodiment of the present application; and Figure 7 The hardware architecture diagram of the computer device according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0021] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.

[0023] To facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: Community data currently presents features such as complex account characteristics, obscure and changeable comment content, and similar text or picture content of data without information being sent in batches. However, the applicant understands that the existing technology relies on rules of professional knowledge or classification models based on training samples, which is difficult to cover existing specific types of data and requires extremely high labor costs. To this end, the embodiment of the present application provides a data recognition technology solution. In this technical solution, it is possible to greatly reduce labor costs and actively perceive the latest cases of specific types of data. See below for details.

[0024] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0025] The technical solutions of the present application are described below through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be construed as being limited to the embodiments described here.

[0026] Embodiment 1 Figure 1 The flowchart of the data identification method according to the first embodiment of the present application is schematically shown.

[0027] like Figure 1 As shown, the data identification method may include steps S200 to S208, wherein: Step S200, obtaining reporting data.

[0028] Step S202: extract multiple feature vectors corresponding to the reported data.

[0029] Step S204: clustering the multiple feature vectors based on text similarity to obtain different groups.

[0030] Step S206, when the propagation degree of any reported data in the same group is greater than a preset threshold, the feature vector of each reported data in the same group is determined as data sample library data; the propagation degree of the reported data is obtained according to the number of data partition areas associated with the reported data.

[0031] Step S208: updating the data sample library based on the data in the data sample library; wherein the data sample library is used to identify target type data.

[0032] The data identification method provided in this embodiment obtains reported data and extracts corresponding multiple feature vectors, and groups the multiple feature vectors by clustering. When the propagation degree of any reported data in a group exceeds a preset threshold, the feature vector of the group is included in the data sample library, and the data sample library is automatically updated accordingly for subsequent identification of target type data. This method can actively perceive the latest target type data cases to automatically update the data sample library, thereby increasing the coverage of target type data identification and reducing labor costs.

[0033] The following combination Figure 1 , each step in steps S200~S208 and other optional steps are explained in detail.

[0034] Step S200 , get the reported data.

[0035] The reported data can be the interactive content that is not in compliance with the regulations reported by the user. The reported data can include reported text and reported pictures. Based on the idea of ​​user autonomy, the reported data can be screened by the user, thereby improving the accuracy of the mining data samples.

[0036] Step S202 , extracting multiple feature vectors corresponding to the reported data.

[0037] In some embodiments, the reported text and the reported image are encoded, and the text encoding method can be a sentence-level vector representation based on a pre-trained BERT model. The BERT model can convert text into a vector representation. For example, a sentence of length 30 can be represented as a matrix of (30*512), where 512 is the feature dimension of each word, and the feature dimension is used to describe the number of numerical values ​​or parameters of data features. The image encoding method can be based on the VIT model, which extracts features from the pixel information of the input image and aggregates them into a high-dimensional feature vector. The VIT model is a model for image processing, which can extract features from images and convert images into vector representations. It should be noted that other models can also be used as needed.

[0038] Step S204 , based on text similarity, the multiple feature vectors are clustered to obtain different groups.

[0039] Use clustering algorithms (such as community clustering) to obtain different groups. Clustering algorithm is an unsupervised learning algorithm that is used to automatically classify data into several groups based on text similarity. Community clustering is a clustering method specifically used for network data or relational data, which can be used to analyze social networks, comment areas and other scenarios. In other words, the more times a sample is reported, the greater the probability that it will be judged as non-standard data. However, in community scenarios, it is easy to mistakenly hurt public opinion-related comments by relying solely on the number of reports. In addition, the reported interactive content is often not exactly the same. For example, some texts may only have a few words changed in the middle, or the pictures may use different filters. Therefore, the cumulative number of reports of these samples may not reach the set storage threshold. In order to avoid these problems as much as possible, it is necessary to use a clustering algorithm to cluster the samples, and then further superimpose the corresponding propagation rules for judgment.

[0040] Step S206 When the propagation degree of any reported data in the same group is greater than a preset threshold, the characteristic vector of each reported data in the same group is determined as the data sample library data; the propagation degree of the reported data is obtained according to the number of data partition areas associated with the reported data.

[0041] In some embodiments, the degree of propagation can be determined by whether the content in the current group appears in different partitions (data partition areas). It should be noted that partition refers to the division of data according to platform characteristics, user interaction methods or logical classifications. If some content only appears in one partition or one comment area, it is possible that similar content is reported due to public discussion. To this end, it is possible to check whether the comments in the current group appear in multiple partitions and comment areas, and whether they are evenly propagated in each partition or comment area. For example, the number of partitions covered should be greater than or equal to 3, and the ratio of the total frequency of the partition with the least comments in the group to the total frequency of the partition with the most comments in the group must be greater than 0.5. For groups that meet the current conditions, the feature vectors in the same group can be used as data sample library data.

[0042] Step S208 , based on the data in the data sample library, updating the data sample library; wherein the data sample library is used to identify target type data.

[0043] The target type data may be non-compliant data, such as data that does not comply with information security specifications or network security specifications. In some embodiments, the image feature vectors and text feature vectors in the data sample library data may be placed in different data sample libraries, respectively. For example, the text feature vectors may be placed in a text vector library, and the image feature vectors may be placed in an image vector library, and both the text vector library and the image vector library may be used as data sample libraries.

[0044] In this embodiment, the reported data is obtained and the corresponding multiple feature vectors are extracted, and the multiple feature vectors are grouped by clustering. When the propagation degree of any reported data in the group exceeds the preset threshold, the feature vector of the group is included in the data sample library, and the data sample library is automatically updated accordingly for subsequent data recognition. This method can actively perceive the latest data cases to automatically update the data sample library, thereby increasing the coverage of target type data recognition and reducing labor costs.

[0045] In an optional embodiment, if Figure 2 As shown, step S208 may include: Step S300: creating multiple processes based on the online service instance.

[0046] Step S302: updating the plurality of processes to update the data sample library and responding to the real-time request of the requesting client according to the updated data sample library.

[0047] The service is deployed on multiple instances at the same time. When a data request is made online, a random instance participates in the calculation. The number of online request data in a unit time is far more than one, and multiple instances are required to participate in the calculation at the same time. The update can be carried out in parallel or the process can be updated one by one. The update of the process one by one makes the update of the data sample library not affect the real-time request in the online service instance. In some embodiments, the data sample library can be updated regularly every day, and the instance needs to reload the updated data sample library. In order not to affect the online data calculation, Gunicorn can be used to create multiple Flask processes. Gunicorn (Green Unicorn) is a Python WSGI HTTP server specifically used to run Python Web applications. It can handle multiple concurrent requests and is often used in production environments to improve the performance and scalability of Web applications. Flask itself is a lightweight Web framework. By default, only one request can be processed through its built-in development server (single thread, single process). In this embodiment, based on multiple processes, the services in the online service instance are updated one by one to avoid blocking the real-time request in the online service instance. As a result, the update of the data sample library and the real-time request can be made independent of each other.

[0048] In an optional embodiment, if Figure 3 As shown, step S208 may include: Step S400, packaging the data sample library data through the reload service to create a data sample library and set the data sample library version; Step S402: The reload service is bound to the online service so that the online service can perceive the data sample library version in real time; wherein the online service is deployed in a plurality of online service instances.

[0049] The reload service organizes the data sample library data into available file packages. Each data sample library can correspond to a text vector storage file and an image vector storage file. The file package can also have a file recording the version information of the generated data sample library, and the version information of the data sample library will be recorded in the database. The recorded database can be Redis, which is a database tool that is good at processing real-time data and can be used to record and manage temporary data.

[0050] In this embodiment, different versions of information can be tracked through the reload service, and the binding of the reload service with the online service can enable the online service to access the version information in the database at any time. Therefore, the online service can perceive the changes in the data sample library version in real time, thereby improving the accuracy and timeliness of identifying the target type data.

[0051] In an optional embodiment, if Figure 4As shown, step S208 may include: Step S500: Record the update status of the online service instance and the current version of the data sample library.

[0052] Step S502: when the current version is inconsistent with the running version of the online service instance, query the update status of the online service instance.

[0053] Step S504: when the update status of the online service instance is an updateable status, update the data sample library online.

[0054] The update status of the online service instance and the current version of the data sample library can be recorded in the database. The online service always has high QPS (query rate per second) requests. After processing each request, the database will be accessed once to obtain the version information in the current database record. If the version information recorded in the database is inconsistent with the version running in the online service instance, and the update status of the online service instance is the update status, the data sample library will be updated. In some embodiments, the update status of the online service instance can be recorded in the Redis database through buzy_signal, and the current version of the data sample library can be recorded through version. If the recorded version information is inconsistent with the version running in the online service instance, buzy_signal is further accessed, and if buzy_signal is 0, it is updated.

[0055] In this embodiment, by comparing the data sample library versions and setting the update status, not only can the data sample library be automatically updated in a timely manner, but also conflicts and resource consumption caused by concurrent updates can be avoided as much as possible.

[0056] In an optional embodiment, step S504 may include: In the process of updating the data sample library, the update state of the online service instance is set to a non-updatable state; wherein the non-updatable state is used to indicate that only one process among multiple online service instances is updating the data sample library.

[0057] In some embodiments, when updating the data sample library, the update status buzy_signal of the record is set to 1. When it is detected that the update status buzy_signal is 1, no update is performed to continue processing the real-time request. After the update is completed, the update status buzy_signal is reset to 0.

[0058] In this embodiment, when the data sample library is updated, the update status is set to non-updatable, so that only one process in multiple online service instances updates the data sample library, thereby avoiding blocking real-time requests of online services.

[0059] In an optional embodiment, if Figure 5 As shown, the data identification method may also include: Step S600, obtaining target detection data.

[0060] Step S602: extracting a target feature vector corresponding to the target detection data.

[0061] Step S604: compare the target feature vector with the vector in the data sample library, and determine whether the target detection data is the target type data according to the comparison result.

[0062] In some embodiments, the target feature vector and the vector in the data sample library can be compared to obtain text similarity. The method for calculating similarity can be cosine similarity, which is a mathematical method for measuring the similarity between two vectors. The cosine value of the angle between the two vectors is calculated to determine their similarity. The vector in the target type data sample library closest to the target feature vector can be retrieved through the vector retrieval library Faiss. Faiss is a vector retrieval library used to quickly retrieve the most similar vectors in massive high-dimensional vector data. Faiss can support searching and positioning in vector sets of any size, and will have different performances on the CPU (central processing unit) and GPU (graphics processing unit). On the GPU, there will be faster retrieval speed and lower memory usage, which can effectively help us withstand more QPS (query rate per second) pressure. When the text similarity exceeds the set threshold, the target detection data is identified as target type data, and the target type data can be made visible (only visible to the sending user) or marked with related users, etc.

[0063] In this embodiment, by comparing the target feature vector with the vectors in the data sample library, the accuracy and efficiency of identifying the target type data can be improved.

[0064] In an optional embodiment, the data identification method may further include: In a preset time window, when the number of times that the detection data is identified as the target type data reaches a threshold, the target type data identification is stopped in the current time window.

[0065] When there is no way to guarantee 100% accuracy of samples in the data sample library, a certain mechanism for preventing accidental damage is required. In some embodiments, different time window recall thresholds can be set for single samples and total samples (target detection data) respectively. When the threshold is triggered, an alarm message is issued and the target type data recall is stopped within the time window. Recall refers to comparing and identifying similar content from a known data sample library to prevent the spread of malicious content and mark or restrict suspicious users. A deletion interface for the data sample library can also be developed separately to remove the accidentally injured target type data samples that are found. The recalled data can also be manually obtained and marked regularly. If it does not meet the requirements, the corresponding data sample data can be deleted or optimized. In addition, a type of data is generally only active for a period of time. After a period of time, the data sample library will inevitably deteriorate. The expiration time of each sample in the data sample library can be set to 90 days by default. All data sample libraries are refreshed once a day, and expired samples are deleted from the data sample library. At the same time, manual intervention can also be performed to forcibly delete samples, or the expiration time can be manually set for short-term data samples. The log data generated during the target type data identification process can also be stored for offline analysis to check the performance of recall data and related models and strategies. At the same time, the content that needs to be monitored (samples entering the data sample library daily, real-time recall data) can be visualized through dashboard monitoring.

[0066] In this embodiment, by setting the time window and the detection number threshold, it is possible to reduce misjudgment in the target type data identification process and improve the fault tolerance and stability.

[0067] In order to make the present application easier to understand, an exemplary application is provided below.

[0068] In this exemplary application, taking online comment content as an example, the recognition process is as follows: S1. The user sends a comment.

[0069] S2. The comment content is transmitted to the server for data identification through the comment service.

[0070] S3. The server calls the interface to compare the encoded comment content (target feature vector) with the encoded data in the data sample library (vector of the data sample library) to obtain text similarity.

[0071] S4. Determine whether the text similarity reaches the set threshold: When the set threshold is reached, the comment content will be processed as self-viewable (only visible to the sending user); If the set threshold is not reached, the comment content will be released for normal display.

[0072] In this exemplary application, by comparing the target feature vector with the vectors in the data sample library, the accuracy and efficiency of identifying the target type data can be improved.

[0073] Embodiment 2 Figure 6 The block diagram of the data recognition device according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium, and are executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Figure 6 As shown, the apparatus 1600 may include: an extraction module 1610, an extraction module 1620, a clustering module 1630, a determination module 1640 and an update module 1650, wherein: The acquisition module 1610 is used to acquire the reported data; An extraction module 1620, configured to extract a plurality of feature vectors corresponding to the reported data; A clustering module 1630, for clustering the plurality of feature vectors based on text similarity to obtain different groups; A determination module 1640 is used to determine the feature vectors of each reported data in the same group as data sample library data when the propagation degree of any reported data in the same group is greater than a preset threshold; the propagation degree of the reported data is obtained according to the number of data partition areas associated with the reported data; The updating module 1650 is used to update the data sample library based on the data in the data sample library; wherein the data sample library is used to identify target type data.

[0074] As an optional embodiment, the data device 1600 may also be used for: Obtain target detection data; Extracting a target feature vector corresponding to the target detection data; The target feature vector is compared with the vector in the data sample library, and whether the target detection data is the target type data is determined according to the comparison result.

[0075] As an optional embodiment, the updating module 1650 is further configured to: Create multiple processes based on online service instances; The multiple processes are updated one by one to update the data sample library and respond to real-time requests from requesting clients.

[0076] As an optional embodiment, the updating module 1650 is further configured to: Packing the data sample library data by overloading the service to create a data sample library and set the data sample library version; The reload service is bound to the online service so that the online service can perceive the data sample library version in real time; wherein the online service is deployed in a plurality of online service instances.

[0077] As an optional embodiment, the updating module 1650 is further configured to: Recording the update status of the online service instance and the current version of the data sample library; In the case where the current version is inconsistent with the running version of the online service instance, querying the update status of the online service instance; and When the update status of the online service instance is an updateable status, the data sample library is updated online.

[0078] As an optional embodiment, the data identification device 1600 is further used for: In a preset time window, when the number of times that the detection data is identified as the target type data reaches a threshold, the target type data identification is stopped in the current time window.

[0079] Embodiment 3 Figure 7 The schematic diagram of the hardware architecture of a computer device 10000 suitable for implementing the data identification method according to the third embodiment of the present application is schematically shown. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. Figure 7 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as a hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk equipped on the computer device 10000, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 10010 can also include both the internal storage module of the computer device 10000 and its external storage device. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed in the computer device 10000, such as program codes of the data identification method, etc. In addition, the memory 10010 can also be used to temporarily store various data that have been output or are to be output.

[0080] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0081] The network interface 10030 may include a wireless network interface or a wired network interface, and the network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0082] It should be pointed out that Figure 7 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementing all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.

[0083] In this embodiment, the data identification method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as processor 10020) to complete the embodiment of the present application.

[0084] Embodiment 4 An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the data identification method in the embodiment are implemented.

[0085] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the computer-readable storage medium can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store an operating system and various application software installed on the computer device, such as the program code of the data identification method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or are to be output.

[0086] Embodiment 5 An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0087] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0088] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A data identification method, characterized in that: The method comprises: Get reported data; Extracting multiple feature vectors corresponding to the reported data; Based on text similarity, clustering the plurality of feature vectors to obtain different groups; When the propagation degree of any reported data in the same group is greater than a preset threshold, the feature vectors of each reported data in the same group are determined as data sample library data; the propagation degree of the reported data is obtained according to the number of data partition areas associated with the reported data; Based on the data in the data sample library, the data sample library is updated; wherein the data sample library is used to identify target type data.

2. The method according to claim 1, characterized in that The method further comprises: Obtain target detection data; Extracting a target feature vector corresponding to the target detection data; The target feature vector is compared with the vector in the data sample library, and whether the target detection data is the target type data is determined according to the comparison result.

3. The method according to claim 1, characterized in that The updating of the data sample library based on the data in the data sample library includes: Create multiple processes based on online service instances; The multiple processes are updated to update the data sample library and respond to real-time requests from requesting clients.

4. The method according to claim 3, characterized in that The updating of the data sample library based on the data in the data sample library includes: Packing the data sample library data by overloading the service to create a data sample library and set the data sample library version; The reload service is bound to the online service so that the online service can perceive the data sample library version in real time; wherein the online service is deployed in a plurality of online service instances.

5. The method according to claim 3, characterized in that: The updating of the data sample library based on the process includes: Recording the update status of the online service instance and the current version of the data sample library; In the case where the current version is inconsistent with the running version of the online service instance, querying the update status of the online service instance; and When the update status of the online service instance is an updateable status, the data sample library is updated online.

6. The method according to claim 5, characterized in that When the update status of the online service instance is updateable, updating the data sample library online includes: In the process of updating the data sample library, setting the update state of the online service instance to a non-updatable state; The non-updatable state is used to indicate that only one process among multiple online service instances is updating the data sample library.

7. The method according to claim 1, characterized in that The method further comprises: In a preset time window, when the number of times that the detection data is identified as the target type data reaches a threshold, the target type data identification is stopped in the current time window.

8. A data identification device, characterized in that: The device comprises: Acquisition module, used to obtain reported data; An extraction module, used to extract multiple feature vectors corresponding to the reported data; A clustering module, used for clustering the plurality of feature vectors based on text similarity to obtain different groups; A determination module, configured to determine the characteristic vectors of each reported data in the same group as data sample library data when the propagation degree of any reported data in the same group is greater than a preset threshold; the propagation degree of the reported data is obtained according to the number of data partition areas associated with the reported data; An updating module is used to update a data sample library based on the data in the data sample library; wherein the data sample library is used to identify target type data.

9. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 7 are implemented.