Information identification method and device
By obtaining the matching results of the information to be identified and the data set of violation words and using the in-depth model to judge its violation, the problem of identification of violation information in the network is solved, and the effective identification of violation information and healthy maintenance of the network environment is achieved.
Patent Information
- Application Number
- CN202111383614.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-11-19
AI Technical Summary
How to effectively identify illegal information spreading on the Internet to prevent it from misleading and affecting netizens and enterprises.
By obtaining the information to be identified, the matching result with the violation word data set is determined, and the depth model is used to determine whether the information containing the target violation word is violated. This method combines preset conditions and deep learning techniques to ensure the accuracy of recognition.
It realizes effective identification of illegal information, creates a healthier network environment, reduces the spread of illegal information, and improves information security.
Smart Images

Figure CN114282097B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an information identification method and device thereof. Background Art
[0002] With the popularization of electronic products, mobile phones, computers and other electronic products have gradually become an indispensable part of people's lives. At the same time, with the rapid development of the Internet industry, various web pages can provide users with more and more information. However, as online information becomes easier to obtain, some lawless elements or people with ulterior motives spread some illegal information on the Internet, which can easily cause irreversible misleading and impact on many netizens and corporate organizations.
[0003] Therefore, how to effectively identify illegal information is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present application discloses an information identification method and device, which can effectively identify illegal information and help create a healthier network environment.
[0005] In a first aspect, an embodiment of the present application provides an information identification method, the method comprising:
[0006] Acquire first information to be identified;
[0007] Determine a matching result between the first information to be identified and the illegal word data set; the matching result includes a target illegal word that exists in both the first information to be identified and the illegal word data set;
[0008] If the matching result satisfies the preset condition, obtaining the second information to be identified including the target illegal word in the first information to be identified;
[0009] The deep model is used to determine whether the second information to be identified is in violation of regulations.
[0010] In an optional implementation, the second information to be identified includes multiple words; the specific implementation of using a deep model to determine whether the second information to be identified is in violation of regulations is: using a deep model to determine the semantic dependency relationship between words in the second information to be identified; and based on the semantic dependency relationship, determine whether the second information to be identified is in violation of regulations.
[0011] In an optional embodiment, the second information to be identified includes a first word, an adjective and a second word; in the second information to be identified, the appearance order of the first word, the adjective and the second word decreases; using a deep model, the specific implementation method of determining the semantic dependency relationship between words in the second information to be identified is: using a deep model, determining from the first word and the second word that the modified object of the adjective is the first word.
[0012] In an optional implementation, the number of target illegal words is one or more; the preset condition includes one or more of the following: the length of the target illegal word is less than a first threshold; the number of the target illegal words is less than a second threshold; the illegal degree value of the first information to be identified is less than a third threshold, and the illegal degree value of the first information to be identified is determined by the part of speech of the target illegal word.
[0013] In an optional implementation, a specific implementation of obtaining the second information to be identified including the target illegal word in the first information to be identified is: determining a position of the target illegal word in the first information to be identified; and segmenting the first information to be identified according to the position to obtain the second information to be identified including the target illegal word; wherein the length of characters included in the second information to be identified is less than a fourth threshold, and / or the second information to be identified has a complete sentence structure.
[0014] In an optional implementation, the first information to be identified is crawled information that does not match an object in a filtering object data set, and the filtering object data set includes blacklist objects and / or whitelist objects.
[0015] In an optional embodiment, the method may further include: crawling and obtaining the crawled information according to a crawling strategy; wherein the crawling strategy includes one or more of the following: during the crawling process, using the first information to send a preset number of requests, and subsequently using the second information to send requests; the first information is identity information and / or address information; if it is detected that the page structure of the crawled web page is not a preset structure, formatting the page structure of the web page; if it is detected that the crawled URL is incomplete, dynamically capturing the page corresponding to the crawled URL.
[0016] In a second aspect, an embodiment of the present application provides an information identification device, which includes a unit for implementing the method described in the first aspect.
[0017] In a third aspect, an embodiment of the present application provides another information identification device, including a processor; the processor is used to execute the method described in the first aspect.
[0018] In an optional implementation, the information identification device may further include a memory; the memory is used to store a computer program; and a processor, specifically used to call the computer program from the memory to execute the method described in the first aspect.
[0019] In a fourth aspect, an embodiment of the present application provides a chip, which is used to execute the method described in the first aspect.
[0020] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of an information identification method provided in an embodiment of the present application;
[0022] Figure 2 This is a schematic diagram of the result of using a deep model to identify illegal information provided by an embodiment of the present application;
[0023] Figure 3 It is a schematic diagram of the architecture of an information identification system provided in an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of a hierarchical structure in a web page provided in an embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of a processing flow of a text analysis module provided in an embodiment of the present application;
[0026] Figure 6 It is a structural schematic diagram of an information identification device provided in an embodiment of the present application;
[0027] Figure 7 It is a structural schematic diagram of another information identification device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0029] In order to better understand the technical solutions provided by the embodiments of the present application, the technical terms involved in the embodiments of the present application are first introduced.
[0030] (1) Deep Model
[0031] The deep model is used to identify whether the information to be identified is illegal. In this application, the deep model can use the deep model architecture biLSTM. Optionally, the deep model can use a combination of biLSTM and textcnn models. Optionally, the deep model can use biLSTM plus textcnn model and virtual adversarial training as the model architecture for text classification, that is, add part of the model after Virtual Adversarial Training perturbation when training biLSTM.
[0032] The idea of adversarial training is to train an opponent (adversarial network) to continuously improve the learning of oneself (generative network). For example, different objectives are used to train the adversarial network and the generative network to compete. The deep model in this application uses Virtual Adversarial Training to improve the generalization ability and robustness of the model.
[0033] Adversarial training is to learn from adversarial samples generated by a certain model. Therefore, the model is bound to be more targeted, so it may have a higher error rate than the original model when facing adversarial sample attacks generated by other models. In addition, each model generally has good robustness to adversarial samples generated by adversarial models. Adversarial training not only fits the perturbations that affect the model, but also weakens the linear assumptions of the model that need to be relied on during single-step attacks, further improving the robustness of the model to black-box attacks.
[0034] Adversarial perturbations typically consist of making small modifications to many real-valued inputs. For text classification, the inputs are discrete and are typically represented as a series of high-dimensional one-hot encoded vectors. Since the set of high-dimensional one-hot encoded vectors does not allow infinitesimal perturbations, the deep model in this application defines perturbations on continuous word embeddings rather than on discrete word inputs. Both traditional adversarial training and virtual adversarial training can be interpreted as regularization strategies as a defense against an adversary providing malicious inputs. Since the perturbed embeddings do not map to any words and the adversary may not have access to the word embedding layer, the training strategy in this application is no longer a defense strategy against an adversary.
[0035] (2) Cloud Services
[0036] The information identification method proposed in the embodiment of the present application can be executed by an information identification device. The existence form of the information identification device can be a virtual device carried on a cloud server. By virtualizing multiple parts similar to independent servers on the physical server (host), each part can be a separate operating system, and the management method is the same as that of the server. Cloud servers can provide elastic cloud technology with adjustable cloud host configurations, cloud host rental services with on-demand use and on-demand instant payment capabilities, and have greatly improved flexibility, controllability, scalability and resource reusability. Its management method is simpler and more efficient than that of physical servers, and it can quickly build more stable and secure applications, reducing the difficulty of development and operation and the overall information technology (IT) cost.
[0037] The information identification method involved in this application can be encapsulated as a cloud service, and an interface can be exposed to the outside. When the information identification method involved in this application needs to be used, by calling the interface, it can be identified whether the information to be identified is illegal.
[0038] (3) Cloud computing
[0039] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". From the user's perspective, the resources in the "cloud" are infinitely scalable and can be obtained at any time, used on demand, expanded at any time, and paid for by use.
[0040] The information identification method provided in the present application involves large-scale calculations and requires large computing power and storage space. Therefore, in a feasible implementation of the present application, sufficient computing power and storage space can be obtained through cloud computing technology.
[0041] In order to effectively identify illegal information and thus create a healthier network environment, the present application embodiment provides an information identification method, such as Figure 1 As shown, the information identification method may include but is not limited to the following steps:
[0042] S101: Obtain first information to be identified.
[0043] The information identification method proposed in the embodiment of the present application can be executed by an information identification device, which can be a server or a terminal device. Among them, the server can be a cloud server, that is, the information identification method can be executed by the cloud server. The terminal device can also be called user equipment (UE), and can also be called terminal, mobile station (MS), mobile terminal (MT), etc. The terminal device can be a mobile phone, a wearable device, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in a smart city, a wireless terminal in a smart home, etc.
[0044] The first information to be identified may be information crawled by a crawler, and may include one or more of the following forms of information: text, image, video, audio, file, etc. For example, the first information to be identified is a web page. Optionally, the information identification device may obtain the first information to be identified from a local database, or the information identification device may obtain the first information to be identified through a cloud service.
[0045] In one implementation, the first information to be identified may be crawled information that does not match an object in a filtering object data set, and the filtering object data set includes a blacklist object and / or a whitelist object. In other words, the first information to be identified may be crawled information filtered by the filtering object data set, and in the case where the filtering object data set includes a blacklist object, the first information to be identified does not match (or is not associated with) the blacklist object; in the case where the filtering object data set includes a whitelist object, the first information to be identified does not match (or is not associated with) the whitelist object; in the case where the filtering object data set includes a blacklist object and a whitelist object, the first information to be identified does not match (or is not associated with) both the blacklist object and the whitelist object. Among them, the first information to be identified does not match (or is not associated with) the blacklist object may mean that the blacklist object is not included in the first information to be identified.
[0046] Objects can be URLs, web pages, words and other information. Blacklist objects can be previously detected objects that violated the rules. Whitelist objects can be some large websites. The data released on these websites will be strictly reviewed, so the information on these websites is generally not in violation of the rules. By filtering the object data set, crawled information can be filtered, and some crawled information can be filtered out without further identification. Only the crawled information obtained after filtering can be further identified, which is conducive to saving computing resources.
[0047] S102: Determine a matching result between the first information to be identified and the illegal word data set; the matching result includes a target illegal word that exists in both the first information to be identified and the illegal word data set.
[0048] The illegal word data set includes multiple illegal words, and the illegal words may be illegal words in one or more fields. For example, illegal words may be words related to pornography, gambling, politics, drugs, gambling, etc.
[0049] The information identification device can retrieve whether the first information to be identified includes the illegal words in the illegal word data set to obtain a matching result. It should be noted that the illegal words mentioned in the embodiment of the present application refer to the words present in the illegal word data set. The matching result may include the target illegal words retrieved in the first information to be identified, and it is understandable that the target illegal words also exist in the illegal word data set. It should be noted that the aforementioned retrieval by the information identification device of whether the first information to be identified includes the illegal words in the illegal word data set to obtain the matching result is used for example. In other implementations, it can also be retrieved by other devices, and the information identification device obtains the matching result from the device. The illegal word data set can store the information identification device, and can also be stored in a cloud server, which is not limited in the embodiment of the present application.
[0050] S103: If the matching result meets a preset condition, obtain the second information to be identified including the target illegal word from the first information to be identified.
[0051] After the information identification device obtains the matching result, it can determine whether the matching result meets the preset condition (such as the violation word match). If the preset condition is met, it can further identify whether it is a violation. If it does not meet the preset condition, it can be determined whether the first information to be identified is a violation without further identification.
[0052] In one implementation, the number of target illegal words is one or more; the preset condition may include one or more of the following: the length of the target illegal word is less than a first threshold; the number of the target illegal words is less than a second threshold; the illegal degree value of the first information to be identified is less than a third threshold, and the illegal degree value of the first information to be identified is determined by the part of speech of the target illegal word.
[0053] The illegal word matching may include one or more of the following processing sub-processes: long word calculation, illegal degree calculation, and part of speech calculation. Among them, the long word calculation sub-process refers to calculating illegal words with relatively long lengths. If the length of the target illegal word is less than the first threshold, it can be indicated that the target illegal word is not an illegal word with a relatively long length. Optionally, if it is detected that the length of the target illegal word is greater than or equal to the first threshold, no further identification is required, and the first information to be identified is also determined to be illegal. Because long words represent more semantic information and certainty in terms of conventional human use of words. In daily language usage habits, if a relatively long character is used as a word, this word greatly represents a certain event or a certain specific idea. Therefore, if a long word is matched in the first information to be identified, the first information to be identified is likely to be illegal. If the number of target illegal words is less than the second threshold, it can be indicated that the number of target illegal words in the first information to be identified is small. Optionally, if it is detected that the number of target illegal words is greater than or equal to the second threshold, no further identification is required, and the first information to be identified is also determined to be illegal. If the number of illegal words is large, the first information to be identified to which the illegal word belongs is likely to be illegal.
[0054] The violation degree calculation sub-process refers to calculating the accumulation of multiple parts of speech of illegal words matched in the first information to be identified. The violation degree value of the first information to be identified is less than the third threshold value, which can indicate that the violation degree value determined by the parts of speech of multiple target illegal words in the first information to be identified is less than the third threshold value. In other words, the violation degree of the first information to be identified is relatively low. Optionally, if the violation degree value of the first information to be identified is greater than or equal to the third threshold value, the first information to be identified can be determined to be illegal without further identification. Violation degree calculation is a quantitative process. For example, a large number or even all of the parts of speech of words appearing in a pornographic web page are pornographic. When the accumulated violation degree is greater than the second threshold value, the web page can be directly determined to be illegal without further identification.
[0055] The part-of-speech calculation subprocess refers to guessing the part-of-speech of a word, and further guessing the semantic information of the information to be identified to which the word belongs. The part-of-speech calculation can be applied to the violation degree calculation subprocess to determine the part-of-speech of each target violation word, and then obtain the violation degree value of the first information to be identified. In one implementation, one part-of-speech can correspond to one violation degree value. After determining the part-of-speech of each target violation word, the violation degree values corresponding to the part-of-speech of each target violation word can be added, and the result obtained can be used as the violation degree value of the first information to be identified.
[0056] If the preset conditions are met, the second information to be identified including the target illegal word can be further obtained from the first information to be identified, and it can be determined whether the second information to be identified is illegal. Whether the second information to be identified is illegal can indicate whether the first information to be identified is illegal. If the second information to be identified is illegal, it means that the first information to be identified is illegal. If the second information to be identified is not illegal, it means that the first information to be identified is not illegal.
[0057] In one implementation, a specific implementation method of obtaining the second information to be identified including the target illegal word in the first information to be identified may be: determining a position of the target illegal word in the first information to be identified; and segmenting the first information to be identified according to the position to obtain the second information to be identified including the target illegal word; wherein the length of characters included in the second information to be identified is less than a fourth threshold, and / or the second information to be identified has a complete sentence structure.
[0058] When the matching result does not meet the preset conditions, the meaning information provided by the first information to be identified is relatively small. The context information and position around the target illegal word can be searched, and then the first information to be identified can be segmented. Optionally, the following segmentation methods can be used: the first is to segment a part of the short sentences around the target illegal word in the first information to be identified, so that the second information to be identified obtained by segmentation may not conform to normal reading habits and logical structure. The length of the characters included in the second information to be identified is less than the fourth threshold, indicating that the second information to be identified is a short sentence cut from the first information to be identified. The second is to segment according to normal reading conditions (that is, segment according to normal semantic conditions). At this time, the segmented second information to be identified has a complete sentence structure. The complete sentence structure can refer to: the second information to be identified includes subject, predicate, and object objects. Among them, the first segmentation method can be called illegal short text segmentation or relative segmentation, and the second segmentation method can be called semantic segmentation or absolute segmentation.
[0059] S104: Use the deep model to determine whether the second information to be identified is in violation of regulations.
[0060] In one implementation, the second information to be identified may include multiple words; using the deep model to determine whether the second information to be identified violates the rules, including: using the deep model to determine the semantic dependency relationship between the words in the second information to be identified; and judging whether the second information to be identified violates the rules based on the semantic dependency relationship. The semantic dependency relationship refers to the degree of correlation between the words in the second information to be identified. For example, if the second information to be identified includes an adjective and two nouns, the semantic dependency relationship may indicate which of the two nouns the adjective is used to modify.
[0061] In one implementation, the second information to be identified includes a first word, an attributive, and a second word; in the second information to be identified, the order of appearance of the first word, the attributive, and the second word decreases. Using a deep model, determining the semantic dependency relationship between words in the second information to be identified includes: using a deep model, determining from the first word and the second word that the modified object of the attributive is the first word, not the second word. In other implementations, the modified object of the attributive may also be the second word. The order of appearance of the first word, the attributive, and the second word decreases, and the modified object of the attributive is the first word, indicating that the attributive that appears later in the second information to be identified is used to modify the first word that appears before the attributive. It can be seen that the deep model in this application can make full use of historical information. For example, in "This restaurant is so dirty, not as good as the one next door", "not so good" is a modification of the degree of "dirty", and the deep model in this application can better capture the two-way semantic dependency.
[0062] In natural language processing, in order to combine the representation of words into the representation of sentences, the addition method can be used, that is, the representation of all words is added up, or the average method is taken, but these methods do not take into account the order of words in the sentence. For example, in the sentence "I don't think he is good". If the deep model in this application is not used, only the representation of all words is added up, so it is impossible to know that the word "no" is a negation of the following "good", and then because of "good", it is mistakenly believed that the emotion of the sentence is commendatory. If the deep model in this application is used, the two-way semantic dependency can be better captured, and it can be known that the word "no" is a negation of the following "good", so that the emotion of the sentence can be accurately determined to be derogatory. The deep model in this application has a very fine-grained classification of sentiment words, such as five categories of strong degree of commendation, weak degree of commendation, neutral, weak degree of derogatory, and strong degree of derogatory. And in the five categories, attention will also be paid to the interaction between sentiment words, degree words, and negative words.
[0063] In one implementation, in the second information to be identified, the number of characters between the first word and the attributive is greater than a preset number, indicating that the first word and the attributive are far apart. That is, the deep model in this application can better capture semantic dependencies over a longer distance, thereby facilitating the accuracy of identifying illegal words.
[0064] Optionally, the learning parameters in the deep model of the present application can be updated. Specifically, the update of the learning parameters in the deep model can use the back-propagation through time (BPTT) algorithm. In this case, the deep model is different from the general model in the forward calculation error (forward) and the reverse update model parameter gradient (backard) stage in that the hidden layer must be calculated for all time steps.
[0065] In the case where the deep model uses a combination of biLSTM and textcnn models, the similarity with the offending words can be calculated during the convolution process, and then the max pooling layer can be used to determine whether the offending words that the model focuses on appear in the information to be identified. Optionally, it can also be determined how large the maximum similarity between the most similar offending words and the convolution kernel is. Assuming that the Chinese output is a word vector, ideally a convolution kernel represents a keyword (such as an offending word). For example, in a 2-classification task, if the entire deep model is used as a black box to detect its output results, it will be found that the model is particularly sensitive to whether the input text contains words such as "like" and "love". This is because if there are one or more of these two words in a large number of training samples, it means that these two words are common features of this type of data, and the convolution kernel can learn these characteristics. In the deep model of this application, a convolution kernel can only learn half of the keyword word vector, and then another convolution kernel learns the other half of the keyword word vector, and finally these feature values are accumulated in the classifier to obtain the final result. Therefore, the deep model of this application can obtain local semantic features of text from multiple dimensions.
[0066] Optionally, in the deep model of the present application, a layer of bidirectional biLSTM can be added before the information to be identified enters the textcnn model to capture the global information of the text, so that a classification model can be formed in which the LSTM layer learns context dependencies and the textcnn captures local important information.
[0067] See also Figure 2 , is a schematic diagram of the results of using a deep model to identify illegal information. Figure 2 Before the improvement, the LSTM model was used to determine whether the text violated the rules. After the improvement, the deep model proposed in this application was used to determine whether the text violated the rules. Figure 2It can be seen that for the text "Let's look at Asian beauty pictures together", the recognition result before improvement was violation, while the recognition result after improvement was no violation. For the text "Dahua Entertainment Hall", since the probability of violating regulations is very high in such places, the recognition result of violation is accurate. For the text "Artists participated in the event at Dahua Entertainment Hall", although the text includes "Dahua Entertainment Hall" which is identified as violation information, it can be seen from the context that the probability of no violation is higher. The deep model of the present application can be used for accurate judgment, because the deep model can capture semantic dependencies over longer distances, while the LSTM model cannot. It can be seen that the use of the deep model of the present application can improve the accuracy of identifying violation information, and can also avoid misidentifying normal information as violation information.
[0068] It should be noted that the thresholds (such as the first threshold, the second threshold, the third threshold, etc.) and preset parameters (such as the preset number, etc.) in the embodiments of the present application can be set or modified by the information identification device.
[0069] In one implementation, the present application also provides an information identification system, the architecture diagram of which is as follows: Figure 3 The information recognition system may include the following modules: online configuration module, crawler module, data analysis, storage and deduplication module, and text analysis module. Figure 3 As shown, the online configuration module can also be called module 1, the crawler module can also be called module 2, the data parsing, storage and deduplication module can also be called module 3, and the text analysis module can also be called module 4.
[0070] It should be noted that the processing and flow of each module is a link, and each link has different processing architectures and storage for data, and the processing flow message queues in each small module are decoupled from each other. These functional small modules are actually atomic capabilities, and their functions are relatively simple and concentrated, so some modules can be upgraded, optimized or transformed separately. The communication between modules depends on the flow of data. Independent module design can solve the problem of strong dependence between modules of the system. Each small module can also provide service capabilities separately. When other projects or systems want to obtain corresponding service capabilities, they can use the service capabilities of the module by connecting to the corresponding module.
[0071] Among them, the online configuration module (i.e., module 1) can have one or more of the following functions: allowing users to flexibly configure crawler seeds, web sites, configure crawling strategies, and set crawling targets for web crawlers. For example, the crawling site level, whether to use a headless browser to render pages, the number of processes and threads used by the crawler, the size of the allocated memory, the control of the network traffic threshold, the restart strategy for program failure, the re-crawling strategy for crawling failure, the response time of each page, the alarm level, etc. Figure 3 Module 1 shown in the figure, when there is a crawling task, loads the seed, and the loaded seed data enters the seed queue so that module 2 can use it for crawling.
[0072] The web page site is the web page site that the crawler needs to crawl. Optionally, the set web page site may have one or more of the following characteristics: the page views are less than the fifth threshold, and the search ranking on the search engine platform is in the top N. N is an integer greater than or equal to 1. Among them, the page views are less than the fifth threshold, which means that the page views are not large, and the search ranking on the search engine platform is in the top N, which means that the search ranking on the search engine platform is high. Exemplarily, the set web page site can be an official website of a medical institution, school education, etc.
[0073] The information identification system can use a headless browser to dynamically load and render pages to ensure the authenticity of the data. Since some web pages in real sites are loaded dynamically, if the crawler does not use a headless browser to load the content of the Uniform Resource Locator (URL), the data obtained will be HyperText Markup Language (html) tags. In dynamic pages, the text information and image information are generally obtained by requesting a remote server through JavaScript (js). The information obtained by using only ordinary crawling strategies will be untrue information. A headless browser refers to a web browser without a graphical user interface (GUI), which is usually controlled by programming or a command line interface.
[0074] The crawling level can be used to control the deepest level of web pages to ensure the number of crawls. The number of processes and threads can be used to control the crawling speed of the crawler. The memory allocation configuration item can be used to ensure the performance of the crawler system. When the amount of data to be crawled is large, setting the memory of the project running appropriately to a larger size can ensure the service capability of the crawler. In addition, this configuration item can also control the memory request of the program so that it will not cause the wild growth of memory and bring negative feedback to the server and other projects. Network traffic refers to the amount of data that can pass through the network in one second, in bits (bit) / second. The network can be compared to a highway. The larger the traffic, the wider the road, and the more data can pass through the network highway (at the same time). When crawling data, in addition to basic network requests, there is also data download. Reasonable setting of network traffic thresholds can control the speed and progress of crawling. The restart strategy for program failure can ensure that the program will not lose its service capability. The re-climbing strategy for crawling failure can ensure that the crawled data is not lost. Page response time refers to the loading time of page content. The content of many web pages is loaded by multiple remote servers, and the response time of such page loading will be very long. The warning level refers to the violation level of the violation information analysis result in the subsequent process. In the embodiment of the present application, the violation information can also be referred to as bad information.
[0075] The crawler module (i.e. module 2) can be used to crawl data according to the strategies and tasks configured in module 1. Figure 3 In the example, the cache in module 2 can be used to temporarily store the loaded seed data and the links and image data after the subsequent URL resolution and deduplication. When the system is running, the data in the cache can be crawled first. If there is no data in the cache or the data in the cache has been loaded, the data in the message queue can be loaded later. In this way, it can be avoided that there is no extra resources to respond to the directly loaded seeds when the system crawling task is heavy.
[0076] The message queue can be a queue of URLs to be crawled or other queues. The crawling strategy can be used to determine the order of URLs in the queue of URLs to be crawled, and the order of URLs can affect the crawling order of pages corresponding to the URLs.
[0077] The crawler system is a multi-process and multi-line support system. The crawled data exists in the form of a tree on the site, which means that the amount of crawled data will increase exponentially as the crawling level goes deeper. Figure 4It is a hierarchical structure in a web page, and the letters A, B, C, D...J in the figure represent hyperlinks. The information identification system in this application can use the following methods to crawl data: depth first search (DFS) or breadth first search (BFS). DFS means that the crawler starts from a certain URL and crawls one link after another until all the lines where a certain link is located are processed, and then switches to other lines. At this time, the crawling order is: A ->B ->D->H ->I ->E ->J ->C ->F ->G. BFS is to insert the links found in the newly downloaded web page directly into the end of the URL queue to be crawled. That is to say, the web crawler will first crawl all the web pages linked in the starting web page, and then select one of the linked web pages, and continue to crawl all the web pages linked in this web page. At this time, the crawling order is: A ->B ->C ->D->E ->F ->G ->H ->I ->J. If the information identification system uses DFS crawling, it will bring a level mark when crawling each page. When this mark does not exceed the level set in the configuration, it will automatically crawl and parse the next level, and mark the parsed data with one to set its level. In this way, as the iteration proceeds, when the final level reaches the set level, it will automatically stop and no longer continue.
[0078] For some specific sites, web crawlers send a large number of requests in a short period of time, consuming a large amount of server bandwidth, which may affect normal user access. In addition, data has become a company's core asset. Enterprises need to protect their core data to maintain or enhance their core competitiveness, so anti-crawler is very important. The information identification system in this application can also solve the anti-crawling problem.
[0079] The anti-crawling problem can be solved by configuring the crawling strategy in the information identification system. In one implementation, the aforementioned information identification method may further include the following steps: crawling and obtaining crawling information according to the crawling strategy; wherein the crawling strategy includes one or more of the following to solve the anti-crawling problem:
[0080] Crawling strategy 1: During the crawling process, a preset number of requests are sent using the first information, and then requests are sent using the second information; the first information is identity information and / or address information. The second information can also be identity information and / or address information. The first information is different from the second information. The address information can be, for example, an IP address. In crawling strategy 1, an IP address is changed every few requests, which can easily bypass multiple anti-crawlers.
[0081] Crawling strategy 2: If it is detected that the page structure of the crawled web page is not the preset structure, the xpath or regular expression is changed according to the source code of the web page. Some anti-crawling strategies change the original html page structure through JavaScript, so that the required content cannot be matched in the program. In crawling strategy 2, if it is detected that the page structure of the crawled web page is not the preset structure, the xpath or regular expression of the web page is changed according to the source code of the web page, and the data in a standard form can be returned, thereby solving the anti-crawling problem. The preset structure can refer to a pre-set standard html page structure. The fact that the page structure of the crawled web page is not the preset structure can indicate that the page structure of the crawled web page has changed. xpath is a third-party library in python that is used to parse web page content.
[0082] Crawling strategy 3: If it is detected that the page structure of the crawled web page is not the preset structure, the page structure of the web page is formatted. Crawling strategy 3 is similar to crawling strategy 2. When it is detected that the page structure of the crawled web page is not the preset structure, the page structure of the web page is formatted, and data in a standard form can be returned, thereby solving the anti-crawling problem.
[0083] Crawling strategy 4: If it is detected that the crawling URL is incomplete, the page corresponding to the crawling URL is dynamically captured. The incomplete crawling URL may indicate that the data of the web page corresponding to the crawling URL is not loaded at one time, so the data crawled by the crawler is incomplete. In this case, crawling strategy 4 dynamically captures the page corresponding to the crawling URL to obtain asynchronously loaded data packets. In this way, each new content loaded by the web page can be captured, thereby solving the anti-crawler problem.
[0084] Crawling strategy 5: If the crawling URL is detected to be incomplete or incorrect, the encrypted file is obtained, the crawled information is encrypted according to the encrypted file, and the encrypted data is returned. The crawling URL is detected to be incomplete or incorrect, which may be because the target site has encrypted some parameters through JavaScript. In this case, crawling strategy 5 obtains the encrypted file, analyzes the encryption algorithm, and encrypts the crawled information according to the encrypted file, and returns the encrypted data, thereby solving the anti-crawling problem.
[0085] Crawling strategy 6: Add Headers to the crawler, copy the browser's User-Agent to the crawler's Headers, or modify the Refer value to the target website domain name. For anti-crawlers that detect Headers, copying the browser's User-Agent to the crawler's Headers can solve the anti-crawler problem. Headers can be the browser's identification data, which can contain the browser's configuration information (such as the kernel version information used, a certain browser, supported network protocols, hypertext protocols, etc.). User-Agent can be used to store the user's own information, such as user identification (Identity, id), user name, user password, user session information, etc. This information will be verified by the access website when visiting the website. Refer is when a page resource is accessed, the browser tells the page which page the access is linked from. Refer can be used to verify the legitimacy of the access.
[0086] The data parsing, storage and deduplication module (i.e. module 3) can be used to: parse the data captured by module 2, and then perform a series of deduplication and persistence operations on the data. Figure 3 As shown, the data captured by module 2 may include pictures, texts, etc., and accordingly, module 3 analyzes the captured data including: analyzing texts, analyzing pictures ( Figure 3 ). After parsing, text, images and links may be obtained. Specifically, the text, images and links obtained after parsing can be placed in the corresponding text parsing queue, image parsing queue and link identification queue. The links in the link identification queue are further deduplicated and hierarchical judged to obtain the links that need to be crawled, and put them into the link crawling queue. The data crawling step in module 2 also includes: crawling the page content corresponding to the links in the link crawling queue. The hierarchical judgment includes: if the hierarchical identifier of the link obtained from the link identification queue does not exceed the hierarchical level set in the configuration, the next level of the link can be further crawled.
[0087] In one implementation, parsing the data captured by the crawler module may refer to: parsing the text information in HTML, which is some plain text, and it is necessary to remove various tags in HTML, because these tags are useless for the subsequent data analysis and judgment, and will also bring huge consumption and pressure to the transmission and storage of data. Optionally, in addition to parsing the plain text information in HTML, the text information in some key tags in HTML can also be parsed, such as parsing the HTML of a web page, and the text information in the meta tag of a web page. By parsing the TITLE and meta tags of a web page, illegal web pages can be detected more efficiently. This is because many bad information web pages may put some key information or content in the TITLE to obtain clicks from potential users. In addition, in order to improve the search ranking of their bad information websites, these bad information will put a large amount of key information or keywords in the meta tag to improve the search ranking of the entire web page. When users search for corresponding keywords, the website will display these bad information in the front position of the browser to facilitate users to find and browse. Meta is an auxiliary tag in the head area of html language, located at the head of the document, and does not contain any content.
[0088] In another implementation, parsing the data captured by the crawler module may refer to: parsing the URL of the hyperlink in HTML. The development modes of websites vary greatly, and the existence of hyperlinks in different sites also varies greatly. In order to parse the correct URL as comprehensively as possible, the information identification system of this application proposes one or more of the following parsing methods to parse the URL of the hyperlink:
[0089] Parsing method 1: Match all URLs with http or https protocol in the web page, and end with any character in [-A-Za-z0-9+&@# / %?=~_|!:,.;] space (including text). This method can be used to parse hyperlinks that conform to the URL format of hypertext links.
[0090] Parsing method 2: Match all URLs on the web page that start with www and end with any of the characters [-A-Za-z0-9+&@# / %?=~_|!:,.;], space (including text). This method can match website domain names that comply with the www protocol on the web page (including hidden links and hyperlinks).
[0091] Parsing method 3: Process special characters such as & \" %3A %2F in the links on the web page. That is, recognize these special characters in the link as normal characters. These symbols are some special browser encoding characters, which are recognized as normal characters. For example, the special character & is recognized as &.
[0092] Parsing method 4: Match the URL in the href="" hyperlink and do corresponding splicing, so that the browser splicing rules can be reproduced, so that the browser can identify and access the spliced link. For example, those starting with .. / .. / , / / , . / , .. / , / , ?. This method handles the case where the href is followed by double quotes. Href stands for hyperlink, which refers to the connection relationship from a web page to a target.
[0093] Parsing method 5: Match the web pages with base href tags and change the splicing rules. Specifically, the sublinks with base href are no longer spliced with the current parent link, but with the parent link's primary domain name.
[0094] Parsing method 6: match the URL in the href='' hyperlink and concatenate accordingly. For example, those starting with .. / .. / , / / , . / , .. / , / , ?. This method handles the case where href is followed by a single quote.
[0095] Parsing method 7: Match the URL in the href hyperlink and concatenate accordingly. For example, those starting with .. / .. / , / / , . / , .. / , / , ?. This method handles the situation where there is no quotation mark after href.
[0096] Parsing method 8: Match the URL in the src="" hyperlink and do the corresponding splicing. For example, those starting with .. / .. / , / / , . / , .. / , / , ?. This method handles the case where src is followed by double quotes.
[0097] Parsing method 9: match the URL in the src='' hyperlink and concatenate accordingly. For example, those starting with .. / .. / , / / , . / , .. / , / , ?. This method handles the case where src is followed by a single quote.
[0098] Parsing method 10: Match the mailbox after the @ symbol, and the mailbox ends with any of the following suffixes: \.edu\.com|\.gov\.cn|\.org\.cn|\.net\.cn|\.com\.cn|\.top\.cn|\.asp|\.com|\.cn|\.top|\.xyz|\.vip|\.net|\.org|\.wang|\.gov|\.mil|\.co|\.biz|\.name|\.info|\.pro|\.int|\.im|\.ltd|\.hk. Among them, | is used to separate specific items, and / is the separator in the specific item. This method can handle the situation where there are illegal links after the mailbox.
[0099] Parsing method 11: Match window.location='. This method can handle new links after the web page jumps.
[0100] Parsing method 12: Match the URL in the option value="" hyperlink and concatenate accordingly. For example, those starting with .. / .. / , / / , . / , .. / , / , ?. This method handles the case where the option value is followed by double quotes.
[0101] Parsing method 13: Match the URL in the option value='' hyperlink and concatenate accordingly. For example, those starting with .. / .. / , / / , . / , .. / , / , ?. This method handles the case where the option value is followed by a single quote.
[0102] Analysis method 14: Process the garbled text in the meta name="keywords" content=" tag on the web page and restore it for link matching.
[0103] Data deduplication means that there are many duplicate links in a website. Module 2 needs to crawl too many web pages. If some duplicate tags are crawled repeatedly, it will not only cause heavy pressure on the crawler system and occupy a lot of software and hardware performance, but also cause the crawler to enter an infinite loop and not know when to end crawling data, which will cause the performance of the entire system to slow down. If there are N websites in the entire network, the complexity of duplication detection is N*log(N), because all web pages need to be traversed once, and each duplication detection requires log(N) complexity. The duplication detection method used by the information identification system in this application is: using Bloom Filter. Its feature is that it can use fixed memory (which does not grow with the number of URLs) to determine whether the URL has been crawled with O(1) efficiency.
[0104] In this application, the amount of data that needs to be processed for information identification is very large, and a lot of the data is image data. This application can obtain sufficient computing power and storage space through cloud computing technology. Optionally, the information identification system may have data recovery and error correction functions. When the storage server crashes or the disk is damaged, the previously stored data can be restored quickly and efficiently. Specifically, erasure codes and checksums can be used to protect data from hardware failures and silent data corruption. Erasure code is a mathematical algorithm that recovers lost and damaged data. Optionally, this application may use the storage database MINIO for data storage, for example, to store crawled information. On standard hardware, MINIO's read / write speeds are as high as 183GB / s and 171GB / s. It should be noted that, Figure 3In the example, the image to be identified is stored in the MINIO database, and can also be stored in other databases, which is not limited in the embodiment of the present application.
[0105] Optionally, module 3 also includes an optical character recognition (OCR) module for recognizing information to be recognized including images. Figure 3 As shown in the figure, the results of OCR recognition can be put into the OCR recognition queue. A large number of illegal contents in many bad information web pages exist in the form of pictures. The OCR module can identify the text information included in the pictures. In this way, illegal pictures can be effectively identified to ensure the comprehensiveness of the information recognition system's ability to capture bad information.
[0106] The text analysis module (i.e. module 4) can be used to determine whether the information to be identified from module 3 violates the regulations. For example, see Figure 3 , determine whether the information to be identified in the text parsing queue and the OCR recognition queue is in violation of the rules. The general processing flow of the text analysis module includes: matching the information to be identified from module 3 with the illegal word data set, and putting the information to be identified whose matching results with the illegal word data set meet the preset conditions into the illegal word filtering queue, and further, using the deep model to analyze and judge the information to be identified in the illegal word filtering queue to obtain the analysis result. The analysis result indicates whether it is in violation of the rules. For the information to be identified whose matching results with the illegal word data set do not meet the preset conditions, the analysis result can be directly obtained, and the details can be referred to the description in the above steps S102 and S103.
[0107] For example, the schematic diagram of the processing flow of the text analysis module can be found in Figure 5 shown. Figure 5In the figure, the wave box represents various data sets, the parallelogram represents a certain processing process, the small long rectangle represents the sub-process included in the processing process, and the large rectangle represents the result. The overall data flow is shown as the solid line in the figure. The data output by module 3 flows into the processing of this link. First, it undergoes data cleaning, and then obtains data without black and white lists (i.e., the first information to be identified mentioned above). Here, the data has two flows, one is to flow to the illegal word matching process, and the other is that the data matched to the blacklist can flow directly to the final result, or the data matched to the external link can flow into the interface or queue, and further use the deep model to perform semantic analysis and judgment on it to obtain the result. After data cleaning, the next processing process, that is, illegal word matching, is carried out. Specifically, the illegal word data set of the full-text matching system without black and white lists is matched, and then the information to be identified that matches the illegal words (such as web page text) is reversed. Full-text search to obtain the position information of the illegal words in the original text, after determining the specific position information, the text can be short-text segmented to obtain short text data related to the illegal words, and finally these data are sent to the interface or queue to use the deep model for judgment. Searching using reverse full-text search is more efficient.
[0108] Figure 5 In the crawler text dataset, the crawler text dataset includes the information crawled by the crawler. In the crawler text dataset, it can be divided into crawling batches or task numbers. The crawler text dataset may include one or more of the following contents: the original web page unstructured data of the crawled relevant web pages, sites, public accounts, news media, regulatory sites and other relevant web pages. Low-confidence text refers to web pages that match relevant illegal words. Since web pages that match illegal words are not necessarily illegal, these data are called low-confidence text. Among them, the data without black and white lists and the content of illegal word matching can be found in the previous description, which will not be repeated here.
[0109] Figure 5The dotted part in the middle refers to the generation of illegal words and the operation of illegal words. In the early stage of the operation of illegal words in the information recognition system, some basic word library resources (i.e., the dictionary data set in the figure) can be found on the Internet, and then added to the word library after operation (such as manual review). The operation of illegal words can include the following methods: adding some attributes to illegal words, such as adding one or more of the following attributes: category, sensitivity level, interception rate, interception accuracy, recall rate, etc. These attributes can be used to calculate the degree of violation of the information to be identified that matches the illegal word. Optionally, the relevant attributes of the illegal words can also be adjusted through feedback in actual production to ensure the accuracy of intercepting bad information. Optionally, the words in the basic word library can also be expanded, such as one or more of the following expansions: synonym expansion, pinyin expansion, heteronymous character expansion, agreement expansion, jump expansion, antonym expansion, etc. The attributes and categories of the expanded words can be the same as the source vocabulary. Optionally, new word discovery and hot word calculation can also be used to automatically obtain new illegal words and corresponding attribute categories, thereby expanding the illegal word data set.
[0110] like Figure 5 As shown, the illegal word matching includes three sub-processes: long word calculation, illegal degree calculation, and part of speech calculation. In the illegal word matching process, the information to be identified that meets the preset conditions is a low-confidence illegal text. Further, the position of the illegal word is searched in the original text, and then the sentence is cut, and the second information to be identified obtained by the sentence cutting is put into the interface or queue, and the deep model alignment is further used for research and judgment to obtain the research and judgment results. For those that do not meet the preset conditions, the research and judgment results are directly obtained. The relevant content of the illegal word matching can be found in the description in step S102, which will not be repeated here.
[0111] In the process of matching the illegal word data set, strong matching based on hash, or filtering based on regular expressions, or using the Deterministic Finite Automaton (DFA) algorithm or the Ahocorasick multi-mode matching algorithm can be used. DFA can achieve efficient filtering of illegal words. The Ahocorasick algorithm is a string search algorithm. The difference between the Ahocorasick algorithm and ordinary string matching is that it matches all dictionary strings at the same time. The algorithm has a time complexity that is approximately linear in the amortized case, which is approximately the length of the string plus the number of all matches.
[0112] It should be noted that, in the present application, “greater than or equal to” can be replaced by “greater than”, and in this case, “less than” can be replaced by “less than or equal to”.
[0113] The use of an information identification system to identify illegal information can have the following beneficial effects: First, comprehensiveness. It can identify illegal information of different dissemination channels and methods, such as text, pictures and other types of illegal information. Second, accuracy. Using the aforementioned deep model to identify illegal information can improve the accuracy of identification. Third, flexibility. The independent design of each module of the information identification system can greatly increase the system's transformability and coupling, making it more flexible. Fourth, the ability to discover and identify new information. The deep model can maintain the ability to discover new information through continuous maintenance and training of the model, and thus can also enhance the ability to identify illegal information. Fifth, high availability and reusability of data. This application uses cloud computing to provide powerful computing power and storage resources. The large amount of stored data can be used for self-learning or updating of the model, thereby achieving high availability and reusability of data.
[0114] See also Figure 6 , Figure 6 Schematic diagram of the structure of an information identification device provided in an embodiment of the present application. Figure 6 As shown, the information identification device 60 includes an acquisition unit 601 and a processing unit 602.
[0115] An acquisition unit 601 is used to acquire first information to be identified;
[0116] A processing unit 602 is used to determine a matching result between the first information to be identified and the illegal word data set; the matching result includes a target illegal word that exists in both the first information to be identified and the illegal word data set;
[0117] The processing unit 602 is further configured to obtain second information to be identified including the target illegal word from the first information to be identified if the matching result satisfies a preset condition;
[0118] The processing unit 602 is further configured to use the deep model to determine whether the second information to be identified violates the regulations.
[0119] In an optional embodiment, the second information to be identified includes multiple words; when the processing unit 602 is used to use a deep model to determine whether the second information to be identified violates the rules, it is specifically used to: use the deep model to determine the semantic dependency relationship between words in the second information to be identified; and determine whether the second information to be identified violates the rules based on the semantic dependency relationship.
[0120] In an optional embodiment, the second information to be identified includes a first word, an adjective and a second word; in the second information to be identified, the appearance order of the first word, the adjective and the second word decreases; when the processing unit 602 is used to use a deep model to determine the semantic dependency relationship between words in the second information to be identified, it is specifically used to: use the deep model to determine from the first word and the second word that the modified object of the adjective is the first word.
[0121] In an optional implementation, the number of target illegal words is one or more; the preset condition includes one or more of the following: the length of the target illegal word is less than a first threshold; the number of the target illegal words is less than a second threshold; the illegal degree value of the first information to be identified is less than a third threshold, and the illegal degree value of the first information to be identified is determined by the part of speech of the target illegal word.
[0122] In an optional implementation, when the processing unit 602 is used to obtain the second information to be identified including the target illegal word in the first information to be identified, it is specifically used to: determine the position of the target illegal word in the first information to be identified; and segment the first information to be identified according to the position to obtain the second information to be identified including the target illegal word; wherein the length of characters included in the second information to be identified is less than a fourth threshold, and / or the second information to be identified has a complete sentence structure.
[0123] In an optional implementation, the first information to be identified is crawled information that does not match an object in a filtering object data set, and the filtering object data set includes blacklist objects and / or whitelist objects.
[0124] In an optional embodiment, the processing unit 602 can also be used to: crawl and obtain the crawl information according to a crawling strategy; wherein the crawling strategy includes one or more of the following: during the crawling process, use the first information to send a preset number of requests, and subsequently use the second information to send requests; the first information is identity information and / or address information; if it is detected that the page structure of the crawled web page is not a preset structure, format the page structure of the web page; if it is detected that the crawled URL is incomplete, dynamically capture the page corresponding to the crawled URL.
[0125] The information recognition device 60 can also be used to implement Figure 1 Other functions of the information identification device in the corresponding embodiment will not be described in detail here.
[0126] See also Figure 7 , Figure 7 Another information identification device 70 provided in the embodiment of the present application can be used to implement the function of the information identification device in the above method embodiment. The information identification device 70 may include a processor 701. Optionally, the information identification device 70 may also include a memory 702. The processor 701 and the memory 702 may be connected via a bus 703 or other means. The bus is Figure 7 The connections between other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0127] The coupling in the embodiment of the present application is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, and is used for information exchange between devices, units or modules. The specific connection medium between the processor 701 and the memory 702 is not limited in the embodiment of the present application.
[0128] The memory 702 may include a read-only memory and a random access memory, and provides instructions and data to the processor 701. A portion of the memory 702 may also include a nonvolatile random access memory.
[0129] The processor 701 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, and optionally, the processor 701 may also be any conventional processor, etc.
[0130] When the information recognition device adopts Figure 7 When the form shown is Figure 7 The processor in can execute the method executed by the information identification device in any of the above method embodiments.
[0131] In an optional implementation, the memory 702 is used to store program instructions; the processor 701 is used to call the program instructions stored in the memory 702 to execute Figure 1 The steps performed by the information identification device in the corresponding embodiment. Specifically, Figure 6 The functions / implementation processes of the acquisition unit and the processing unit can be achieved through Figure 7 The processor 701 in the embodiment calls the computer execution instructions stored in the memory 702 to implement.
[0132] In the embodiment of the present application, the method provided in the embodiment of the present application can be implemented by running a computer program (including program code) capable of executing each step involved in the above method on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above computing device through the computer-readable recording medium and run therein.
[0133] Based on the same inventive concept, the principle and beneficial effects of solving the problem by the information identification device 70 provided in the embodiment of the present application are similar to the principle and beneficial effects of solving the problem by the information identification device in the method embodiment of the present application. Please refer to the principle and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.
[0134] The embodiment of the present application also provides a chip, which can execute the relevant steps of the information identification device in the aforementioned method embodiment. In a possible implementation, the chip includes at least one processor, at least one first memory and at least one second memory; wherein the aforementioned at least one first memory and the aforementioned at least one processor are interconnected through a line, and the aforementioned first memory stores instructions; the aforementioned at least one second memory and the aforementioned at least one processor are interconnected through a line, and the aforementioned second memory stores data that needs to be stored in the aforementioned method embodiment.
[0135] For each device or product applied to or integrated in a chip, each module contained therein may be implemented in the form of hardware such as circuits, or at least some of the modules may be implemented in the form of software programs, which run on a processor integrated inside the chip, and the remaining (if any) modules may be implemented in the form of hardware such as circuits.
[0136] An embodiment of the present application also provides a computer-readable storage medium, in which one or more instructions are stored, and the one or more instructions are suitable for being loaded by a processor and executing the method provided by the above method embodiment.
[0137] The embodiment of the present application also provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the method provided by the above method embodiment.
[0138] It should be noted that, for the above-mentioned various method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0139] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0140] The modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0141] A person skilled in the art may understand that all or part of the steps in the various methods of the above-mentioned embodiments may be completed by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, which may include a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.
[0142] The above disclosure is only a preferred embodiment of the present application, which is only a part of the embodiments of the present application and cannot be used to limit the scope of rights of the present application.
Claims
1. An information identification method, characterized in that: The method comprises: Acquire first information to be identified; Determine a matching result between the first information to be identified and the illegal word data set; the matching result includes a target illegal word that exists in both the first information to be identified and the illegal word data set; Determining whether the matching result satisfies a preset condition, wherein the number of the target illegal words is one or more; the preset condition includes: one or more of the following: the length of the target illegal word is less than a first threshold, the number of the target illegal words is less than a second threshold, and the illegal degree value of the first information to be identified is less than a third threshold, and the illegal degree value of the first information to be identified is determined by calculating the cumulative illegal degree values corresponding to the parts of speech of the multiple target illegal words matched in the first information to be identified, and one part of speech corresponds to one illegal degree value; If the length of the target illegal word is greater than or equal to the first threshold, or the number of the target illegal words is greater than or equal to the second threshold, or the illegal degree value of the first information to be identified is greater than or equal to the third threshold, then the first information to be identified is determined to be illegal; If the matching result meets the preset condition, then obtaining the second information to be identified including the target illegal word in the first information to be identified; the second information to be identified includes a first word, an attributive and a second word; in the second information to be identified, the first word, the attributive and the second word appear in descending order; the number of characters between the first word and the attributive is greater than a preset number; Determine the semantic dependency relationship between the words in the second information to be identified by using the deep model; wherein the semantic dependency relationship refers to the degree of correlation between the words in the second information to be identified, and determine the modification relationship between the attributive and the first word and the second word by using the semantic dependency relationship; Whether the second information to be identified violates the rules is determined based on the semantic dependency, so as to determine whether the first information to be identified violates the rules based on whether the second information to be identified violates the rules.
2. The method according to claim 1, characterized in that: The step of acquiring the second information to be identified including the target illegal word from the first information to be identified includes: Determining a position of the target illegal word in the first information to be identified; According to the position, the first information to be recognized is segmented to obtain second information to be recognized including the target illegal word; wherein the length of characters included in the second information to be recognized is less than a fourth threshold, and / or the second information to be recognized has a complete sentence structure.
3. The method according to claim 1, characterized in that: The first information to be identified is crawled information that does not match an object in a filtering object data set, and the filtering object data set includes blacklist objects and / or whitelist objects.
4. The method according to claim 3, characterized in that The method further comprises: According to the crawling strategy, crawling obtains the crawling information; The crawling strategy includes one or more of the following: During the crawling process, a preset number of requests are sent using the first information, and requests are subsequently sent using the second information; the first information is identity information and / or address information; If it is detected that the page structure of the crawled web page is not a preset structure, formatting the page structure of the web page; If it is detected that the crawled uniform resource locator URL is incomplete, the page corresponding to the crawled URL is dynamically captured.
5. An information recognition device, characterized in that: The method comprises means for executing the method according to any one of claims 1 to 4.
6. An information recognition device, characterized in that: Including processors; The processor is configured to execute the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed, the method according to any one of claims 1 to 4 is executed.
Citation Information
Patent Citations
Illegal text recognition method and device, storage medium and electronic device
CN111738011A