Looking for breakthrough ideas for innovation challenges? Try Patsnap Eureka!
Method, device and equipment for identifying risk content
What is Al technical title?
Al technical title is built by PatSnap Al team. It summarizes the technical point description of the patent document.
A risk and content technology, applied in the field of identifying risk content, can solve the problems of risk content identification, manufacturing, difficulty, etc.
Active Publication Date: 2021-07-13
ADVANCED NEW TECH CO LTD
View PDF11 Cites 0 Cited by
Summary
Abstract
Description
Claims
Application Information
AI Technical Summary
This helps you quickly interpret patents by identifying the three key elements:
Problems solved by technology
Method used
Benefits of technology
Problems solved by technology
However, in order to prevent the published content from being identified as risky content, information publishers create more and more variant information, which makes it difficult to identify risky content
Method used
the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more
Image
Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
Click on the blue label to locate the original text in one second.
Reading with bidirectional positioning of images and text.
Smart Image
Examples
Experimental program
Comparison scheme
Effect test
Embodiment Construction
[0025] At present, in order to prevent the published content from being identified as risky content (risky content refers to information that has a negative impact on the purity of the network environment), information publishers create more and more variant information, which makes the identification of risky content difficult. difficulty. Generally, the methods of variation include but are not limited to: using homophones, or English, or abbreviated letters, or deformed Chinese characters, etc. to replace sensitive words in sentences that originally belonged to risky content. For example: "I lost 41,000, dog, fart, chicken, gold, cut the meat", "Interested + I'm Weixin sd232433, Huakoubei will help you mention it", etc. Among them, replacing "fund" with "jijin", "WeChat" with "weixin", and "huabei" with "huakoubei" all make identification more difficult. On the other hand, if only homonyms are considered, but the context of the sentence is not considered, there is a possibi...
the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More
PUM
Login to View More
Abstract
The embodiments of this specification provide a method, device, and equipment for identifying risky content. The method includes: for the target text to be identified, identifying whether the target text contains a risk text segment, and the pinyin corresponding to the risk text segment belongs to the set risk pinyin; if the target text contains the risk text segment, generating A feature sequence corresponding to the target text, the feature sequence including: the text features contained in the target text and the pinyin features corresponding to the risk text fragment; based on the feature sequence and a model obtained through machine learning algorithm training , to determine whether the target text belongs to risk content, the input of the model is a feature vector corresponding to each feature in the feature sequence, and the output of the model represents the possibility that the target text belongs to risk content.
Description
technical field [0001] One or more embodiments of this specification relate to the technical field of machine learning, and in particular to a method, device, and device for identifying risky content. Background technique [0002] At present, in order to maintain the purity of the Internet environment, it is necessary to identify and deal with undesirable risky content (such as tyranny, terrorism, pornography, gambling and drugs, and spam advertisements, etc.). Generally, an identification engine can be built for different risk content to identify risk content. However, in order to prevent the published content from being identified as risky content, information publishers create more and more variant information, which makes it difficult to identify risky content. For example, use homonyms to replace sensitive words in sentences that originally belonged to risk content, for example: "I lost 41,000, dog, fart, chicken, gold, cut the meat", among them, use "chicken gold" ins...
Claims
the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More
Application Information
Patent Timeline
Application Date:The date an application was filed.
Publication Date:The date a patent or application was officially published.
First Publication Date:The earliest publication date of a patent with the same application number.
Issue Date:Publication date of the patent grant document.
PCT Entry Date:The Entry date of PCT National Phase.
Estimated Expiry Date:The statutory expiry date of a patent right according to the Patent Law, and it is the longest term of protection that the patent right can achieve without the termination of the patent right due to other reasons(Term extension factor has been taken into account ).
Invalid Date:Actual expiry date is based on effective date or publication date of legal transaction data of invalid patent.