Family network fraud risk early warning method, device, equipment, product and medium

By analyzing home network communication data and utilizing preset scenario recognition rules and text semantic classification models, fraud risk scenarios and applications in home networks are identified, and detailed early warning information is generated. This solves the problem of incomplete early warning information in existing technologies and achieves accurate fraud risk early warning for home networks.

CN121125132APending Publication Date: 2025-12-12CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411399638.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing technologies generate incomplete warning information for home network fraud risk alerts, making it difficult for users to clearly understand the types and sources of fraud risks, thus hindering the implementation of effective preventative measures.

Method used

By acquiring home network communication data, analyzing web page text information and access records, using preset scenario recognition rules and text semantic classification models to identify fraud risk scenarios, marking fraud risk domains, and generating detailed warning information to push to family members.

Benefits of technology

It provides accurate identification and early warning of specific fraud risks in home networks, helping users take timely preventive measures and improve their network security protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125132A_ABST
    Figure CN121125132A_ABST
Patent Text Reader

Abstract

The invention provides a home network fraud risk early warning method, device, equipment, product and medium, and relates to the technical field of network security. The method comprises the following steps: determining webpage text information of a website to be identified and an access record of the website to be identified based on home network communication data; identifying the webpage text information based on a preset scene identification rule, and determining a fraud risk scene of the website to be identified based on an identification result; marking the website to be identified as a fraud risk domain name based on a comparison result between the webpage text information and each piece of sensitive application information in a preset sensitive application information set; determining a fraud application name based on the webpage text information; and generating early warning information based on the fraud risk scene, the fraud application name and the access record, and pushing the early warning information to family members in the home network. According to the method, the fraud risk type and source of the website to be identified are identified, comprehensive early warning information is generated, and the home network security protection capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a method, apparatus, equipment, product, and medium for early warning of home network fraud risks. Background Technology

[0002] With the rapid development of the internet, home networks have become an indispensable part of people's daily lives. Family members access the internet through various devices such as computers, mobile phones, and tablets. While this brings increased convenience, it also brings security risks. Fraudsters frequently target home users with deceptive tactics such as spoofing websites, applications, and even fake advertisements, causing unnecessary losses. Current technologies, when dealing with fraud in home networks, typically only generate fraud warnings at the single-website level for the entire family. Specifically, existing technologies identify potential fraud risks and generate warnings based on information from visited web pages or applications. However, these warnings are incomplete; users do not clearly understand the types and sources of fraud risks they face. This lack of specificity makes it difficult for home users to take effective preventative measures, thus limiting their network security protection capabilities. Summary of the Invention

[0003] This invention provides a method, device, equipment, product, and medium for early warning of home network fraud risks, in order to solve the shortcomings of existing technologies in generating incomplete early warning information, which prevents users from clearly understanding the types and sources of fraud risks they face, and thus makes it difficult to take effective preventive measures.

[0004] This invention provides a method for early warning of home network fraud risks, comprising: Acquire home network communication data; determine the webpage text information and access records of the URL to be identified based on the home network communication data; The webpage text information is identified based on preset scene recognition rules to obtain recognition results, and the fraud risk scenario of the URL to be identified is determined based on the recognition results; Based on the comparison results between the webpage text information and each sensitive application information in the preset sensitive application information set, the URL to be identified is marked as a fraud risk domain. Based on the webpage text information of the URL to be identified, which is marked as the domain name of the fraud risk, the name of the fraudulent application is determined; Based on the fraud risk scenario, the name of the fraudulent application, and the access records, an early warning message is generated and pushed to family members in the home network.

[0005] According to the present invention, a method for early warning of home network fraud risks includes identifying the webpage text information based on preset scenario identification rules to obtain identification results, and determining the fraud risk scenario of the URL to be identified based on the identification results, comprising: If the identification result is successful, then the fraud risk scenario of the URL to be identified is determined based on the identification result; If the identification result is a failure, the webpage text information is input into the text semantic classification model to obtain the classification result output by the text semantic classification model, and the fraud risk scenario is determined based on the classification result; The text semantic classification model is trained based on the sample webpage text information and the corresponding sample classification result labels.

[0006] According to the present invention, a method for early warning of family network fraud risks is provided, wherein each sensitive application information in the preset sensitive application information set includes full-text information and sensitive word information; the comparison result includes a first similarity result and a second similarity result; The step of marking the URL to be identified as a fraud risk domain based on the comparison results between the webpage text information and each sensitive application information in a preset sensitive application information set includes: Determine the first similarity result between the webpage text information and the full text information of each sensitive application in the preset sensitive application information set; Based on the first similarity results that are greater than the first preset threshold, a set of similar sensitive application information is determined in the preset sensitive application information set; the URL to be identified corresponding to the webpage text information is marked as a high-risk domain name; Based on the second similarity results between the webpage text information and the sensitive word information of each sensitive application in the similar sensitive application information set, the URL to be identified corresponding to the webpage text information is marked as a fraud risk domain.

[0007] According to the present invention, a method for early warning of home network fraud risks includes marking the URL to be identified corresponding to the webpage text information as a fraud risk domain based on the second similarity result between the webpage text information and the sensitive word information of each sensitive application in the similar sensitive application information set. If any of the second similarity results is greater than the second preset threshold, then the URL to be identified corresponding to the webpage text information is marked as a fraud risk domain. If all the second similarity results are less than the second preset threshold, then the webpage content corresponding to the URL to be identified is snapshotted to determine the snapshot webpage text information; Based on the third similarity results between the text information of the snapshot webpage and the sensitive word information of each sensitive application in the similar sensitive application information set, the URL to be identified is marked as a fraud risk domain.

[0008] According to a method for early warning of home network fraud risk provided by the present invention, the web page text information of the URL to be identified, which is marked as the fraud risk domain name, is input into the encoding layer of the application name extraction model to obtain the web page text vector output by the encoding layer; The webpage text vector is input into the feature extraction layer of the application name extraction model to obtain the webpage text sequence features output by the feature extraction layer; The webpage text sequence features are input into the annotation layer of the application name extraction model to obtain the label sequence output by the annotation layer; the annotation layer is used to annotate the features of individual characters in the webpage text sequence features. The name of the fraudulent application is determined based on the tag sequence; The application name extraction model is trained based on sample webpage text information and the corresponding sample label sequence labels.

[0009] According to a method for early warning of home network fraud risk provided by the present invention, the access record includes the access time of the URL to be identified, and the device information for accessing the URL to be identified; The generation of early warning information based on the fraud risk scenario, the name of the fraudulent application, and the access records includes: Based on the device information that accessed the URL to be identified, determine the family member information that accessed the URL to be identified; Based on the fraud risk scenario, the name of the fraudulent application, the access time, and the family member information, an early warning message is generated.

[0010] The present invention also provides a home network risk warning device, comprising: An acquisition module is used to acquire home network communication data; and based on the home network communication data, determine the webpage text information of the URL to be identified and the access records of the URL to be identified; The scene recognition module is used to recognize the webpage text information based on preset scene recognition rules, obtain recognition results, and determine the fraud risk scenario of the URL to be identified based on the recognition results; The marking module is used to mark the URL to be identified as a fraud risk domain based on the comparison results between the webpage text information and each sensitive application information in the preset sensitive application information set; The name determination module is used to determine the name of the fraudulent application based on the webpage text information of the URL to be identified, which is marked as the fraud risk domain name; The early warning module is used to generate early warning information based on the fraud risk scenario, the name of the fraudulent application, and the access records, and push the early warning information to family members in the home network.

[0011] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the home network fraud risk warning method as described above.

[0012] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the home network fraud risk warning method as described above.

[0013] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the home network fraud risk warning method as described above.

[0014] This invention provides a method, apparatus, device, product, and medium for early warning of home network fraud risks. The method involves: acquiring home network communication data; determining the webpage text information and access records of a URL to be identified based on the home network communication data; identifying the webpage text information based on preset scenario identification rules to obtain identification results; determining the fraud risk scenario of the URL to be identified based on the identification results; marking the URL to be identified as a fraud risk domain based on comparison results between the webpage text information and various sensitive application information in a preset sensitive application information set; determining the name of a fraudulent application based on the webpage text information of the URL marked as a fraud risk domain; generating early warning information based on the fraud risk scenario, the name of the fraudulent application, and the access records; and pushing the early warning information to family members in the home network. This invention extracts [data / information] by analyzing home network communication data. The process involves analyzing the webpage text information and access records of the URL to be identified. Webpage text information is crucial for identifying potential fraud risks, while access records help understand the user's actual access behavior and timing. By analyzing the extracted webpage text information, possible fraud risk scenarios are identified. The webpage text information is compared with the information of various sensitive applications in a pre-set sensitive application information set. If the comparison results are similar, the URL to be identified is marked as a fraud risk domain, ensuring accurate identification of potential fraud domains. Furthermore, when the URL to be identified is marked as a fraud risk domain, the name of the fraudulent application is extracted from its webpage text information. This generates comprehensive early warning information, which includes the specific type of fraud risk and the applications involved, and is pushed to relevant family members in the home network. Through such early warnings, home users can receive timely and targeted alerts and take effective preventative measures. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 This is one of the flowcharts of the home network fraud risk warning method provided by the present invention.

[0017] Figure 2 This is a schematic diagram of the application name extraction model of the home network fraud risk early warning method provided by the present invention.

[0018] Figure 3 This is a schematic diagram of the structure of the home network fraud risk warning device provided by the present invention.

[0019] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0021] The patent applicant has discovered significant shortcomings in existing technologies for fraud risk warnings within home networks. Specifically, existing technologies typically only provide fraud warnings for a single website across the entire household. While such warnings can identify potential fraudulent websites, they lack detailed descriptions of specific fraud scenarios and applications, preventing home users from developing a clear understanding of the fraud risks they face. For example, existing technologies cannot accurately identify specific fraud scenarios, such as investment scams, fake apps, or false advertising, nor can they clearly identify specific fraudulent applications. This generalized approach limits home users' actual ability to prevent online fraud risks, making it difficult for them to take effective countermeasures.

[0022] While existing technologies have been improved in some aspects, such as increasing the number or types of features or identifying variant fraudulent websites through the calculation of indicators like information matching degree, these technologies still have shortcomings. Existing technologies often only generate a general warning for the entire household, lacking specific, targeted information. Therefore, even if household users receive a warning, they cannot clearly know the specific fraud scenario and application, which limits their ability to intervene. For example, users cannot know which device, which application, or in what specific scenario the fraud risk is present. This vague warning information makes it difficult for users to respond effectively to online fraud, thus limiting their cybersecurity protection capabilities.

[0023] To address the above problems, the present invention proposes the following embodiments.

[0024] Figure 1 This is a flowchart illustrating the home network fraud risk early warning method provided by the present invention, as follows: Figure 1 As shown, the method includes the following: Step 110: Obtain home network communication data; based on the home network communication data, determine the webpage text information of the URL to be identified and the access records of the URL to be identified.

[0025] Here, webpage text information refers to the text content on a webpage, including the title and body. This information is crucial for identifying whether a webpage involves fraudulent activities. Access records include information such as the time, device, and frequency of user visits to the URL to be identified, serving as the data basis for generating subsequent warning information. Home network communication data refers to home network traffic.

[0026] In one embodiment, home network communication data is acquired, and preprocessing operations are performed on the home network communication data, including data cleaning and data filtering, to obtain preprocessed home network communication data. Further, access records of the URL to be identified are extracted from the preprocessed data. The preprocessed home network communication data is then subjected to web crawling, whereby the crawler captures web page content to obtain the web page text information of the URL to be identified, including the title and body of the web page. Based on the web page text information and access records of the URL to be identified, detailed information support is provided for subsequent fraud risk identification.

[0027] It should be noted that home network communication data may include multiple URLs to be identified. This embodiment only provides an example explanation for one URL to be identified. The operations for other URLs to be identified are the same as those in this embodiment.

[0028] It should be noted that data filtering includes whitelist filtering, blacklist filtering, webpage request status code filtering, and server address attribution filtering.

[0029] Specifically, whitelisting refers to excluding URLs from a whitelist. Whitelists typically contain trusted URLs that usually require no further analysis. Blacklisting identifies and excludes known malicious or fraudulent URLs, preventing them from being mistakenly identified as targets. Web request status code filtering only processes web requests with a 200 status code, indicating a successful data return; other status codes, such as 404 (Not Found) or 500 (Server Error), are ignored. Server address attribution filtering performs initial screening based on the server's geographical location or region of origin, helping to identify potentially suspicious or unfamiliar servers.

[0030] For example, data cleaning refers to removing irrelevant or noisy data to improve the accuracy of subsequent analysis. This includes, for instance, removing abnormal or erroneous network requests.

[0031] For example, web crawling refers to web crawling of home network traffic that has undergone preliminary preprocessing; the crawler will grab web page content to obtain the web page text information of the URL to be identified, including the title and body of the web page.

[0032] Step 120: Recognize the webpage text information based on preset scene recognition rules to obtain recognition results, and determine the fraud risk scenario of the URL to be identified based on the recognition results.

[0033] Here, the pre-defined scenario recognition rules refer to a series of rules or standards defined in advance in the fraud risk recognition system. These rules are designed based on known fraud methods, patterns and characteristics, and aim to determine whether a webpage conforms to a specific fraud scenario by analyzing the webpage text information. These rules may include specific keywords, patterns or content features.

[0034] Here, fraud risk scenarios refer to specific types or patterns of fraud. Each fraud risk scenario has its own unique characteristics and manifestations, so each scenario requires one or more specific rules for identification.

[0035] Specifically, there is a direct correspondence between the preset scenario identification rules and fraud risk scenarios. Each identification rule is designed for one or more specific fraud scenarios; when webpage text information is analyzed through these rules, the system determines whether the webpage matches a particular fraud risk scenario based on the matching results. This means that the matching results of the rules directly determine which risk scenario the webpage belongs to. In other words, by identifying the webpage text information according to the preset scenario identification rules, the identification results can further determine the fraud risk scenario corresponding to the URL to be identified.

[0036] Step 130: Based on the comparison results between the webpage text information and each sensitive application information in the preset sensitive application information set, the URL to be identified is marked as a fraud risk domain.

[0037] Here, the preset sensitive application information set is a predefined database containing sensitive information about known fraudulent applications or related information. This information set is used to compare with web page text information to identify potential fraud risks.

[0038] For example, the sensitive application information in the preset sensitive application information set can be relevant information extracted from each confirmed fraudulent application. The extracted content includes the fraudulent application name, HTML files, and fraud evidence. HTML files refer to the webpage structure and content related to the fraudulent application; these files contain information such as the webpage's text, layout, and style, which can help identify similar fraudulent websites. Fraud evidence refers to specific evidence related to fraudulent activities, such as user reports, identified fraud methods and patterns. This evidence supports further analysis and judgment.

[0039] Specifically, the information of each fraudulent application is further analyzed and processed to generate a unique "fingerprint". These fingerprints can help quickly identify similar or variant fraudulent applications. That is, the preset sensitive application information set includes multiple sensitive application information; each sensitive application information corresponds to a fraudulent application; each sensitive application information includes the name of the fraudulent application, the characteristics of the HTML file (such as text features, structural features, etc.), and related fraud evidence.

[0040] The advantage of this approach is that by comparing webpage text information with a sensitive application information set, the system can more accurately identify fraudulent URLs that may disguise themselves as legitimate websites. This dual verification process effectively reduces false positives and improves overall identification accuracy. The use of the sensitive application information set helps the system detect fraudulent websites that attempt to evade identification through subtle modifications or covert methods. Even if the webpage text contains only partial sensitive information, the comparison process can still identify potential risks and prevent these websites from being missed. The results of marking domains as fraudulent risk will further drive the subsequent generation of warning information, providing users with more accurate and specific risk alerts and enhancing the security capabilities of home users.

[0041] The specific comparison process between the webpage text information and each sensitive application information in the preset sensitive application information set can be referred to in the following embodiments, which will not be elaborated on here.

[0042] Step 140: Determine the name of the fraudulent application based on the webpage text information of the URL to be identified, which is marked as the fraud risk domain.

[0043] The fact that the URL to be identified is marked as a fraud risk domain means that, based on the detection and comparison in the previous steps, the URL is determined to be highly likely to be related to online fraud. Therefore, based on the webpage text information of the URL to be identified, the name of the fraudulent application can be further determined.

[0044] It's important to note that a URL refers to the address of a resource used to access a specific webpage. It typically includes a protocol (such as HTTP / HTTPS), domain name, path, and query parameters. A URL can point to a webpage or a specific online resource, such as a file or application. A domain name is usually used to identify the owner of a website or the location of its server.

[0045] In one embodiment, by analyzing the webpage text information of the URL to be identified, which is marked as a fraud risk domain, it is possible to determine whether the URL poses a fraud risk and further infer the specific application name that the URL is attempting to impersonate. This method leverages the close relationship between URLs, domain names, and applications, enabling the system to more accurately identify and mark specific fraudulent applications, thereby providing users with more targeted security warnings.

[0046] For example, when a URL is marked as a fraud risk domain, it indicates that the URL has been confirmed to be related to potential fraudulent activities. Based on this judgment, further analysis of the webpage text content corresponding to the URL is conducted. The webpage text typically contains identification information related to a specific application, such as the application name, version number, and feature description.

[0047] By parsing the text content of these web pages, the system can extract the explicit application name. This is because most applications mention their name on their official websites or related web pages, whether through page titles, navigation menus, or copyright information. This information is publicly displayed in the web page text, especially on pages that offer downloads, logins, or user guidance.

[0048] Step 150: Based on the fraud risk scenario, the name of the fraudulent application, and the access records, generate early warning information and push the early warning information to family members in the home network.

[0049] In one embodiment, fraud risk scenarios, fraudulent application names, and access records are integrated. This allows for a comprehensive consideration of the risk scenario, the specific application, and its actual situation within the home network; based on the integrated information, detailed warning information is generated; and the generated warning information is pushed to family members within the home network, typically via notifications or messages, enabling family members to be aware of potential risks in a timely manner.

[0050] The specific methods for pushing early warning information to family members in the home network can be found in the following examples, which will not be elaborated on here.

[0051] The home network fraud risk early warning method provided in this invention analyzes home network communication data to extract webpage text information and access records of the URL to be identified. Webpage text information is key data for identifying potential fraud risks, while access records help understand the user's actual access behavior and time. By analyzing the extracted webpage text information, possible fraud risk scenarios are identified. The webpage text information is compared with the information of various sensitive applications in a preset sensitive application information set. If the comparison results are similar, the URL to be identified is marked as a fraud risk domain, ensuring accurate identification of potential fraud domains. Furthermore, when the URL to be identified is marked as a fraud risk domain, the name of the fraudulent application is extracted from its webpage text information, thereby generating comprehensive early warning information. This early warning information includes the specific type of fraud risk and the applications involved, and is pushed to relevant family members in the home network. Through such early warnings, home users can receive timely and targeted warnings and take effective preventative measures.

[0052] Based on any of the above embodiments, in this method, the step of recognizing the webpage text information based on preset scenario recognition rules to obtain a recognition result, and determining the fraud risk scenario of the URL to be identified based on the recognition result, includes: If the identification result is successful, then the fraud risk scenario of the URL to be identified is determined based on the identification result; If the identification result is a failure, the webpage text information is input into the text semantic classification model to obtain the classification result output by the text semantic classification model, and the fraud risk scenario is determined based on the classification result; The text semantic classification model is trained based on the sample webpage text information and the corresponding sample classification result labels.

[0053] Here, the text semantic classification model refers to a machine learning-based model that, after training, can classify web page text based on its semantic features. During training, sample web page texts and their corresponding classification labels are used to improve the model's ability to classify new web page texts.

[0054] In one embodiment, if the identification result is a failure, that is, when the preset scenario identification rules cannot identify the fraud risk scenario corresponding to the webpage text information of the website to be identified, the webpage text information is input into the text semantic classification model, and the webpage content is further identified through a more detailed semantic classification model, thereby ensuring that no potential fraud risk scenarios are missed.

[0055] Specifically, when the preset scenario identification rules fail to accurately classify the webpage text information into known fraud risk scenarios, the webpage text information is passed to a text semantic classification model. This model is trained using a large amount of labeled webpage text information and corresponding classification labels, and can classify webpages based on the semantic features of the text. Even if the preset rules fail to identify risks, the text semantic classification model can further identify potential fraud scenarios by analyzing the semantic content of the webpage, reducing false negatives.

[0056] The home network fraud risk early warning method provided in this invention firstly uses preset scene recognition rules to quickly identify known fraud scenarios, providing efficient initial screening. Secondly, for cases where identification fails, a text semantic classification model is used for further processing, which not only supplements the shortcomings of the preset rules but also identifies new or complex fraud scenarios. The text semantic classification model is trained on a large amount of labeled data, ensuring high accuracy and robustness of the classification results. Therefore, the entire process not only improves the coverage and accuracy of identification but also reduces the false negative rate, ensuring that users receive comprehensive and accurate fraud risk warnings, thereby effectively preventing potential network fraud risks.

[0057] Based on any of the above embodiments, in this method, each sensitive application information in the preset sensitive application information set includes full-text information and sensitive word information; the comparison result includes a first similarity result and a second similarity result; The step of marking the URL to be identified as a fraud risk domain based on the comparison results between the webpage text information and each sensitive application information in a preset sensitive application information set includes: Determine the first similarity result between the webpage text information and the full text information of each sensitive application in the preset sensitive application information set; Based on the first similarity results that are greater than the first preset threshold, a set of similar sensitive application information is determined in the preset sensitive application information set; the URL to be identified corresponding to the webpage text information is marked as a high-risk domain name; Based on the second similarity results between the webpage text information and the sensitive word information of each sensitive application in the similar sensitive application information set, the URL to be identified corresponding to the webpage text information is marked as a fraud risk domain.

[0058] Here, full-text information refers to the text content of each sensitive application in the sensitive application information set, including the text and titles on web pages. Full-text information is used to calculate the similarity with the web page text information of the URL to be identified. By comparing the full-text information, it can be determined whether the web page text information of the URL to be identified contains content similar to known fraudulent applications.

[0059] Here, sensitive word information refers to predefined keywords, phrases, or sensitive words in each sensitive application information set. These words are typically associated with fraudulent activities. Sensitive word information is used for further similarity calculations. By identifying sensitive word information in webpage text, it is possible to confirm the presence of indicators of fraudulent activities, thereby improving the accuracy of identification.

[0060] Here, the first preset threshold is a threshold set in the similarity calculation to determine whether the similarity result is high enough to determine whether to mark the URL to be identified as a high-risk domain.

[0061] Here, the first similarity result refers to the similarity score between the webpage text information of the URL to be identified and the full-text information of each sensitive application in the sensitive application information set. The second similarity result refers to the similarity score between the webpage text information of the URL to be identified and the sensitive word information of each sensitive application in the similar sensitive application information set.

[0062] Specifically, the first similarity result is used for initial screening of potentially fraudulent domain names. If the first similarity result is greater than a preset threshold, it indicates that the webpage text highly matches information from certain sensitive applications and is initially marked as a high-risk domain name. The second similarity result is used for refined screening to ensure that only those webpages that match not only the full-text information but also exhibit similarities in sensitive words are ultimately marked as fraudulent domain names.

[0063] In one embodiment, the webpage text information of the URL to be identified is compared with the full-text information of each sensitive application in a preset sensitive application information set, and the calculated similarity result is called the first similarity result. Based on each of the first similarity results, sensitive application information with a similarity greater than a first preset threshold is selected, and this information constitutes a similar sensitive application information set. The URL to be identified is initially marked as a high-risk domain name. Further, the webpage text information is compared with the sensitive word information of each sensitive application in the similar sensitive application information set to obtain a second similarity result. Based on the high-risk domain name, the second similarity result is used for further screening. If the similarity between the webpage text information and the sensitive word information in the similar sensitive application information set also reaches a preset threshold, then the URL is finally marked as a fraud risk domain name.

[0064] The advantage of this approach is that it improves the comprehensiveness and accuracy of identification through two similarity calculation steps. The first step, full-text information similarity filtering, can quickly eliminate most non-fraudulent web pages, while the second step, sensitive word similarity filtering, further refines the identification, ensuring that only domains with genuine fraud risks are marked.

[0065] For example, the full-text information can be obtained by crawling the sensitive application information in a preset sensitive application information set using web crawling technology.

[0066] The collective operation of marking the URLs to be identified corresponding to the webpage text information as fraud risk domains based on each of the second similarity results can be referred to in the following embodiments, which will not be elaborated on here.

[0067] The home network fraud risk early warning method provided in this invention performs a preliminary similarity comparison between the full-text information of a preset sensitive application information set and the text information of the webpage to be identified. This comparison provides a first similarity result, which is used to filter out webpage text information with high similarity and mark it as a high-risk domain. Furthermore, by performing a more detailed second similarity comparison between the webpage text information of these high-risk domains and sensitive word information in the sensitive application information set, domains that truly pose a fraud risk can be accurately marked. Through the dual comparison of full-text information and sensitive word information, the accuracy of identifying fraudulent application domains is improved, and the possibility of false positives and false negatives is reduced.

[0068] Based on any of the above embodiments, in this method, marking the URL to be identified corresponding to the webpage text information as a fraud risk domain based on the second similarity result between the webpage text information and the sensitive word information of each sensitive application information in the similar sensitive application information set includes: If any of the second similarity results is greater than the second preset threshold, then the URL to be identified corresponding to the webpage text information is marked as a fraud risk domain. If all the second similarity results are less than the second preset threshold, then the webpage content corresponding to the URL to be identified is snapshotted to determine the snapshot webpage text information; Based on the third similarity results between the text information of the snapshot webpage and the sensitive word information of each sensitive application in the similar sensitive application information set, the URL to be identified is marked as a fraud risk domain.

[0069] Here, the second preset threshold is a threshold set in the similarity calculation to determine whether the similarity result is high enough to determine whether the URL to be identified should be marked as a fraud risk domain. The third similarity result is the similarity calculation result between the snapshot webpage text information and the sensitive application information, used to conduct a more in-depth similarity comparison to confirm the fraud risk.

[0070] Here, snapshot processing refers to archiving the current state of the webpage to be identified for subsequent analysis; this step generates snapshot webpage text information, which is a static version that records the content of the webpage at that time.

[0071] In one embodiment, when all second similarity results are less than a second preset threshold, that is, the webpage text information does not match the sensitive word information in each sensitive application information, it indicates that the current text information cannot directly confirm the fraud risk; at this time, snapshot processing is adopted; to determine the snapshot webpage text information; based on the comparison between the snapshot webpage text information and the sensitive word information in the preset sensitive application information set, a third similarity result is calculated; furthermore, if there is a third similarity result greater than a certain threshold in the third similarity result, it indicates that the snapshot webpage text information is related to the sensitive application, and then the URL is marked as a fraud risk domain.

[0072] Specifically, the webpage text information is extracted from the URL to be identified using web crawling technology. Web crawling typically captures the HTML content of the webpage at the target URL, including the page title and body. Furthermore, some fraudulent websites employ masking techniques, meaning the webpage text information captured by the crawler may be incomplete or obscured, potentially causing the initial comparison of sensitive words to fail to accurately identify fraud risks. When the webpage text information obtained through web crawling fails to reach a preset similarity threshold, it indicates that the crawled data may be incomplete or inaccurate. In this case, snapshot processing is required, and OCR technology is used to extract the snapshot webpage text information. This step captures the actual displayed content of the webpage at the target URL, overcoming any information that the crawler might miss. Further, the snapshot webpage text information is re-compared with the sensitive word information in the sensitive application information to calculate a third similarity result. If the similarity value in the snapshot text exceeds the third preset threshold, it can more accurately confirm whether the target URL is a fraudulent domain.

[0073] The home network fraud risk early warning method provided in this invention compensates for the deficiencies in the initial crawling by performing secondary analysis of the text information of snapshot web pages, ensuring comprehensive detection of hidden or obscured content, thereby improving the accuracy of fraud risk identification. Snapshot processing can effectively deal with the obscuration techniques of fraudulent websites, enabling accurate extraction and analysis of key content even when information is hidden, thus improving anti-interference capabilities. The use of snapshots ensures that even if the initial similarity analysis fails to confirm the fraud risk, the web page content of the URL to be identified can still be thoroughly evaluated, improving the overall reliability and accuracy of detection.

[0074] Based on any of the above embodiments, in this method, determining the fraudulent application name based on the webpage text information of the URL to be identified, which is marked as the fraudulent risk domain name, includes: The webpage text information of the URL to be identified, which is marked as the fraud risk domain, is input into the encoding layer in the application name extraction model to obtain the webpage text vector output by the encoding layer; The webpage text vector is input into the feature extraction layer of the application name extraction model to obtain the webpage text sequence features output by the feature extraction layer; The webpage text sequence features are input into the annotation layer of the application name extraction model to obtain the label sequence output by the annotation layer; the annotation layer is used to annotate the features of individual characters in the webpage text sequence features. The name of the fraudulent application is determined based on the tag sequence; The application name extraction model is trained based on sample webpage text information and the corresponding sample label sequence labels.

[0075] To better understand the application name extraction model of this invention, as shown in Figure 2, the application name extraction model includes an encoding layer, a feature extraction layer, and a labeling layer.

[0076] Here, the encoding layer, the first part of the application name extraction model, transforms webpage text information into webpage text vectors. These vectors represent the semantic features of the webpage text for subsequent processing. The feature extraction layer, the second part of the model, is responsible for extracting webpage text sequence features from the webpage text vectors. These features reflect important information and patterns in the webpage content for further analysis. The annotation layer, the third part of the model, annotates the features of individual characters in the webpage text sequence features. The annotation layer outputs a sequence of labels used to identify key entities in the webpage text, such as the application name. Training is performed based on training sample webpage text information and corresponding sample label sequences. During training, the model learns how to map webpage text information to the corresponding label sequences.

[0077] For example, the encoding layer can be a BERT model. The web page text information is first encoded by a pre-trained BERT model to generate corresponding feature vectors. These vectors capture the semantic information of the characters or words.

[0078] For example, the feature extraction layer can be an LSTM layer (Long Short-Term Memory). LSTM is a special type of recurrent neural network (RNN) that can effectively capture long-short-term dependencies when processing sequential data. Compared to traditional RNNs, LSTM can alleviate the problem of long-term dependencies by using "memory units" to determine which information to retain and discard. In this layer, the BERT-encoded vector sequence is input into the LSTM layer. LSTM processes the data at each time step through its internal memory mechanism, extracting web page text sequence features related to the context.

[0079] For example, the annotation layer can be a CRF layer (Conditional Random Field). CRF is a model for sequence labeling, suitable for tasks where there are dependencies between labels (such as named entity recognition). It can consider contextual information, thereby globally optimizing the labels for the entire sequence. The feature vector sequence output by the LSTM layer is input into the CRF layer. The CRF layer learns the optimal label sequence based on these features. For each input sequence, the CRF calculates all possible label paths and selects the path that results in the highest global score as the final annotation sequence. For example, it might assign labels such as "B-APP_NAME" (the beginning of the application name), "I-APP_NAME" (the middle of the application name), or "O" (not belonging to any entity) to certain characters. If "ExampleApp" appears in the webpage text, the label sequence might be "B-APP_NAME I-APP_NAME I-APP_NAME". The system will combine these three characters labeled as application names to finally output "ExampleApp" as the application name.

[0080] The home network fraud risk warning method provided in this embodiment of the invention can effectively extract fraudulent application names from web page text information through a hierarchical processing method of application name extraction model, thereby improving the accuracy and efficiency of detecting fraudulent content.

[0081] Based on any of the above embodiments, in this method, the access record includes the access time of the URL to be identified, and the device information for accessing the URL to be identified; The generation of early warning information based on the fraud risk scenario, the name of the fraudulent application, and the access records includes: Based on the device information that accessed the URL to be identified, determine the family member information that accessed the URL to be identified; Based on the fraud risk scenario, the name of the fraudulent application, the access time, and the family member information, an early warning message is generated.

[0082] Here, access records include access time and device information, used to record and track specific website visits by users. Identifying specific family members through device information helps to accurately deliver alerts to relevant individuals.

[0083] In one embodiment, family member information that accesses the URL to be identified is determined from a device personnel mapping table based on the device information that accesses the URL to be identified.

[0084] It should be noted that the device information includes the home MAC address and the MAC address of the connected devices. The device-person mapping table includes multiple device information entries and multiple personnel information entries, and these multiple device information entries have a mapping relationship with the multiple personnel information entries.

[0085] In another embodiment, a warning message is generated based on the fraud risk scenario, the name of the fraudulent application, the access time, and family member information. This warning message enables more refined risk alerts.

[0086] The home network fraud risk early warning method provided in this invention generates accurate early warning information by comprehensively considering access records, fraud risk scenarios, and fraudulent application names. Specifically, it first extracts access time and device information from access records to identify specific family members. Then, it combines this information with fraud risk scenarios and fraudulent application names to generate personalized early warning alerts. In this way, users can clearly understand the time of the fraud incident, the family members involved, the specific scenario of the fraud, and the application name, thus enabling them to take appropriate measures in a timely manner. This detailed early warning method not only improves the accuracy of risk identification but also enhances users' ability to protect their home network security.

[0087] The following describes the home network fraud risk warning device provided by the present invention. The home network fraud risk warning device described below can be referred to in correspondence with the home network fraud risk warning method described above.

[0088] Figure 3 is a schematic diagram of the structure of the home network fraud risk warning device provided by the present invention. As shown in Figure 3, the home network fraud risk warning device includes: The acquisition module 310 is used to acquire home network communication data; and based on the home network communication data, determine the webpage text information of the URL to be identified and the access records of the URL to be identified. The scene recognition module 320 is used to recognize the webpage text information based on preset scene recognition rules, obtain recognition results, and determine the fraud risk scenario of the URL to be identified based on the recognition results; The marking module 330 is used to mark the URL to be identified as a fraud risk domain based on the comparison results between the webpage text information and each sensitive application information in the preset sensitive application information set. Name determination module 340 is used to determine the name of a fraudulent application based on the webpage text information of the URL to be identified, which is marked as the fraud risk domain name; The early warning module 350 is used to generate early warning information based on the fraud risk scenario, the name of the fraudulent application, and the access records, and push the early warning information to family members in the home network.

[0089] Figure 4An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a home network fraud risk warning method. The method includes: acquiring home network communication data; determining the webpage text information and access records of the URL to be identified based on the home network communication data; identifying the webpage text information based on preset scenario identification rules to obtain identification results; determining the fraud risk scenario of the URL to be identified based on the identification results; marking the URL to be identified as a fraud risk domain name based on the comparison results of the webpage text information with each sensitive application information in a preset sensitive application information set; determining the fraudulent application name based on the webpage text information of the URL to be identified marked as the fraud risk domain name; generating warning information based on the fraud risk scenario, the fraudulent application name, and the access records; and pushing the warning information to family members in the home network.

[0090] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0091] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the home network fraud risk warning method provided by the above methods. The method includes: acquiring home network communication data; determining the webpage text information of the URL to be identified and the access records of the URL to be identified based on the home network communication data; identifying the webpage text information based on a preset scenario identification rule to obtain an identification result; determining the fraud risk scenario of the URL to be identified based on the identification result; marking the URL to be identified as a fraud risk domain name based on the comparison results of the webpage text information with each sensitive application information in a preset sensitive application information set; determining the fraud application name based on the webpage text information of the URL to be identified marked as the fraud risk domain name; generating warning information based on the fraud risk scenario, the fraud application name, and the access records; and pushing the warning information to family members in the home network.

[0092] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a home network fraud risk warning method provided by the above methods. The method includes: acquiring home network communication data; determining webpage text information and access records of a URL to be identified based on the home network communication data; identifying the webpage text information based on preset scenario identification rules to obtain identification results; determining a fraud risk scenario of the URL to be identified based on the identification results; marking the URL to be identified as a fraud risk domain name based on comparison results between the webpage text information and each sensitive application information in a preset sensitive application information set; determining a fraudulent application name based on the webpage text information of the URL to be identified marked as the fraud risk domain name; generating warning information based on the fraud risk scenario, the fraudulent application name, and the access records; and pushing the warning information to family members in the home network.

[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for early warning of home-based online fraud risks, characterized in that, include: Obtain home network communication data; Based on the aforementioned home network communication data, determine the webpage text information of the URL to be identified and the access records of the URL to be identified; The webpage text information is identified based on preset scene recognition rules to obtain recognition results, and the fraud risk scenario of the URL to be identified is determined based on the recognition results; Based on the comparison results between the webpage text information and each sensitive application information in the preset sensitive application information set, the URL to be identified is marked as a fraud risk domain. Based on the webpage text information of the URL to be identified, which is marked as the domain name of the fraud risk, the name of the fraudulent application is determined; Based on the fraud risk scenario, the name of the fraudulent application, and the access records, an early warning message is generated and pushed to family members in the home network.

2. The method for early warning of home network fraud risks according to claim 1, characterized in that, The process of identifying the webpage text information based on preset scenario recognition rules to obtain recognition results, and determining the fraud risk scenario of the URL to be identified based on the recognition results, includes: If the identification result is successful, then the fraud risk scenario of the URL to be identified is determined based on the identification result; If the identification result is a failure, the webpage text information is input into the text semantic classification model to obtain the classification result output by the text semantic classification model, and the fraud risk scenario is determined based on the classification result; The text semantic classification model is trained based on the sample webpage text information and the corresponding sample classification result labels.

3. The method for early warning of family online fraud risks according to claim 1, characterized in that, Each sensitive application information in the preset sensitive application information set includes full-text information and sensitive word information; the comparison results include a first similarity result and a second similarity result; The step of marking the URL to be identified as a fraud risk domain based on the comparison results between the webpage text information and each sensitive application information in a preset sensitive application information set includes: Determine the first similarity result between the webpage text information and the full text information of each sensitive application in the preset sensitive application information set; Based on the first similarity results that are greater than the first preset threshold, a set of similar sensitive application information is determined in the preset sensitive application information set; the URL to be identified corresponding to the webpage text information is marked as a high-risk domain name; Based on the second similarity results between the webpage text information and the sensitive word information of each sensitive application in the similar sensitive application information set, the URL to be identified corresponding to the webpage text information is marked as a fraud risk domain.

4. The method for early warning of home network fraud risks according to claim 3, characterized in that, The second similarity result based on the webpage text information and the sensitive word information of each sensitive application in the similar sensitive application information set marks the URL to be identified corresponding to the webpage text information as a fraud risk domain, including: If any of the second similarity results is greater than the second preset threshold, then the URL to be identified corresponding to the webpage text information is marked as a fraud risk domain. If all the second similarity results are less than the second preset threshold, then the webpage content corresponding to the URL to be identified is snapshotted to determine the snapshot webpage text information; Based on the third similarity results between the text information of the snapshot webpage and the sensitive word information of each sensitive application in the similar sensitive application information set, the URL to be identified is marked as a fraud risk domain.

5. The method for early warning of home network fraud risks according to claim 1, characterized in that, The process of determining the fraudulent application name based on the webpage text information of the URL to be identified, marked as the fraudulent risk domain, includes: The webpage text information of the URL to be identified, which is marked as the fraud risk domain, is input into the encoding layer in the application name extraction model to obtain the webpage text vector output by the encoding layer; The webpage text vector is input into the feature extraction layer of the application name extraction model to obtain the webpage text sequence features output by the feature extraction layer; The webpage text sequence features are input into the annotation layer of the application name extraction model to obtain the label sequence output by the annotation layer; the annotation layer is used to annotate the features of individual characters in the webpage text sequence features. The name of the fraudulent application is determined based on the tag sequence; The application name extraction model is trained based on sample webpage text information and the corresponding sample label sequence labels.

6. The method for early warning of family online fraud risks according to claim 1, characterized in that, The access record includes the access time of the URL to be identified, as well as the device information for accessing the URL to be identified; The generation of early warning information based on the fraud risk scenario, the name of the fraudulent application, and the access records includes: Based on the device information that accessed the URL to be identified, determine the family member information that accessed the URL to be identified; Based on the fraud risk scenario, the name of the fraudulent application, the access time, and the family member information, an early warning message is generated.

7. A home-based online fraud risk early warning device, characterized in that, include: The acquisition module is used to acquire home network communication data. Based on the aforementioned home network communication data, determine the webpage text information of the URL to be identified and the access records of the URL to be identified; The scene recognition module is used to recognize the webpage text information based on preset scene recognition rules, obtain recognition results, and determine the fraud risk scenario of the URL to be identified based on the recognition results; The marking module is used to mark the URL to be identified as a fraud risk domain based on the comparison results between the webpage text information and each sensitive application information in the preset sensitive application information set; The name determination module is used to determine the name of the fraudulent application based on the webpage text information of the URL to be identified, which is marked as the fraud risk domain name; The early warning module is used to generate early warning information based on the fraud risk scenario, the name of the fraudulent application, and the access records, and push the early warning information to family members in the home network.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the home network fraud risk warning method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the home network fraud risk warning method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the home network fraud risk warning method as described in any one of claims 1 to 6.