Website identification method and device, electronic equipment and storage medium
By obtaining the key website phrases and contextual statements of the website to be identified, combining semantic weights and position weights, the problem of inaccurate website recognition in the prior art is solved, and a more detailed and comprehensive and accurate website recognition is achieved.
Patent Information
- Application Number
- CN202510389030.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, when identifying whether a website is an asset based on the correlation between the company name and the local content of the website, it is easy to lead to inaccurate identification, ignore the semantic importance and location structure of the local content in the overall website, resulting in misjudgment and deviation.
By obtaining the key website phrases and contextual statements of the website to be identified, combining semantic weights and position weights, the target feature vectors are determined to improve the recognition accuracy.
It effectively alleviates the misjudgment problem caused by simple correlation, improves the accuracy and reliability of website recognition, and ensures the comprehensiveness and accuracy of the identification results.
Smart Images

Figure CN120257984A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security technology, and in particular, to a website recognition method, device, electronic device and storage medium. Background Art
[0002] With the digital development of enterprises, websites have become an important platform for carrying enterprise data and business operations. Comprehensively identifying all websites of an enterprise is a crucial task.
[0003] In the prior art, keyword matching or statistical analysis techniques are usually adopted to identify whether a website is an enterprise's website asset based on the relevance between the enterprise name and the partial content of the website.
[0004] However, simply identifying whether a website is an enterprise's website asset based on the relevance between the enterprise name and the partial content of the website may lead to inaccurate identification. For example, if some partial content is highly relevant to the enterprise name but is not the core theme of the website, then it is inaccurate to regard it as the enterprise's website asset. Summary of the Invention
[0005] Embodiments of the present application provide a website recognition method, device, electronic device and storage medium to improve the accuracy of website recognition.
[0006] The specific technical solutions provided by the embodiments of the present application are as follows:
[0007] In a first aspect, the present application provides a website recognition method, including:
[0008] Obtain each key website phrase related to the enterprise name of the target enterprise in the website content of the website to be recognized, and extract the partial content corresponding to each key website phrase. The partial content is the context statement of the corresponding key website phrase in the website content;
[0009] Based on the first importance degree of each key website phrase in the corresponding partial content and the second importance degree of each obtained partial content in the website content, determine the semantic weight of each partial content, and based on the position information of each partial content, determine the position weight of each partial content;
[0010] Based on the feature vector, semantic weight and position weight of each partial content, determine the target feature vector corresponding to the website to be recognized, and based on the target feature vector, determine the recognition result of the website to be recognized.
[0011] In a second aspect, the present application provides a website recognition device, and the device includes:
[0012] An acquisition module, configured to acquire each key website phrase related to the enterprise name of a target enterprise in the website content of a website to be recognized, and extract the respective local content corresponding to each key website phrase, where the local content is the context statement of the corresponding key website phrase in the website content;
[0013] A processing module, configured to determine the semantic weight of each local content based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content, and determine the position weight of each local content based on the position information of each local content;
[0014] A determination module, configured to determine the target feature vector corresponding to the website to be recognized based on the feature vector, semantic weight, and position weight of each local content, and determine the recognition result of the website to be recognized based on the target feature vector.
[0015] In a possible embodiment, when acquiring each key website phrase related to a target enterprise in the website content of a website to be recognized, the acquisition module is further configured to:
[0016] Acquire each initial website phrase included in the website content of the website to be recognized and each name phrase included in the enterprise name of the target enterprise;
[0017] Based on the correlation degree between the feature vector of each initial website phrase and the feature vector of each name phrase, determine each key website phrase in each initial website phrase whose correlation degree with the target enterprise meets the correlation degree condition.
[0018] In a possible embodiment, when determining the semantic weight of each local content based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content, the processing module is further configured to:
[0019] Respectively, based on the similarity between the local content corresponding to each key website phrase and the first mask content and the enterprise name, determine the first importance degree of each key website phrase in the corresponding local content, where the first mask content is the remaining content after removing the corresponding key website phrase from the corresponding local content;
[0020] Respectively, based on the similarity between each local content and the website content and the similarity between each local content and the corresponding second mask content, determine the second importance degree of each local content in the website content, where the second mask content is the remaining content after removing the corresponding local content from the website content;
[0021] Based on the first importance degree and the second importance degree of each local content, determine the semantic weight of each local content.
[0022] In a possible embodiment, when determining the first importance degree of each key website phrase in the corresponding local content based on the similarity between the local content and the first mask content corresponding to each key website phrase and the enterprise name respectively, the processing module is further configured to:
[0023] For each key website phrase, perform the following operations respectively:
[0024] Calculate the first similarity between the feature vector of the local content corresponding to a key website phrase and the feature vector of the enterprise name, and calculate the second similarity between the feature vector of the first mask content corresponding to a key website phrase and the feature vector of the enterprise name;
[0025] Calculate the difference between the first similarity and the second similarity to obtain the first importance degree of a key website phrase in its corresponding local content.
[0026] In a possible embodiment, when determining the second importance degree of each local content in the website content based on the similarity between each local content and the website content and the similarity with the corresponding second mask content respectively, the processing module is further configured to:
[0027] For each local content, perform the following operations respectively:
[0028] Calculate the third similarity between the feature vector of a local content and the feature vector of the website content, and calculate the fourth similarity between the feature vector of a local content and the feature vector of its corresponding second mask content;
[0029] Calculate the difference between the third similarity and the fourth similarity to obtain the second importance degree of a local content in the website content.
[0030] In a possible embodiment, when determining the position weight of each local content based on the position information of each local content, the processing module is further configured to:
[0031] Determine the position weight of each local content based on the length and the starting character position of each local content and the total length of the website content.
[0032] In a possible embodiment, when determining the target feature vector corresponding to the website to be recognized based on the feature vector, semantic weight and position weight of each local content, the determining module is further configured to:
[0033] Determine the fusion weight of each local content based on the semantic weight and position weight of each local content respectively;
[0034] Perform weighted summation on the feature vectors of each local content based on the fusion weight of each local content to obtain the target feature vector.
[0035] In a possible embodiment, the recognition result is obtained by inputting the website content of the website to be recognized and the enterprise name of the target enterprise into the target recognition model. The device further includes a training module, and the training module is configured to:
[0036] Based on a training sample set, perform iterative training on the recognition model to be trained to obtain a target recognition model. Each initial training sample in the training sample set at least includes: the sample enterprise name of the sample enterprise and the sample website content of the sample website. During one iterative training process, perform the following operations:
[0037] Obtain each key sample website phrase related to the sample enterprise name in the sample website content, and extract the sample local content corresponding to each key sample website phrase;
[0038] Based on the first sample importance degree of each key sample website phrase in the corresponding sample local content and the second sample importance degree of each obtained sample local content in the sample website content, determine the sample semantic weight of each sample local content, and based on the position information of each sample local content, determine the sample position weight of each sample local content;
[0039] Based on the sample feature vector, sample semantic weight, and sample position weight of each sample local content, determine the sample target feature vector corresponding to the sample website, and based on the sample target feature vector, determine the sample recognition result of the sample website, and perform parameter adjustment based on the loss value corresponding to the sample recognition result.
[0040] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in any one of the first aspects above are implemented.
[0041] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the first aspects above are implemented.
[0042] In a fifth aspect, the present application provides a computer program product, which includes: computer program code. When the computer program code runs on a computer, the computer is made to execute the method described in any one of the first aspects.
[0043] In the embodiments of the present application, various key website phrases related to the enterprise name of the target enterprise are obtained from the website content of the website to be recognized, and the local content corresponding to each key website phrase is extracted. The local content is the context statement of the corresponding key website phrase in the website content. Then, based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content, the semantic weight of each local content is determined, and based on the position information of each local content, the position weight of each local content is determined. Finally, based on the feature vector, semantic weight, and position weight of each local content, the target feature vector corresponding to the website to be recognized is determined, and based on the target feature vector, the recognition result of the website to be recognized is determined. In this way, by combining the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content to obtain the semantic weight of each local content, the misjudgment problem caused by simple association is effectively alleviated, and the accuracy of website recognition is improved. In addition, the position weight of the local content in the website content is also considered, and the position weight is combined with the semantic information, which can not only provide more detailed website recognition at the semantic level, but also improve the comprehensiveness of the recognition result at the structural level, effectively improving the accuracy and reliability of asset recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic diagram of a possible application scenario in the embodiments of the present application;
[0045] Figure 2 It is a schematic flowchart of the training of the target recognition model in the embodiments of the present application;
[0046] Figure 3 It is a schematic diagram of the architecture of the target recognition model in the embodiments of the present application;
[0047] Figure 4 It is a schematic flowchart of the implementation of a website recognition method provided in the embodiments of the present application;
[0048] Figure 5 It is a schematic diagram of determining the semantic weight in the embodiments of the present application;
[0049] Figure 6 It is a schematic diagram of determining the recognition result in the embodiments of the present application;
[0050] Figure 7 It is a schematic diagram of the structure of a website recognition device provided in the embodiments of the present application;
[0051] Figure 8 It is a schematic diagram of the structure of an electronic device in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following will, in conjunction with the accompanying drawings in the embodiments of this application, clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application. Without conflict, the embodiments in this application and the features in the embodiments can be combined arbitrarily with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0053] In the description and claims of this application and the above accompanying drawings, the terms "first" and "second" are used to distinguish different objects rather than to describe a specific order. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. "Multiple" in this application can mean at least two, for example, it can be two, three, or more, and there is no limitation in the embodiments of this application.
[0054] The following makes an explanation of the exemplary embodiments of this application in conjunction with the accompanying drawings, including various details of the embodiments of this application to facilitate understanding. It should be considered that they are only exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described here without departing from the scope of disclosure of this application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below. It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be considered exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0055] In the technical solution of this application, the acquisition, transmission, storage, use, etc. of data all comply with the requirements of relevant national laws and regulations.
[0056] Before introducing the website recognition method provided by the embodiments of this application, for the convenience of understanding, first, the technical background of the embodiments of this application will be introduced in detail below.
[0057] With the digital development of enterprises, websites have become important platforms for carrying enterprise data and business operations. However, with the replacement of website versions and personnel changes, the management of enterprise website assets has gradually got out of control, accumulating a large number of "three-none and seven-edge" website assets. "Three-none" website assets refer to websites without filing, without operation, and without maintenance; "seven-edge" website assets refer to websites without primary and secondary, without association, without management, without monitoring, without responsibility, without planning, and without clear definition. These website assets are prone to contain potential security vulnerabilities and system defects, providing potential entrances for malicious attacks and data leakage. Therefore, comprehensively identifying all the websites of an enterprise is a crucial task.
[0058] In the prior art, keyword matching or statistical analysis techniques are usually adopted to identify whether a website is an enterprise's website asset based on the relevance between the enterprise name and the local content of the website.
[0059] However, the above methods ignore the semantic importance of local content in the overall website and are vulnerable to interference from invalid information. Even if some local content is highly relevant to the enterprise name, if these local contents are not the core themes of the website, then it is inaccurate to regard them as the enterprise's website assets. In addition, the layout structure of website content reflects the core theme of the website to a certain extent. The above methods' neglect of the content position leads to errors and biases in asset determination.
[0060] In view of this, in the embodiments of the present application, a website recognition method, device, equipment and medium are provided. Each key website phrase related to the enterprise name of the target enterprise is obtained from the website content of the website to be recognized, and the local content corresponding to each key website phrase is extracted. The local content is the context statement of the corresponding key website phrase in the website content. Then, based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content, the semantic weight of each local content is determined, and based on the position information of each local content, the position weight of each local content is determined. Finally, based on the feature vector, semantic weight and position weight of each local content, the target feature vector corresponding to the website to be recognized is determined, and based on the target feature vector, the recognition result of the website to be recognized is determined. In this way, by combining the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content, the semantic weight of each local content is obtained, effectively alleviating the misjudgment problem caused by simply making judgments based on associations, and improving the accuracy of website recognition. In addition, the position weight of the local content in the website content is also considered, and the position weight is combined with the semantic information, which can not only provide more detailed website recognition at the semantic level, but also improve the comprehensiveness of the recognition result at the structural level, avoiding over-reliance on semantic intensity and ignoring the impact of content layout on website determination, and effectively improving the accuracy and reliability of asset recognition.
[0061] The preferred embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0062] Refer to Figure 1 As shown, it is a schematic diagram of a possible application scenario in the embodiments of the present application. In this application scenario diagram, it includes a server 110 and terminal devices 120 (including terminal devices 1201, 1202... 120n).
[0063] The server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal devices 120 and the server 110 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0064] The terminal device 120 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, etc.; various software can be installed on the terminal device, such as application programs and mini-programs, etc.
[0065] It should be noted that the website recognition method in the embodiments of the present application can be deployed in a computing device, and the computing device can be a server or a terminal device, where the server can be the Figure 1 server 110 shown in Figure 1 , or can be other servers other than the Figure 1 server 110 shown in Figure 1 ; the terminal device can be the
[0066] terminal device 120 shown in
[0067] , or can be other terminal devices other than the
[0068] terminal device 120 shown in Figure 2 . That is, this method can be executed independently by the server or the terminal device, or can be jointly executed by the server and the terminal device.
[0069] Among them, each initial training sample in the training sample set includes at least: the sample enterprise name of the sample enterprise and the sample website content of the sample website. The training sample set includes positive samples and negative samples. In the positive samples, the sample enterprise and the sample website are associated, and in the negative samples, the sample enterprise and the sample website are not associated.
[0070] Optionally, in the embodiments of the present application, website crawler data is collected to construct a website database and a training sample set.
[0071] Specifically, when constructing the website database, a large number of domain names are collected using the Domain Name System (DNS) database. The http: / / and https: / / are added to the domain name prefix to construct the Uniform Resource Locator (URL) based on the domain name. A large number of IPs and suspicious web ports (such as 80, 8000, 8888, 8889, 443, 6443, etc.) are concatenated and prefixed to obtain URLs in the form of http: / / ip:port and https: / / ip:port. Then, web crawlers are used to obtain website data for these URLs and store the website data in the website database.
[0072] Specifically, when constructing the training sample set, the main domain name of the sample enterprise is obtained using the Internet Content Provider (ICP). The sample websites crawled based on the main domain name of the sample enterprise and the current sample enterprise form positive samples, and the sample websites crawled based on the main domain name of the sample enterprise and other sample enterprises form negative samples.
[0073] In the embodiments of the present application, there is no limit on the number of positive and negative samples. To ensure the effectiveness of training, the number of positive and negative samples can be in a 1:1 relationship.
[0074] During one iteration training process, the following operations are performed:
[0075] Step 20: Obtain each key sample website phrase related to the sample enterprise name in the sample website content, and extract the corresponding sample local content for each key sample website phrase.
[0076] Among them, the sample local content is the context statement of the corresponding key sample website phrase in the sample website content, and the sample website content includes the text and title of the sample website.
[0077] In the embodiments of the present application, each initial sample website phrase included in the sample website content is obtained, each key sample website phrase related to the sample enterprise name among the initial sample website phrases is determined, and the corresponding sample local content for each key sample website phrase is extracted from the website content.
[0078] Optionally, in the embodiments of the present application, a possible embodiment is provided to determine each key sample website phrase related to the sample enterprise name. Specifically, the following operations are performed:
[0079] Step 200: Obtain each initial sample website phrase included in the sample website content and each sample name phrase included in the sample enterprise name.
[0080] Step 201: Based on the correlation between the sample feature vectors of each initial sample website phrase and the sample feature vectors of each sample name phrase, determine each key sample website phrase in the initial sample website phrases whose correlation with the sample enterprise name meets the correlation condition.
[0081] In the embodiment of the present application, for each initial sample website phrase, the following operations are respectively performed: Calculate the sample correlation between the sample feature vector of an initial sample website phrase and the sample feature vectors of each sample name phrase, and use the maximum sample correlation obtained as the sample correlation between the initial sample website phrase and the sample enterprise name. Then, compare the sample correlations between each initial sample website phrase and the sample enterprise name, and use the top k initial sample website phrases with the highest sample correlations as the k key sample website phrases, where k is the total number of key sample website phrases.
[0082] Step 21: Based on the first sample importance degree of each key sample website phrase in the corresponding sample local content and the second sample importance degree of each obtained sample local content in the sample website content, determine the sample semantic weight of each sample local content, and based on the position information of each sample local content, determine the sample position weight of each sample local content.
[0083] Optionally, in the embodiment of the present application, a possible embodiment is provided to determine the semantic weight of each sample local content, and the following operations are specifically performed:
[0084] Step 210: Based on the similarity between the sample local content and the first sample mask content corresponding to each key sample website phrase and the sample enterprise name respectively, determine the first sample importance degree of each key sample website phrase in the corresponding sample local content.
[0085] Wherein, the first sample mask content is the remaining content of the corresponding sample local content after removing the corresponding key sample website phrase, and the first sample mask content corresponding to a key sample website phrase is obtained by performing a mask operation on the key sample website phrase in the sample local content corresponding to the key sample website phrase.
[0086] In the embodiment of the present application, for each key sample website phrase, the following operations are respectively performed: Based on the sample similarity between the sample local content corresponding to a key sample website phrase and the sample enterprise name, and the sample similarity between the first sample mask content corresponding to the key sample website phrase and the sample enterprise name, determine the first sample importance degree of the key sample website phrase in its corresponding sample local content.
[0087] In the embodiments of the present application, when determining the first sample importance degree of a key sample website phrase in its corresponding sample local content, the specific steps are as follows:
[0088] Step 2100: Calculate the first sample similarity between the sample feature vector of the sample local content corresponding to a key sample website phrase and the sample feature vector of the sample enterprise name, and calculate the second sample similarity between the sample feature vector of the first sample masked content corresponding to a key sample website phrase and the sample feature vector of the sample enterprise name.
[0089] In the embodiments of the present application, extract the sample feature vectors of the sample local content, the first sample masked content, and the sample enterprise name respectively, then calculate the first sample similarity between the sample feature vector of the sample local content corresponding to a key sample website phrase and the sample feature vector of the sample enterprise name, and calculate the second sample similarity between the sample feature vector of the first sample masked content corresponding to a key sample website phrase and the sample feature vector of the sample enterprise name.
[0090] Exemplarily, use the encoder BERT to encode the sample local content sq1 corresponding to a key sample website phrase, the first sample masked content sq2 corresponding to the key sample website phrase, and the sample enterprise name sq3, to obtain the sample feature vector sv1 of the sample local content corresponding to the key sample website phrase, the sample feature vector sv2 of the first sample masked content corresponding to the key sample website phrase, and the sample feature vector sv3 of the sample enterprise name. The first sample similarity between the sample feature vector of the sample local content and the sample feature vector of the sample enterprise name is ssimA = sv1 * sv3, and the second sample similarity between the sample feature vector of the first sample masked content and the sample feature vector of the sample enterprise name is ssimB = sv2 * sv3.
[0091] Step 2101: Calculate the difference between the first sample similarity and the second sample similarity to obtain the first sample importance degree of a key sample website phrase in its corresponding sample local content.
[0092] Exemplarily, the first sample importance degree sw1 of a key sample website phrase in its corresponding sample local content is sw1 = ssimA - ssimB = sv1 * sv3 - sv2 * sv3.
[0093] Step 211: Determine the second sample importance degree of each sample local content in the sample website content respectively based on the similarity between each sample local content and the sample website content, and the similarity between each sample local content and the corresponding second sample masked content.
[0094] Among them, the second sample mask content is the remaining content after removing the corresponding sample local content from the sample website content, and the second sample mask content corresponding to a sample local content is obtained by performing a mask operation on the sample local content in the sample website content.
[0095] In the embodiments of the present application, for each sample local content, the following operations are respectively performed: Based on the sample similarity between a sample local content and the sample website content, and the sample similarity between the sample local content and its corresponding second sample mask content, determine the second sample importance degree of the sample local content in the sample website content.
[0096] In the embodiments of the present application, when determining the second sample importance degree of a sample local content in the sample website content, the specific steps are as follows:
[0097] Step 2110: Calculate the third sample similarity between the sample feature vector of a sample local content and the sample feature vector of the sample website content, and calculate the fourth sample similarity between the sample feature vector of a sample local content and the sample feature vector of its corresponding second sample mask content.
[0098] In the embodiments of the present application, extract the sample feature vectors of a sample local content, the sample website content, and the second sample mask content corresponding to the sample local content respectively, and then calculate the third sample similarity between the sample feature vector of a sample local content and the sample feature vector of the sample website content, and calculate the fourth sample similarity between the sample feature vector of a sample local content and the sample feature vector of its corresponding second sample mask content.
[0099] Exemplarily, use the encoder BERT to encode a sample local content sq, the sample website content sd1, and the second sample mask content sd2 corresponding to the sample local content to obtain the sample feature vector se of the sample local content, the sample feature vector se1 of the sample website content, and the sample feature vector se2 of the second sample mask content corresponding to the sample local content. The third sample similarity between the sample feature vector of the sample local content and the sample feature vector of the sample website content is ssimC = se * sd1, and the fourth sample similarity between the sample feature vector of the sample local content and the sample feature vector of its corresponding second sample mask content is ssimD = se * sd2.
[0100] Step 2111: Calculate the difference between the third sample similarity and the fourth sample similarity to obtain the second sample importance degree of a sample local content in the sample website content.
[0101] Exemplarily, the second sample importance degree sw2 of a sample local content in the sample website content is sw2 = ssimC - ssimD = se * sd1 - se * sd2.
[0102] Step 212: Determine the sample semantic weights of each sample local content based on the first sample importance degree and the second sample importance degree of each sample local content.
[0103] In the embodiments of the present application, for each sample local content, the following operations are respectively performed: Combine the first sample importance degree and the second sample importance degree of a sample local content, and pass through the sigmoid function to obtain the sample semantic weight of this sample local content.
[0104] In the embodiments of the present application, the sample semantic weight of a sample local content can be expressed as:
[0105]
[0106] Among them, sw1 is the first sample importance degree of the sample local content, sw2 is the second sample importance degree of the sample local content, and λ is a weight parameter, which can be preset or a parameter that the model needs to learn.
[0107] In the embodiments of the present application, the position information of each sample local content includes: the length and the starting character position of the sample local content. Based on the length and the starting character position of each sample local content, and the total length of the sample website content, determine the sample position weights of each sample local content.
[0108] In the embodiments of the present application, the sample position weight of a sample local content can be expressed as:
[0109]
[0110] Among them, α and β are parameters that the model can learn, len sq is the length of this sample local content, len sall is the total length of the sample website content, and spos is the starting character position of this sample local content.
[0111] Step 22: Determine the sample target feature vector corresponding to the sample website based on the sample feature vectors, sample semantic weights, and sample position weights of each sample local content, and determine the sample recognition result of the sample website based on the sample target feature vector, and perform parameter adjustment based on the loss value corresponding to the sample recognition result.
[0112] In the embodiments of the present application, based on the sample feature vectors, sample semantic weights, and sample position weights of each sample's local content, the sample target feature vector corresponding to the sample website is determined, and the sample target feature vector is passed through a fully connected layer and a sigmoid layer to obtain a logit value, that is, the sample recognition result of the sample website. Based on the true label of the sample and the logit value, the loss value corresponding to the sample recognition result is calculated, and the parameters are adjusted based on the loss value to optimize the model.
[0113] Optionally, in the embodiments of the present application, to determine the sample target feature vector corresponding to the sample website, a possible embodiment is provided, and the following operations are specifically performed:
[0114] Step 220: Based on the sample semantic weights and sample position weights of each sample's local content, determine the sample fusion weights of each sample's local content.
[0115] In the embodiments of the present application, the sample semantic weights of each sample's local content are passed through the softmax activation function to obtain the new sample semantic weights of each sample's local content, and the position weights of each sample's local content are passed through the softmax activation function to obtain the new sample position weights of each sample's local content. The new sample semantic weights and new sample position weights of each sample's local content are fused to obtain the sample fusion weights of each sample's local content.
[0116] Among them, the sample fusion weight of each sample's local content can be expressed as:
[0117]
[0118] k is the total number of sample local contents, sposition1 is the sample position weight of a sample local content, and swight' k is the sample semantic weight of a sample local content.
[0119] Step 221: Based on the sample fusion weights of each sample's local content, perform a weighted sum on the sample feature vectors of each sample's local content to obtain the sample target feature vector.
[0120] In the embodiments of the present application, the sample target feature vector can be expressed as:
[0121]
[0122] Among them, k is the total number of sample local contents, st i is the sample fusion weight of the i-th sample local content, is the sample feature vector of the i-th sample local content.
[0123] In the embodiments of the present application, the logit value of the sample website can be expressed as:
[0124] In the embodiments of the present application, the loss value corresponding to the sample recognition result can be expressed as:
[0125]
[0126] where y is the true label of the sample website, is the logit value of the sample website.
[0127] Participate Figure 3 As shown, it is a schematic diagram of the architecture of the target recognition model in the embodiments of the present application, including a semantic analysis module and an aggregation module based on location information.
[0128] In this way, the obtained target recognition model improves the accuracy and efficiency of website recognition.
[0129] After obtaining the trained target recognition model, select a website to be recognized of the target enterprise from the website database, input the enterprise name of the target enterprise and the website content of the website to be recognized into the target recognition model, and obtain the recognition result of the website to be recognized.
[0130] Refer to Figure 4 As shown, it is a flowchart of the implementation of a website recognition method provided by the embodiments of the present application. The specific implementation process of this method is as follows:
[0131] Step 40: Obtain each key website phrase related to the enterprise name of the target enterprise in the website content of the website to be recognized, and extract the local content corresponding to each key website phrase.
[0132] Among them, the local content is the context statement of the corresponding key website phrase in the website content, and the website content of the website to be recognized includes the website text and the title.
[0133] In the embodiments of the present application, obtain each initial website phrase included in the website content of the website to be recognized, determine each key website phrase related to the enterprise name of the target enterprise among the initial website phrases, and extract the local content corresponding to each key website phrase from the website content.
[0134] Optionally, in the embodiments of the present application, a possible embodiment is provided for determining each key website phrase related to the enterprise name of the target enterprise, and the following operations are specifically performed:
[0135] Step 400: Obtain each initial website phrase included in the website content of the website to be recognized and each name phrase included in the enterprise name of the target enterprise.
[0136] Step 401: Based on the correlation between the feature vectors of each initial website phrase and the feature vectors of each name phrase, determine each key website phrase in the initial website phrases whose correlation with the target enterprise meets the correlation condition.
[0137] In the embodiments of the present application, for each initial website phrase, the following operations are respectively performed: Calculate the correlation between the feature vector of an initial website phrase and the feature vectors of each name phrase, take the maximum correlation obtained as the correlation between the initial website phrase and the target enterprise, and then compare the correlations between each initial website phrase and the target enterprise, and take the top k initial website phrases with the highest correlations as the k key website phrases, where k is the total number of key website phrases.
[0138] Exemplarily, assume that the name phrases obtained from the enterprise name are [c1, c2,..., c m , use word2vec to obtain the feature vector of the i-th name phrase c i as v ci , and the feature vector of the j-th website phrase o j of the website to be identified is v oj . The correlation between an initial website phrase and the target enterprise can be expressed as:
[0139]
[0140] where δ is a hyperparameter, and the correlations less than the correlation threshold are filtered out by the English. The mathematical representation of relu is:
[0141]
[0142] In this way, the initial website phrases extracted from the website content are sorted and filtered, and only the context statements containing each key website phrase are extracted from the website content, and other statements are discarded, so as to achieve the filtering of redundant text, reduce the interference of invalid content on the website, and improve the data quality.
[0143] Step 41: Based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content, determine the semantic weight of each local content, and based on the position information of each local content, determine the position weight of each local content.
[0144] Participate Figure 5 As shown in
[0145] Optionally, in the embodiments of the present application, a possible embodiment is provided for determining the semantic weight of each local content, and the following operations are specifically performed:
[0146] Step 410: Determine the first importance level of each key website phrase in the corresponding local content based on the similarity between each key website phrase and the enterprise name with respect to their respective corresponding local content and the first mask content.
[0147] Among them, the first mask content is the remaining content of the corresponding local content after removing the corresponding key website phrase. The first mask content corresponding to a key website phrase is obtained by performing a masking operation on the key website phrase in the local content corresponding to the key website phrase.
[0148] In the embodiments of the present application, for each key website phrase, the following operations are respectively performed: Determine the first importance level of the key website phrase in the corresponding local content based on the similarity between the local content corresponding to a key website phrase and the enterprise name, and the similarity between the first mask content corresponding to the key website phrase and the enterprise name.
[0149] In the embodiments of the present application, when determining the first importance level of a key website phrase in the corresponding local content, the specific steps are as follows:
[0150] Step 4100: Calculate the first similarity between the feature vector of the local content corresponding to a key website phrase and the feature vector of the enterprise name, and calculate the second similarity between the feature vector of the first mask content corresponding to a key website phrase and the feature vector of the enterprise name.
[0151] In the embodiments of the present application, extract the feature vectors of the local content, the first mask content, and the enterprise name respectively, and then calculate the first similarity between the feature vector of the local content corresponding to a key website phrase and the feature vector of the enterprise name, and calculate the second similarity between the feature vector of the first mask content corresponding to a key website phrase and the feature vector of the enterprise name.
[0152] Exemplarily, use the encoder BERT to encode the local content q1 corresponding to a key website phrase, the first mask content q2 corresponding to the key website phrase, and the enterprise name q3 to obtain the feature vector v1 of the local content corresponding to the key website phrase, the feature vector v2 of the first mask content corresponding to the key website phrase, and the feature vector v3 of the enterprise name. The first similarity between the feature vector of the local content and the feature vector of the enterprise name is simA = v1 * v3, and the second similarity between the feature vector of the first mask content and the feature vector of the enterprise name is simB = v2 * v3.
[0153] Step 4101: Calculate the difference between the first similarity and the second similarity to obtain the first importance level of a key website phrase in the corresponding local content.
[0154] Exemplarily, the first importance degree w1 of a key website phrase in its corresponding local content is w1 = simA - simB = v1 * v3 - v2 * v3.
[0155] Since the more important the key website phrase is to the context statement (local content), the greater the semantic change of the context statement after extracting this key website phrase. Therefore, the greater the first importance degree w1, the greater the semantic gap between the local content q1 corresponding to the key website phrase and the first masked content q2 corresponding to the key website phrase, and the greater the impact of the key website phrase on the context statement.
[0156] In this way, through the first masked content corresponding to each key website phrase, the first importance degree of each key website phrase in the corresponding local content can be accurately determined.
[0157] Step 411: Based on the similarity between each local content and the website content, and the similarity between each local content and the corresponding second masked content, determine the second importance degree of each local content in the website content respectively.
[0158] Wherein, the second masked content is the remaining content of the website content after removing the corresponding local content, and the second masked content corresponding to a local content is obtained by performing a mask operation on the local content in the website content.
[0159] In the embodiments of the present application, for each local content, the following operations are respectively performed: Based on the similarity between a local content and the website content, and the similarity between the local content and its corresponding second masked content, determine the second importance degree of the local content in the website content.
[0160] In the embodiments of the present application, when determining the second importance degree of a local content in the website content, the specific steps are as follows:
[0161] Step 4110: Calculate the third similarity between the feature vector of a local content and the feature vector of the website content, and calculate the fourth similarity between the feature vector of a local content and the feature vector of its corresponding second masked content.
[0162] In the embodiments of the present application, the feature vectors of a local content, the website content, and the second masked content corresponding to the local content are extracted respectively, and then the third similarity between the feature vector of a local content and the feature vector of the website content is calculated, and the fourth similarity between the feature vector of a local content and the feature vector of its corresponding second masked content is calculated.
[0163] Exemplarily, use the encoder BERT to encode a local content q, a website content d1, and a second masked content d2 corresponding to the local content, to obtain a feature vector e of the local content, a feature vector e1 of the website content, and a feature vector e2 of the second masked content corresponding to the local content. The third similarity between the feature vector of the local content and the feature vector of the website content is simC = e * d1, and the fourth similarity between the feature vector of the local content and the feature vector of its corresponding second masked content is simD = e * d2.
[0164] Step 4111: Calculate the difference between the third similarity and the fourth similarity to obtain the second importance degree of a local content in the website content.
[0165] Exemplarily, the second importance degree w2 of a local content in the website content = simC - simD = e * d1 - e * d2.
[0166] Since the more important the context statement (local content) is to the website content, the greater the semantic change of the website content after extracting this context statement. Therefore, the greater the second importance degree w2, the greater the semantic gap between the website content d1 and the second masked content d2 corresponding to the local content, and the greater the impact of the local content on the website.
[0167] In this way, through the second masked content corresponding to each local content, the second importance degree of each local content in the website content can be accurately determined.
[0168] Step 412: Based on the first importance degree and the second importance degree of each local content, determine the semantic weight of each local content.
[0169] In the embodiments of the present application, for each local content, the following operations are respectively performed: combine the first importance degree and the second importance degree of a local content, and pass through the sigmoid function to obtain the semantic weight of this local content.
[0170] In the embodiments of the present application, the semantic weight of a local content can be expressed as:
[0171]
[0172] where w1 is the first importance degree of the local content, w2 is the second importance degree of the local content, and λ is a weight parameter, which can be preset or learned by the model.
[0173] In this way, by analyzing the semantic changes between the enterprise name and the local content and the website content, calculating the importance of the key website phrases in the local content and the importance of the local content in the entire website, a comprehensive semantic weight is obtained, which improves the accuracy of content relevance determination, effectively alleviates the misjudgment problem caused by simple association-based determination, and makes website recognition more accurate and comprehensive.
[0174] In the embodiment of the present application, the position information of each local content includes: the length and the starting character position of the local content. Based on the length and the starting character position of each local content and the total length of the website content, the position weight of each local content is determined.
[0175] In the embodiment of the present application, the position weight of a local content can be expressed as:
[0176]
[0177] where α and β are parameters that can be learned by the model, len q is the length of this local content, len all is the total length of the website content, pos is the starting character position of this local content. Based on the local content, the website can be divided into len all / len q blocks, and the local content is in the pos / len q +1 block.
[0178] In this way, the position weight fully considers the length of the local content, the total length of the website content, and the position index of the previous local content, and accurately quantifies the position weight of the local content.
[0179] Step 42: Based on the feature vectors, semantic weights, and position weights of each local content, determine the target feature vector corresponding to the website to be recognized, and based on the target feature vector, determine the recognition result of the website to be recognized.
[0180] Among them, the recognition result represents whether the website to be recognized is the website of the target enterprise, that is, whether the website to be recognized is associated with the target enterprise.
[0181] In the embodiment of the present application, based on the feature vectors, semantic weights, and position weights of each local content, determine the target feature vector corresponding to the website to be recognized, and pass the target feature vector through a fully connected layer and a sigmoid layer to obtain the recognition result of the website to be recognized.
[0182] Optionally, in the embodiment of the present application, to determine the target feature vector corresponding to the website to be recognized, a possible embodiment is provided, and the following operations are specifically performed:
[0183] Step 420: Determine the fusion weights of each local content based on the semantic weight and position weight of each local content respectively.
[0184] In the embodiments of the present application, the semantic weights of each local content are passed through the softmax activation function to obtain the new semantic weights of each local content, and the position weights of each local content are passed through the softmax activation function to obtain the new position weights of each local content. The new semantic weights and new position weights of each local content are fused to obtain the fusion weights of each local content.
[0185] Among them, the fusion weight of each local content can be expressed as:
[0186]
[0187] k is the total number of local contents, position1 is the position weight of a local content, and wight' k is the semantic weight of a local content.
[0188] Step 421: Based on the fusion weights of each local content, perform weighted summation on the feature vectors of each local content to obtain the target feature vector.
[0189] In the embodiments of the present application, the target feature vector can be expressed as:
[0190]
[0191] Among them, k is the total number of local contents, t i is the fusion weight of the i-th local content, is the feature vector of the i-th local content.
[0192] Exemplarily, as shown in Figure 6 is a schematic diagram for determining the recognition result in the embodiments of the present application. The semantic weights wight of each local content are passed through the softmax activation function to obtain the new semantic weights of each local content, and the position weights position of each local content are passed through the softmax activation function to obtain the new position weights of each local content. The new semantic weights and new position weights of each local content are fused to obtain the fusion weight t of each local content. Then, based on the fusion weights of each local content, perform weighted summation on the feature vectors of each local content to obtain the target feature vector. Finally, the target feature vector passes through the fully connected layer and the sigmoid layer to obtain the logit value as: Determine the recognition result.
[0193] In this way, by fusing the position weight with the semantic weight, it is ensured that the semantic importance of local content can be comprehensively evaluated in combination with its significance in the website, avoiding over-reliance on semantic intensity and ignoring the impact of content layout on website recognition, thereby greatly improving the accuracy and reliability of the recognition results.
[0194] Furthermore, in the embodiments of the present application, after traversing the website database and identifying the website set associated with the target enterprise in the above manner, the changes of the websites in the website set are tracked, abnormal websites ("three-none and seven-edge" websites) are detected, and potential vulnerabilities are patched, reducing the risk of being attacked, enhancing the overall network security protection ability, and effectively supporting the enterprise's rapid response ability in the face of security threats.
[0195] Furthermore, in the embodiments of the present application, after traversing the website database and identifying the website set associated with the target enterprise in the above manner, non-compliance issues in the website set (such as non-compliance of ICP filing, etc.) are identified to ensure that the website operates within the scope of compliance, reducing the legal and financial risks that may be faced due to violations, and conducting regular audits and reports on its Internet assets.
[0196] Based on the same inventive concept, an embodiment of the present application also provides a website recognition device. Refer to Figure 7 As shown in the figure, it is a schematic structural diagram of a website recognition device in an embodiment of the present application, which specifically includes
[0197] An acquisition module 701, configured to acquire each key website phrase related to the enterprise name of the target enterprise in the website content of the website to be recognized, and extract the local content corresponding to each key website phrase. The local content is the context statement of the corresponding key website phrase in the website content;
[0198] A processing module 702, configured to determine the semantic weight of each local content based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each local content in the website content, and determine the position weight of each local content based on the position information of each local content;
[0199] A determination module 703, configured to determine the target feature vector corresponding to the website to be recognized based on the feature vector, semantic weight, and position weight of each local content, and determine the recognition result of the website to be recognized based on the target feature vector.
[0200] In a possible embodiment, when acquiring each key website phrase related to the target enterprise in the website content of the website to be recognized, the acquisition module 701 is further configured to:
[0201] Acquire each initial website phrase included in the website content of the website to be recognized and each name phrase included in the enterprise name of the target enterprise;
[0202] Based on the correlation between the feature vectors of each initial website phrase and the feature vectors of each name phrase, determine each key website phrase in each initial website phrase whose correlation with the target enterprise meets the correlation condition.
[0203] In a possible embodiment, when determining the semantic weight of each local content based on the first importance level of each key website phrase in the corresponding local content and the second importance level of each obtained local content in the website content, the processing module 702 is further configured to:
[0204] Based on the similarity between the local content corresponding to each key website phrase and the first masked content and the enterprise name respectively, determine the first importance level of each key website phrase in the corresponding local content, where the first masked content is the remaining content after removing the corresponding key website phrase from the corresponding local content;
[0205] Based on the similarity between each local content and the website content and the similarity between each local content and the corresponding second masked content respectively, determine the second importance level of each local content in the website content, where the second masked content is the remaining content after removing the corresponding local content from the website content;
[0206] Based on the first importance level and the second importance level of each local content, determine the semantic weight of each local content.
[0207] In a possible embodiment, when determining the first importance level of each key website phrase in the corresponding local content based on the similarity between the local content corresponding to each key website phrase and the first masked content and the enterprise name respectively, the processing module 702 is further configured to:
[0208] For each key website phrase, perform the following operations respectively:
[0209] Calculate the first similarity between the feature vector of the local content corresponding to a key website phrase and the feature vector of the enterprise name, and calculate the second similarity between the feature vector of the first masked content corresponding to a key website phrase and the feature vector of the enterprise name;
[0210] Calculate the difference between the first similarity and the second similarity to obtain the first importance level of a key website phrase in its corresponding local content.
[0211] In a possible embodiment, when determining the second importance level of each local content in the website content based on the similarity between each local content and the website content and the similarity between each local content and the corresponding second masked content respectively, the processing module 702 is further configured to:
[0212] For each local content, perform the following operations respectively:
[0213] Calculate a third similarity between the feature vector of a local content and the feature vector of the website content, and calculate a fourth similarity between the feature vector of a local content and the feature vector of its corresponding second masked content;
[0214] Calculate the difference between the third similarity and the fourth similarity to obtain a second importance degree of a local content in the website content.
[0215] In a possible embodiment, when determining the position weights of the local contents based on the position information of the local contents, the processing module 702 is further configured to:
[0216] Determine the position weights of the local contents based on the lengths and starting character positions of the local contents, and the total length of the website content.
[0217] In a possible embodiment, when determining the target feature vector corresponding to the website to be recognized based on the feature vectors, semantic weights, and position weights of the local contents, the determining module 703 is further configured to:
[0218] Determine the fusion weights of the local contents based on the semantic weights and position weights of the local contents respectively;
[0219] Perform weighted summation on the feature vectors of the local contents based on the fusion weights of the local contents to obtain the target feature vector.
[0220] In a possible embodiment, the recognition result is obtained by inputting the website content of the website to be recognized and the enterprise name of the target enterprise into the target recognition model. The device further includes a training module 704, and the training module 704 is configured to:
[0221] Iteratively train the recognition model to be trained based on the training sample set to obtain the target recognition model. Each initial training sample in the training sample set at least includes: the sample enterprise name of the sample enterprise and the sample website content of the sample website. In one iterative training process, perform the following operations:
[0222] Obtain each key sample website phrase related to the sample enterprise name in the sample website content, and extract the corresponding sample local content of each key sample website phrase;
[0223] Determine the sample semantic weights of the sample local contents based on the first sample importance degree of each key sample website phrase in the corresponding sample local content and the second sample importance degree of each obtained sample local content in the sample website content, and determine the sample position weights of the sample local contents based on the position information of the sample local contents;
[0224] Based on the sample feature vectors, sample semantic weights, and sample position weights of the local content of each sample, determine the sample target feature vector corresponding to the sample website, and based on the sample target feature vector, determine the sample recognition result of the sample website, and perform parameter tuning based on the loss value corresponding to the sample recognition result.
[0225] Based on the above embodiments, refer to Figure 8 The following is a schematic structural diagram of an electronic device in an embodiment of the present application.
[0226] An embodiment of the present application provides an electronic device, which may include a processor 810 (Central Processing Unit, CPU), a memory 820, an input device 830, and an output device 840, etc. The input device 830 may include a keyboard, a mouse, a touch screen, etc. The output device 840 may include a display device, such as a liquid crystal display (Liquid Crystal Display, LCD), a cathode ray tube (Cathode Ray Tube, CRT), etc.
[0227] The memory 820 may include a read-only memory (ROM) and a random access memory (RAM), and provide program instructions and data stored in the memory 820 to the processor 810. In an embodiment of the present application, the memory 820 may be used to store a program of any website recognition method in an embodiment of the present application.
[0228] By calling the program instructions stored in the memory 820, the processor 810 is configured to execute any website recognition method in an embodiment of the present application according to the obtained program instructions.
[0229] Based on the above embodiments, in an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the website recognition method in any of the above method embodiments.
[0230] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0231] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in a process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0232] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in a process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0233] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0234] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. A website recognition method, characterized in that, Including: Obtain each key website phrase related to the enterprise name of the target enterprise in the website content of the website to be recognized, and extract the respective local content corresponding to each key website phrase, where the local content is the context statement of the corresponding key website phrase in the website content; Based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content, determine the semantic weight of each local content, and based on the position information of each local content, determine the position weight of each local content; Based on the feature vectors, semantic weights, and position weights of each local content, determine the target feature vector corresponding to the website to be recognized, and based on the target feature vector, determine the recognition result of the website to be recognized.
2. The method according to claim 1, wherein The obtaining of each key website phrase related to the target enterprise in the website content of the website to be recognized includes: Obtain each initial website phrase included in the website content of the website to be recognized and each name phrase included in the enterprise name of the target enterprise; Based on the correlation degree between the feature vectors of each initial website phrase and the feature vectors of each name phrase, determine each key website phrase among the initial website phrases whose correlation degree with the target enterprise meets the correlation degree condition.
3. The method according to claim 1, wherein The determining of the semantic weight of each local content based on the first importance degree of each key website phrase in the corresponding local content and the second importance degree of each obtained local content in the website content includes: Respectively, based on the similarity between the local content corresponding to each key website phrase and the first mask content and the enterprise name, determine the first importance degree of each key website phrase in the corresponding local content, where the first mask content is the remaining content after removing the corresponding key website phrase from the corresponding local content; Respectively, based on the similarity between each local content and the website content and the similarity between each local content and the corresponding second mask content, determine the second importance degree of each local content in the website content, where the second mask content is the remaining content after removing the corresponding local content from the website content; Based on the first importance degree and the second importance degree of each local content, determine the semantic weight of each local content.
4. The method according to claim 3, wherein The respectively determining the first importance degree of each key website phrase in the corresponding local content based on the similarity between the local content corresponding to each key website phrase and the first mask content and the enterprise name includes: For each key website phrase, respectively perform the following operations: Calculate the first similarity between the feature vector of the local content corresponding to a key website phrase and the feature vector of the enterprise name, and calculate the second similarity between the feature vector of the first mask content corresponding to the key website phrase and the feature vector of the enterprise name; Calculate the difference between the first similarity and the second similarity to obtain the first importance degree of a key website phrase in its corresponding local content.
5. The method according to claim 3, characterized in that, Determining the second importance level of each local content in the website content respectively based on the similarity between each local content and the website content, as well as the similarity between each local content and the corresponding second mask content, includes: For each of the local contents, the following operations are performed respectively: Calculating a third similarity between the feature vector of a local content and the feature vector of the website content, and calculating a fourth similarity between the feature vector of the local content and the feature vector of its corresponding second mask content; Calculating the difference between the third similarity and the fourth similarity to obtain the second importance level of the local content in the website content.
6. The method according to claim 1, wherein Determining the position weights of the local contents based on the position information of the local contents, includes: Determining the position weights of the local contents based on the lengths and starting character positions of the local contents, and the total length of the website content.
7. The method according to claim 1, characterized in that Determining the target feature vector corresponding to the website to be recognized based on the feature vectors, semantic weights, and position weights of the local contents, includes: Respectively determining the fusion weights of the local contents based on the semantic weights and position weights of the local contents; Performing weighted summation on the feature vectors of the local contents based on the fusion weights of the local contents to obtain the target feature vector.
8. The method according to any one of claims 1 to 7, characterized in that The recognition result is obtained by inputting the website content of the website to be recognized and the enterprise name of the target enterprise into a target recognition model, where the target recognition model is trained in the following manner: Based on a training sample set, performing iterative training on a recognition model to be trained to obtain a target recognition model, where each initial training sample in the training sample set at least includes: the sample enterprise name of a sample enterprise and the sample website content of a sample website. During one iterative training process, the following operations are performed: Obtaining each key sample website phrase in the sample website content related to the sample enterprise name, and extracting the corresponding sample local content for each key sample website phrase; Determining the sample semantic weights of the sample local contents based on the first sample importance level of each key sample website phrase in the corresponding sample local content and the second sample importance level of each obtained sample local content in the sample website content, and determining the sample position weights of the sample local contents based on the position information of the sample local contents; Determining the sample target feature vector corresponding to the sample website based on the sample feature vectors, sample semantic weights, and sample position weights of the sample local contents, determining the sample recognition result of the sample website based on the sample target feature vector, and adjusting the parameters based on the loss value corresponding to the sample recognition result.
9. A website recognition device, characterized in that, Includes: An acquisition module, configured to acquire each key website phrase in the website content of the website to be recognized related to the enterprise name of the target enterprise, and extract the corresponding local content for each key website phrase, where the local content is the context statement of the corresponding key website phrase in the website content. A processing module, configured to determine the semantic weights of the respective local contents based on the first importance degrees of the respective key website phrases in the corresponding local contents and the second importance degrees of the obtained respective local contents in the website content, and determine the position weights of the respective local contents based on the position information of the respective local contents; A determination module, configured to determine a target feature vector corresponding to the website to be recognized based on the feature vectors, semantic weights and position weights of the respective local contents, and determine a recognition result of the website to be recognized based on the target feature vector.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the method described in any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 8 are implemented.