A page recognition method, related device and storage medium
By extracting information and semantically representing advertising pages, and using clustering algorithms to identify duplicate pages, the problem of low accuracy in duplicate page identification based on advertising IDs in existing technologies has been solved, achieving efficient and accurate duplicate page identification.
Patent Information
- Application Number
- CN202211131329.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-09-16
AI Technical Summary
Existing deduplication methods based on ad IDs have low accuracy in identifying duplicate ad pages and cannot effectively identify similar or identical ad pages, resulting in poor retrieval and ranking performance of the ad system.
By extracting information from the target page, textual and semantic representation information is obtained. Clustering algorithms are used to classify the page into clusters. Based on the semantic representation information, clustering processing is performed to determine whether the page is a duplicate page.
It enables convenient and accurate identification of duplicate pages, improves the accuracy of clustering results, reduces the need for manual labeling of category data, and improves the recognition efficiency of the advertising system.
Smart Images

Figure CN117009612B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and in particular to a page identification method, related equipment and a storage medium. BACKGROUND
[0002] In order to improve the exposure of advertisements, commodity advertisers often open multiple accounts and submit similar or identical advertisements. These operations cause many problems for the retrieval and sorting of the advertisement system, such as invalidation of diversity strategies, underestimation of a single advertisement behavior, and the like. In response, in addition to strengthening the education and guidance of advertisers, identifying these similar or identical advertisement pages plays an important role in the entire advertisement system, such as reuse of historical data in the model, determination of new advertisements, de-duplication in the advertisement playing stage, and user experience problems.
[0003] Currently, when performing advertisement page de-duplication, an advertisement id-based de-duplication method is usually used. However, the advertisement id-based de-duplication method is relatively simple, but if an advertiser repeatedly creates an advertisement page or multiple commodities exist under the advertisement page, the accuracy of de-duplication is low and the effect is poor.
[0004] Therefore, how to conveniently and accurately identify duplicate pages has become a problem to be solved. SUMMARY
[0005] Embodiments of the present application provide a page identification method, related equipment and a storage medium, which can conveniently and accurately identify duplicate pages.
[0006] In a first aspect, embodiments of the present application provide a page identification method, which includes:
[0007] Performing information extraction on a target page to obtain text information contained in the target page, the target page including description information of a target object.
[0008] Determining feature information of the target page based on the text information contained in the target page, the feature information including a first page category and semantic representation information.
[0009] Obtaining a plurality of clusters corresponding to the first page category, and performing first clustering processing on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category to obtain a clustering result, the plurality of clusters being obtained by performing second clustering processing on a plurality of pages corresponding to the first page category in a page library.
[0010] Determining whether the target page is a duplicate page based on the clustering result.
[0011] In a second aspect, embodiments of the present application provide a page identification device, which includes:
[0012] The acquisition module is configured to perform information extraction on the target page to obtain text information contained in the target page, the target page including description information of the target object.
[0013] The determination module is configured to determine feature information of the target page based on the text information contained in the target page, the feature information including a first page category and semantic representation information.
[0014] The acquisition module is further configured to acquire a plurality of clusters corresponding to the first page category.
[0015] The processing module is configured to perform first clustering processing on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category, to obtain a clustering result, the plurality of clusters being obtained by performing second clustering processing on a plurality of pages corresponding to the first page category in a page library.
[0016] The determination module is further configured to determine whether the target page is a duplicate page based on the clustering result.
[0017] In a third aspect, an embodiment of the present application provides a computer device, the computer device including a processor, a network interface and a storage device, the processor, the network interface and the storage device being connected to each other, wherein the network interface is controlled by the processor to receive and send data, the storage device is configured to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to execute the page identification method according to the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium storing a computer program, the computer program including program instructions, and the program instructions being executed by a processor to execute the page identification method according to the first aspect.
[0019] In a fifth aspect, an embodiment of the present application provides a computer program product including a computer program, and the computer program is executed by a computer processor to implement the page identification method according to the first aspect.
[0020] In the embodiments of the present application, the computer device can perform information extraction on the target page to obtain text information contained in the target page, the target page including description information of the target object; the computer device determines feature information of the target page based on the text information contained in the target page, the feature information including a first page category and semantic representation information of the target page; the computer device obtains a plurality of clusters corresponding to the first page category, and performs first clustering processing on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category, to obtain a clustering result, the plurality of clusters corresponding to the first page category being obtained by performing second clustering processing on a plurality of pages corresponding to the first page category in the page library; and the computer device determines whether the target page is a duplicate page based on the clustering result. As can be seen, the embodiments of the present application do not need a large amount of manually labeled category data, but automatically extract comprehensive page information to accurately determine the category and semantic representation information of the page, and then cluster in the plurality of clusters included in the page category according to the semantic representation information, which helps to improve the accuracy of the clustering result, so that the duplicate page can be conveniently and accurately identified. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0022] Figure 1a is a schematic diagram of the architecture of a page recognition system provided by the embodiments of the present application;
[0023] Figure 1b is a schematic diagram of the overall implementation framework of the page recognition provided by the embodiments of the present application;
[0024] Figure 2 is a schematic diagram of the flow of a page recognition method provided by the embodiments of the present application;
[0025] Figure 3 is a schematic diagram of the flow of a multi-modal page content extraction provided by the embodiments of the present application;
[0026] Figure 4 is a schematic diagram of the flow of another page recognition method provided by the embodiments of the present application;
[0027] Figure 5a is a schematic diagram of the structure of a classification model provided by the embodiments of the present application;
[0028] Figure 5bis a structural schematic diagram of another classification model provided by an embodiment of the present application.
[0029] Figure 5c is a flow schematic diagram of page clustering and identification provided by an embodiment of the present application.
[0030] Figure 6 is a structural schematic diagram of a page identification apparatus provided by an embodiment of the present application.
[0031] Figure 7 is a structural schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0033] Key terms and related definitions involved in the embodiments of the present application are as follows:
[0034] H5: HTML5 is a language description method for building web content. HTML5 is the next generation standard of the Internet, a language method for building and presenting Internet content, and is considered one of the core technologies of the Internet.
[0035] BERT: Bidirectional Encoder Representation from Transformers, is a pre-trained language representation model.
[0036] DSSM: Deep Structured Semantic Models, is a deep network semantic model, the core idea of which is to map query and doc into a common dimensional semantic space, and to train the implicit semantic model by maximizing the cosine similarity between the query and doc semantic vectors.
[0037] SKU: Stock Keeping Unit, the basic unit of inventory measurement, which can be in pieces, boxes, pallets, etc. SKU is a necessary method for large chain supermarket DC (distribution center) logistics management. Now it has been extended to the abbreviation of product uniform number, each product corresponds to a unique SKU number. Single product: for a product, when any of its brand, model, configuration, grade, color, packaging capacity, unit, production date, shelf life, purpose, price, origin and other attributes are different from other products, it can be called a single product.
[0038] SPU: Standard Product Unit, the smallest unit of product information aggregation, a set of reusable and easily searchable standardized information that describes the characteristics of a product. In simple terms, goods with the same attribute values and characteristics can be called an SPU. For example: brand + model is an SPU, plus white color and size 4.0, which means a SKU. SPU + color + size is a SKU, and SKU is subordinate to SPU.
[0039] OCR: (Optical Character Recognition) is the process of electronic devices (such as scanners or digital cameras) checking characters printed on paper, detecting light and dark patterns to determine their shape, and then translating the shape into computer text using character recognition methods.
[0040] Product_id, the ID of the advertising target, which is filled in by the advertiser and used to identify the advertising object, such as appid as the product id under app advertising.
[0041] Customer, i.e. advertiser, provides the target and advertising conversion data for the advertising system.
[0042] pCVR (predicted CVR): predicted conversion rate.
[0043] pCTR (predicted CTR): estimated click-through rate.
[0044] The scheme provided in the present application belongs to machine learning technology under artificial intelligence basic technology, and simultaneously relates to cloud computing and big data of cloud basic technology. The technologies involved in the page recognition scheme provided in the present application are briefly described below.
[0045] Artificial Intelligence (AI) is the theory, method, technology and application system of using digital computer or digital computer controlled machine to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision making. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies.
[0046] Machine Learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a special study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.
[0047] Cloud technology refers to a kind of hosting technology that unifies a series of resources such as hardware, software and network in a wide area network or local area network to realize data calculation, storage, processing and sharing. Cloud infrastructure technology includes cloud computing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology and other technologies based on cloud computing business model applications, which can form a resource pool, be used on demand and be flexible and convenient. Cloud computing technology will become an important support. The background service of the technical network system needs a large amount of computing and storage resources, such as video websites, picture websites and more portals. With the high development and application of the Internet industry, every item may have its own identification mark in the future, which needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and various industry data will need strong system support, which can only be realized through cloud computing.
[0048] Big data refers to a collection of data that cannot be captured, managed and processed within a certain time range by conventional software tools, and is a massive, high-growth and diversified information asset that needs new processing mode to have stronger decision-making, insight discovery and process optimization capabilities. With the advent of the cloud era, big data has attracted more and more attention, and big data needs special technology to effectively process large amounts of data over time. The technology suitable for big data includes massive parallel processing database, data mining, distributed file system, distributed database, cloud computing platform, Internet and scalable storage system.
[0049] The embodiment of the application can be applied to various scenes such as cloud technology, artificial intelligence, intelligent transportation and auxiliary driving.
[0050] Please refer to Figure 1a , which is a schematic diagram of the architecture of a page recognition system provided by the embodiment of the application. The page recognition system comprises a computer device 101 and a terminal device 102, wherein:
[0051] The computer device 101 can manage various page resources, such as storing HTML5 page resources. The page involved in the embodiment of the application can be a landing page of an object such as an advertisement, and the advertisement page launched by an advertiser can also be distributed to each terminal device 102.
[0052] The terminal device 102 can communicate with the computer device 101, receive data sent by the computer device 101, and also can show various types of data to the user, such as the advertisement page launched by the advertiser, which can be specifically a page showing the description information of a certain commodity.
[0053] The computer device 101 can be specifically an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms; the terminal device 102 can be specifically a smart phone, a tablet computer, a notebook computer, a desktop computer, a vehicle-mounted intelligent terminal, etc., and the embodiment of the application is not limited.
[0054] In some possible implementation manners, the computer device 101 can perform information extraction on the target page to obtain text information contained in the target page. The target page can be an advertisement page to be put by an advertiser or the like, which includes description information of a target object, for example, promotion information of a certain commodity, and the content can be any one or combination of text, image, video, audio, and the like. The computer device determines feature information of the target page based on the text information contained in the target page, and the feature information includes a first page category of the target page and semantic representation information. The computer device 101 obtains a plurality of clusters corresponding to the first page category, and performs first clustering processing (incremental clustering) on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category, to obtain a clustering result. The plurality of clusters corresponding to the first page category are obtained by performing second clustering processing (initial clustering) on a plurality of pages corresponding to the first page category in a page library. The computer device 101 can determine whether the target page is a duplicate page based on the clustering result. It can be seen that, compared with a duplicate detection method based on a Standard Product Unit (SPU) or a Stock Keeping Unit (SKU), the duplicate detection method based on the SPU or the SKU needs a large amount of manual labeling on commodity advertisements, so as to generate a standard and usable SPU or SKU, which is high in cost and difficult to implement. The embodiment of the present application does not need a large amount of manual labeling of category data, but automatically extracts comprehensive page information, accurately determines the category and semantic representation information of the page, and then performs clustering in a plurality of clusters included in the page category according to the semantic representation information, which helps to improve the accuracy of the clustering result, so as to conveniently and accurately identify a duplicate page.
[0055] See Figure 1b is a schematic diagram of an overall implementation framework of page recognition provided by the embodiment of the present application. The framework mainly includes three parts of landing page content grabbing, page semantic representation, and category-sensitive incremental clustering.
[0056] Specifically, after the landing page (corresponding to the target page described above) of the advertisement is captured, text extraction is performed on the landing page, the page title, meta, picture, and picture-text relationship can be parsed, and the image data in the landing page is subjected to OCR recognition processing, so that comprehensive text information in the page is obtained (111); the extracted text information is subjected to text filtering to reduce the amount of information of the text, eliminate words or symbols without actual meaning, and speed up the processing efficiency of the subsequent process. Specifically, the text filtering can be performed by using a stop sentence template and a stop word list (112); the filtered text information can be subjected to two aspects of processing. On the one hand, page classification is performed to detect the category system to which the landing page belongs, and in order to perform fine classification, multiple levels of categories can be set, for example, four levels of categories can be set, including a first-level category, a second-level category, a third-level category, and a fourth-level category. For a certain electric toothbrush product, it can be classified as digital electrical machinery (first-level category) - household appliances (second-level category) - personal care and health (third-level category) - electric toothbrush (fourth-level category), so as to obtain the page category (114); on the other hand, the semantic information of the page is represented. The intent words in the text information can be mined first, and then the intent words are taken as input to extract feature vectors by using a related model (such as BERT+DSSM), and then semantic feature vectors are obtained (115); then, according to the category to which the landing page belongs, specifically, the category corresponding to the most fine-grained level, for example, the fourth-level category - electric toothbrush, the landing page is clustered (116) to analyze the similarity between the landing page and other pages in the category, whether it can be divided into the same cluster, so as to obtain the identification result of whether the landing page is a repeated page (117). The clustering method only performs clustering calculation on the landing pages in the same group (i.e., belonging to the same page category), and does not perform clustering calculation on the landing pages under different groups. Through the above process, repeated pages can be conveniently and accurately identified.
[0057] The implementation details of the technical solutions of the embodiments of the present application are described in detail as follows:
[0058] Please refer to Figure 2 The page identification method provided by the page identification system shown in Figure 1a The page identification method provided by the page identification system shown in Figure 1a The page identification method provided by the page identification system shown in
[0059] 201, information extraction is performed on the target page to obtain text information contained in the target page, and the target page includes description information of a target object.
[0060] The target object is an object that needs to be advertised or promoted, such as a product or a service, and the target page can be a landing page of an advertisement of a certain product / service, which includes various types of description information of the target object, such as an identification, a function, an instruction for use, an efficacy, etc. of the target object. The content form of the description information can be any one of text, image, video, and audio, or a combination of multiple types of text, image, video, and audio, and the embodiments of the present application are not limited thereto.
[0061] Specifically, after obtaining the target page, the computer device can perform multi-dimensional information extraction on the target page, that is, can extract pure text information in the target page, or can extract multimedia information such as images, videos, and audios in the target page and convert them into text information, so that comprehensive text information can be extracted from the target page.
[0062] In some possible embodiments, the computer device can obtain source data of the target page from a web page database based on the identification information of the target page, perform analysis and processing on the source data to obtain first text information and multimedia feature data, determine second text information based on the multimedia feature data, and determine text information contained in the target page based on the first text information and the second text information, specifically, the first text information and the second text information can be taken as the text information contained in the target page.
[0063] In some possible embodiments, the multimedia feature data can include multimedia address data and image-text relationship data, and the specific implementation of the computer device for determining the second text information based on the multimedia feature data can be as follows:
[0064] The data extraction tool is called to extract corresponding multimedia data based on the multimedia address data, the multimedia data can include one or both of an image and a video, description information of the multimedia data is determined based on a specified field in the image-text relationship data, and the multimedia data and the description information of the multimedia data are identified to obtain the second text information, for example, the image and the video in the multimedia data are identified by OCR to obtain corresponding text information.
[0065] In some possible embodiments, the multimedia data can also include audio data embedded in the target page, and the computer device can convert the audio in the multimedia data to obtain corresponding text information.
[0066] As Figure 3As shown, taking images in multimedia data as an example, the data scraped all originate from the landing pages of advertisements. The scraped data is divided into two types: webpage HTML data and image data. The specific process of extracting information from the page can include: querying the corresponding HTML source code from the full webpage database based on the identification information of the H5 landing page (such as URL links), converting the HTML source code into binary protocol packets, and then calling a data analyzer to obtain image-text relationship data. The URL data (i.e., image address data) of the images in the landing page is obtained through the image-text relationship data. After deduplication, the corresponding image data is scraped by the scraping platform, and the text information in the image is obtained through OCR recognition. Simultaneously, the descriptive information of the multimedia data can be determined based on specified fields in the image-text relationship data, such as the alt attribute information (i.e., the specified field mentioned above) in the HTML source code. For example... <img src="smiley-2.gif"alt="Smiley face"width="42"height="42"> The alt information, "Smiley face," is the descriptive information corresponding to the image. This information, along with the text obtained from image OCR recognition, serves as the image's text representation. This allows for the comprehensive extraction of image-related text information, not only recognizing the information inherent in the image itself but also incorporating the descriptive information from the image-text relationship data. Furthermore, after obtaining the landing page's HTML source code, the embedded hyperlinks can be parsed. If the landing page contains links to other pages, deduplication and selection operations are performed to extract the links that need to be crawled again. The HTML source code of the corresponding webpage can be queried from the full webpage database. If it doesn't exist, the link is submitted to the crawling platform, which then retrieves the data from other webpages embedded in the landing page and adds it to the full webpage database. The same analysis and processing can be performed on the HTML source code data of these other webpages embedded in the landing page, and the obtained text information can be used as the corresponding text information for the landing page.
[0067] 202. Determine the feature information of the target page based on the text information contained in the target page, wherein the feature information includes a first page category and semantic representation information.
[0068] Specifically, the computer device can determine the first page category to which the target page belongs based on the text information contained in the target page, that is, determine which category the target page belongs to among the defined multiple page categories; the computer device can also determine the semantic representation information of the target page based on the text information contained in the target page, such as a 64-dimensional feature vector.
[0069] 203、obtaining a plurality of clusters corresponding to the first page category, and performing first clustering processing on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category to obtain a clustering result, the plurality of clusters being obtained by performing second clustering processing on a plurality of pages corresponding to the first page category in a page library.
[0070] Specifically, the computer device can first perform initial clustering on each page category, which can specifically include:
[0071] obtaining a plurality of pages corresponding to a second page category from the page library, the second page category being any one of a plurality of page categories, and the plurality of page categories including the first page category; wherein the computer device can classify each page according to the feature information of each page to determine the page category to which each page belongs.
[0072] performing second clustering processing (i.e., initial clustering) on a seed page in the plurality of pages corresponding to the second page category according to a preset number of clusters to obtain a plurality of clusters corresponding to the second page category. The clustering can be performed using a K-means clustering algorithm or other feasible clustering algorithm, and the seed page can be a certain number of pages selected from the plurality of pages corresponding to the second page category, i.e., the seed page is first subjected to initial clustering to obtain a plurality of clusters corresponding to each page category, and each cluster can also be assigned a cluster identifier (e.g., cluster ID).
[0073] Further, the computer can obtain a plurality of clusters corresponding to the first page category, and perform first clustering processing on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category to obtain a clustering result, i.e., on the basis of the previous initial clustering, incremental clustering (or secondary clustering) is performed on the target page to determine whether the target page can be divided into an existing cluster or a new cluster needs to be added for the target page, thereby obtaining a clustering result.
[0074] In some feasible embodiments, for each page category, the initial clustering performs the following steps:
[0075] (1) selecting K seed advertisement pages (referred to as seed pages) as cluster centers for initial clustering (e.g., K = 100);
[0076] (2) calculating the distance (e.g., Euclidean distance or cosine distance) from the semantic representation information (64-dimensional feature vector) of each seed page to the K cluster centers, respectively, finding the cluster center closest to the seed page, and attributing it to the corresponding cluster;
[0077] (3) after all seed pages are attributed to clusters, K clusters are formed. Then, the center of gravity (e.g., average distance center) of each cluster is recalculated and defined as a new "cluster center".
[0078] (4) iteratively performing steps (2)-(3) until each seed page is divided into a corresponding cluster.
[0079] In some possible implementations, the seed pages are a certain number of advertisement pages manually selected or randomly selected by a machine, and for each page category, a part of pages can be selected from the corresponding multiple pages as the seed pages of the page category.
[0080] In some possible implementations, the target page can be an advertisement page newly launched by an advertiser, or can be a page other than the seed pages, that is, the target page refers to a page that has not been divided into a cluster and needs to be divided into an existing cluster or a new cluster through clustering, so as to determine whether it is a duplicate page.
[0081] In some possible implementations, before identifying whether the target page is a duplicate page through clustering, the computer device can first analyze whether the target page is a duplicate page according to address information (such as a URL link address) of the page, for example, the computer can obtain the address information of the target page and the address information of each page of the multiple pages corresponding to the first page category in the page library; by comparing the address information of the target page with the address information of the multiple pages corresponding to the first page category, it is determined whether there is a page in the multiple pages corresponding to the first page category whose address information matches the address information of the target page; if there is, it can be directly determined that the target page is a duplicate page, for example, the URL of the target page is the same as the URL of a page in the multiple pages corresponding to the first page category, then it can be determined that the target page is a duplicate page; if not, then the related steps of identifying whether the target page is a duplicate page through clustering need to be performed.
[0082] 204、based on the clustering result, determining whether the target page is a duplicate page.
[0083] Specifically, the computer device can determine whether the target page can be divided into the multiple clusters corresponding to the first page category according to the clustering result, if the target page can be divided into any one of the existing clusters, the target page is a duplicate page, if the target page cannot be divided into the existing clusters, the target page is not a duplicate page, for example, the multiple clusters corresponding to the first page category are cluster 1, cluster 2, cluster 3, and cluster 4, if the target page can be divided into any one of cluster 1, cluster 2, cluster 3, and cluster 4, for example, cluster 3, through clustering, the target page is a duplicate page; if the target page cannot be divided into the multiple clusters corresponding to the first page category, the target page is not a duplicate page.
[0084] In some possible implementation manners, for the identified duplicate pages (i.e., similar or identical advertisement pages), the computer device can add a mark to the duplicate advertisement pages (e.g., target pages), and for the advertisement pages added with the mark, the advertisements can be filtered or de-weighted in advertisement retrieval and diversity sorting.
[0085] In the embodiment of the present application, the computer device can perform information extraction on the target page to obtain text information contained in the target page, the target page including description information of a target object; the computer device determines feature information of the target page based on the text information contained in the target page, the feature information including a first page category of the target page and semantic representation information; the computer device obtains a plurality of clusters corresponding to the first page category, and performs first clustering processing on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category, to obtain a clustering result, the plurality of clusters corresponding to the first page category being obtained by performing second clustering processing on a plurality of pages corresponding to the first page category in a page library; and the computer device determines whether the target page is a duplicate page based on the clustering result. As can be seen, the embodiment of the present application does not need a large amount of manually labeled category data, but automatically extracts comprehensive page information to accurately determine the category and semantic representation information of the page, and then clusters in a plurality of clusters included in the page category according to the semantic representation information, which helps to improve the accuracy of the clustering result, so that the duplicate page can be conveniently and accurately identified.
[0086] Please refer to Figure 4 is another page identification method provided by the page identification system shown in Figure 1a The page identification method of the embodiment of the present application is mainly described from the side of the computer device shown in Figure 1a The page identification method includes the following steps:
[0087] 401. Perform information extraction on a target page to obtain text information contained in the target page, the target page including description information of a target object.
[0088] The specific implementation of step 401 can refer to the related description of step 201 in the foregoing embodiments, which will not be described here again.
[0089] 402. Call a first classification model to process the text information contained in the target page to obtain a first page category of the target page, the first classification model being obtained by training an initial neural network based on sample pages and page category labels.
[0090] However, considering that different categories of goods, such as shoes, bags, mobile phones, and necklaces, do not require deduplication, a category (also known as class) classification is needed to improve the accuracy of deduplication. Only goods belonging to the same category will be clustered for deduplication.
[0091] Specifically, the structure and processing flow of the first classification model can be as follows: Figure 5a As shown, the first classification model can employ neural network structures such as FastText, RNN, LSTM, CNN, and Transformer. The computer extracts and filters the text information of the target page, then inputs it into the input layer of the first classification model. The embedding layer extracts features from the text information, converting it into vector form. This vector is then processed through hidden and representation layers, and finally, a normalization layer outputs the classification result. For multi-level classification systems, such as a 4-level category system, the model can specifically output: the estimated probability of the 4-level category system, with the category with the highest score being taken as the category of the target page (i.e., the first page category).
[0092] In some feasible implementations, the computer device can perform supervised training on the initial neural network based on sample pages and page category labels to obtain a first classification model. The sample construction method can be to manually annotate the web pages with multi-level categories, and the amount of labeled data can be flexibly set.
[0093] 403. Extract keywords from the text information contained in the target page based on the target vocabulary to obtain the keyword set corresponding to the target page.
[0094] Specifically, computer devices can perform intent word recognition on text information on a target page based on a target vocabulary. The target vocabulary may include specific brand information and other keywords, thereby extracting keywords and obtaining a set of keywords corresponding to the target page.
[0095] 404. Call the semantic representation module in the second classification model to process the keywords included in the keyword set to obtain the semantic representation information of the target page. The second classification model includes the semantic representation module and the semantic discrimination module. The second classification model is trained on the semantic representation module and the semantic discrimination module based on sample page pairs and annotation information. The annotation information is used to indicate whether the two sample pages included in the sample page pair are duplicate pages.
[0096] The second classification model is mainly used to extract semantic information from the page and generate semantic representation information, such as a 64-dimensional feature vector. Specifically, the semantic representation module in the second classification model can be called to process the keywords included in the keyword set to obtain the semantic representation information of the target page.
[0097] Specifically, the structure and processing flow of the second classification model can be as shown in the following table. Figure 5b As shown, the second classification model includes a semantic representation module and a semantic discrimination module. The semantic representation module can employ a BERT neural network, and the semantic discrimination module can employ a DSSM neural network. When training the second classification model, the computer device can input the labeled sample page pair (e.g., including page 1 and page 2) into a semantic representation module respectively, output respective 64-dimensional feature vectors, input the two 64-dimensional feature vectors into a semantic discrimination module, and output a discrimination result of whether the two pages are similar, i.e., a binary classification result, i.e., the two pages are similar or not. Based on the prediction result output by the semantic discrimination module and the actual labeled information of whether the two sample pages are duplicate pages, the network parameters of the semantic representation module and the semantic discrimination module are adjusted until the prediction result and the actual labeled information satisfy a convergence condition, and the training of the second classification model is completed.
[0098] Further, after the model training is completed, when the semantic information of a page needs to be extracted, the computer device can directly use the output of the semantic representation module in the second classification model as the semantic representation information of the target page, i.e., a 64-dimensional feature vector, so as to accurately determine the semantic representation information of the page.
[0099] In some possible implementation manners, the sample page pair can be manually labeled or screened based on an account structure rule. The account structure rule refers to that if two advertisements belong to the same promotion unit (e.g., the same advertiser) and the landing page intent words are completely identical, it is considered that the advertisement landing page content is consistent.
[0100] In some possible implementation manners, the sample pages and the sample page pairs involved in the above first classification model and second classification model can be regarded as seed pages of various page categories.
[0101] 405、Obtain the plurality of clusters corresponding to the first page category, and include the pages included in the plurality of clusters corresponding to the first page category as candidate pages, to obtain a plurality of candidate pages.
[0102] 406、Obtain the semantic representation information of each candidate page in the plurality of candidate pages.
[0103] Specifically, when performing incremental clustering, the computer device can first obtain the pages belonging to the first page category that have been divided into clusters through initial clustering (including the aforementioned seed pages), and take these pages as candidate pages that need to perform distance calculation with the target page, obtain the semantic representation information of each candidate page, such as a 64-dimensional feature vector, and the specific generation manner of the semantic representation information of each candidate page can refer to the related description in step 404.
[0104] 407. Based on the semantic representation information of the target page and the semantic representation information of each candidate page, performing first clustering processing on the target page to obtain a target cluster to which the target page belongs, and taking the target cluster as the clustering result of the target page.
[0105] Specifically, when performing secondary clustering (or incremental clustering), the semantic representation information of the target page and the semantic representation information of each candidate page in the plurality of clusters corresponding to the first page category are used to perform first clustering processing on the target page to obtain a target cluster to which the target page belongs, that is, to determine which cluster the target page can be divided into or to add a cluster for the target page. Incremental clustering refers to a process of continuously aggregating pages not divided into clusters into clusters based on the result of initial clustering for each page category.
[0106] In some possible implementation manners, the computer device can obtain the distance (such as Euclidean distance or cosine distance) between the semantic representation information of the target page and the semantic representation information of each candidate page, determine a first candidate page corresponding to the minimum distance from the plurality of candidate pages, if the distance corresponding to the first candidate page is less than or equal to a preset distance threshold, add the target page to the cluster to which the first candidate page belongs, and take the cluster to which the first candidate page belongs as the target cluster to which the target page belongs, and if the distance corresponding to the first candidate page is greater than the preset distance threshold, generate a new cluster, add the target page to the new cluster, and take the new cluster as the target cluster to which the target page belongs.
[0107] In some possible implementation manners, different clusters can be distinguished by using cluster IDs, that is, each cluster (including a new cluster) can be assigned a mutually different ID. By taking the cluster ID as a diversity and freshness judgment ID, the precision of the judgment is improved, and by generating the semantic expression of the landing page of the advertisement, the estimation accuracy of various models (such as pCVR / pCTR) is improved.
[0108] 408、If the clustering result indicates that the target cluster to which the target page belongs is one of the multiple clusters corresponding to the first page category, it is determined that the target page is a duplicate page; if the clustering result indicates that the target cluster to which the target page belongs is a new cluster, it is determined that the target page is not a duplicate page.
[0109] Specifically, the computer device can determine, according to the clustering result, whether the target cluster to which the target page is divided is one of the multiple clusters corresponding to the first page category or a new cluster. If the target cluster to which the target page belongs is one of the multiple clusters corresponding to the first page category, it is determined that the target page is a duplicate page; if the target cluster to which the target page belongs is a new cluster, it is determined that the target page is not a duplicate page.
[0110] In some possible implementations, as shown in FIG. 7, which is a flow diagram of a page clustering and identification method provided by an embodiment of the present application, the method can include the following steps. Figure 5c
[0111] 501、Determine whether the URL address of the landing page meets the business rule. If yes, execute 508 to determine that it is a duplicate page.
[0112] The business rule refers to that the address information (such as URL link address) of the page is the same. If the address information of the landing page (such as the target page described above) is the same as the address information of the existing page in the page library, it is determined that the URL address of the landing page meets the business rule.
[0113] 502、If not, obtain a candidate page set in which multiple pages corresponding to the page category to which the landing page belongs have been allocated cluster IDs, and the candidate page set includes multiple candidate pages.
[0114] 503、Recall the candidate page in the candidate page set that is closest to the landing page.
[0115] 504、Determine whether the distance is less than a threshold value.
[0116] 505、If not, the landing page is independently clustered. 506, It is determined to be a new page (i.e., not a duplicate page).
[0117] 507、If yes, the landing page is divided into the cluster in which the closest candidate page is located. 508, It is determined to be a duplicate page.
[0118] In the embodiments of the present application, the computer device can perform information extraction on a target page to obtain text information contained in the target page, the target page including description information of a target object; the computer device determines a first page category to which the target page belongs by calling a first classification model; after performing keyword extraction on the text information contained in the target page to obtain a keyword set corresponding to the target page, the computer device processes keywords included in the keyword set by calling a semantic representation module in a second classification model to obtain semantic representation information of the target page, and then compares the semantic representation information of the target page with semantic representation information of each page that has been divided into a cluster to determine whether the target page can be divided into a plurality of clusters corresponding to the first page category to which the target page belongs, and obtain a target cluster to which the target page belongs. If the target cluster to which the target page belongs is one of the plurality of clusters corresponding to the first page category, it is determined that the target page is a duplicate page; if the target cluster to which the target page belongs is a new cluster, it is determined that the target page is not a duplicate page. By automatically extracting comprehensive page information, accurately determining the category and semantic representation information of the page based on a plurality of classification models, and then clustering in a plurality of clusters included in the page category according to the semantic representation information, the accuracy of the clustering result can be improved, so that the duplicate page can be conveniently and accurately identified.
[0119] It can be understood that in the specific embodiments of the present application, the information related to the page and other related data are involved. When the above embodiments of the present application are applied to specific products or technologies, the user's permission or consent is required, and the collection, use and processing of the related data need to comply with the relevant laws, regulations and standards of the country and region.
[0120] Please refer to Figure 6 is a structural schematic diagram of a page recognition device according to an embodiment of the present application. The device comprises:
[0121] The acquisition module 601 is configured to perform information extraction on a target page to obtain text information contained in the target page, the target page including description information of a target object.
[0122] The determination module 602 is configured to determine feature information of the target page based on the text information contained in the target page, the feature information including a first page category and semantic representation information.
[0123] The acquisition module 601 is further configured to acquire a plurality of clusters corresponding to the first page category.
[0124] The processing module 603 is configured to perform first clustering processing on the target page based on the semantic representation information of the target page and a plurality of clusters corresponding to the first page category, to obtain a clustering result, wherein the plurality of clusters are obtained by performing second clustering processing on a plurality of pages corresponding to the first page category in a page library.
[0125] The determination module 602 is further configured to determine whether the target page is a duplicate page based on the clustering result.
[0126] Optionally, the acquisition module 601 is specifically configured to:
[0127] acquire source data of the target page from a page database based on identification information of the target page.
[0128] perform parsing processing on the source data to obtain first text information and multimedia feature data.
[0129] determine second text information based on the multimedia feature data.
[0130] determine text information contained in the target page based on the first text information and the second text information.
[0131] Optionally, the multimedia feature data includes multimedia address data and image-text relationship data, and the acquisition module 601 is specifically configured to:
[0132] invoke a data crawling tool to crawl corresponding multimedia data based on the multimedia address data, wherein the multimedia data includes one or both of pictures and videos.
[0133] determine description information of the multimedia data based on a specified field in the image-text relationship data.
[0134] perform identification processing on the multimedia data and the description information of the multimedia data to obtain the second text information.
[0135] Optionally, the determination module 602 is specifically configured to:
[0136] invoke a first classification model to process the text information contained in the target page to obtain a first page category of the target page, wherein the first classification model is obtained by training an initial neural network based on sample pages and page category labels.
[0137] perform keyword extraction on the text information contained in the target page based on a target word table to obtain a keyword set corresponding to the target page.
[0138] The semantic representation module in the second classification model is called to process the keywords included in the keyword set, to obtain semantic representation information of the target page. The second classification model includes the semantic representation module and a semantic discrimination module, and is obtained by training the semantic representation module and the semantic discrimination module based on a sample page pair and label information. The label information is used to indicate whether the two sample pages included in the sample page pair are duplicate pages.
[0139] Optionally, the processing module 603 is specifically configured to:
[0140] Pages included in the plurality of clusters corresponding to the first page category are taken as candidate pages, to obtain a plurality of candidate pages.
[0141] Semantic representation information of each candidate page in the plurality of candidate pages is obtained.
[0142] The target page is subjected to first clustering processing based on the semantic representation information of the target page and the semantic representation information of each candidate page, to obtain a target cluster to which the target page belongs.
[0143] The target cluster is taken as a clustering result of the target page.
[0144] Optionally, the processing module 603 is specifically configured to:
[0145] A distance between the semantic representation information of the target page and the semantic representation information of each candidate page is obtained.
[0146] A first candidate page corresponding to a minimum distance is determined from the plurality of candidate pages.
[0147] If the distance corresponding to the first candidate page is less than or equal to a preset distance threshold, the target page is added to a cluster to which the first candidate page belongs, and the cluster to which the first candidate page belongs is taken as a target cluster to which the target page belongs.
[0148] If the distance corresponding to the first candidate page is greater than the preset distance threshold, a new cluster is generated, the target page is added to the new cluster, and the new cluster is taken as a target cluster to which the target page belongs.
[0149] Optionally, the determination module 602 is specifically configured to:
[0150] If the clustering result indicates that a target cluster to which the target page belongs is one of the plurality of clusters corresponding to the first page category, the target page is determined to be a duplicate page.
[0151] If the clustering result indicates that the target cluster to which the target page belongs is a newly added cluster, it is determined that the target page is not a duplicate page.
[0152] Optionally, the obtaining module 601 is further configured to obtain a plurality of pages corresponding to a second page category from the page library, the second page category being any one of a plurality of page categories, and the plurality of page categories including the first page category.
[0153] The processing module 603 is further configured to perform second clustering processing on the seed page in the plurality of pages corresponding to the second page category according to a preset clustering quantity, to obtain a plurality of clusters corresponding to the second page category.
[0154] Optionally, the obtaining module 601 is further configured to obtain address information of the target page and address information of the plurality of pages corresponding to the first page category in the page library.
[0155] The processing module 603 is further configured to determine, by comparing the address information of the target page with the address information of the plurality of pages corresponding to the first page category, whether there is a page in the plurality of pages corresponding to the first page category whose address information matches the address information of the target page.
[0156] The processing module 603 is further configured to, if there is, determine that the target page is a duplicate page.
[0157] The processing module 603 is further configured to, if there is not, trigger the obtaining module 601 to obtain the plurality of clusters corresponding to the first page category.
[0158] It should be noted that the functions of each functional module of the page identification apparatus of the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the related description of the above method embodiments, which will not be described here.
[0159] Please refer to Figure 7 is a structural schematic diagram of a computer device according to an embodiment of the present application. The computer device according to the embodiment of the present application includes a power supply module and the like, and includes a processor 701, a storage device 702, and a network interface 703. The processor 701, the storage device 702, and the network interface 703 can interact with each other.
[0160] The storage device 702 can include volatile memory, such as random-access memory (RAM); the storage device 702 can also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; the storage device 702 can also include a combination of the above-mentioned kinds of memory.
[0161] The processor 701 can be a central processing unit (CPU). In an embodiment, the processor 701 can also be a graphics processing unit (GPU). The processor 701 can also be a combination of CPU and GPU. In an embodiment, the storage device 702 is configured to store program instructions, and the processor 701 can invoke the program instructions to perform the following operations:
[0162] Performing information extraction on a target page to obtain text information contained in the target page, the target page including description information of a target object.
[0163] Determining feature information of the target page based on the text information contained in the target page, the feature information including a first page category and semantic representation information.
[0164] Obtaining a plurality of clusters corresponding to the first page category.
[0165] Performing first clustering processing on the target page based on the semantic representation information of the target page and the plurality of clusters corresponding to the first page category to obtain a clustering result, the plurality of clusters being obtained by performing second clustering processing on a plurality of pages corresponding to the first page category in a page library.
[0166] Determining whether the target page is a duplicate page based on the clustering result.
[0167] Optionally, the processor 701 is specifically configured to:
[0168] Obtaining source data of the target page from a page database based on identification information of the target page.
[0169] Performing parsing processing on the source data to obtain first text information and multimedia feature data.
[0170] Determining second text information based on the multimedia feature data.
[0171] determine text information contained in the target page based on the first text information and the second text information.
[0172] Optionally, the multimedia feature data includes multimedia address data and image-text relationship data, and the processor 701 is specifically configured to:
[0173] invoke a data crawling tool to crawl corresponding multimedia data based on the multimedia address data, the multimedia data including one or both of a picture and a video.
[0174] determine description information of the multimedia data based on a specified field in the image-text relationship data.
[0175] perform identification processing on the multimedia data and the description information of the multimedia data to obtain second text information.
[0176] Optionally, the processor 701 is specifically configured to:
[0177] invoke a first classification model to process the text information contained in the target page to obtain a first page category of the target page, the first classification model being obtained by training an initial neural network based on sample pages and page category labels.
[0178] perform keyword extraction on the text information contained in the target page based on a target vocabulary to obtain a keyword set corresponding to the target page.
[0179] invoke a semantic representation module in a second classification model to process keywords included in the keyword set to obtain semantic representation information of the target page, the second classification model including the semantic representation module and a semantic discrimination module, the second classification model being obtained by training the semantic representation module and the semantic discrimination module based on sample page pairs and annotation information, the annotation information being used to indicate whether two sample pages included in the sample page pairs are duplicate pages.
[0180] Optionally, the processor 701 is specifically configured to:
[0181] include pages included in a plurality of clusters corresponding to the first page category as candidate pages to obtain a plurality of candidate pages.
[0182] obtain semantic representation information of each candidate page in the plurality of candidate pages.
[0183] perform first clustering processing on the target page based on the semantic representation information of the target page and the semantic representation information of each candidate page to obtain a target cluster to which the target page belongs.
[0184] The target cluster is taken as a clustering result of the target page.
[0185] Optionally, the processor 701 is specifically configured to:
[0186] Obtain a distance between the semantic representation information of the target page and the semantic representation information of each candidate page.
[0187] Determine a first candidate page corresponding to the minimum distance from the plurality of candidate pages.
[0188] If the distance corresponding to the first candidate page is less than or equal to a preset distance threshold, the target page is added to a cluster to which the first candidate page belongs, and the cluster to which the first candidate page belongs is taken as a target cluster to which the target page belongs.
[0189] If the distance corresponding to the first candidate page is greater than the preset distance threshold, a new cluster is generated, the target page is added to the new cluster, and the new cluster is taken as the target cluster to which the target page belongs.
[0190] Optionally, the processor 701 is specifically configured to:
[0191] If the clustering result indicates that the target cluster to which the target page belongs is one of the plurality of clusters corresponding to the first page category, it is determined that the target page is a duplicate page.
[0192] If the clustering result indicates that the target cluster to which the target page belongs is a new cluster, it is determined that the target page is not a duplicate page.
[0193] Optionally, the processor 701 is further configured to obtain a plurality of pages corresponding to a second page category from the page library, the second page category being any one of a plurality of page categories, and the plurality of page categories including the first page category.
[0194] The processor 701 is further configured to perform second clustering processing on a seed page in the plurality of pages corresponding to the second page category according to a preset clustering number, to obtain a plurality of clusters corresponding to the second page category.
[0195] Optionally, the processor 701 is further configured to obtain address information of the target page and address information of the plurality of pages corresponding to the first page category in the page library.
[0196] The processor 701 is further configured to determine whether there is a page in the plurality of pages corresponding to the first page category whose address information matches the address information of the target page by comparing the address information of the target page with the address information of the plurality of pages corresponding to the first page category.
[0197] The processor 701 is further configured to determine, if present, that the target page is a duplicate page.
[0198] The processor 701 is further configured to, if not, obtain multiple clusters corresponding to the first page category.
[0199] In specific implementations, the processor 701, storage device 702, and network interface 703 described in the embodiments of this application can execute the embodiments of this application. Figure 2 , Figure 4 The implementation methods described in the relevant embodiments of the provided method can also be used to execute the embodiments of this application. Figure 6 The implementation methods described in the relevant embodiments of the provided device will not be repeated here.
[0200] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The technical solutions of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which can be a personal computer, server, or network device, specifically a processor in the computer device) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium may include: a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), and other media capable of storing program code.
[0201] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A page recognition method, characterized in that, The method includes: Information is extracted from the target page to obtain the text information contained in the target page, wherein the target page includes descriptive information of the target object; The feature information of the target page is determined based on the text information contained in the target page, and the feature information includes a first page category and semantic representation information; Multiple clusters corresponding to the first page category are obtained, and the target page is subjected to a first clustering process based on the semantic representation information of the target page and the multiple clusters corresponding to the first page category to obtain a clustering result. The multiple clusters are obtained by performing a second clustering process on multiple pages corresponding to the first page category in the page library. Based on the clustering results, determine whether the target page is a duplicate page; The step of extracting information from the target page to obtain the text information contained in the target page includes: The source data of the target page is obtained from the page database based on the target page's identifier information; The source data is parsed to obtain first text information and multimedia feature data, wherein the multimedia feature data includes multimedia address data and image-text relationship data. The data scraping tool is invoked to scrape the corresponding multimedia data based on the multimedia address data. The multimedia data includes one or both of images and videos. Determine the descriptive information of the multimedia data based on specified fields in the text-image relationship data; The multimedia data and its descriptive information are identified and processed to obtain second text information; The text information contained in the target page is determined based on the first text information and the second text information.
2. The method according to claim 1, characterized in that, The step of determining the feature information of the target page based on the text information contained in the target page includes: The first classification model is called to process the text information contained in the target page to obtain the first page category of the target page. The first classification model is obtained by training an initial neural network based on sample pages and page category labels. Based on the target vocabulary, keywords are extracted from the text information contained in the target page to obtain the keyword set corresponding to the target page; The semantic representation module in the second classification model is invoked to process the keywords included in the keyword set to obtain the semantic representation information of the target page. The second classification model includes the semantic representation module and the semantic discrimination module. The second classification model is trained on the semantic representation module and the semantic discrimination module based on sample page pairs and annotation information. The annotation information is used to indicate whether the two sample pages included in the sample page pair are duplicate pages.
3. The method according to claim 1, characterized in that, The first clustering process, which involves performing a first clustering operation on the target page based on its semantic representation information and multiple clusters corresponding to the first page category, yields the clustering results, including: Multiple candidate pages are obtained by taking the pages included in the multiple clusters corresponding to the first page category as candidate pages; Obtain the semantic representation information of each candidate page among the multiple candidate pages; Based on the semantic representation information of the target page and the semantic representation information of each candidate page, the target page is subjected to a first clustering process to obtain the target cluster to which the target page belongs; The target cluster is used as the clustering result of the target page.
4. The method according to claim 3, characterized in that, The first clustering process, based on the semantic representation information of the target page and the semantic representation information of each candidate page, is performed on the target page to obtain the target cluster to which the target page belongs, including: Obtain the distance between the semantic representation information of the target page and the semantic representation information of each candidate page; The first candidate page with the smallest distance is determined from the plurality of candidate pages; If the distance to the first candidate page is less than or equal to a preset distance threshold, the target page is added to the cluster to which the first candidate page belongs, and the cluster to which the first candidate page belongs is taken as the target cluster to which the target page belongs. If the distance to the first candidate page is greater than the preset distance threshold, a new cluster is generated, the target page is added to the new cluster, and the new cluster is used as the target cluster to which the target page belongs.
5. The method according to claim 1, 3, or 4, characterized in that, Determining whether the target page is a duplicate page based on the clustering results includes: If the clustering result indicates that the target page belongs to one of the multiple clusters corresponding to the first page category, then the target page is determined to be a duplicate page; If the clustering result indicates that the target page belongs to a newly added cluster, then the target page is determined not to be a duplicate page.
6. The method according to claim 1, 3, or 4, characterized in that, The method further includes: Obtain multiple pages corresponding to the second page category from the page library, wherein the second page category is any one of the multiple page categories, and the multiple page categories include the first page category; The seed pages among the multiple pages corresponding to the second page category are subjected to a second clustering process according to a preset number of clusters to obtain multiple clusters corresponding to the second page category.
7. The method according to claim 1, characterized in that, Before obtaining the multiple clusters corresponding to the first page category, the method further includes: Obtain the address information of the target page and the address information of multiple pages corresponding to the first page category in the page library; By comparing the address information of the target page with the address information of multiple pages corresponding to the first page category, it is determined whether there is a page among the multiple pages corresponding to the first page category whose address information matches the address information of the target page; If it exists, then the target page is determined to be a duplicate page; If it does not exist, then proceed with the step of obtaining multiple clusters corresponding to the first page category.
8. A page recognition device, characterized in that, The device includes: The acquisition module is used to extract information from the target page to obtain the text information contained in the target page, wherein the target page includes description information of the target object; The determining module is used to determine the feature information of the target page based on the text information contained in the target page, wherein the feature information includes a first page category and semantic representation information; The acquisition module is further configured to acquire multiple clusters corresponding to the first page category; The processing module is used to perform a first clustering process on the target page based on the semantic representation information of the target page and multiple clusters corresponding to the first page category to obtain a clustering result. The multiple clusters are obtained by performing a second clustering process on multiple pages corresponding to the first page category in the page library. The determining module is further configured to determine whether the target page is a duplicate page based on the clustering results; The acquisition module is specifically used for: The source data of the target page is obtained from the page database based on the target page's identifier information; The source data is parsed to obtain first text information and multimedia feature data, wherein the multimedia feature data includes multimedia address data and image-text relationship data. The data scraping tool is invoked to scrape the corresponding multimedia data based on the multimedia address data. The multimedia data includes one or both of images and videos. Determine the descriptive information of the multimedia data based on specified fields in the text-image relationship data; The multimedia data and its descriptive information are identified and processed to obtain second text information; The text information contained in the target page is determined based on the first text information and the second text information.
9. A computer device, characterized in that, The computer device includes a processor, a network interface, and a storage device, which are interconnected. The network interface is controlled by the processor to send and receive data. The storage device is used to store a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute the page recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions, which are executed by a processor to perform the page recognition method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a computer processor, it implements the page recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Commodity image semantic annotation method based on domain ontology
CN108537240A
Method and device for detecting repeated page content
CN108764352A