Webpage information acquisition method and device, equipment, medium and program product
By calculating the similarity between the web topic words and the emotional lexicon, filtering out the target web topic words, the problem of insufficient user demand positioning in the existing technology is solved, and the efficiency and accuracy of web information collection is improved.
Patent Information
- Application Number
- CN202510656049.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-26
AI Technical Summary
The existing methods of collecting web page information cannot accurately locate user needs, resulting in excessive data acquisition and poor user experience.
By obtaining the unified resource locator of the target web page, using the preset emotional vocabulary to calculate the similarity between the web page topic words and the emotional vocabulary, filtering out the target web page topic words, and realizing automatic filtering of web page content data.
Improve the efficiency and accuracy of web information collection and improve user experience.
Smart Images

Figure CN120541315A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, medium and program product for collecting web page information. Background Art
[0002] Internal employee forums have become one of the important channels for corporate employees to exchange information, share resources, maintain relationships, and share emotions. Collecting and analyzing information posted on internal employee forums is of great significance to improving management standardization.
[0003] Currently, existing web information collection methods typically use crawler technology to obtain specific web content. However, this method has certain limitations. Typically, it obtains too much data, lacks analysis and screening of the obtained data, and cannot accurately identify the actual needs of users, resulting in a poor user experience. Summary of the Invention
[0004] The present invention provides a web page information collection method, apparatus, device, medium and program product, which can realize automatic screening of acquired web page content data, improve the efficiency and accuracy of web page information collection, and enhance user experience.
[0005] According to one aspect of the present invention, a method for collecting web page information is provided, comprising:
[0006] Obtaining a uniform resource locator (URL) of a target webpage, and obtaining webpage content data corresponding to the target webpage based on the URL;
[0007] Acquire a plurality of initial webpage keywords based on the webpage content data, and calculate the initial similarity between each initial webpage keyword and each stored emotion word in each preset emotion word library;
[0008] According to the initial similarity between each of the initial webpage keywords and each stored emotion word in each preset emotion word library, and the weight value corresponding to each preset emotion word library, the target similarity between each of the initial webpage keywords and each stored emotion word is obtained;
[0009] According to the target similarity between each of the initial web page keywords and each of the existing sentiment words, a target web page keyword is obtained from each of the initial web page keywords.
[0010] According to another aspect of the present invention, there is provided a device for collecting web page information, comprising:
[0011] A web page content data acquisition module, configured to acquire a uniform resource locator (URL) of a target web page and, based on the URL, acquire web page content data corresponding to the target web page;
[0012] An initial similarity calculation module is used to obtain a plurality of initial web page keywords based on the web page content data, and calculate the initial similarity between each of the initial web page keywords and each stored emotion word in each preset emotion word library;
[0013] A target similarity acquisition module is used to acquire a target similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library according to the initial similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library, and the weight value corresponding to each preset emotion word library;
[0014] The target webpage theme word acquisition module is used to acquire the target webpage theme word from each of the initial webpage theme words according to the target similarity between each of the initial webpage theme words and each of the stored sentiment words.
[0015] According to another aspect of the present invention, an electronic device is provided, comprising:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the webpage information collection method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to enable a processor to implement the webpage information collection method according to any embodiment of the present invention when executed.
[0020] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method for collecting webpage information according to any embodiment of the present invention is implemented.
[0021] The technical solution of the embodiment of the present invention obtains the uniform resource locator of the target web page, and obtains the web page content data corresponding to the target web page according to the uniform resource locator; obtains multiple initial web page keywords according to the web page content data, and calculates the initial similarity between each initial web page keyword and each existing emotional word in each preset emotional word library; obtains the target similarity between each initial web page keyword and each existing emotional word according to the initial similarity between each initial web page keyword and each existing emotional word in each preset emotional word library, and the weight value corresponding to each preset emotional word library; obtains the target web page keyword from each initial web page keyword according to the target similarity between each initial web page keyword and each existing emotional word; after obtaining the web page content data, the similarity between the initial web page keyword and the existing emotional word is evaluated based on the weight value corresponding to each preset emotional word library, and the target web page keyword is obtained according to the similarity screening, thereby realizing automatic screening of the obtained web page content data, improving the efficiency and accuracy of web page information collection, and improving user experience.
[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 This is a flow chart of a method for collecting web page information provided according to the first embodiment of the present invention;
[0025] Figure 2 This is a flow chart of a method for collecting web page information provided according to the second embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of the structure of a web page information collection device provided in accordance with a third embodiment of the present invention;
[0027] Figure 4 It is a structural diagram of an electronic device for implementing the webpage information collection method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first", "second", "target", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] It is worth noting that for the convenience of understanding the technical solution of the present invention, the structure of the system to which the web page information collection method in the present invention is applied will be introduced first. Specifically, the overall framework of the distributed platform is built using the MapReduce architecture. MapReduce consists of a JobTracher and several TaskTrackers. The JobTracher runs on the master node and is the core of the entire MapReduce architecture. It is used to group and allocate the Uniform Resource Locators (URLs) in the theme library according to the principle of load balancing to the TaskTrackers. Among them, the URL addresses of the enterprise internal forums can be collected in advance to form a theme library.
[0031] The TaskTracher runs on the slave nodes. It is used to start the Map distributed download of web page data after receiving the instructions from the JobTracher. After acting alone and ending, it saves them in the format of <URL, web page data> key-value pairs respectively, and starts the Reduce to merge the key-value pairs and extract relevant information. Secondly, if the parsed web page contains other links, the links are de-duplicated and analyzed. If the format is different from the existing URLs in the theme library, they are recorded in the theme library.
[0032] Embodiment 1
[0033] Figure 1A flowchart of a web page information collection method is provided for the first embodiment of the present invention. This embodiment is applicable to the case of collecting web page information from internal employee forums. The method can be executed by a web page information collection device, which can be implemented in the form of hardware and / or software. Typically, the web page information collection device can be configured in an electronic device, such as a computer device or a server. Figure 1 As shown, the method includes:
[0034] S110: Obtain a uniform resource locator (URL) of a target webpage, and obtain webpage content data corresponding to the target webpage according to the URL.
[0035] It should be noted that the information collected in this embodiment is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0036] The target webpage can be a webpage for information collection and analysis, such as an internal employee forum. Specifically, the TaskTracher running on the slave node receives the URL of the target webpage sent by the JobTracher, and downloads and parses the webpage content based on the URL to obtain the webpage content data corresponding to the target webpage.
[0037] S120 , obtaining a plurality of initial web page keywords based on the web page content data, and calculating an initial similarity between each initial web page keyword and each stored emotion word in each preset emotion word library.
[0038] Specifically, the web page content data can be segmented to obtain a plurality of initial web page keywords. This embodiment does not specifically limit the type of word segmentation method. Afterwards, a pre-trained word vector model, such as Word2vec, can be used to perform vector conversion on the initial web page keywords and each existing emotional word in the preset emotional vocabulary to obtain the word vector corresponding to the initial web page keyword and the word vector corresponding to the existing emotional word; further, methods such as cosine similarity and Euclidean distance can be used to calculate the similarity between the word vector corresponding to the initial web page keyword and the word vector corresponding to the existing emotional word, as the initial similarity between the initial web page keyword and the existing emotional word.
[0039] The preset emotional vocabulary can include a basic emotional vocabulary, a hot emotional vocabulary, and a temporary emotional vocabulary. The emotional words in the basic emotional vocabulary are words that can clearly express the user's emotions. The emotional words in the hot emotional vocabulary are trending words on social platforms or recently popular emotional words. The emotional words in the temporary emotional vocabulary are new and unverified words. They have a set expiration period. If they do not appear repeatedly within a certain frequency during this period, they will be removed.
[0040] S130 , obtaining a target similarity between each of the initial webpage keywords and each stored emotion word in each preset emotion word library according to the initial similarity between each of the initial webpage keywords and each stored emotion word in each preset emotion word library, and the weight value corresponding to each preset emotion word library.
[0041] In this embodiment, a corresponding weight value is pre-set for each preset sentiment vocabulary; for example, the weight values gradually decrease in the order of the basic sentiment vocabulary, the hot sentiment vocabulary, and the temporary sentiment vocabulary. Specifically, the initial similarity between the current initial webpage topic word and each stored sentiment word in the current preset sentiment vocabulary can be multiplied by the weight value corresponding to the current preset sentiment vocabulary, and the product obtained is used as the corresponding target similarity.
[0042] S140 , obtaining a target web page theme word from each of the initial web page theme words according to the target similarity between each of the initial web page theme words and each of the stored sentiment words.
[0043] Specifically, the maximum target similarity corresponding to each initial web page theme word can be determined based on the target similarity between each initial web page theme word and each existing sentiment word; then, if it is detected that the maximum target similarity corresponding to a certain initial web page theme word is greater than or equal to a preset similarity threshold, the initial web page theme word can be determined as the target web page theme word.
[0044] Optionally, the technical solution of this embodiment further includes:
[0045] Obtaining a data type corresponding to the webpage content data, and obtaining a data compression strategy corresponding to the webpage content data based on the data type;
[0046] The webpage content data is compressed based on the data compression strategy to obtain compressed content data, and the compressed content data is sent to a target data center for storage.
[0047] Among them, the data type may include text data, structured data, multimedia resource data, etc. In this embodiment, after obtaining the web page content data, the data type corresponding to the web page data content can be obtained first, and then the data compression strategy corresponding to the current web page content data can be obtained based on the current data type and the mapping relationship between the preset data type and the data compression strategy. For example, when the data type is text data, the data compression strategy is to use a specified compression library for real-time compression; or, when the data type is structured data, the data compression strategy is to first convert it into a columnar storage format, then encode the numerical metadata, and finally optimize the storage of repeated fields by combining run-length encoding and dictionary encoding; or, when the data type is multimedia resource data, the data compression strategy is to perform WebP / AVIF format transcoding.
[0048] Specifically, after determining the data compression strategy, the webpage content data can be compressed according to the data compression strategy to obtain compressed content data. The compressed content data and the identification of the data compression strategy are then sent to the target data center. After receiving the compressed content data, the target data center can decompress it according to the identification of the data compression strategy to obtain the original webpage content data, and store the webpage content data in a designated database. The target data center can be pre-specified or selected according to preset rules.
[0049] In this embodiment, by constructing an intelligent data compression pipeline and implementing differentiated data compression strategies according to data types, data transmission efficiency can be improved.
[0050] Optionally, during data transmission, an adaptive chunking mechanism can be employed to dynamically adjust the data chunk size based on real-time bandwidth. Zero-copy transmission technology can also be employed to combine metadata and content data transmission, eliminating memory copy overhead. Secondly, a layered metadata tagging system can be developed to attach descriptive information such as compression algorithm identifiers and data versions to each data chunk, ensuring that data centers can quickly reconstruct the original data.
[0051] Optionally, sending the compressed content data to a target data center for storage may include:
[0052] Obtain the location information of this node and the location information of each data center, and based on the distance priority rule, determine the target data center among the data centers according to the location information of this node and the location information of each data center, and send the compressed content data to the target data center for storage.
[0053] In this embodiment, geolocation-aware routing can be established to prioritize data centers located in the same region for data transmission and storage. Specifically, the distances between the node and each data center can be determined based on the node's location information and the location information of each data center, and the closest data center can be selected as the target data center.
[0054] In this embodiment, by selecting the target data center based on the distance priority rule, network jump delay can be reduced and the time spent on data transmission can be shortened.
[0055] Optionally, in this embodiment, a local cache system can also be deployed on each slave node to automatically retain high-frequency access data within a preset period of time (such as 3 hours, etc.), and combine the least recently used algorithm to achieve rapid response to hot data.
[0056] The technical solution of the embodiment of the present invention obtains the uniform resource locator of the target web page, and obtains the web page content data corresponding to the target web page according to the uniform resource locator; obtains multiple initial web page keywords according to the web page content data, and calculates the initial similarity between each initial web page keyword and each existing emotional word in each preset emotional word library; obtains the target similarity between each initial web page keyword and each existing emotional word according to the initial similarity between each initial web page keyword and each existing emotional word in each preset emotional word library, and the weight value corresponding to each preset emotional word library; obtains the target web page keyword from each initial web page keyword according to the target similarity between each initial web page keyword and each existing emotional word; after obtaining the web page content data, the similarity between the initial web page keyword and the existing emotional word is evaluated based on the weight value corresponding to each preset emotional word library, and the target web page keyword is obtained according to the similarity screening, thereby realizing automatic screening of the obtained web page content data, improving the efficiency and accuracy of web page information collection, and improving user experience.
[0057] Example 2
[0058] Figure 2 This is a flow chart of a web page information collection method provided by Example 2 of the present invention. This embodiment is a further refinement of the above technical solution. The technical solution in this embodiment can be combined with one or more of the above implementation methods. Figure 2 As shown, the method includes:
[0059] S210: Obtain a uniform resource locator (URL) of a target webpage, and collect webpage data corresponding to the target webpage based on the URL.
[0060] Specifically, after obtaining the URL corresponding to the target web page, TaskTracher can perform distributed web page downloading based on the URL to obtain web page data corresponding to the target web page. This embodiment does not specifically limit the web page downloading method.
[0061] S220: Obtain a webpage type corresponding to the target webpage, and obtain a target resolution strategy corresponding to the target webpage based on the webpage type.
[0062] Among them, web page types can include static pages, dynamic pages, etc. In this embodiment, a modular parsing engine can be deployed in TaskTracher, and Docker containerization technology can be used to achieve lightweight deployment (image volume <50 megabytes). The engine has a built-in intelligent parsing strategy selector that can dynamically adapt to different web page types. Specifically, the target parsing strategy corresponding to the target web page can be determined based on the web page type corresponding to the target web page and the mapping relationship between the preset web page type and the parsing strategy.
[0063] Optionally, obtaining a target resolution strategy corresponding to the target webpage according to the webpage type may include:
[0064] If it is detected that the webpage type is a static page, determining that the target parsing strategy corresponding to the target webpage is Hypertext Markup Language parsing;
[0065] If it is detected that the webpage type is a dynamic page, it is determined that the target parsing strategy corresponding to the target webpage is headless rendering parsing.
[0066] In this embodiment, the parsing strategy used for static pages is Hypertext Markup Language (HTML) parsing, while the parsing strategy used for dynamic pages is headless rendering parsing. Specifically, for static pages, an HTML parser is used to directly extract structured data, while for dynamic pages, a headless browser is automatically called to perform Document Object Model rendering to ensure complete parsing of dynamic content.
[0067] In this embodiment, by performing HTML parsing on static pages and performing headless rendering parsing on dynamic pages, the content parsing efficiency of static and dynamic pages can be improved.
[0068] S230: Perform content analysis on the webpage data based on the target analysis strategy to obtain webpage content data corresponding to the target webpage.
[0069] Specifically, the web page data may be parsed using a modular parsing engine and a target parsing strategy to obtain the web page content data specifically contained in the target web page.
[0070] In this embodiment, by selecting differentiated parsing strategies for different web page types, the efficiency and accuracy of web page parsing can be improved.
[0071] S240: Acquire a plurality of initial web page keywords according to the web page content data.
[0072] S250: Using a pre-trained bidirectional transformer model, obtain a first word vector corresponding to the current initial webpage topic word and a second word vector corresponding to the currently stored sentiment word.
[0073] In this embodiment, a pre-trained bidirectional transformer model (Bidirectional Encoder Representations from Transformers, BERT) can be used to analyze the contextual information of web page keywords and stored sentiment words, capture the "same words with different meanings" and "different words with the same meanings", and obtain the first word vector corresponding to the initial web page keyword and the second word vector corresponding to the stored sentiment words. BERT can be composed of multiple layers of Transformer encoders stacked together.
[0074] S260. Use the word frequency-inverse document frequency method to obtain the first word weight value corresponding to the first word vector and the second word weight value corresponding to the second word vector, and use the cosine similarity method to obtain the similarity between the first word weight value and the second word weight value as the initial similarity between the current initial web page topic word and the currently stored sentiment word.
[0075] Specifically, the term frequency-inverse document frequency (TF-IDF) method can be used to transform the first term vector to obtain the corresponding TF-IDF vector as the first term weight value. At the same time, the second term vector can be transformed to obtain the corresponding TF-IDF vector as the second term weight value. For example, the formula ω can be used. ij =tf ij ×idf j and idf j =log(N / df j ) Transform the word vector to obtain the corresponding word weight value ω ij , where tf ij Represents word vector t j In the web data i The word frequency in , N represents the number of web page data, df j Indicates that all web page data contains word vector t j The number of web page data, idf jFurthermore, the cosine similarity method can be used to calculate the similarity between the first word weight value and the second word weight value to serve as the initial similarity between the initial web page topic words and the stored sentiment words.
[0076] In this embodiment, a hybrid computing framework is constructed by combining the BERT deep semantic model with the traditional TF-IDF statistical model, which not only retains the word frequency statistical advantages of the traditional method and ensures the stable recognition of basic terms, but also takes into account the contextual meaning of the text and improves the accuracy of word segmentation similarity evaluation.
[0077] Optionally, this embodiment can also be provided with dynamic incremental training, for example, adopting a "rolling window" update strategy, retaining the data of the last 30 days for each training, and dynamically eliminating obsolete corpus; or, automatically collecting new forum posts every day to build an incremental data set, and setting a semantic drift warning, triggering full model retraining when the new word recognition accuracy drops to a threshold.
[0078] S270 , obtaining a target similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library according to the initial similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library, and the weight value corresponding to each preset emotion word library.
[0079] S280 , obtaining a target web page theme word from each of the initial web page theme words according to the target similarity between each of the initial web page theme words and each of the stored sentiment words.
[0080] The solution of this embodiment introduces a three-level resource management mechanism during the parsing process. It uses virtual machine parameters to limit memory usage to a threshold of the total node capacity, and uses the control group function to control the occupancy rate of the central processing unit to below a certain percentage. When the system load exceeds the threshold, it automatically switches to lightweight mode, such as disabling specific energy-consuming functions such as rendering, to reduce end-to-end processing latency and thus improve parsing efficiency. Secondly, in the semantic understanding process, a dual-engine hybrid architecture is constructed. The BERT model is used to capture complex contextual associations, and TF-IDF is combined to ensure the stability of basic terms, achieving a multi-dimensional breakthrough in semantic understanding. Moreover, the three-level vocabulary management is combined with semantic anchor technology, and dynamic incremental training controls update efficiency, effectively solving the problems of polysemy and emerging concept recognition, and providing continuous and accurate semantic parsing capabilities for different emotional scenarios. Finally, the establishment of the URL theme library and the emotional vocabulary library is a long-term optimization process. The results of each information collection are re-trained into the relevant theme library, providing strong support for the accuracy and comprehensiveness of information collection.
[0081] The technical solution of the embodiment of the present invention collects web page data corresponding to the target web page according to the uniform resource locator; obtains the web page type corresponding to the target web page, and obtains the target parsing strategy corresponding to the target web page according to the web page type; performs content parsing on the web page data based on the target parsing strategy to obtain web page content data corresponding to the target web page; by selecting differentiated parsing strategies for different web page types, the efficiency and accuracy of web page parsing can be improved; a pre-trained bidirectional transformer model is used to obtain a first word vector corresponding to the current initial web page theme word and a second word vector corresponding to the current stored sentiment word; a word frequency-inverse document frequency method is used to obtain a first word weight value corresponding to the first word vector and a second word weight value corresponding to the second word vector, and a cosine similarity method is used to obtain the similarity between the first word weight value and the second word weight value as the initial similarity between the current initial web page theme word and the current stored sentiment word; by combining BERT and TF-IDF to perform similarity evaluation between web page keywords and stored sentiment words, the accuracy of the similarity evaluation can be improved.
[0082] Example 3
[0083] Figure 3 This is a schematic diagram of the structure of a web page information collection device provided by the third embodiment of the present invention. Figure 3 As shown, the device includes: a web page content data acquisition module 310, an initial similarity calculation module 320, a target similarity acquisition module 330 and a target web page keyword acquisition module 340; wherein,
[0084] The webpage content data acquisition module 310 is used to acquire the uniform resource locator of the target webpage and acquire the webpage content data corresponding to the target webpage according to the uniform resource locator;
[0085] An initial similarity calculation module 320 is used to obtain a plurality of initial web page keywords based on the web page content data, and calculate the initial similarity between each of the initial web page keywords and each stored emotion word in each preset emotion word library;
[0086] A target similarity acquisition module 330 is configured to acquire a target similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library based on the initial similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library, and the weight value corresponding to each preset emotion word library;
[0087] The target webpage keyword acquisition module 340 is configured to acquire a target webpage keyword from each of the initial webpage keywords according to target similarities between each of the initial webpage keywords and each of the stored sentiment words.
[0088] The technical solution of the embodiment of the present invention obtains the uniform resource locator of the target web page, and obtains the web page content data corresponding to the target web page according to the uniform resource locator; obtains multiple initial web page keywords according to the web page content data, and calculates the initial similarity between each initial web page keyword and each existing emotional word in each preset emotional word library; obtains the target similarity between each initial web page keyword and each existing emotional word according to the initial similarity between each initial web page keyword and each existing emotional word in each preset emotional word library, and the weight value corresponding to each preset emotional word library; obtains the target web page keyword from each initial web page keyword according to the target similarity between each initial web page keyword and each existing emotional word; after obtaining the web page content data, the similarity between the initial web page keyword and the existing emotional word is evaluated based on the weight value corresponding to each preset emotional word library, and the target web page keyword is obtained according to the similarity screening, thereby realizing automatic screening of the obtained web page content data, improving the efficiency and accuracy of web page information collection, and improving user experience.
[0089] Optionally, the webpage content data acquisition module 310 is specifically configured to acquire webpage data corresponding to the target webpage according to the uniform resource locator;
[0090] Obtaining a webpage type corresponding to the target webpage, and obtaining a target parsing strategy corresponding to the target webpage based on the webpage type;
[0091] The webpage data is subjected to content analysis based on the target analysis strategy to obtain webpage content data corresponding to the target webpage.
[0092] Optionally, the webpage content data acquisition module 310 is specifically configured to determine that the target parsing strategy corresponding to the target webpage is Hypertext Markup Language parsing if it is detected that the webpage type is a static page;
[0093] If it is detected that the webpage type is a dynamic page, it is determined that the target parsing strategy corresponding to the target webpage is headless rendering parsing.
[0094] Optionally, the initial similarity calculation module 320 is specifically configured to use a pre-trained bidirectional transformer model to obtain a first word vector corresponding to the current initial webpage topic word and a second word vector corresponding to the currently stored sentiment word;
[0095] The word frequency-inverse document frequency method is used to obtain the first word weight value corresponding to the first word vector and the second word weight value corresponding to the second word vector, and the cosine similarity method is used to obtain the similarity between the first word weight value and the second word weight value as the initial similarity between the current initial web page topic word and the currently stored sentiment word.
[0096] Optionally, the webpage information collection device further includes:
[0097] A data compression strategy acquisition module is used to acquire a data type corresponding to the webpage content data, and acquire a data compression strategy corresponding to the webpage content data according to the data type;
[0098] The data compression module is used to perform data compression processing on the webpage content data based on the data compression strategy, obtain compressed content data, and send the compressed content data to a target data center for storage.
[0099] Optionally, a data compression module is specifically used to obtain the location information of this node and the location information of each data center, and based on the distance priority rule, determine the target data center among each data center according to the location information of this node and the location information of each data center, and send the compressed content data to the target data center for storage.
[0100] The web page information collection device provided in the embodiment of the present invention can execute the web page information collection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0101] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0102] Example 4
[0103] Figure 4 A schematic diagram of the structure of an electronic device 40 that can be used to implement an embodiment of the present invention is shown. The electronic device 40 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 40 can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0104] like Figure 4As shown, the electronic device 40 includes at least one processor 41, and a memory connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory 42 or the computer program loaded from the storage unit 48 to the random access memory 43. Various programs and data required for the operation of the electronic device 40 can also be stored in the RAM 43. The processor 41, ROM 42 and RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0105] Multiple components in the electronic device 40 are connected to the I / O interface 45, including an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0106] Processor 41 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit, a graphics processing unit, various specialized artificial intelligence computing chips, various processors running machine learning model algorithms, a digital signal processor, and any other suitable processor, controller, microcontroller, etc. Processor 41 executes the various methods and processes described above, such as the method for collecting webpage information.
[0107] In some embodiments, the web page information collection method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the web page information collection method described above can be performed. Alternatively, in other embodiments, processor 41 can be configured to perform the web page information collection method by any other appropriate means (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays, application specific integrated circuits, application specific standard products, system-on-a-chip systems, on-load programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device or any suitable combination of the foregoing.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device 40 having: a display device (e.g., a cathode ray tube or a liquid crystal display) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device 40. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0112] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks, wide area networks, blockchain networks, and the Internet.
[0113] A computing system may include clients and servers. The client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server.
[0114] This embodiment may also include a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the webpage information collection method provided by any embodiment of the present invention.
[0115] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0116] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for collecting web page information, characterized in that: include: Obtaining a uniform resource locator (URL) of a target webpage, and obtaining webpage content data corresponding to the target webpage based on the URL; Acquire a plurality of initial webpage keywords based on the webpage content data, and calculate the initial similarity between each initial webpage keyword and each stored emotion word in each preset emotion word library; According to the initial similarity between each of the initial webpage keywords and each stored emotion word in each preset emotion word library, and the weight value corresponding to each preset emotion word library, the target similarity between each of the initial webpage keywords and each stored emotion word is obtained; According to the target similarity between each of the initial web page keywords and each of the existing sentiment words, a target web page keyword is obtained from each of the initial web page keywords.
2. The method according to claim 1, characterized in that Acquiring webpage content data corresponding to the target webpage according to the uniform resource locator, including: According to the uniform resource locator, web page data corresponding to the target web page is collected; Obtaining a webpage type corresponding to the target webpage, and obtaining a target parsing strategy corresponding to the target webpage based on the webpage type; The webpage data is subjected to content analysis based on the target analysis strategy to obtain webpage content data corresponding to the target webpage.
3. The method according to claim 2, characterized in that Obtaining a target parsing strategy corresponding to the target webpage according to the webpage type includes: If it is detected that the webpage type is a static page, determining that the target parsing strategy corresponding to the target webpage is Hypertext Markup Language parsing; If it is detected that the webpage type is a dynamic page, it is determined that the target parsing strategy corresponding to the target webpage is headless rendering parsing.
4. The method according to claim 1, wherein Calculating the initial similarity between each of the initial web page keywords and each stored emotion word in each preset emotion word library includes: Using a pre-trained bidirectional transformer model, obtain the first word vector corresponding to the current initial webpage topic word and the second word vector corresponding to the currently stored sentiment word; The word frequency-inverse document frequency method is used to obtain the first word weight value corresponding to the first word vector and the second word weight value corresponding to the second word vector, and the cosine similarity method is used to obtain the similarity between the first word weight value and the second word weight value as the initial similarity between the current initial web page topic word and the currently stored sentiment word.
5. The method according to claim 1, wherein Also includes: Obtaining a data type corresponding to the webpage content data, and obtaining a data compression strategy corresponding to the webpage content data based on the data type; The webpage content data is compressed based on the data compression strategy to obtain compressed content data, and the compressed content data is sent to a target data center for storage.
6. The method according to claim 5, characterized in that Sending the compressed content data to a target data center for storage includes: Obtain the location information of this node and the location information of each data center, and based on the distance priority rule, determine the target data center among the data centers according to the location information of this node and the location information of each data center, and send the compressed content data to the target data center for storage.
7. A web page information collection device, characterized in that: include: A web page content data acquisition module, configured to acquire a uniform resource locator (URL) of a target web page and, based on the URL, acquire web page content data corresponding to the target web page; An initial similarity calculation module is used to obtain a plurality of initial web page keywords based on the web page content data, and calculate the initial similarity between each of the initial web page keywords and each stored emotion word in each preset emotion word library; A target similarity acquisition module is used to acquire a target similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library according to the initial similarity between each of the initial webpage keywords and each of the stored emotion words in each preset emotion word library, and the weight value corresponding to each preset emotion word library; The target webpage theme word acquisition module is used to acquire the target webpage theme word from each of the initial webpage theme words according to the target similarity between each of the initial webpage theme words and each of the stored sentiment words.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor, and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the webpage information collection method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is used to enable a processor to implement the webpage information collection method according to any one of claims 1 to 6 when executed.
10. A computer program product, characterized in that The method comprises a computer program, which implements the web page information collection method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Cited By
Webpage data collection method, electronic equipment and computer readable storage medium
CN121167008A