Web page text extraction method, device, equipment and storage medium based on image and text
Through graphic and text recognition and conflict screening technology, the web page text is extracted, which solves the problem that traditional methods are difficult to adapt to different web page structures, and improves the efficiency and accuracy of data crawling.
Patent Information
- Application Number
- CN202310300637.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-03-24
AI Technical Summary
Traditional web page text extraction methods require understanding of the web page structure and formulating professional templates, which are difficult to adapt to different web page structures, resulting in a decrease in data crawling speed and quality.
Through graphic and text recognition and positioning, the pre-trained picture text recognition model is used to identify the web page images, and the text image area is extracted based on the graphic and text conflict screening rules, and the text image data is obtained through text sequence number query.
It improves the efficiency and universality of web page text crawling, can extract web page text more accurately, and reduces interference from advertisements and irregular pictures.
Smart Images

Figure CN116386068B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a method, device, equipment and computer-readable storage medium for extracting text from a web page based on images and texts. Background Art
[0002] With the development of computer networks, big data has received increasing attention, and many companies have begun to obtain massive data from the Internet for analysis to help them make decisions and improve.
[0003] Traditional web page text extraction methods require understanding the structure of web pages and developing professional templates to implement data crawling. However, the number of web pages on the Internet has exploded, and the structures of web pages are also different. It is difficult to develop data crawling templates that are suitable for each network, and the speed and quality of obtaining network data are greatly reduced. Summary of the invention
[0004] The present invention provides a method, device, equipment and storage medium for extracting web page text based on images and texts, the main purpose of which is to extract the web page text through image and text recognition and positioning, thereby increasing the efficiency and universality of web text crawling.
[0005] To achieve the above object, the present invention provides a method for extracting web page text based on images and texts, comprising:
[0006] Retrieving a target web page according to preset keywords, crawling the hypertext markup data of the target web page, and intercepting the web page image of the target web page;
[0007] Using a pre-trained image text recognition model, perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and an image area set;
[0008] According to the preset image-text conflict screening rule and the text block set, the image area set is screened for text areas to obtain a text image area set;
[0009] Performing text recognition on each body text block in the body text block set to obtain a body text sequence, and performing a text sequence number marking operation on each body text sequence;
[0010] According to the text sequence number of the text sequence adjacent to the context of each text image region in the text image region set, the hypertext markup data is queried to obtain the text image data;
[0011] Determining the sentence integrity of each of the body text sequences;
[0012] According to the incomplete text sequence of the main text, the hypertext markup data is queried to obtain the main text data;
[0013] The complete sentence body text sequence, the body text data and the body image data are stored in a pre-constructed database.
[0014] Optionally, the performing text area screening on the picture area set according to the preset text-image conflict screening rule and the text block set to obtain the text picture area set includes:
[0015] According to the preset text-image conflict screening rules, construct the maximum external cut frame of the text block set;
[0016] Identifying the positional relationship between each image region in the image region set and the maximum circumscribed frame, and extracting a built-in image set from the image region set according to the positional relationship between each image region;
[0017] The image-text coverage conflict relationship of each built-in image in the built-in image set is identified, and based on the image-text coverage conflict relationship, the built-in images that do not cover the text are extracted from the included image set to obtain a text image area set.
[0018] Optionally, the performing object recognition on the webpage image using a pre-trained picture text recognition model includes:
[0019] Performing Gaussian filtering on the webpage image to obtain a noise-reduced image;
[0020] Using a pre-trained picture text recognition model, a feature extraction operation is performed on the denoised image to obtain an image feature sequence;
[0021] The image feature sequence is subjected to fully connected image-text classification based on the text font and text background to obtain an object recognition result.
[0022] Optionally, before performing object recognition on the webpage image using the pre-trained picture text recognition model, the method further includes:
[0023] Use preset embedding points to capture massive data texts and classify them to obtain image samples and text samples, and mark the image samples and text samples with type labels;
[0024] The marked image samples and the text samples are mixed as an image and text sample set, and the image and text sample set is randomly grouped according to a preset ratio to obtain a training set and a test set;
[0025] Extracting a training sample from the training set in turn, and using a pre-built picture and text recognition model to classify the training sample based on pictures and texts to obtain a test category;
[0026] Using a cross entropy loss algorithm, the loss value between the type label of the training sample and the test category is calculated, and the loss value is minimized to obtain the model parameters of the image text recognition model when the loss value is minimized, and the model parameters are updated inversely according to the gradient descent method to obtain an updated image text recognition model;
[0027] Determine whether all training samples in the training set participate in the training;
[0028] When not all the training samples in the training set participate in the training, returning to the above step of extracting one training sample from the training set in turn, and iteratively training the updated picture text recognition model;
[0029] When all the training samples in the training set participate in the training, the updated image text recognition model is tested using the test set to obtain the test accuracy;
[0030] Determining whether the test accuracy is greater than or equal to a preset pass threshold;
[0031] When the test accuracy is less than the qualified threshold, return to the above step of randomly grouping the image and text sample set according to a preset ratio, regroup the image and text sample set, and update the updated image and text recognition model;
[0032] When the test accuracy is greater than or equal to the qualified threshold, a trained picture text recognition model is obtained.
[0033] Optionally, the crawling of the hypertext markup data of the target webpage includes:
[0034] Send a request to the target webpage using a pre-built request library, obtain a status code fed back by the target webpage, and determine whether the status code indicates a successful request;
[0035] When the status code request is successful, the encrypted hypertext markup data of the target webpage is captured according to the response text attribute, and the encrypted hypertext markup data is parsed using a pre-built BeautifulSoup library to obtain decrypted hypertext markup data;
[0036] The decrypted hypertext markup data is formatted and outputted using a pre-built beautification function to obtain the hypertext markup data of the target web page.
[0037] Optionally, querying the hypertext markup data according to the text sequence number of the context-adjacent text sequence of each text image region in the text image region set to obtain the text image data includes:
[0038] Searching the hypertext markup data for the text sequence number of the body text sequence adjacent to the context, obtaining the file data position of the body text block corresponding to the text sequence number of the body text sequence adjacent to the context, and searching for the img tag of the embedded image in the file data position;
[0039] The hypertext markup data is queried according to the img tag to obtain the text image data.
[0040] Optionally, the determining the sentence integrity of each of the body text sequences includes:
[0041] Identifying the grammatical completeness of the body text sequence to obtain a first completeness score;
[0042] Identifying the semantic integrity of the body text sequence to obtain a second integrity score;
[0043] Identifying the contextual knowledge completeness of the text body sequence to obtain a third completeness score;
[0044] According to a preset weight configuration rule, the first integrity score, the second integrity score and the third integrity score are weightedly calculated to obtain the sentence integrity of the main text sequence.
[0045] In order to solve the above problems, the present invention also provides a web page text extraction device based on images and texts, the device comprising:
[0046] A data capture module, used to retrieve a target web page according to preset keywords, capture the hypertext markup data of the target web page, and capture the web page image of the target web page;
[0047] The image text classification module is used to use a pre-trained image text recognition model to perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and a picture area set;
[0048] A text image query module, used to perform text recognition on each text block in the text block set to obtain a text sequence, and to perform a text sequence number marking operation on each text sequence, and to query the hypertext markup data according to the text sequence number of the text sequence adjacent to the context of each text image region in the text image region set to obtain text image data;
[0049] The main text query module is used to determine the sentence integrity of each of the main text sequences, and to query the hypertext markup data based on the main text sequences with incomplete sentences to obtain the main text data, and to store the main text sequences with complete sentences, the main text data and the main text image data in a pre-constructed database.
[0050] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:
[0051] at least one processor; and,
[0052] a memory communicatively connected to the at least one processor; wherein,
[0053] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned image-text based web page text extraction method.
[0054] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned image-based web page text extraction method.
[0055] The embodiment of the present invention intercepts the web page image of the target web page, identifies the main text block and the picture area in the web page image, and then extracts the main text picture area through the image-text conflict screening rule, wherein the image-text conflict screening rule can eliminate a large number of non-standard pictures such as advertisements according to the maximum circumscribed boundary of the text, and can also delete pictures with image-text overlap conflicts to obtain a more accurate main text picture area; data is extracted from the incomplete main text according to the sentence integrity to obtain the entire web page text. Therefore, the embodiment of the present invention provides a web page text extraction method, device, equipment and storage medium based on image-text, which can extract the web page text through image-text recognition and positioning, thereby increasing the efficiency and universality of network text crawling. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A schematic diagram of a flow chart of a method for extracting text from a web page based on images and texts provided by an embodiment of the present invention;
[0057] Figure 2 A detailed flow chart of a step in a method for extracting web page text based on images and texts provided in one embodiment of the present invention;
[0058] Figure 3 A detailed flow chart of a step in a method for extracting web page text based on images and texts provided in one embodiment of the present invention;
[0059] Figure 4 A detailed flow chart of a step in a method for extracting web page text based on images and texts provided in one embodiment of the present invention;
[0060] Figure 5 A functional module diagram of a web page text extraction device based on images and texts provided in one embodiment of the present invention;
[0061] Figure 6 A schematic diagram of the structure of an electronic device for implementing the image-text based web page text extraction method provided by an embodiment of the present invention.
[0062] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0063] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0064] The embodiment of the present application provides a method for extracting web page text based on images and texts. In the embodiment of the present application, the execution subject of the method for extracting web page text based on images and texts includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the web page text extraction method based on images and texts can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0065] Reference Figure 1 FIG. 1 is a flow chart of a method for extracting web page text based on images and texts according to an embodiment of the present invention. In this embodiment, the method for extracting web page text based on images and texts includes:
[0066] S1. Retrieve a target web page according to preset keywords, capture hypertext markup data of the target web page, and capture a web page image of the target web page.
[0067] In the embodiments of the present invention, the Internet can be searched according to platform keywords such as *Chacha, *Bo, *Du, etc. and event keywords such as enterprise, management, and user portrait to obtain a list of target web pages. Then, the main text is crawled from each target web page in the list to obtain hypertext markup data and capture web page images.
[0068] Among them, the hypertext markup data (HyperText Mark-up Language, abbreviated as HTML) is a standard language for making web pages on the World Wide Web. Through various tags, data can be displayed. The hypertext markup data in the embodiments of the present invention is the decrypted data.
[0069] Specifically, in the embodiments of the present invention, the crawling of the hypertext markup data of the target web page includes:
[0070] Send a request to the target web page using a pre-built request library, obtain the status code feedback by the target web page, and determine whether the status code indicates a successful request;
[0071] When the status code indicates a successful request, capture the encrypted hypertext markup data of the target web page according to the response text attribute, and use the pre-built BeautifulSoup library to parse the encrypted hypertext markup data to obtain decrypted hypertext markup data;
[0072] Use a pre-built beautification function to format and output the decrypted hypertext markup data to obtain the hypertext markup data of the target web page.
[0073] Among them, both the request library (requests library) and the BeautifulSoup library are libraries in Python.
[0074] For example, in the embodiments of the present invention, first send a GET request to an enterprise web page on the *Chacha platform through the requests library to obtain the status code of the response. When the status code is 200, it indicates that the request is passed. Then, the HTML code of the enterprise web page can be captured through the response text (response.text) attribute. Then, use the BeautifulSoup library to parse the HTML code to obtain the decrypted HTML code. Finally, use the beautification function (prettify()) to format and output the decrypted HTML code to obtain hypertext markup data.
[0075] In addition, when the embodiments of the present invention crawl hypertext markup data, it is also necessary to capture the web page image of the enterprise web page. Among them, the web page image is an image obtained by fully expanding and scrolling and capturing the enterprise web page.
[0076] S2. Use a pre-trained image text recognition model to perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and a picture area set.
[0077] In an embodiment of the present invention, the picture text recognition model is an image convolutional neural network model for recognizing text and pictures, including the functions of text recognition, text intent recognition and text background range recognition.
[0078] In detail, in an embodiment of the present invention, the use of a pre-trained image text recognition model to perform object recognition on the web page image includes:
[0079] Performing Gaussian filtering on the webpage image to obtain a noise-reduced image;
[0080] Using a pre-trained picture text recognition model, a feature extraction operation is performed on the denoised image to obtain an image feature sequence;
[0081] The image feature sequence is subjected to fully connected image-text classification based on the text font and text background to obtain an object recognition result.
[0082] In the embodiment of the present invention, Gaussian filtering is performed on the web page image through a filter, wherein the Gaussian filtering refers to smoothing processing between pixels of the image.
[0083] The embodiment of the present invention performs convolution, pooling and flattening on the denoised image through the feature extraction network in the image text recognition model to obtain a reduced-dimensional image feature sequence, wherein the number of elements in the image feature sequence is the same as the number of convolution kernels in the convolution process, and each convolution kernel is responsible for identifying a feature. The pooling and flattening processes are both feature dimensionality reduction processes, which will not be described in detail here.
[0084] Furthermore, an embodiment of the present invention utilizes the fully connected layer of the image text recognition model to perform fully connected classification judgment on the image feature sequence to obtain an object recognition result, and utilizes the marking function in the final output network to mark and select the object recognition result to obtain a main text block set and a picture area set.
[0085] For further information, see Figure 2 As shown, in the embodiment of the present invention, before using the pre-trained picture text recognition model to perform object recognition on the web page image, the method further includes:
[0086] S201, using preset embedding points to capture massive data texts and classify them to obtain image samples and text samples, and mark the image samples and the text samples with type labels;
[0087] S202, mixing the marked image samples and the text samples as an image and text sample set, and randomly grouping the image and text sample set according to a preset ratio to obtain a training set and a test set;
[0088] S203, extracting a training sample from the training set in turn, and using a pre-built picture and text recognition model to classify the training sample based on pictures and texts to obtain a test category;
[0089] S204, using a cross entropy loss algorithm to calculate the loss value between the type label of the training sample and the test category, and minimize the loss value to obtain the model parameters of the image text recognition model when the loss value is minimized, and perform network reverse update on the model parameters according to a gradient descent method to obtain an updated image text recognition model;
[0090] S205, determining whether all training samples in the training set participate in the training;
[0091] When not all the training samples in the training set participate in the training, return to the above step S203 to iteratively train the updated image text recognition model;
[0092] When all the training samples in the training set participate in the training, S206, the updated image text recognition model is tested using the test set to obtain a test accuracy rate;
[0093] S207, determining whether the test accuracy is greater than or equal to a preset qualified threshold;
[0094] When the test accuracy is less than the qualified threshold, return to the above step S202, regroup the image and text sample set, and update the updated image and text recognition model;
[0095] When the test accuracy is greater than or equal to the qualified threshold, S208, a trained image text recognition model is obtained.
[0096] Among them, the image samples and text samples obtained in the embodiment of the present invention are all in the form of images, and preprocessing operations are performed, namely: operations such as image, cropping, normalization, and operations such as text segmentation and removal of stop words.
[0097] During the training process, the present invention controls the training direction through the cross entropy loss algorithm and the gradient descent method, and controls the training effect of the model by testing the accuracy. When the accuracy reaches the preset qualified threshold N, such as 95%, it can be considered that the training process of the model is completed, and the image text recognition model is obtained. Among them, the cross entropy loss algorithm is used to calculate the difference between the type label (true label) and the test category (prediction label) to measure the prediction accuracy of the model; the gradient descent method is a method of minimizing the loss function, which is used to find the model parameters at the minimum loss value, and then realize reverse update.
[0098] S3. According to the preset text-image conflict screening rule and the text block set, the text area screening is performed on the picture area set to obtain a text picture area set.
[0099] For details, see Figure 3 As shown, in the embodiment of the present invention, the text area of the picture area set is screened according to the preset text-image conflict screening rule and the text block set to obtain the text picture area set, including:
[0100] S31, constructing the maximum circumscribed frame of the text block set according to a preset image-text conflict screening rule;
[0101] S32, identifying the positional relationship between each image region in the image region set and the maximum circumscribed frame, and extracting a built-in image set from the image region set according to the positional relationship between each image region;
[0102] S33, identifying the image-text coverage conflict relationship of each built-in image in the built-in image set, and extracting the built-in images that do not cover the text from the built-in image set according to the image-text coverage conflict relationship, to obtain a text image area set.
[0103] In the embodiment of the present invention, the image-text conflict screening rule refers to determining whether the positions of the image and text are standardized, whether the text and the image, or the image and the image are embedded or overlapped, and further distinguishing non-text images such as advertisements.
[0104] The embodiment of the present invention constructs the maximum circumscribed frame of the main text block set, and then treats the image area that intersects with or is outside the maximum circumscribed frame as an invalid image, and retains the built-in images that are completely contained in the maximum circumscribed frame to obtain a built-in image set. For example, when a person's frontal photo appears in a corporate webpage, it may be located between paragraphs of the main text, or it may be embedded in the main text like in newspaper or paper layout, but most images such as advertisements exceed the article frame and interfere with reading. Therefore, most advertising interference can be eliminated through the built-in image set. Finally, the built-in image set is tested for conflicts to see if they overlap with each other, and a main text image area set is obtained.
[0105] S4. Perform text recognition on each body text block in the body text block set to obtain a body text sequence, and perform a text sequence number marking operation on each body text sequence.
[0106] In the embodiment of the present invention, after identifying each body text block, text recognition is performed on each body text block, image text data is converted into text data, a body text sequence is obtained, and each body text sequence is marked with a text sequence number. The text sequence number may include web page display information such as a digital sequence number and a paragraph position.
[0107] S5. Query the hypertext markup data according to the text sequence number of the context-adjacent text sequence of each text image region in the text image region set to obtain text image data.
[0108] For details, see Figure 4 As shown, in the embodiment of the present invention, the method of querying the hypertext markup data according to the text sequence number of the context-adjacent text sequence of each text image region in the text image region set to obtain the text image data includes:
[0109] S51, searching the hypertext markup data for the text sequence number of the body text sequence adjacent to the context, obtaining the file data position of the body text block corresponding to the text sequence number of the body text sequence adjacent to the context, and searching for the img tag of the embedded image in the file data position;
[0110] S52: query the hypertext markup data according to the img tag to obtain text image data.
[0111] In the embodiment of the present invention, the developer tool can automatically lock the mouse position at a fixed position on the web page according to the text sequence number, and then the corresponding text tag in the HTML file will be highlighted, and then the highlighted part can be queried to obtain the corresponding Tag, so as to query and obtain the text image data.
[0112] S6. Determine the sentence integrity of each of the main text sequences.
[0113] In detail, in the embodiment of the present invention, the step of determining the sentence integrity of each of the body text sequences includes:
[0114] Identifying the grammatical completeness of the body text sequence to obtain a first completeness score;
[0115] Identifying the semantic integrity of the body text sequence to obtain a second integrity score;
[0116] Identifying the contextual knowledge completeness of the text body sequence to obtain a third completeness score;
[0117] According to a preset weight configuration rule, the first integrity score, the second integrity score and the third integrity score are weightedly calculated to obtain the sentence integrity of the main text sequence.
[0118] In the embodiment of the present invention, the sentence integrity of the main text sequence is comprehensively identified through grammatical integrity, semantic integrity and contextual knowledge integrity. Among them, grammatical integrity is to check whether the phrase sentence conforms to the grammatical rules. If the grammatical structure of the phrase sentence is incomplete, it may indicate that it is incomplete or has grammatical errors; semantic integrity refers to identifying whether the intention of the text is complete. If the intention is ambiguous or the wording is inappropriate, it can be considered that the semantic integrity is insufficient; contextual knowledge integrity refers to judging whether the phrase sentence is complete based on domain knowledge and common sense reasoning. For example, if the phrase sentence describes an event or behavior, it may need to contain information such as the subject, predicate and object of the event or behavior, otherwise it may be incomplete.
[0119] In the embodiment of the present invention, an enterprise can assign appropriate weights to each integrity level according to different scenarios, and then perform weighted calculation to obtain statement integrity.
[0120] S7. According to the incomplete main text sequence of the sentence, the hypertext markup data is queried to obtain the main text data.
[0121] In the embodiment of the present invention, when the sentence is complete, the body text data can be directly obtained by recognizing the image, and when the sentence is incomplete, it is still necessary to extract the corresponding text through the hypertext markup data to obtain the corresponding body text data.
[0122] S8. Store the complete text sequence of the sentence, the text data of the sentence and the text image data into a pre-constructed database.
[0123] The embodiment of the present invention stores a complete text sequence of a sentence, the text data of the text and the text image data in a pre-constructed database to complete the extraction process of the network text.
[0124] Afterwards, data can be extracted from the database through front-end design and other methods to construct character portraits, corporate structure portraits, etc.
[0125] The embodiment of the present invention intercepts the web page image of the target web page, identifies the main text block and the picture area in the web page image, and then extracts the main text picture area through the image-text conflict screening rule, wherein the image-text conflict screening rule can eliminate a large number of non-standard pictures such as advertisements according to the maximum circumscribed boundary of the text, and can also delete pictures with image-text overlap conflicts to obtain a more accurate main text picture area; then, according to the text sequence number of the main text, extract the main text image data, and extract the incomplete main text according to the sentence integrity to obtain the entire web page text. Therefore, the web page text extraction method based on images and texts provided by the embodiment of the present invention can extract the web page text through image-text recognition and positioning, thereby increasing the efficiency and universality of network text crawling.
[0126] like Figure 5 , which is a functional module diagram of a web page text extraction device based on images and texts provided in one embodiment of the present invention.
[0127] The web page text extraction device 100 based on images and texts of the present invention can be installed in an electronic device. According to the functions to be implemented, the web page text extraction device 100 based on images and texts can include a data capture module 101, an image text classification module 102, a text image query module 103 and a text query module 104. The module of the present invention can also be called a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0128] In this embodiment, the functions of each module / unit are as follows:
[0129] The data capture module 101 is used to retrieve a target web page according to a preset keyword, capture the hypertext markup data of the target web page, and capture the web page image of the target web page;
[0130] The image text classification module 102 is used to use a pre-trained image text recognition model to perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and an image area set;
[0131] The text image query module 103 is used to perform text recognition on each text block in the text text block set to obtain a text sequence, and to perform a text sequence number marking operation on each text sequence, and to query the hypertext markup data according to the text sequence number of the text text sequence adjacent to the context of each text image region in the text image region set to obtain text image data;
[0132] The main text query module 104 is used to determine the sentence completeness of each of the main text sequences, and to query the hypertext markup data according to the main text sequences with incomplete sentences to obtain the main text data, and to store the main text sequences with complete sentences, the main text data and the main text image data in a pre-constructed database.
[0133] In detail, each module described in the web page text extraction device 100 based on images and texts in the embodiment of the present application is used in the same manner as described above. Figures 1 to 4 The same technical means as the web page text extraction method based on images and texts described in the text extraction method can produce the same technical effects, which will not be repeated here.
[0134] like Figure 6 , which is a schematic diagram of the structure of an electronic device 1 for implementing a method for extracting web page text based on images and texts provided by an embodiment of the present invention.
[0135] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a graphic-based web page text extraction program.
[0136] In some embodiments, the processor 10 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 10 is the control core (ControlUnit) of the electronic device 1, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules stored in the memory 11 (for example, executing a web page text extraction program based on graphics and text, etc.), and calls data stored in the memory 11 to execute various functions of the electronic device and process data.
[0137] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit of the electronic device and an external storage device. The memory 11 may not only be used to store application software and various types of data installed in the electronic device, such as the code of a web page text extraction program based on graphics and text, but may also be used to temporarily store data that has been output or is to be output.
[0138] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0139] The communication interface 13 is used for communication between the above-mentioned electronic device 1 and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface.
[0140] Figure 6 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 6The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0141] For example, although not shown, the electronic device 1 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The electronic device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0142] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0143] The image-based webpage text extraction program stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve:
[0144] Retrieving a target web page according to preset keywords, crawling the hypertext markup data of the target web page, and intercepting the web page image of the target web page;
[0145] Using a pre-trained image text recognition model, perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and an image area set;
[0146] According to the preset image-text conflict screening rule and the text block set, the image area set is screened for text areas to obtain a text image area set;
[0147] Performing text recognition on each body text block in the body text block set to obtain a body text sequence, and performing a text sequence number marking operation on each body text sequence;
[0148] According to the text sequence number of the text sequence adjacent to the context of each text image region in the text image region set, the hypertext markup data is queried to obtain the text image data;
[0149] Determining the sentence integrity of each of the body text sequences;
[0150] According to the incomplete text sequence of the main text, the hypertext markup data is queried to obtain the main text data;
[0151] The complete sentence body text sequence, the body text data and the body image data are stored in a pre-constructed database.
[0152] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to the description of the relevant steps in the corresponding embodiment of the accompanying drawings, which will not be repeated here.
[0153] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0154] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, the computer program can implement:
[0155] Retrieving a target web page according to preset keywords, crawling the hypertext markup data of the target web page, and intercepting the web page image of the target web page;
[0156] Using a pre-trained image text recognition model, perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and an image area set;
[0157] According to the preset image-text conflict screening rule and the text block set, the image area set is screened for text areas to obtain a text image area set;
[0158] Performing text recognition on each body text block in the body text block set to obtain a body text sequence, and performing a text sequence number marking operation on each body text sequence;
[0159] According to the text sequence number of the text sequence adjacent to the context of each text image region in the text image region set, the hypertext markup data is queried to obtain the text image data;
[0160] Determining the sentence integrity of each of the body text sequences;
[0161] According to the incomplete text sequence of the main text, the hypertext markup data is queried to obtain the main text data;
[0162] The complete sentence body text sequence, the body text data and the body image data are stored in a pre-constructed database.
[0163] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0164] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0166] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0167] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0168] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0169] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0170] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0171] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for extracting web page text based on images and texts. It is characterized in that The method comprises: Retrieving a target web page according to preset keywords, crawling the hypertext markup data of the target web page, and intercepting the web page image of the target web page; Using a pre-trained image text recognition model, perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and an image area set; According to the preset image-text conflict screening rule and the text block set, the image area set is screened for text areas to obtain a text image area set; Performing text recognition on each body text block in the body text block set to obtain a body text sequence, and performing a text sequence number marking operation on each body text sequence; According to the text sequence number of the text sequence adjacent to the context of each text image region in the text image region set, the hypertext markup data is queried to obtain the text image data; Determining the sentence integrity of each of the body text sequences; According to the incomplete text sequence of the main text, the hypertext markup data is queried to obtain the main text data; The complete sentence body text sequence, the body text data and the body image data are stored in a pre-constructed database.
2. The method for extracting web page text based on images and texts as claimed in claim 1, It is characterized in that The method of filtering the image area set according to the preset image-text conflict screening rule and the text block set to obtain the text image area set includes: According to the preset text-image conflict screening rules, construct the maximum external cut frame of the text block set; Identifying the positional relationship between each image region in the image region set and the maximum circumscribed frame, and extracting a built-in image set from the image region set according to the positional relationship between each image region; The image-text coverage conflict relationship of each built-in image in the built-in image set is identified, and based on the image-text coverage conflict relationship, the built-in images that do not cover the text are extracted from the built-in image set to obtain a text image area set.
3. The method for extracting web page text based on images and texts as claimed in claim 1, It is characterized in that The using a pre-trained picture text recognition model to perform object recognition on the web page image includes: Performing Gaussian filtering on the webpage image to obtain a noise-reduced image; Using a pre-trained picture text recognition model, a feature extraction operation is performed on the denoised image to obtain an image feature sequence; The image feature sequence is subjected to fully connected image-text classification based on the text font and text background to obtain an object recognition result.
4. The method for extracting web page text based on images and texts as claimed in claim 3, It is characterized in that Before performing object recognition on the webpage image using the pre-trained picture text recognition model, the method further includes: Use preset embedding points to capture massive data texts and classify them to obtain image samples and text samples, and mark the image samples and text samples with type labels; The marked image samples and the text samples are mixed as an image and text sample set, and the image and text sample set is randomly grouped according to a preset ratio to obtain a training set and a test set; Extracting a training sample from the training set in turn, and using a pre-built picture and text recognition model to classify the training sample based on pictures and texts to obtain a test category; Using a cross entropy loss algorithm, the loss value between the type label of the training sample and the test category is calculated, and the loss value is minimized to obtain the model parameters of the image text recognition model when the loss value is minimized, and the model parameters are updated inversely according to the gradient descent method to obtain an updated image text recognition model; Determine whether all training samples in the training set participate in the training; When not all the training samples in the training set participate in the training, returning to the above step of extracting one training sample from the training set in turn, and iteratively training the updated picture text recognition model; When all the training samples in the training set participate in the training, the updated image text recognition model is tested using the test set to obtain the test accuracy; Determining whether the test accuracy is greater than or equal to a preset pass threshold; When the test accuracy is less than the qualified threshold, return to the above step of randomly grouping the image and text sample set according to a preset ratio, regroup the image and text sample set, and update the updated image and text recognition model; When the test accuracy is greater than or equal to the qualified threshold, a trained picture text recognition model is obtained.
5. The method for extracting web page text based on images and texts as claimed in claim 1, It is characterized in that The step of crawling the hypertext markup data of the target webpage includes: Using a pre-built request library, a request is sent to the target web page, a status code is fed back by the target web page, and it is determined whether the status code indicates a successful request; When the status code request is successful, the encrypted hypertext markup data of the target webpage is captured according to the response text attribute, and the encrypted hypertext markup data is parsed using a pre-built BeautifulSoup library to obtain decrypted hypertext markup data; The decrypted hypertext markup data is formatted and outputted using a pre-built beautification function to obtain the hypertext markup data of the target web page.
6. The method for extracting web page text based on images and texts as claimed in claim 1, It is characterized in that The step of querying the hypertext markup data according to the text sequence number of the context-adjacent text sequence of each text image region in the text image region set to obtain the text image data includes: Searching the hypertext markup data for the text sequence number of the body text sequence adjacent to the context, obtaining the file data position of the body text block corresponding to the text sequence number of the body text sequence adjacent to the context, and searching for the img tag of the embedded image in the file data position; The hypertext markup data is queried according to the img tag to obtain the text image data.
7. The method for extracting web page text based on images and texts as claimed in claim 1, It is characterized in that The determining of the sentence integrity of each of the body text sequences includes: Identifying the grammatical completeness of the body text sequence to obtain a first completeness score; Identifying the semantic integrity of the body text sequence to obtain a second integrity score; Identifying the contextual knowledge completeness of the text body sequence to obtain a third completeness score; According to a preset weight configuration rule, the first integrity score, the second integrity score and the third integrity score are weightedly calculated to obtain the sentence integrity of the main text sequence.
8. A web page text extraction device based on images and texts, It is characterized in that The device comprises: A data capture module, used to retrieve a target web page according to preset keywords, capture the hypertext markup data of the target web page, and capture the web page image of the target web page; The image text classification module is used to use a pre-trained image text recognition model to perform object recognition on the web page image, and perform a marking and box selection operation on the object recognition result to obtain a body text block set and a picture area set; A text image query module, for performing text area screening on the image area set according to a preset text-image conflict screening rule and the text text block set to obtain a text image area set, performing text recognition on each text block in the text text block set to obtain a text sequence, and performing a text sequence number marking operation on each text sequence, and querying the hypertext markup data according to the text sequence numbers of the text text sequences adjacent to the context of each text image area in the text image area set to obtain text image data; The main text query module is used to determine the sentence integrity of each of the main text sequences, and to query the hypertext markup data based on the main text sequences with incomplete sentences to obtain the main text data, and to store the main text sequences with complete sentences, the main text data and the main text image data in a pre-constructed database.
9. An electronic device, It is characterized in that The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image-based web page text extraction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, the method for extracting web page text based on images and texts as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Training data set generation method and device
CN110147817A
Method and related device for acquiring webpage text content
CN110309392A