Theme extraction method and system for webpage content

Through the SAC-KG language model and NLP large model, the knowledge graph is automatically constructed, and the problem of inaccurate titles of web page content is solved, clear theme refinement results are generated, and user experience and search engine rankings are improved.

CN120354848APending Publication Date: 2025-07-22中科天玑数据科技股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510411054.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the title of the web page content is usually formulated by humans. Due to cultural qualities and personal writing skills, it is difficult to clearly and accurately summarize the core and full text of the web page content.

Method used

The SAC-KG language model and NLP large model are used to automatically build a knowledge graph, extract different types of data from the web pages into text data, generate natural language text describing triples, and compress it into the preset word count range through a generative summary to form clear and accurate theme refinement results.

Benefits of technology

It realizes the automation and accurate theme refinement of web page content, improves user experience and search engine ranking, and enhances the attractiveness of content and the market competitiveness of the website.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354848A_ABST
    Figure CN120354848A_ABST
Patent Text Reader

Abstract

The invention provides a webpage content theme extraction method and system. The method comprises the steps that a crawler obtains original text data, picture data, audio data and / or video data in a webpage; extracting text data based on the picture data, the audio data and / or the video data, and merging the extracted text data with the original text data to obtain comprehensive text data; analyzing the comprehensive text data by using an SAC-KG language model and automatically constructing a knowledge graph, and generating a natural language text describing a triple contained in the knowledge graph by using an NLP large model; the natural language text is reduced to a preset word number range to serve as a generated theme extraction result about webpage content; wherein key frame images and audio data contained in the video data are extracted, the audio data are converted into text data by using a voice recognition model, and the text data contained in the key frame images and / or picture data are extracted based on an OCR technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of natural language processing, and particularly relates to a method and system for theme extraction of web page content. Background Art

[0002] Theme extraction of web page content is a key step in enhancing user experience and information transmission efficiency. In the era of explosive Internet information, users often need to quickly obtain the required information in a short time, and clear and accurate theme extraction can help users quickly understand the core content of the web page, save their browsing time, and improve satisfaction. At the same time, for search engines, a clear theme also helps the web page to obtain a higher ranking in relevant search results, thereby increasing exposure and traffic. Therefore, carefully refining the web page theme can not only enhance the attractiveness and readability of the content, but also effectively improve the overall performance and market competitiveness of the website.

[0003] In the prior art, usually the title of the web page content is directly extracted, but usually this title is manually formulated by the author of the web page content. Limited by cultural qualities and personal writing skills, it may not be able to clearly and accurately summarize the core and full text of the web page content in the title.

[0004] Therefore, how to provide an automated method for extracting web page content to generate a clear and accurate theme extraction result that can summarize the core content of the web page is a technical problem to be solved urgently. Summary of the Invention

[0005] In order to overcome the problems existing in the prior art, this application proposes a method and system for theme extraction of web page content. This solution combines the SAC-KG language model and the NLP large model to achieve accurate and comprehensive theme extraction and generalization.

[0006] In a first aspect, an embodiment of the present invention provides a method for theme extraction of web page content, the method includes: the crawler obtains the original text data, picture data, audio data, and / or video data in the web page; extracts text data based on the picture data, audio data, and / or video data, and merges the extracted text data with the original text data to obtain comprehensive text data; uses the SAC-KG language model to parse the comprehensive text data and automatically construct a knowledge graph, and uses the NLP large model to generate natural language text describing the triples included in the knowledge graph; reduces the natural language text to within a preset number of words as the generated theme extraction result for the web page content; wherein, extracts the key frame images and audio data included in the video data, uses a speech recognition model to convert the audio data into text data, and extracts the text data included in the key frame images and / or picture data based on the OCR technology.

[0007] With this embodiment, different types of data contained in a web page are all converted into or extracted as text data. The SAC-KG language model is used to construct a knowledge graph based on the comprehensive text data. During this process, key entities are extracted as nodes, and the relationships between the nodes are stored in triples. Then, a large NLP model is used to generate natural language text describing the triple structure contained in the knowledge graph. Finally, the natural language text is compressed to within a preset number of words as the theme extraction result.

[0008] In some embodiments of the present invention, nodes for recording the positions of picture data, audio data, and / or video data are preset in the original text data; in the step of merging the extracted text data with the original text data, it includes: adding the text data extracted based on the picture data, audio data, and / or video data to the positions of the corresponding nodes.

[0009] In some embodiments of the present invention, the SAC-KG language model includes a generator, a validator, and a pruner; the step of using the SAC-KG language model to parse the comprehensive text data and automatically construct a knowledge graph includes: inputting the comprehensive text data into the generator of the SAC-KG language model to generate a preliminary knowledge graph; verifying and evaluating the quality of the triples in the preliminary knowledge graph based on preset rules, and filtering out the triples that do not meet the preset rule range; determining the key entities and non-key entities in the filtered first-level knowledge graph, expanding the key entities, and pruning the non-key entities to obtain the final knowledge graph.

[0010] In some embodiments of the present invention, the step of using a large NLP model to generate natural language text describing the triples contained in the knowledge graph includes: extracting multiple subgraphs from the knowledge graph, and extracting the triples contained in the subgraphs; using natural language to express the triples extracted from each subgraph; merging all the natural language expressions, and merging synonymous sentences based on the vector similarity of the text to generate the natural language text.

[0011] In some embodiments of the present invention, the step of reducing the natural language text to within a preset number of words as the generated theme extraction result of the web page content includes: using the method of generative summarization to compress the number of words of the natural language text, setting the preset number of words range as the target generated number of words range, and generating the theme extraction result of the web page content.

[0012] In some embodiments of the present invention, the step of converting audio data into text data using a speech recognition model includes: preprocessing the audio data by including noise reduction processing, silence excision, and audio format conversion; extracting feature data of the preprocessed audio data using the method of perceptual linear prediction; and generating corresponding text data based on the feature data using a pre-trained Transformer model for speech recognition.

[0013] In some embodiments of the present invention, the method further includes: using a large NLP model to extract the text data included in the key frame images and / or picture data, and correcting the text data extracted based on the OCR technology using the text data extracted by the large NLP model.

[0014] In some embodiments of the present invention, the data structure type for storing the text data is a string, a text file, an array, a list, or an inverted index.

[0015] In a second aspect, an embodiment of the present application provides a system for refining the theme of web page content, including a processor, a memory, and a computer program / instruction stored on the memory. The processor is configured to execute the computer program / instruction, and when the computer program / instruction is executed, the system implements the steps described in the above embodiments.

[0016] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the steps described in the above embodiments are implemented.

[0017] Additional advantages, objects, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objects and other advantages of the present invention can be pointed out and obtained specifically in the description and the accompanying drawings.

[0018] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. Description of the Drawings

[0019] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure.

[0020] In the drawings:

[0021] Figure 1 It is a flowchart of the method for refining the theme of web page content in an embodiment of the present invention.

[0022] Figure 2 This is a schematic diagram of the triple structure in the knowledge graph in an embodiment of the present invention.

[0023] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Detailed implementation manners

[0024] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0025] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0026] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0027] In the prior art, usually the title of the web page content is directly extracted, but usually this title is artificially formulated by the author of the web page content. Limited by cultural qualities and personal writing skills, it may not be able to clearly and accurately summarize the core and full text of the web page content in the title.

[0028] Therefore, the present invention designs an automated method for refining web page content to generate a theme refinement result that can clearly and accurately summarize the core content of the web page.

[0029] Figure 1 This is a flowchart of the method for refining the theme of web page content in an embodiment of the present invention. The method includes the following steps:

[0030] Step S110: The crawler obtains the original text data, picture data, audio data and / or video data in the web page.

[0031] Step S120: Extract text data based on the picture data, audio data, and / or video data, and merge the extracted text data with the original text data to obtain comprehensive text data.

[0032] Step S130: Parse the comprehensive text data using the SAC-KG language model and automatically construct a knowledge graph, and use the NLP large model to generate natural language text describing the triples included in the knowledge graph.

[0033] Step S140: Reduce the natural language text to within a preset number of words as the refined result of the theme of the web page content generated.

[0034] Among them, extract the key frame images and audio data included in the video data, use a speech recognition model to convert the audio data into text data, and extract the text data included in the key frame images and / or picture data based on OCR technology.

[0035] Adopting the embodiment of the present invention, different types of data included in a web page can be converted into or extracted as text data, and the SAC-KG language model is used to construct a knowledge graph based on the comprehensive text data. In this process, key entities are extracted as nodes, and the relationships between the nodes are stored in triples. Then, the NLP large model is used to generate natural language text describing the triple structure included in the knowledge graph, and finally, the natural language text is compressed to within a preset number of words as the refined result of the theme.

[0036] In some embodiments of the present invention, nodes for recording the positions of picture data, audio data, and / or video data are preset in the original text data.

[0037] Further, in the step of merging the extracted text data with the original text data in Step S120, it includes: adding the text data extracted based on the picture data, audio data, and / or video data to the positions of the corresponding nodes.

[0038] Adopting this embodiment, all the content included in the web page can be converted into text data for subsequent processing based on the NLP large model.

[0039] In some embodiments of the present invention, the SAC-KG language model includes a generator, a validator, and a pruner.

[0040] Further, the step of parsing the comprehensive text data using the SAC-KG language model and automatically constructing a knowledge graph in step S130 includes: (1) inputting the comprehensive text data into the generator of the SAC-KG language model to generate a preliminary knowledge graph; (2) verifying and evaluating the quality of the triples in the preliminary knowledge graph based on preset rules, and filtering out the triples that do not meet the scope of the preset rules; (3) determining the key entities and non-key entities in the filtered first-level knowledge graph, expanding the key entities, and pruning the non-key entities to obtain the final knowledge graph.

[0041] By adopting this implementation manner, automatic generation of a knowledge graph based on the SAC-KG language model can be realized to further refine the key path in the comprehensive text data. The key path identification of the target can be achieved through weighted calculation.

[0042] Figure 2 It is a schematic diagram of the triple structure in the knowledge graph in an embodiment of the present invention. Figure 2 The "entity-relationship-entity" triple structure is proposed, and the triple structure included in the key path is transformed into a natural language expression, and then polished and word count calibrated by the nlp model to obtain the output theme refinement result.

[0043] In some embodiments of the present invention, the step of using the NLP large model to generate a natural language text describing the triples included in the knowledge graph in step S130 includes: (1) extracting multiple subgraphs from the knowledge graph and extracting the triples included in the subgraphs; (2) using natural language to express the triples extracted from each subgraph; (3) merging all the natural language expressions and merging synonymous sentences based on the vector similarity of the texts to generate the natural language text.

[0044] By adopting this implementation manner, the NLP large model can be used to generate natural language to describe the triple structure in the key path, so as to generate a theme that may be slightly longer in length, and after compression, the theme refinement result of the web page content with the word count length meeting the requirements is obtained.

[0045] In some embodiments of the present invention, the step of reducing the natural language text to within a preset word count range as the generated theme refinement result of the web page content includes: using the method of generative summarization to compress the word count of the natural language text, setting the preset word count range as the target generated word count range, and generating the theme refinement result of the web page content.

[0046] Further, the generative summary model can be trained with the original text and compressed text on a large data scale.

[0047] By adopting this implementation manner, the generated theme extraction result can be compressed within a preset number of words.

[0048] In some embodiments of the present invention, the step of converting audio data into text data using a speech recognition model includes: (1) preprocessing the audio data including noise reduction processing, silence excision, and audio format conversion; (2) extracting feature data of the preprocessed audio data using the method of perceptual linear prediction; (3) generating corresponding text data based on the feature data using a pre-trained Transformer model for speech recognition.

[0049] In some embodiments of the present invention, the method further includes: using an NLP large model to extract the text data included in the key frame images and / or picture data, and correcting the text data extracted based on the OCR technology using the text data extracted by the NLP large model.

[0050] In some embodiments of the present invention, the data structure type for storing the text data is a string, a text file, an array, a list, or an inverted index.

[0051] In a second aspect, an embodiment of the present application provides a theme extraction system for web page content. The device includes a computer device, the computer device includes a processor and a memory, computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps implemented by the theme extraction system for web page content.

[0052] In a third aspect, an embodiment of the present application provides a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the above-mentioned theme extraction system for web page content is implemented.

[0053] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device is a hardware entity for implementing the theme extraction system for web page content.

[0054] As Figure 3 shown, an embodiment of the present application provides an electronic device. The device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the above-mentioned theme extraction system for web page content is implemented.

[0055] The electronic device may include a processor 1201 and a memory 1202 storing computer program instructions.

[0056] Specifically, the above-mentioned processor 1201 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0057] The memory 1202 may include a mass storage device for data or instructions. By way of example and not limitation, the memory 1202 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 1202 may include removable or non-removable (or fixed) media. In a suitable case, the memory 1202 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 1202 is a non-volatile solid-state memory.

[0058] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0059] The processor 1201 reads and executes the computer program instructions stored in the memory 1202 to implement any one of the methods for determining battery thermal runaway parameters in the above embodiments.

[0060] In an example, the electronic device may further include a communication interface 1203 and a bus 1210. Among them, as Figure 3 shown, the processor 1201, the memory 1202, and the communication interface 1203 are connected through the bus 1210 and complete communication with each other.

[0061] The communication interface 1203 is mainly used to implement communication between the modules, devices, units, and / or devices in the embodiments of the present application.

[0062] Bus 1210 includes hardware, software, or both, and couples components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, bus 1210 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0063] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0064] It should also be noted that the functional blocks shown in the above-described block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an Application Specific Integrated Circuit (ASIC), appropriate firmware, a plug-in, a function card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, Erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, Radio Frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0065] It also needs to be noted that in the exemplary embodiments mentioned in the present application, some methods or systems are described based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0066] As described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0067] As described above, the above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present application.

Claims

1. A method for extracting the theme of web page content, characterized in that, The method includes: A crawler obtains original text data, image data, audio data, and / or video data in a web page; Based on the image data, audio data, and / or video data, text data is extracted, and the extracted text data is merged with the original text data to obtain comprehensive text data; The SAC-KG language model is used to parse the comprehensive text data and automatically construct a knowledge graph, and an NLP large model is used to generate natural language text describing the triples included in the knowledge graph; The natural language text is reduced to within a preset number of words as the generated topic refinement result of the web page content; Among them, key frame images and audio data included in the video data are extracted, the audio data is converted into text data using a speech recognition model, and the text data included in the key frame images and / or image data is extracted based on OCR technology.

2. The method according to claim 1, wherein Nodes for recording the positions of image data, audio data, and / or video data are preset in the original text data; In the step of merging the extracted text data with the original text data, it includes: adding the text data extracted based on the image data, audio data, and / or video data to the positions of the corresponding nodes.

3. The method according to claim 1, characterized in that, The SAC-KG language model includes a generator, a validator, and a pruning unit; The step of using the SAC-KG language model to parse the comprehensive text data and automatically construct a knowledge graph includes: Inputting the comprehensive text data into the generator of the SAC-KG language model to generate a preliminary knowledge graph; Based on preset rules, the quality of the triples in the preliminary knowledge graph is verified and evaluated, and the triples that do not meet the preset rule range are filtered out; Determine the key entities and non-key entities in the filtered first-level knowledge graph, expand the key entities, and perform pruning processing on the non-key entities to obtain the final knowledge graph.

4. The method according to claim 1, wherein The step of using the NLP large model to generate natural language text describing the triples included in the knowledge graph includes: Extract multiple subgraphs from the knowledge graph and extract the triples included in the subgraphs; Use natural language to express the triples extracted from each subgraph; Merge all the natural language expressions, and merge the synonymous sentences based on the vector similarity of the text to generate the natural language text.

5. The method according to claim 1, characterized in that The step of reducing the natural language text to within a preset number of words as the generated topic refinement result of the web page content includes: Compress the number of words of the natural language text in a generative summary manner, set the preset number of words as the target generated number of words range, and generate the topic refinement result of the web page content.

6. The method according to claim 1, characterized in that, The step of using the speech recognition model to convert the audio data into text data includes: Preprocess the audio data in advance, including noise reduction processing, silence excision, and audio format conversion; Use the perceptually linear prediction method to extract the feature data of the preprocessed audio data; Use a pre-trained Transformer model for speech recognition to generate corresponding text data based on the feature data.

7. The method according to claim 1, characterized in that The method further includes: using an NLP large model to extract text data included in key frame images and / or picture data, and correcting the text data extracted based on OCR technology by using the text data extracted by the NLP large model.

8. The method according to claim 1, characterized in that, The data structure type for storing the text data is a string, a text file, an array, a list, or an inverted index.

9. A theme extraction system for web page content, comprising a processor, a memory, and computer programs / instructions stored on the memory, characterized in that, The processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Methods, devices, computer equipment, and storage media for generating audio and video mind maps

    CN121351943B